TPU Kernel Engineer
Designs and optimizes TPU kernels to address performance issues in ML research, training, and inference systems. Provides feedback on model impacts and solves large-scale systems problems, requiring deep accelerator expertise.
About the job
You may be a good fit if you:
- Have significant experience optimizing ML systems for TPUs, GPUs, or other accelerators
- Are results-oriented, with a bias towards flexibility and impact
- Pick up slack, even if it goes outside your job description
- Enjoy pair programming (we love to pair!)
- Want to learn more about machine learning research
- Care about the societal impacts of your work
Strong candidates may also have experience with:
- High performance, large-scale ML systems
- Designing and implementing kernels for TPUs or other ML accelerators
- Understanding accelerators at a deep level, e.g. a background in computer architecture
- ML framework internals
- Language modeling with transformers
Representative projects:
- Implement low-latency, high-throughput sampling for large language models
- Adapt existing models for low-precision inference
- Build quantitative models of system performance
- Design and implement custom collective communication algorithms
- Debug kernel performance at the assembly level
Logistics
Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
Annual Salary: $280,000 — $850,000 USD
Skills
Tpu, GPU, Kernel Optimization, Ml Systems, Computer Architecture, Ml Frameworks, Transformers, Collective Communication, Low-Precision Inference, Assembly
Similar jobs
ML Engineering jobsBuild and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.