Software Engineer, Trainium
Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
About the job
Responsibilities
- Build and optimize OpenAI's inference stack for AWS Trainium.
- Develop high-performance kernels for critical model operations and workloads.
- Extend and improve compiler support to efficiently target Trainium hardware.
- Build systems to execute and optimize model forward passes on Trainium.
- Profile workloads and identify bottlenecks across kernels, compiler-generated code, runtime, and model execution.
- Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
- Work across the hardware/software boundary to unlock performance from specialized AI accelerators.
- Own complex performance and systems problems end-to-end, from investigation through production deployment.
Requirements
- 3+ years of relevant engineering experience, ideally in ML systems, compilers, kernels, runtimes, or performance engineering.
- Strong systems programming fundamentals and experience writing performance-critical software.
- Experience with GPU, TPU, Trainium, or other specialized accelerator architectures.
- Ability to reason about performance across hardware, kernels, compilers, and ML frameworks.
- Ability to own technically ambiguous problems end-to-end and learn new hardware and software domains.
Nice-to-Haves
- AWS Trainium or AWS Neuron SDK experience.
- Contributions to PyTorch or JAX.
- Experience with LLVM, MLIR, XLA, or Triton.
- Experience developing kernels for specialized accelerators.
Skills
Aws Trainium, Aws Neuron Sdk, Systems Programming, GPU, Tpu, Performance Engineering, Compilers, Kernels, PyTorch, JAX, Llvm, Mlir, Xla, Triton, Inference Systems
Similar jobs
ML Engineering jobsBuild and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.
Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.