Skip to content
OpenAIOpenAI

Software Engineer, Trainium

Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.

About the job

Responsibilities

  • Build and optimize OpenAI's inference stack for AWS Trainium.
  • Develop high-performance kernels for critical model operations and workloads.
  • Extend and improve compiler support to efficiently target Trainium hardware.
  • Build systems to execute and optimize model forward passes on Trainium.
  • Profile workloads and identify bottlenecks across kernels, compiler-generated code, runtime, and model execution.
  • Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
  • Work across the hardware/software boundary to unlock performance from specialized AI accelerators.
  • Own complex performance and systems problems end-to-end, from investigation through production deployment.

Requirements

  • 3+ years of relevant engineering experience, ideally in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Strong systems programming fundamentals and experience writing performance-critical software.
  • Experience with GPU, TPU, Trainium, or other specialized accelerator architectures.
  • Ability to reason about performance across hardware, kernels, compilers, and ML frameworks.
  • Ability to own technically ambiguous problems end-to-end and learn new hardware and software domains.

Nice-to-Haves

  • AWS Trainium or AWS Neuron SDK experience.
  • Contributions to PyTorch or JAX.
  • Experience with LLVM, MLIR, XLA, or Triton.
  • Experience developing kernels for specialized accelerators.

Skills

Aws Trainium, Aws Neuron Sdk, Systems Programming, GPU, Tpu, Performance Engineering, Compilers, Kernels, PyTorch, JAX, Llvm, Mlir, Xla, Triton, Inference Systems

Anthropic

Anthropic

San Francisco, CA
Applied AI Engineer, Beneficial Deployments
$280k+/yrHybridML Engineering

Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Model Runtime
$266k+/yrHybridML Engineering

Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.

Hyperbound

Hyperbound

San Francisco, CA

Machine Learning Engineer
$260k+/yrOn-siteML Engineering

Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.

OpenAI

OpenAI

San Francisco, CA

Machine Learning Engineer, Multimodal Perception and Authentication
$342k+/yrHybridML Engineering

Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.