Software Engineer, Model Runtime
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
About the job
Responsibilities
- Design and implement the LLM inference runtime for frontier models running on custom silicon.
- Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
- Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
- Optimize latency, throughput, memory efficiency, and hardware utilization across model architectures and serving workloads.
- Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks.
- Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
- Create profiling, observability, benchmarking, and performance-modeling tools.
- Debug correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
- Translate workload insights into requirements for future silicon and system architecture.
Requirements
- Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- Experience building or optimizing runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
- Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- Ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
- Experience profiling and debugging performance across multiple layers of a hardware-software stack.
- Ability to design clean abstractions while retaining low-level control for specialized hardware.
- Ability to work across model, systems, compiler, kernel, and hardware teams on ambiguous technical problems.
- Focus on correctness, observability, reliability, maintainability, and graceful behavior at scale.
- Candidates may need to meet legal status requirements under U.S. export control laws and regulations.
Compensation
- ATS-listed salary range: $266,000–$445,000 annually.
Skills
C++, Rust, Python, Llm Inference, Distributed Systems, Compilers, Kernels, Continuous Batching, Kv Cache, Model Parallelism, Memory Management, Performance Profiling, Model Serving, Custom Silicon
Similar jobs
ML Engineering jobsBuild research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.
Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.