# Software Engineer, Model Runtime

**Company:** [OpenAI](https://hotfix.jobs/companies/openai)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $266k – $445k/yr
**Skills:** C++, Rust, Python, Llm Inference, Distributed Systems, Compilers, Kernels, Continuous Batching, Kv Cache, Model Parallelism, Memory Management, Performance Profiling, Model Serving, Custom Silicon
**Posted:** 2026-08-24

> Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.

## Job Description

## Responsibilities
- Design and implement the LLM inference runtime for frontier models running on custom silicon.
- Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
- Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
- Optimize latency, throughput, memory efficiency, and hardware utilization across model architectures and serving workloads.
- Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks.
- Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
- Create profiling, observability, benchmarking, and performance-modeling tools.
- Debug correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
- Translate workload insights into requirements for future silicon and system architecture.

## Requirements
- Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- Experience building or optimizing runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
- Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- Ability to reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
- Experience profiling and debugging performance across multiple layers of a hardware-software stack.
- Ability to design clean abstractions while retaining low-level control for specialized hardware.
- Ability to work across model, systems, compiler, kernel, and hardware teams on ambiguous technical problems.
- Focus on correctness, observability, reliability, maintainability, and graceful behavior at scale.
- Candidates may need to meet legal status requirements under U.S. export control laws and regulations.

## Compensation
- ATS-listed salary range: $266,000–$445,000 annually.

## Similar jobs

- [Software Engineer, AI for Chip Design](https://hotfix.jobs/jobs/abfb017d-1ce2-4d24-901d-16084eb7b3bc) - OpenAI - San Francisco, CA - $266k – $468k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/22efd49a-a0e5-48d7-aea3-8a7bdd7d7337) - Hyperbound - San Francisco, CA - $260k – $300k/yr
- [Applied AI Engineer, Beneficial Deployments](https://hotfix.jobs/jobs/c38bca51-e1c4-4f65-8790-b09d52d8169c) - Anthropic - San Francisco, CA - $280k – $320k/yr
- [Software Engineer, Trainium](https://hotfix.jobs/jobs/5e1a8341-f0e9-45c8-8492-12782b38f079) - OpenAI - San Francisco, CA - $295k – $380k/yr
- [Applied Scientist III](https://hotfix.jobs/jobs/00c9850b-9666-4032-aa7c-5fd2253c5ef7) - Garner Health - New York, NY - $236k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/7f549613-3d04-42b8-b691-8ed7c2d9ecd1
**Canonical:** https://hotfix.jobs/jobs/7f549613-3d04-42b8-b691-8ed7c2d9ecd1