Performance Engineer, Inference Engine
Build and optimize a high-scale LLM inference engine spanning accelerator programming, host-device coordination, and distributed systems. The role requires strong systems programming, performance analysis, and an understanding of LLM inference across compute, memory, and interconnects.
About the job
Responsibilities
- Build and optimize Anthropic’s inference engine across accelerator and cloud platforms.
- Improve throughput, cost, reliability, and latency at scale.
- Keep accelerator utilization high by reducing host, device, and systems overhead.
- Optimize model-state caching and reuse versus recomputation.
- Measure and profile systems, model performance bottlenecks, test hypotheses, implement changes, and validate improvements.
- Build observability and infrastructure that preserve model quality and support reproducible systems.
- Collaborate with safeguards and safety teams to maintain robust production inference.
- Pair program and contribute beyond the immediate job scope.
Requirements
- Working mental model of LLM inference, including prefill and decode behavior across accelerator compute, memory, interconnect, and host systems.
- Strong systems programming skills in Rust, C++, or a similar language.
- Analytical, measurement-driven approach to performance optimization.
- Demonstrated ability to learn unfamiliar, complex systems quickly and ship consequential changes.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- Experience with an LLM serving engine.
- GPU or accelerator programming.
- Operating-system internals.
- Transformer-based language modeling.
- Experience building allocators, caches, schedulers, or high-bandwidth transports.
- Fluency in Rust.
- Experience with determinism, replay, and property-based testing.
Compensation
- Annual salary: $350,000–$850,000 USD.
Skills
Rust, C++, Llm Inference, Gpu Programming, Accelerator Programming, Operating Systems, Transformers, Allocators, Caching, Schedulers, Rdma, Property-Based Testing
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.
Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.
Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.
Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.