# Performance Engineer, Inference Engine

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA, New York, NY
**Role:** ML Engineering
**Salary:** $350k – $850k/yr
**Skills:** Rust, C++, Llm Inference, Gpu Programming, Accelerator Programming, Operating Systems, Transformers, Allocators, Caching, Schedulers, Rdma, Property-Based Testing
**Posted:** 2026-09-09

> Build and optimize a high-scale LLM inference engine spanning accelerator programming, host-device coordination, and distributed systems. The role requires strong systems programming, performance analysis, and an understanding of LLM inference across compute, memory, and interconnects.

## Job Description

## Responsibilities
- Build and optimize Anthropic’s inference engine across accelerator and cloud platforms.
- Improve throughput, cost, reliability, and latency at scale.
- Keep accelerator utilization high by reducing host, device, and systems overhead.
- Optimize model-state caching and reuse versus recomputation.
- Measure and profile systems, model performance bottlenecks, test hypotheses, implement changes, and validate improvements.
- Build observability and infrastructure that preserve model quality and support reproducible systems.
- Collaborate with safeguards and safety teams to maintain robust production inference.
- Pair program and contribute beyond the immediate job scope.

## Requirements
- Working mental model of LLM inference, including prefill and decode behavior across accelerator compute, memory, interconnect, and host systems.
- Strong systems programming skills in Rust, C++, or a similar language.
- Analytical, measurement-driven approach to performance optimization.
- Demonstrated ability to learn unfamiliar, complex systems quickly and ship consequential changes.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- Experience with an LLM serving engine.
- GPU or accelerator programming.
- Operating-system internals.
- Transformer-based language modeling.
- Experience building allocators, caches, schedulers, or high-bandwidth transports.
- Fluency in Rust.
- Experience with determinism, replay, and property-based testing.

## Compensation
- Annual salary: **$350,000–$850,000 USD**.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [AI Infrastructure Engineer](https://hotfix.jobs/jobs/0d5aa4cf-a861-427b-8d8f-cba8ac94104e) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, General Agents](https://hotfix.jobs/jobs/e075e229-9ae8-46ba-92a5-20a0a6d4f0db) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, RL Scaling](https://hotfix.jobs/jobs/bd1b548c-7c82-4627-84a9-149150dad06d) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Machine Learning Engineer, Multimodal Perception and Authentication](https://hotfix.jobs/jobs/26314e6e-423f-4a3f-b400-ee0b56429f64) - OpenAI - San Francisco, CA - $342k – $399k/yr

**Apply:** https://hotfix.jobs/jobs/adc277be-252d-483e-9d49-cbc5716291f9
**Canonical:** https://hotfix.jobs/jobs/adc277be-252d-483e-9d49-cbc5716291f9