# Engineering Manager, Inference Infrastructure

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA, New York, NY, Seattle, WA
**Role:** Engineering Management
**Salary:** $405k – $625k/yr
**Experience:** 5+ years
**Skills:** Load Balancing, Scheduling, Cluster Orchestration, Autoscaling, Distributed Systems, High-Performance Networking, Llm Inference, Kubernetes, Service Meshes, Kv Caching, Continuous Batching, Request Scheduling, Multi-Cloud, Accelerator Fleets, Incident Response
**Posted:** 2026-09-01

> Leads teams building and operating the control plane for Anthropic’s large-scale inference fleet, improving routing, capacity, performance, reliability, and cost. Requires deep production-systems expertise, engineering management experience, and strong cross-functional leadership.

## Job Description

## Responsibilities
- Own the technical roadmap for coordinating the inference fleet, including traffic routing, capacity placement, cache placement, demand responsiveness, and control-plane/inference-engine synchronization.
- Partner with product, inference engine, performance, and capacity teams to identify throughput, latency, utilization, and cost improvements and deliver measurable results.
- Establish quantitative modeling practices for evaluating system changes and expected impact.
- Set technical strategy for control-plane evolution across heterogeneous hardware, multiple cloud providers, and serving surfaces.
- Lead operational practices including on-call rotations, incident response, postmortems, and deployment safety.
- Coordinate across API, inference engine, capacity planning, and cloud deployment teams.
- Develop, retain, and hire high-caliber engineers; coach engineers and shape team structure as scope grows.
- Unblock critical initiatives and synthesize design decisions when needed.

## Requirements
- Experience managing engineering teams responsible for critical-path production infrastructure at scale.
- Deep systems expertise in areas such as load balancing, scheduling, cluster orchestration, autoscaling, cache-coherent distributed state, or high-performance networking.
- Experience shipping performance or efficiency improvements in large-scale systems and quantifying latency, cost, and other impacts.
- Experience operating production infrastructure with on-call, incident response, capacity events, and deployment discipline.
- Results-oriented approach and comfort balancing throughput, latency, cost, stability, timelines, and feature velocity.
- Strong cross-functional collaboration and relationship-building skills.
- Curiosity about machine-learning systems and transformer inference.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- 5+ years of engineering management experience.
- Experience with LLM inference serving, including KV caching, continuous batching, request scheduling, or prefill/decode disaggregation.
- Experience with cluster schedulers, autoscalers, load balancers, service meshes, or fleet control planes at scale, including Kubernetes internals or equivalent systems.
- Experience operating workloads across multiple clouds or partner platforms.
- Familiarity with heterogeneous accelerator fleets and hardware-dependent workload placement and rollout sequencing.
- Experience leading teams at supercomputing or hyperscaler infrastructure scale.
- Experience leading multiple teams or groups through rapid growth, hiring, onboarding, and team restructuring.

## Compensation
- Annual salary: **$405,000–$625,000 USD**

## Similar jobs

- [Engineering Manager, ChatGPT Search Infrastructure](https://hotfix.jobs/jobs/8e74e886-d9d3-4e13-9609-23da6400dcb3) - OpenAI - San Francisco, CA - $401k – $445k/yr
- [Engineering Manager, Client Platform Engineering](https://hotfix.jobs/jobs/5f2c78d0-1072-48ee-8a17-0daa662d561a) - OpenAI - San Francisco, CA - $401k – $445k/yr
- [Engineering Manager, Artifacts](https://hotfix.jobs/jobs/2d300ca6-e7a9-4a12-bb65-66e373f0cae1) - OpenAI - San Francisco, CA - $347k – $405k/yr
- [Manager, Engineering](https://hotfix.jobs/jobs/79720669-7b3d-48cb-952e-292b11eed5f2) - Garner Health - New York, NY - $310k – $385k/yr
- [Engineering Manager](https://hotfix.jobs/jobs/c9006d33-4b7c-4cce-9af7-898b64c08af4) - Rain - New York, NY - $300k – $400k/yr

**Apply:** https://hotfix.jobs/jobs/c7fb9f4a-0b98-48a5-b88d-496ae7e51214
**Canonical:** https://hotfix.jobs/jobs/c7fb9f4a-0b98-48a5-b88d-496ae7e51214