Skip to content
AnthropicAnthropic

Engineering Manager, Inference Infrastructure

Leads teams building and operating the control plane for Anthropic’s large-scale inference fleet, improving routing, capacity, performance, reliability, and cost. Requires deep production-systems expertise, engineering management experience, and strong cross-functional leadership.

About the job

Responsibilities

  • Own the technical roadmap for coordinating the inference fleet, including traffic routing, capacity placement, cache placement, demand responsiveness, and control-plane/inference-engine synchronization.
  • Partner with product, inference engine, performance, and capacity teams to identify throughput, latency, utilization, and cost improvements and deliver measurable results.
  • Establish quantitative modeling practices for evaluating system changes and expected impact.
  • Set technical strategy for control-plane evolution across heterogeneous hardware, multiple cloud providers, and serving surfaces.
  • Lead operational practices including on-call rotations, incident response, postmortems, and deployment safety.
  • Coordinate across API, inference engine, capacity planning, and cloud deployment teams.
  • Develop, retain, and hire high-caliber engineers; coach engineers and shape team structure as scope grows.
  • Unblock critical initiatives and synthesize design decisions when needed.

Requirements

  • Experience managing engineering teams responsible for critical-path production infrastructure at scale.
  • Deep systems expertise in areas such as load balancing, scheduling, cluster orchestration, autoscaling, cache-coherent distributed state, or high-performance networking.
  • Experience shipping performance or efficiency improvements in large-scale systems and quantifying latency, cost, and other impacts.
  • Experience operating production infrastructure with on-call, incident response, capacity events, and deployment discipline.
  • Results-oriented approach and comfort balancing throughput, latency, cost, stability, timelines, and feature velocity.
  • Strong cross-functional collaboration and relationship-building skills.
  • Curiosity about machine-learning systems and transformer inference.
  • Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

Nice-to-haves

  • 5+ years of engineering management experience.
  • Experience with LLM inference serving, including KV caching, continuous batching, request scheduling, or prefill/decode disaggregation.
  • Experience with cluster schedulers, autoscalers, load balancers, service meshes, or fleet control planes at scale, including Kubernetes internals or equivalent systems.
  • Experience operating workloads across multiple clouds or partner platforms.
  • Familiarity with heterogeneous accelerator fleets and hardware-dependent workload placement and rollout sequencing.
  • Experience leading teams at supercomputing or hyperscaler infrastructure scale.
  • Experience leading multiple teams or groups through rapid growth, hiring, onboarding, and team restructuring.

Compensation

  • Annual salary: $405,000–$625,000 USD

Skills

Load Balancing, Scheduling, Cluster Orchestration, Autoscaling, Distributed Systems, High-Performance Networking, Llm Inference, Kubernetes, Service Meshes, Kv Caching, Continuous Batching, Request Scheduling, Multi-Cloud, Accelerator Fleets, Incident Response

OpenAI

OpenAI

San Francisco, CA

Engineering Manager, ChatGPT Search Infrastructure
$401k+/yrOn-site5+ YOEEngineering Management

Leads the engineering team responsible for ChatGPT Search Infrastructure, setting technical direction for scalable, low-latency search platforms and integrations. Requires experience managing senior engineers and deep expertise in distributed systems, online serving, experimentation, and AI-powered products.

OpenAI

OpenAI

San Francisco, CA
Engineering Manager, Client Platform Engineering
$401k+/yrHybrid5+ YOEEngineering Management

Leads the client platform engineering organization responsible for secure, reliable endpoint services across major operating systems. The role combines people management, platform strategy, architecture oversight, operational excellence, and cross-functional partnership.

OpenAI

OpenAI

San Francisco, CA

Engineering Manager, Artifacts
$347k+/yrHybrid5+ YOEEngineering Management

Leads and grows the engineering team building AI-native artifacts such as documents, spreadsheets, slides, and dashboards. The role sets technical direction across product, infrastructure, rendering, storage, reliability, and model integration while partnering closely with research and product teams.

Garner Health

Garner Health

New York, NY

Manager, Engineering
$310k+/yrHybrid5+ YOEEngineering Management

Leads the Data Platform Engineering team, building scalable batch and streaming infrastructure for complex healthcare data while guiding architecture, operations, and team development. Requires 5+ years of engineering experience and 2+ years managing engineering teams.

Rain

Rain

New York, NY

Engineering Manager
$300k+/yrHybrid5+ YOEEngineering Management

Leads and develops senior backend and platform engineers while owning architecture, reliability, and technical standards for systems moving money globally. The role requires at least five years of leadership-level experience, current hands-on engineering, distributed systems expertise, and high-stakes payments or fintech experience.