Skip to content
AnthropicAnthropic

Staff + Senior Software Engineer, Inference Deployment

Build and maintain large-scale distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and deployment pipelines across diverse AI accelerators while maximizing compute efficiency for production and research workloads. Requires significant distributed systems experience.

About the job

Key Responsibilities

  • Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide.
  • Develop resilient, flexible systems that adapt in real time to real world events.
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
  • Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
  • Build and operate production-grade deployment pipelines for releasing new models to users.
  • Provide high-performance inference infrastructure that enables researchers to develop next-generation models.
  • Integrate new AI accelerator platforms and support inference for new model architectures.

Minimum Qualifications

  • Significant software engineering experience, particularly with distributed systems.
  • Results-oriented, with a bias towards flexibility and impact.
  • Willingness to pick up slack, even if it goes outside your job description.
  • Desire to learn more about machine learning systems and infrastructure.
  • Thrive in environments where technical excellence directly drives both business results and research breakthroughs.
  • Care about the societal impacts of your work.

Preferred Qualifications

  • Experience with high-performance, large-scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience with load balancing, request routing, or traffic management systems.
  • Familiarity with LLM inference optimization, batching, and caching strategies.
  • Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
  • Proficiency in Python or Rust.

Representative Projects

  • Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments.
  • Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
  • Building production-grade deployment pipelines for releasing new models to millions of users reliably.
  • Contributing to new inference features.
  • Supporting inference for new model architectures.
  • Analyzing observability data to tune performance based on real-world production workloads.
  • Managing multi-region deployments and geographic routing for global customers.

Education

  • Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role.

Skills

Distributed Systems, Kubernetes, Python, Rust, AWS, GCP, Azure, Load Balancing, Request Routing, Llm Inference, Machine Learning Systems

Anthropic

Anthropic

San Francisco, CA

Staff+ Software Engineer, ML Inference Path
$320k+/yrHybrid7+ YOEML Engineering

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

Garner Health

Garner Health

New York, NY

Staff Applied Scientist
$300k+/yrHybrid7+ YOEML Engineering

Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Reddit

Reddit

United States

Senior Staff Machine Learning Systems Engineer, Ads ML Platform
$293k+/yrRemote8+ YOEML Engineering

Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.

Square

Square

San Francisco, CA

Staff Machine Learning Engineer, Fraud & Abuse
$277k+/yrRemote12+ YOEML Engineering

Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.