# Staff + Senior Software Engineer, Inference

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** Unspecified
**Role:** Backend Engineering
**Experience:** 7+ years
**Skills:** Distributed Systems, Machine Learning Systems, Llm Inference, Load Balancing, Request Routing, Traffic Management, Kubernetes, AWS, GCP, Azure, Python, Rust, Autoscaling, Caching, Observability
**Posted:** 2026-08-28

> Build and operate high-performance distributed inference infrastructure serving Claude at global scale, including routing, orchestration, autoscaling, deployment pipelines, and accelerator integration. The role requires significant software engineering experience with distributed systems; experience in large-scale ML infrastructure is preferred.

## Job Description

## Responsibilities
- Design, build, and maintain distributed systems serving Claude to millions of users.
- Develop resilient systems that adapt to real-world events in real time.
- Build intelligent request routing, load balancing, and traffic management across thousands of accelerators.
- Maximize fleet compute efficiency through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure for next-generation model development.
- Integrate AI accelerator platforms and support inference for new model architectures.
- Analyze observability data to tune performance using production workloads.
- Manage multi-region deployments and geographic routing.

## Requirements
- Significant software engineering experience, particularly with distributed systems.
- Flexibility, strong ownership, and a results-oriented approach.
- Willingness to work across responsibilities and learn machine learning systems and infrastructure.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure, including AWS, Google Cloud, or Azure.
- Proficiency in Python or Rust.

## Similar jobs

- [Staff Backend Engineer, Database Automation](https://hotfix.jobs/jobs/ba3b028d-95ae-4f7f-9f85-ea5553f732a9) - GitLab - Remote - $153k – $259k/yr
- [Staff Backend Engineer - Grafana App Platform](https://hotfix.jobs/jobs/b9fae390-8767-4142-a84f-8ea457c4580f) - Grafana Labs - Remote - CA$186k – CA$224k/yr
- [Staff Backend Engineer](https://hotfix.jobs/jobs/127bb501-4fd5-435d-92d9-d47c9bf13f75) - GitLab - Remote - $153k – $259k/yr
- [Senior Staff Software Engineer, Finhub](https://hotfix.jobs/jobs/1951727b-0f94-41ad-a9e8-bd7a603a152d) - Coinbase - Remote - $254k – $299k/yr
- [Staff Software Engineer, Backend](https://hotfix.jobs/jobs/0971f827-acba-45c6-b582-7ed16106c518) - Nango - Remote - $140k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/59f7201d-51ad-430c-b235-0239ae7227b0
**Canonical:** https://hotfix.jobs/jobs/59f7201d-51ad-430c-b235-0239ae7227b0