Staff + Senior Software Engineer, Inference
Build and operate high-performance distributed inference infrastructure serving Claude at global scale, including routing, orchestration, autoscaling, deployment pipelines, and accelerator integration. The role requires significant software engineering experience with distributed systems; experience in large-scale ML infrastructure is preferred.
About the job
Responsibilities
- Design, build, and maintain distributed systems serving Claude to millions of users.
- Develop resilient systems that adapt to real-world events in real time.
- Build intelligent request routing, load balancing, and traffic management across thousands of accelerators.
- Maximize fleet compute efficiency through autoscaling and orchestration of production, research, and experimental workloads.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure for next-generation model development.
- Integrate AI accelerator platforms and support inference for new model architectures.
- Analyze observability data to tune performance using production workloads.
- Manage multi-region deployments and geographic routing.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Flexibility, strong ownership, and a results-oriented approach.
- Willingness to work across responsibilities and learn machine learning systems and infrastructure.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure, including AWS, Google Cloud, or Azure.
- Proficiency in Python or Rust.
Skills
Distributed Systems, Machine Learning Systems, Llm Inference, Load Balancing, Request Routing, Traffic Management, Kubernetes, AWS, GCP, Azure, Python, Rust, Autoscaling, Caching, Observability
Similar jobs
Backend Engineering jobsBuild and operate Go-based automation and a PostgreSQL-as-a-service platform for high-throughput, always-on production systems. The role requires deep PostgreSQL production experience, backend development expertise, infrastructure automation, and staff-level technical leadership.
Staff Backend Engineer responsible for evolving Grafana into a scalable, multi-tenant observability application platform. The role requires production SaaS experience, distributed-systems expertise, strong backend coding skills, and familiarity with or willingness to learn Golang.
Provides cross-team technical leadership for large backend initiatives, modular architecture, production operations, and AI-assisted engineering adoption. The role requires deep backend architecture experience, monolith modernization expertise, strong written influence, and mentoring ability.
Senior Staff Software Engineer defining the technical vision and architecture for FinHub’s financial ledger and money-movement platform. The role requires 12+ years of backend distributed-systems experience and deep expertise in high-consequence transaction systems.
Staff Backend Software Engineer responsible for designing and scaling developer infrastructure, APIs, and backend services while shaping product direction and collaborating directly with customers. Requires extensive backend experience, Node.js proficiency, and a strong focus on developer experience.