Staff + Senior Software Engineer, Scaling
Build and scale distributed inference infrastructure serving Claude across accelerators and cloud providers. The role requires substantial software engineering experience, with expertise in large-scale systems, routing, orchestration, and production machine learning infrastructure.
About the job
Responsibilities
- Design, build, and maintain distributed systems serving Claude to millions of users worldwide.
- Develop resilient systems that adapt in real time to operational events.
- Build intelligent request routing, load balancing, and traffic management across accelerators and cloud providers.
- Improve compute efficiency and fleet cost through autoscaling and workload orchestration.
- Build and operate production-grade deployment pipelines for releasing new models.
- Provide high-performance inference infrastructure for model research.
- Integrate new AI accelerator platforms and support new model architectures.
- Analyze observability data to tune production performance.
- Manage multi-region deployments and geographic routing.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Results-oriented, flexible, and impact-focused approach.
- Willingness to take on work outside the formal job description.
- Interest in machine learning systems and infrastructure.
- Ability to thrive where technical excellence drives business and research outcomes.
- Care for the societal impacts of the work.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Experience with load balancing, request routing, or traffic management.
- Familiarity with LLM inference optimization, batching, and caching strategies.
- Experience with Kubernetes and cloud infrastructure.
- Proficiency in Python or Rust.
Compensation and benefits
- Annual salary: $320,000–$485,000 USD.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office collaboration space.
Skills
Distributed Systems, Machine Learning Systems, Load Balancing, Request Routing, Traffic Management, Llm Inference, Kubernetes, AWS, GCP, Microsoft Azure, Python, Rust, Autoscaling, Caching, Observability
Similar jobs
Backend Engineering jobsDesign and operate high-QPS backend systems on Claude’s token-generation path, owning latency, reliability, safe deployments, and incident response. The role requires 8+ years of software engineering experience, strong distributed-systems expertise, and production ownership of mission-critical services.
This Senior Staff Software Engineer will define and lead the architecture of EarnIn’s Identity & Trust platform, spanning authentication, authorization, verification, privacy, fraud prevention, and secure data processing. The role requires 8+ years of software engineering experience and deep expertise in secure distributed backend systems.
Sets architectural direction and remains hands-on in building ultra-low-latency institutional trading systems spanning matching, risk, order management, and market data. Requires 10+ years of software engineering experience and deep expertise in distributed systems and backend performance.
Leads architecture and technical strategy for Vanta’s government compliance product, building scalable backend systems and new workflows in a 0-to-1 environment. Requires 10+ years of backend experience, strong systems thinking, cross-functional influence, and a track record of mentoring engineers.
Senior Staff Software Engineer defining the technical vision and architecture for FinHub’s financial ledger and money-movement platform. The role requires 12+ years of backend distributed-systems experience and deep expertise in high-consequence transaction systems.