Staff Software Engineer, Inference
Build and optimize large-scale distributed inference infrastructure serving Claude across diverse AI accelerators and cloud platforms. The role requires significant software engineering experience with distributed systems; expertise in LLM inference, Kubernetes, cloud infrastructure, Python, or Rust is advantageous.
About the job
Responsibilities
- Build and maintain critical systems serving Claude to millions of users worldwide.
- Address infrastructure blockers across the inference stack, from intelligent request routing to fleet-wide orchestration across diverse AI accelerators.
- Maximize compute efficiency while enabling high-performance inference infrastructure for AI research.
- Design intelligent routing algorithms that optimize request distribution across thousands of accelerators.
- Autoscale compute fleets to match supply and demand across production, research, and experimental workloads.
- Build production-grade deployment pipelines for releasing new models.
- Integrate new AI accelerator platforms.
- Contribute to inference features such as structured sampling and prompt caching.
- Support inference for new model architectures.
- Analyze observability data to tune performance based on production workloads.
- Manage multi-region deployments and geographic routing.
Requirements
- Significant software engineering experience, particularly with distributed systems.
- Experience with high-performance, large-scale distributed systems.
- Experience implementing and deploying machine learning systems at scale.
- Results-oriented approach with flexibility and a focus on impact.
- Willingness to take on work outside the formal job description.
- Interest in machine learning systems and infrastructure.
- Ability to thrive in environments where technical excellence drives business and research outcomes.
- Care about the societal impacts of the work.
- Bachelor's degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- Load balancing, request routing, or traffic management systems.
- LLM inference optimization, batching, and caching strategies.
- Kubernetes and cloud infrastructure, including AWS or GCP.
- Python or Rust.
- Multi-accelerator deployments.
Compensation and benefits
- Annual salary: €295,000–€355,000 EUR.
- Hybrid policy requiring staff to work from an office at least 25% of the time; some roles may require more.
- Visa sponsorship may be available.
- Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Skills
Distributed Systems, Machine Learning Systems, Llm Inference, Load Balancing, Request Routing, Traffic Management, Kubernetes, AWS, GCP, Python, Rust, Caching, Autoscaling, Multi-Accelerator Deployments, Observability
Similar jobs
Backend Engineering jobsLeads architecture and technical strategy for Vanta’s government compliance product, building scalable backend systems and new workflows in a 0-to-1 environment. Requires 10+ years of backend experience, strong systems thinking, cross-functional influence, and a track record of mentoring engineers.
Leads the design and delivery of shared distributed-platform capabilities for cost intelligence, operational tooling, reliability, and developer experience. The role requires deep backend and systems expertise with Java, Go, Kafka, and Kubernetes, plus technical leadership across multiple engineering teams.
Leads the architecture, implementation, and operation of MongoDB’s large-scale observability infrastructure, covering metrics, logs, traces, and alerts. Requires 10+ years building highly concurrent distributed systems, strong systems programming and performance expertise, and technical leadership across teams.
Build and lead large-scale distributed backend features for MongoDB Atlas and Atlas Data Federation & Archiving. The role requires 10+ years of software development experience, compiled-language expertise, cloud-provider experience, and production ownership.
Technical leader for a security-focused team rearchitecting MongoDB server ingress networking. The role designs and operates high-performance Rust services, drives technical strategy, handles production incidents, and mentors engineers; it requires 10+ years of systems software experience and strong networking and security fundamentals.