Skip to content
AnthropicAnthropic

Staff Software Engineer, Inference

Build and operate high-performance distributed inference infrastructure serving Claude across large-scale accelerator fleets. The role requires strong software engineering experience with production distributed systems, Kubernetes, cloud platforms, and machine learning infrastructure.

About the job

Responsibilities

  • Design, build, and maintain distributed systems that serve Claude to millions of users worldwide.
  • Develop resilient, flexible systems that adapt in real time to real-world events.
  • Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
  • Maximize compute efficiency across the fleet through autoscaling and orchestration of production, research, and experimental workloads.
  • Build and operate production-grade deployment pipelines for releasing new models to users.
  • Provide high-performance inference infrastructure for next-generation model development.
  • Integrate new AI accelerator platforms and support inference for new model architectures.

Minimum Qualifications

  • Proficiency in Python or Rust.
  • Software engineering experience building and operating distributed systems in production.
  • Working knowledge of containerized infrastructure, such as Kubernetes, and at least one major cloud platform: AWS, GCP, or Azure.
  • Results-oriented, flexible, and impact-focused.
  • Willingness to take on work beyond the formal job description.
  • Desire to learn more about machine learning systems and infrastructure.
  • Ability to thrive where technical excellence drives business results and research breakthroughs.
  • Care about the societal impacts of the work.

Preferred Qualifications

  • Significant experience with high-performance, large-scale distributed systems.
  • Experience implementing and deploying machine learning systems at scale.
  • Experience building load balancing, request routing, or traffic management systems.
  • Familiarity with LLM inference optimization, batching, and caching strategies.
  • Deep experience operating Kubernetes and cloud infrastructure at scale.
  • Experience with AI accelerator platforms, including GPUs, TPUs, or emerging hardware.

Representative Projects

  • Design intelligent routing algorithms to optimize request distribution across accelerators and environments.
  • Autoscale the compute fleet to match supply with demand across production, research, and experimental workloads.
  • Build production-grade deployment pipelines for reliably releasing new models to millions of users.
  • Contribute to new inference features and support inference for new model architectures.
  • Analyze observability data to tune performance based on real-world production workloads.
  • Manage multi-region deployments and geographic routing for global customers.

Compensation and Benefits

  • Annual salary: £325,000–£390,000 GBP.
  • Competitive compensation and benefits.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Office collaboration space.
  • Minimum education: bachelor’s degree or an equivalent combination of education, training, and experience.
  • Visa sponsorship is available for eligible roles and candidates.

Skills

Python, Rust, Distributed Systems, Kubernetes, AWS, GCP, Microsoft Azure, Load Balancing, Request Routing, Autoscaling, Llm Inference, Machine Learning Systems, Gpus, Tpus, Caching

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Datadog

Datadog

Bordeaux, France
Senior AI Engineer – Notebooks
No salary listedHybrid6+ YOEML Engineering

Build and operate AI-powered, customer-facing workflows for Datadog Notebooks, combining reliable backend systems with LLM capabilities. The role requires 6+ years of engineering experience, Go or Python expertise, and experience delivering production AI products.

MongoDB

MongoDB

Cork, Ireland
Senior Data Scientist
No salary listedHybrid5+ YOEML Engineering

Research, prototype, and ship statistical and machine learning features that improve MongoDB’s fleet stability, release safety, resource efficiency, and operational automation. The role requires 5+ years of hands-on ML development, strong Python and systems-design skills, and a master’s degree or equivalent quantitative experience.

Reddit

Reddit

United Kingdom
Senior Machine Learning Engineer, Ads Foundational Representations
No salary listedRemote5+ YOEML Engineering

Build and deploy embedding, sequence, and language-model representations for Reddit Ads, taking ML projects from requirements and experimentation through production. The role requires 5+ years of end-to-end industry ML experience, with expertise in NLP or computer vision and deep-learning frameworks.

Hudl

Hudl

Barcelona, Spain
Senior MLOps Engineer - Edge
No salary listedRemote5+ YOEML Engineering

Build and operate edge MLOps infrastructure for smart-camera machine-learning systems, including model deployment, TensorRT compilation, fleet updates, telemetry, and reliability. The role requires production MLOps experience, embedded inference optimization, and strong collaboration with data-science and embedded-engineering teams.