Build and maintain large-scale distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and deployment pipelines across diverse AI accelerators while maximizing compute efficiency for production and research workloads. Requires significant distributed systems experience.
320k – 485k/yr
Hybrid7+ YOEML Engineering
About the role
Key Responsibilities
Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide.
Develop resilient, flexible systems that adapt in real time to real world events.
Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators.
Maximize compute efficiency across the fleet by autoscaling and orchestrating production, research, and experimental workloads.
Build and operate production-grade deployment pipelines for releasing new models to users.
Provide high-performance inference infrastructure that enables researchers to develop next-generation models.
Integrate new AI accelerator platforms and support inference for new model architectures.
Minimum Qualifications
Significant software engineering experience, particularly with distributed systems.
Results-oriented, with a bias towards flexibility and impact.
Willingness to pick up slack, even if it goes outside your job description.
Desire to learn more about machine learning systems and infrastructure.
Thrive in environments where technical excellence directly drives both business results and research breakthroughs.
Care about the societal impacts of your work.
Preferred Qualifications
Experience with high-performance, large-scale distributed systems.
Experience implementing and deploying machine learning systems at scale.
Experience with load balancing, request routing, or traffic management systems.
Familiarity with LLM inference optimization, batching, and caching strategies.
Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure).
Proficiency in Python or Rust.
Representative Projects
Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments.
Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads.
Building production-grade deployment pipelines for releasing new models to millions of users reliably.
Contributing to new inference features.
Supporting inference for new model architectures.
Analyzing observability data to tune performance based on real-world production workloads.
Managing multi-region deployments and geographic routing for global customers.
Education
Bachelor’s degree or an equivalent combination of education, training, and/or experience in a field relevant to the role.
Skills
Distributed SystemsKubernetesPythonRustAWSGCPAzureload balancingrequest routingllm inferencemachine learning systems
Build and own validation pipelines, CI/CD infrastructure, and platform integrations to launch frontier models and inference features reliably across AWS, GCP, and Azure. Requires strong large-scale distributed systems experience and track record improving release velocity.
320k – 485k/yr
Hybrid7+ YOEML Engineering
Staff Software Engineer, Inference
AnthropicSan Francisco, CA +2
Build and maintain distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and high-performance infrastructure across diverse AI accelerators.
320k – 485k/yr
Hybrid7+ YOEML Engineering
Staff Software Engineer, Machine Learning
AttentiveNew York, NY +1
Builds, scales, and operates production-grade ML systems for real-time personalization on Attentive's platform. Requires 6+ years experience with Python, PyTorch/TensorFlow, and scalable ML pipelines in a fast-paced environment.
320k – 360k/yr
Remote6+ YOEML Engineering
Member of Technical Staff
xAIPalo Alto, CA
Member of Technical Staff at xAI performing cutting-edge AI research and hands-on engineering to advance large language models, quantitative reasoning, and language understanding. Requires Bachelor's in CS/ML-related field plus 2 years experience with big data tools, Kubernetes, deep learning, NLP/CV/speech, and languages like Rust/C++/Python.
324k – 396k/yr
On-site2+ YOEML Engineering
Member of Technical Staff
xAIPalo Alto, CA
Hands-on technical leader building and scaling large language models and AI systems. Requires 3-5+ years of AI/ML experience with strong Python and deep learning frameworks.