# Staff+ Software Engineer, ML Inference Path

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $320k – $485k/yr
**Experience:** 7+ years
**Skills:** Python, PyTorch, TensorFlow, JAX, Distributed Systems, ML Infrastructure, A/B Testing, Deployment Pipelines, Monitoring, Data Drift, LLMs, Transformers, Inference Optimization, Automated Testing, Rollback Systems
**Posted:** 2026-09-09

> Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

## Job Description

## Responsibilities
- Design and build scalable ML infrastructure for real-time safety deployments across classifier and model ecosystems.
- Build monitoring and observability tools for classifier performance, data quality, and system health.
- Collaborate with research teams to productionize safety research and translate experimental techniques into robust, scalable systems.
- Optimize inference latency and throughput for real-time safety evaluations while maintaining reliability.
- Implement automated testing, deployment, and rollback systems for production ML models.
- Partner with Safeguards, Security, and Alignment teams to deliver infrastructure meeting safety and production requirements.
- Develop internal tools and frameworks that accelerate safety research and deployment.

## Requirements
- Proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or JAX.
- Understanding of distributed systems principles and experience building high-throughput, low-latency systems.
- Experience building automated or self-service deployment pipelines and evaluation infrastructure for independent researcher rollouts.
- Experience implementing A/B testing frameworks and experimentation infrastructure for ML systems.
- Results-oriented approach with a focus on reliability and impact in safety-critical systems.
- Strong collaboration skills and interest in translating research into production systems.
- Care about AI safety and its societal impacts.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- 5+ years of experience building production ML infrastructure, ideally in safety-critical domains such as fraud detection, content moderation, or risk assessment.
- Experience with large language models and modern transformer architectures.
- Experience developing monitoring and alerting systems for ML model performance and data drift.
- Experience in trust and safety, fraud prevention, or content moderation.
- Knowledge of privacy-preserving ML techniques and compliance requirements.

## Compensation
- Annual salary: $320,000–$485,000 USD.

## Similar jobs

- [Staff Applied Scientist](https://hotfix.jobs/jobs/d518c7a7-07d4-4f06-af79-9c478df769ca) - Garner Health - New York, NY - $300k – $390k/yr
- [Staff Machine Learning Operations Engineer](https://hotfix.jobs/jobs/b5d46280-8221-4b6e-b0b2-033cc6f2c8d7) - Garner Health - New York, NY - $298k – $351k/yr
- [Senior Staff Machine Learning Systems Engineer, Ads ML Platform](https://hotfix.jobs/jobs/efc443bf-db64-4fa3-b123-eb399c454a93) - Reddit - Remote - $293k – $410k/yr
- [Staff Machine Learning Engineer, Fraud & Abuse](https://hotfix.jobs/jobs/2068a66a-3fd2-42f6-ab53-d199be305035) - Square - Remote - $277k – $415k/yr
- [Senior Staff Machine Learning Engineer, Feed Relevance](https://hotfix.jobs/jobs/a2f9b178-c6cf-49bf-ba93-00688d1027fb) - Reddit - Remote - $266k – $372k/yr

**Apply:** https://hotfix.jobs/jobs/90cf7427-87de-41e1-af88-0d4e1187b94b
**Canonical:** https://hotfix.jobs/jobs/90cf7427-87de-41e1-af88-0d4e1187b94b