Staff+ Software Engineer, ML Inference Path
Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
About the job
Responsibilities
- Design and build scalable ML infrastructure for real-time safety deployments across classifier and model ecosystems.
- Build monitoring and observability tools for classifier performance, data quality, and system health.
- Collaborate with research teams to productionize safety research and translate experimental techniques into robust, scalable systems.
- Optimize inference latency and throughput for real-time safety evaluations while maintaining reliability.
- Implement automated testing, deployment, and rollback systems for production ML models.
- Partner with Safeguards, Security, and Alignment teams to deliver infrastructure meeting safety and production requirements.
- Develop internal tools and frameworks that accelerate safety research and deployment.
Requirements
- Proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or JAX.
- Understanding of distributed systems principles and experience building high-throughput, low-latency systems.
- Experience building automated or self-service deployment pipelines and evaluation infrastructure for independent researcher rollouts.
- Experience implementing A/B testing frameworks and experimentation infrastructure for ML systems.
- Results-oriented approach with a focus on reliability and impact in safety-critical systems.
- Strong collaboration skills and interest in translating research into production systems.
- Care about AI safety and its societal impacts.
- Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.
Nice-to-haves
- 5+ years of experience building production ML infrastructure, ideally in safety-critical domains such as fraud detection, content moderation, or risk assessment.
- Experience with large language models and modern transformer architectures.
- Experience developing monitoring and alerting systems for ML model performance and data drift.
- Experience in trust and safety, fraud prevention, or content moderation.
- Knowledge of privacy-preserving ML techniques and compliance requirements.
Compensation
- Annual salary: $320,000–$485,000 USD.
Skills
Python, PyTorch, TensorFlow, JAX, Distributed Systems, ML Infrastructure, A/B Testing, Deployment Pipelines, Monitoring, Data Drift, LLMs, Transformers, Inference Optimization, Automated Testing, Rollback Systems
Similar jobs
ML Engineering jobsLeads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.
Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.