Skip to content
AnthropicAnthropic

Staff+ Software Engineer, ML Inference Path

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

About the job

Responsibilities

  • Design and build scalable ML infrastructure for real-time safety deployments across classifier and model ecosystems.
  • Build monitoring and observability tools for classifier performance, data quality, and system health.
  • Collaborate with research teams to productionize safety research and translate experimental techniques into robust, scalable systems.
  • Optimize inference latency and throughput for real-time safety evaluations while maintaining reliability.
  • Implement automated testing, deployment, and rollback systems for production ML models.
  • Partner with Safeguards, Security, and Alignment teams to deliver infrastructure meeting safety and production requirements.
  • Develop internal tools and frameworks that accelerate safety research and deployment.

Requirements

  • Proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or JAX.
  • Understanding of distributed systems principles and experience building high-throughput, low-latency systems.
  • Experience building automated or self-service deployment pipelines and evaluation infrastructure for independent researcher rollouts.
  • Experience implementing A/B testing frameworks and experimentation infrastructure for ML systems.
  • Results-oriented approach with a focus on reliability and impact in safety-critical systems.
  • Strong collaboration skills and interest in translating research into production systems.
  • Care about AI safety and its societal impacts.
  • Bachelor’s degree or equivalent combination of education, training, and experience in a relevant field.

Nice-to-haves

  • 5+ years of experience building production ML infrastructure, ideally in safety-critical domains such as fraud detection, content moderation, or risk assessment.
  • Experience with large language models and modern transformer architectures.
  • Experience developing monitoring and alerting systems for ML model performance and data drift.
  • Experience in trust and safety, fraud prevention, or content moderation.
  • Knowledge of privacy-preserving ML techniques and compliance requirements.

Compensation

  • Annual salary: $320,000–$485,000 USD.

Skills

Python, PyTorch, TensorFlow, JAX, Distributed Systems, ML Infrastructure, A/B Testing, Deployment Pipelines, Monitoring, Data Drift, LLMs, Transformers, Inference Optimization, Automated Testing, Rollback Systems

Garner Health

Garner Health

New York, NY

Staff Applied Scientist
$300k+/yrHybrid7+ YOEML Engineering

Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Reddit

Reddit

United States

Senior Staff Machine Learning Systems Engineer, Ads ML Platform
$293k+/yrRemote8+ YOEML Engineering

Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.

Square

Square

San Francisco, CA

Staff Machine Learning Engineer, Fraud & Abuse
$277k+/yrRemote12+ YOEML Engineering

Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.

Reddit

Reddit

United States

Senior Staff Machine Learning Engineer, Feed Relevance
$266k+/yrRemote10+ YOEML Engineering

Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.