Skip to content
AnthropicAnthropic

Staff + Senior Software Engineer, Cloud Inference Launch Engineering

Build and own validation pipelines, CI/CD infrastructure, and platform integrations to launch frontier models and inference features reliably across AWS, GCP, and Azure. Requires strong large-scale distributed systems experience and track record improving release velocity.

About the job

Key Responsibilities

  • Be on the critical path for frontier model launches, bringing up inference for new model architectures and shipping them to cloud platforms in lockstep with our first-party platform.
  • Work with the core inference team to bring new inference features (e.g. structured sampling, prompt caching) to cloud platforms, owning the platform-specific integration that gets them to production.
  • Identify and dive deep on the gaps that make inference behave differently across first-party and CSPs (config drift, observability, deployment patterns, hard cross-platform bugs) and fix them at the source.
  • Design, build, and own the CI/CD infrastructure for the inference server and load balancer across cloud platforms, with shadow traffic, performance baselines (throughput and latency), and correctness checks that catch regressions before production.
  • Drive down merge-to-production cycle time by making validation faster, more parallel, and cost-effective enough to run on the same constrained accelerator pool that serves customers, without trading away reliability.
  • Analyze observability data across providers to identify performance bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads.

Minimum Qualifications

  • Strong interest in LLM serving (prior inference or ML experience not required).
  • Significant software engineering experience with a strong background in high-performance, large-scale distributed systems serving millions of users.
  • Track record of building automation or test infrastructure that measurably improved release velocity or reliability.
  • Experience building or operating services on at least one major cloud platform (AWS, GCP, or Azure), with exposure to Kubernetes, Infrastructure as Code, or container orchestration.
  • Thrive in cross-functional collaboration with both internal teams and external partners.
  • Fast learner who can quickly ramp up on new technologies, hardware platforms, and provider ecosystems.
  • Highly autonomous and take ownership of problems end-to-end.

Preferred Qualifications

  • LLM inference optimization, batching, and caching strategies.
  • Capacity-constrained scheduling or shared-resource test infrastructure.
  • Solid understanding of multi-region deployments, request routing, load balancing, global traffic management.
  • Working with CSP partner teams to scale infrastructure across multiple platforms, navigating differences in networking, security, privacy, and managed services.
  • Proficiency in Python or Rust.

Compensation and Benefits

  • Annual compensation range: $320,000–$485,000 USD (total compensation).
  • Minimum education: Bachelor’s degree or equivalent.
  • Location-based hybrid policy: expect staff to be in office at least 25% of the time.
  • Competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours.

Skills

Distributed Systems, Kubernetes, Infrastructure As Code, Python, Rust, AWS, GCP, Azure, CI/CD, Llm Inference, Load Balancing, Observability

Anthropic

Anthropic

San Francisco, CA

Staff+ Software Engineer, ML Inference Path
$320k+/yrHybrid7+ YOEML Engineering

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

Garner Health

Garner Health

New York, NY

Staff Applied Scientist
$300k+/yrHybrid7+ YOEML Engineering

Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Reddit

Reddit

United States

Senior Staff Machine Learning Systems Engineer, Ads ML Platform
$293k+/yrRemote8+ YOEML Engineering

Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.

Square

Square

San Francisco, CA

Staff Machine Learning Engineer, Fraud & Abuse
$277k+/yrRemote12+ YOEML Engineering

Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.