Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.
Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.
Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.
405k – 485k/yr
Hybrid10+ YOEDevOps / SRE
Staff Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +1
Staff-level IC role owning end-to-end CI/CD, merge queue, and deploy pipelines for Anthropic's engineering org. Focus on AI-assisted review, test reliability, and progressive delivery at monorepo scale.
405k – 485k/yr
Hybrid7+ YOEDevOps / SRE
Staff Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +1
Staff-level engineer to own end-to-end development environments at Anthropic, focusing on container lifecycle, cold-start optimization, environment isolation, and pre-push validation for AI researchers and engineers.
405k – 485k/yr
Hybrid7+ YOEDevOps / SRE
Staff Software Engineer, Node Infra
AnthropicSan Francisco, CA +2
Own technical strategy and roadmap for node lifecycle management, health automation, and scaling AI clusters across clouds and accelerators. Requires deep distributed systems expertise, ML accelerator experience, and 12+ years leading complex multi-team infrastructure initiatives.
405k – 485k/yr
Hybrid12+ YOEDevOps / SRE
Staff Software Engineer, Kubernetes Platform
AnthropicSan Francisco, CA +2
Senior-level engineer to own and scale Anthropic's massive Kubernetes control plane and scheduler for training frontier AI models across hundreds of thousands of nodes. Requires deep Kubernetes internals experience and 12+ years building production distributed systems.
Own technical strategy and roadmap for agent-driven cluster lifecycle management across cloud providers and datacenters. Lead complex multi-quarter infrastructure initiatives and mentor engineers on large-scale compute systems.
405k – 485k/yr
Hybrid12+ YOEDevOps / SRE
Performance Engineer, Inference Systems
AnthropicSan Francisco, CA +2
Performance engineer focused on cross-layer investigations of Anthropic's inference fleet for Claude, optimizing throughput, latency, reliability, and correctness while building observability and partnering with kernel and serving teams.
350k – 850k/yr
HybridDevOps / SRE
Staff+ Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +2
Leads technical strategy and builds scalable developer infrastructure including build systems, CI/CD pipelines, and tooling for large monorepo environments. Requires 3+ years leading complex projects, proficiency in Python/Rust/Go, and experience with container orchestration.
405k – 625k/yr
HybridDevOps / SRE
Incident Response Manager - Product & Engineering
AnthropicNew York, NY +2
Leads incident response operations for product and engineering, serving as on-call commander to coordinate cross-functional teams, manage communications, and improve processes during high-stakes incidents. Requires 5+ years in incident management with technical depth in infrastructure and cloud systems.
290k – 365k/yr
Hybrid5+ YOEDevOps / SRE
Staff Engineer, Datacenter Server Lifecycle
AnthropicSan Francisco, CA +1
Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.
320k – 405k/yr
Hybrid8+ YOEDevOps / SRE
Research Engineer, RL Infrastructure and Reliability (Knowledge Work)
AnthropicSan Francisco, CA
Owns reliability, observability, and infrastructure for Knowledge Work team's RL training environments and evaluations. Ensures stability at scale through proactive hardening, SLOs, load testing, and incident response for ML systems.
350k – 850k/yr
HybridDevOps / SRE
Staff+ Software Engineer, Platform
AnthropicSan Francisco, CA +2
Staff-level software engineer builds and scales platform infrastructure across teams, including dev tools, service infra, multicloud, auth, connectivity, API distributability, and ML adaptation systems. Requires 8+ years full-stack experience with Staff leadership, focusing on robust, scalable solutions in fast-paced AI environment.
405k – 485k/yr
Hybrid8+ YOEDevOps / SRE
Performance Engineer
AnthropicSan Francisco, CA +2
Performance Engineer optimizes throughput and robustness of large-scale ML distributed systems by solving novel performance issues. Requires significant software engineering experience at supercomputing scale and interest in ML.
280k – 850k/yr
HybridDevOps / SRE
Search
Location
15 jobs
Job results
Staff+ Software Engineer, Caching
AnthropicSan Francisco, CA +2
Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.
Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.
Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.
405k – 485k/yr
Hybrid10+ YOEDevOps / SRE
Staff Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +1
Staff-level IC role owning end-to-end CI/CD, merge queue, and deploy pipelines for Anthropic's engineering org. Focus on AI-assisted review, test reliability, and progressive delivery at monorepo scale.
405k – 485k/yr
Hybrid7+ YOEDevOps / SRE
Staff Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +1
Staff-level engineer to own end-to-end development environments at Anthropic, focusing on container lifecycle, cold-start optimization, environment isolation, and pre-push validation for AI researchers and engineers.
405k – 485k/yr
Hybrid7+ YOEDevOps / SRE
Staff Software Engineer, Node Infra
AnthropicSan Francisco, CA +2
Own technical strategy and roadmap for node lifecycle management, health automation, and scaling AI clusters across clouds and accelerators. Requires deep distributed systems expertise, ML accelerator experience, and 12+ years leading complex multi-team infrastructure initiatives.
405k – 485k/yr
Hybrid12+ YOEDevOps / SRE
Staff Software Engineer, Kubernetes Platform
AnthropicSan Francisco, CA +2
Senior-level engineer to own and scale Anthropic's massive Kubernetes control plane and scheduler for training frontier AI models across hundreds of thousands of nodes. Requires deep Kubernetes internals experience and 12+ years building production distributed systems.
Own technical strategy and roadmap for agent-driven cluster lifecycle management across cloud providers and datacenters. Lead complex multi-quarter infrastructure initiatives and mentor engineers on large-scale compute systems.
405k – 485k/yr
Hybrid12+ YOEDevOps / SRE
Get new-job notifications on iOS
Hotfix on iOS
Get a push summary when new jobs match your saved alerts.
Performance Engineer, Inference Systems
AnthropicSan Francisco, CA +2
Performance engineer focused on cross-layer investigations of Anthropic's inference fleet for Claude, optimizing throughput, latency, reliability, and correctness while building observability and partnering with kernel and serving teams.
350k – 850k/yr
HybridDevOps / SRE
Staff+ Software Engineer, Developer Productivity
AnthropicSan Francisco, CA +2
Leads technical strategy and builds scalable developer infrastructure including build systems, CI/CD pipelines, and tooling for large monorepo environments. Requires 3+ years leading complex projects, proficiency in Python/Rust/Go, and experience with container orchestration.
405k – 625k/yr
HybridDevOps / SRE
Incident Response Manager - Product & Engineering
AnthropicNew York, NY +2
Leads incident response operations for product and engineering, serving as on-call commander to coordinate cross-functional teams, manage communications, and improve processes during high-stakes incidents. Requires 5+ years in incident management with technical depth in infrastructure and cloud systems.
290k – 365k/yr
Hybrid5+ YOEDevOps / SRE
Staff Engineer, Datacenter Server Lifecycle
AnthropicSan Francisco, CA +1
Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.
320k – 405k/yr
Hybrid8+ YOEDevOps / SRE
Research Engineer, RL Infrastructure and Reliability (Knowledge Work)
AnthropicSan Francisco, CA
Owns reliability, observability, and infrastructure for Knowledge Work team's RL training environments and evaluations. Ensures stability at scale through proactive hardening, SLOs, load testing, and incident response for ML systems.
350k – 850k/yr
HybridDevOps / SRE
Staff+ Software Engineer, Platform
AnthropicSan Francisco, CA +2
Staff-level software engineer builds and scales platform infrastructure across teams, including dev tools, service infra, multicloud, auth, connectivity, API distributability, and ML adaptation systems. Requires 8+ years full-stack experience with Staff leadership, focusing on robust, scalable solutions in fast-paced AI environment.
405k – 485k/yr
Hybrid8+ YOEDevOps / SRE
Performance Engineer
AnthropicSan Francisco, CA +2
Performance Engineer optimizes throughput and robustness of large-scale ML distributed systems by solving novel performance issues. Requires significant software engineering experience at supercomputing scale and interest in ML.