Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
320k – 485k/yr
HybridDevOps / SRE
About the role
Responsibilities
Design, build, and own shared infrastructure for Interpretability, including research environments, data systems, and compute tooling.
Lead cross-team efforts with agentic engineering, security, compute, and storage platform teams so company-wide solutions serve research needs.
Discover and resolve major organization-wide developer experience issues.
Help take interpretability methods from research code to dependable audit pipelines.
Requirements
High proficiency in at least one programming language, such as Python, Rust, Go, or Java, and productivity with Python.
Significant experience building and operating secure, scalable software infrastructure, including cloud systems, distributed systems, or developer tooling.
Strong cross-functional communication skills, working effectively with researchers, platform teams, and security teams.
Curiosity about unfamiliar domains and interpretability research’s role in AI safety.
Ability to prioritize impactful work, operate with ambiguity, and question assumptions.
Care for the societal impacts and ethics of technical work.
Nice-to-Haves
Experience with cloud infrastructure, including Google Cloud or AWS, Kubernetes, networking, and infrastructure as code.
Security engineering experience, including identity, authentication, access management, sandboxing, or red teaming.
Experience with data warehousing, large-scale storage systems, and data lifecycle management.
Experience with compute schedulers and accelerator fleet management.
Experience building developer productivity tooling and observability stacks.
Experience building tooling to accelerate research teams.
Representative Projects
Design and stand up a hardened research environment for experimentation with frontier model weights.
Build lifecycle management for petabytes of research data, including visibility, retention, and cost efficiency.
Build self-service scheduling and capacity tooling.
Create observability that catches infrastructure regressions before they cost researchers valuable time.
Compensation and Benefits
Annual salary: $320,000–$485,000 USD.
Minimum education: Bachelor’s degree or an equivalent combination of education, training, and experience.
The role is based in the San Francisco office, with exceptional remote candidates considered case by case.
Staff are currently expected to work from one of the company’s offices at least 25% of the time.
Visa sponsorship may be available depending on the role and candidate.
Skills
PythonRustGoJavaGCPAWSKubernetesNetworkingInfrastructure As Codeidentity and access managementsandboxingData Warehousingstorage systemsObservability
Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.
310k – 460k/yrOn-siteDevOps / SRE
Workload Porting & Performance Engineer
OpenAISan Francisco, CA +1
Evaluates new hardware platforms by porting benchmarks and workloads, analyzes performance across compute/memory/networking, identifies bottlenecks, and optimizes for AI systems. Requires expertise in performance analysis, system architecture, and debugging across hardware/software boundaries.
342k – 555k/yrHybridDevOps / SRE
3P Architect
OpenAISan Francisco, CA +1
Defines rack- and cluster-level reference architectures for AI infrastructure, translates workload requirements into designs, collaborates with partners and modeling teams to evaluate tradeoffs, and drives vendor roadmaps to address technology gaps.
342k – 555k/yrHybridDevOps / SRE
Performance & Systems Engineer, Codex
OpenAISan Francisco, CA
Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.
295k – 445k/yrHybridDevOps / SRE
Systems Integration Engineer, Build Systems | Consumer Devices
OpenAISan Francisco, CA
Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.