Senior Staff engineer building CI/CD pipelines, developer workflows, and internal platforms to accelerate engineering velocity for OpenAI's consumer device and cloud software. Requires 10+ years experience including 5+ years in dev productivity or platform engineering.
347k – 385k/yr
Hybrid10+ YOEDevOps / SRE
About the role
Responsibilities
Design, build, and operate CI/CD systems and software delivery pipelines for software that runs both on device and in the cloud.
Lead the design and architecture of internal platform capabilities across build, test, deployment, workflow automation, and developer tooling.
Make hands-on technical decisions about platform design, abstraction boundaries, and system tradeoffs based on the team’s current stage, scale, and operational needs.
Improve developer productivity by shortening feedback loops across build, test, debugging, environment setup, and release-adjacent workflows.
Build self-serve paved-road workflows that reduce manual effort and make common engineering tasks fast, reliable, and easy to adopt.
Use AI-native tooling and automation to improve engineering workflows, failure triage, and developer output quality.
Build secure-by-default platform capabilities, including access controls, secrets and credential handling, artifact permissions, auditability, and policy enforcement in software delivery workflows.
Partner closely with product, systems, release, quality, and infrastructure teams to understand pain points and turn them into durable platform improvements.
Define and track metrics for platform health and engineering effectiveness, including build times, queue times, failure rates, flaky failures, and deployment lead time.
Guide engineering teams on pragmatic observability, reliability, and scalability choices for the systems they build.
Participate in the on-call rotation for the systems owned by the team.
Requirements
10+ years of software engineering experience, including 5+ years building developer productivity, CI/CD, internal platform, or engineering systems.
Deep experience designing and operating robust CI/CD pipelines, build systems, and software delivery infrastructure for complex products.
Strong track record of personally building platform features and workflow improvements that materially increased engineering velocity, reliability, or developer experience.
Experience making architectural decisions for internal platforms, including when to standardize, when to abstract, and when to keep systems simple.
Experience adapting technical decisions to the maturity and scaling stage of an organization, balancing speed, reliability, maintainability, and adoption.
Working knowledge of secure software delivery practices such as least-privilege access, secrets management, policy enforcement, auditability, or software supply chain hardening.
Strong empathy for the tools, workflows, and frustrations that create toil or slow engineering teams down.
Comfortable operating in ambiguous, fast-changing environments and bringing structure where needed.
Communicate clearly and work effectively across teams with different priorities, constraints, and technical needs.
Nice-to-Haves
Experience with large-scale deployment of CPU/GPU nodes running in Kubernetes clusters across regions.
Familiarity with Terraform, Buildkite, Bazel, Postgres, Cosmos DB, Kafka.
Compensation
Total compensation range: $347,000 - $385,000 USD.
Build and lead Anthropic's managed caching infrastructure as a foundational service, including a scalable Redis fleet, client libraries, and CDC-driven invalidation. Requires deep distributed systems and caching expertise to optimize latency and consistency across hot paths for Claude.
320k – 485k/yr
Hybrid10+ YOEDevOps / SRE
Staff Engineer, Datacenter Server Lifecycle
AnthropicSan Francisco, CA +1
Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.
320k – 405k/yr
Hybrid8+ YOEDevOps / SRE
Staff Software Engineer, Infrastructure
DecagonSan Francisco, CA
Designs, builds, and operates high-scale, low-latency production infrastructure services, owning SLOs and end-to-end reliability. Partners with teams to optimize performance, evolve CI/CD, and support diverse deployments; requires 8+ years experience with strong observability and cloud expertise.
Staff+ Software Engineer owning the strategy, architecture, and development of Anthropic's configuration management, feature flagging, and large-scale experimentation platforms to enable safe, data-driven changes and boost developer productivity.
Staff+ Software Engineer building Anthropic's Agent Runtime Platform and knowledge infrastructure to enable thousands of employees to be highly productive with AI agents. Requires 10+ years large-scale distributed systems experience and agent expertise to define agentic productivity, build runtimes, write evals, and drive operational excellence.