Skip to content
OpenAIOpenAI

Tech Lead, Deployment & Operations — Custom Infrastructure

Lead deployment and operations for OpenAI’s custom silicon and systems into data center environments. Drive hardware bring-up, validation, production deployment, and fleet reliability at scale while leading a technical team.

About the job

Responsibilities

  • Lead a team responsible for deployment and operations of OpenAI’s custom silicon and systems in data center environments
  • Own the path from hardware bring-up and validation through production deployment, operational readiness, and sustained fleet support
  • Partner closely with silicon, systems, software, infrastructure, networking, data center, supply chain, and external partner teams to ensure successful deployment at scale
  • Define deployment processes, operational playbooks, technical readiness criteria, escalation paths, and reliability practices for new hardware platforms
  • Drive cross-functional execution across lab bring-up, rack/system integration, data center deployment, fleet monitoring, debugging, and issue resolution
  • Stay hands-on technically through architecture reviews, deployment planning, failure analysis, operational debugging, and critical system-level decision-making
  • Identify gaps in tooling, observability, automation, validation coverage, and operational processes, and build plans to close them
  • Establish clear metrics for deployment readiness, reliability, performance, maintainability, and operational health
  • Build a strong engineering culture grounded in ownership, technical rigor, operational excellence, and high-velocity execution
  • Be a contributor and technical driver for the architecture and design of future ML systems

Requirements

  • 8+ years of engineering experience in hardware systems, infrastructure, data center deployment, production operations, systems engineering, silicon bring-up, or related technical domains
  • Strong technical depth in one or more of: hardware deployment, data center operations, rack-scale systems, silicon bring-up, systems validation, fleet operations, reliability engineering, infrastructure automation, or hardware/software integration
  • Experience bringing complex hardware systems from development or validation into production environments
  • Experience working closely with silicon, systems, software, infrastructure, networking, or data center teams
  • Experience with deployment planning, operational readiness, incident response, debugging, and root-cause analysis for production systems
  • Experience building tooling, automation, observability, or operational processes that improve deployment quality and fleet reliability
  • Demonstrated ability to hire, develop, and lead senior technical talent
  • Ability to move fluidly between people leadership, technical strategy, and hands-on operational problem solving
  • Strong written and verbal communication skills, especially in high-urgency, cross-functional technical environments
  • Experience working in fast-moving environments

Nice-to-Haves

  • Enjoy mentoring and developing engineers while staying deeply engaged in technical execution
  • Excited by the challenge of bringing new custom hardware platforms into real-world production data center environments
  • Comfortable operating across silicon, systems, software, infrastructure, and data center operations
  • Comfortable leading through ambiguity, especially when the hardware, tooling, and operational model are still being built
  • Strong judgment around deployment sequencing, technical risk, operational readiness, and when to escalate
  • Care deeply about building practical systems, tools, and processes that work reliably at scale
  • Bias toward ownership and comfortable jumping into urgent technical issues when needed

Skills

Hardware Deployment, Data Center Operations, Silicon Bring-Up, Systems Validation, Fleet Operations, Reliability Engineering, Infrastructure Automation, Hardware/Software Integration, Deployment Planning, Incident Response, Root-Cause Analysis, Observability, Tooling, Automation, Rack-Scale Systems

Anthropic

Anthropic

Austin, TX
Data Center Operations Lead - Partner Site Operations
$320k+/yrHybrid8+ YOEDevOps / SRE

Leads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Release Engineer
$235k+/yrHybrid7+ YOEDevOps / SRE

Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Developer Experience
$225k+/yrHybrid5+ YOEDevOps / SRE

Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.

Descript

Descript

San Francisco, CA

Software Engineer, Infrastructure
$220k+/yrRemote8+ YOEDevOps / SRE

Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.