Skip to content
OpenAIOpenAI

Software Engineer, Full-Stack — Developer Experience

Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.

About the job

In This Role, You Will

  • Design, build, and operate CI infrastructure that gives engineers fast, reliable feedback on every change.
  • Improve Bazel-based build and test workflows across a large, polyglot codebase, including dependency modeling, remote caching, remote execution, and build/test performance.
  • Build systems that reduce unnecessary CI work through affected-target detection, test selection, caching, batching, and smarter scheduling.
  • Partner closely with product and infrastructure teams to understand their workflows, pain points, and reliability needs, then turn those into practical platform improvements.
  • Improve the observability and debuggability of build and CI failures, making it easier for engineers to distinguish product regressions, infrastructure failures, and flakes.
  • Use modern AI tools to rethink how engineers interact with CI: failure explanation, fix suggestions, automatic retries, and agent-assisted debugging.
  • Own the reliability of the systems you build, including participating in an on-call rotation for critical developer infrastructure.

Technologies Commonly Used

  • Bazel and Starlark for large-scale build and test workflows
  • Buildkite for CI orchestration
  • Python and FastAPI for internal services
  • Kubernetes for large-scale infrastructure
  • Terraform for infrastructure as code
  • Postgres, Kafka, and other systems used to power internal engineering platforms

You May Be A Strong Fit If You

  • Have 5+ years of software engineering experience, including significant experience building infrastructure or tooling for developers.
  • Have hands-on experience with Bazel, Buck, Pants, Gradle, or similar build systems, and understand the tradeoffs of hermetic builds, dependency graphs, caching, and remote execution.
  • Have built or operated CI systems at scale, especially in environments where build time, queue time, test flakiness, and developer trust materially affect engineering velocity.
  • Care deeply about developer experience and have empathy for the small sources of friction that slow teams down or create operational toil.
  • Are comfortable debugging distributed systems and using metrics, logs, traces, and structured data to understand reliability and performance problems.
  • Can work across teams, communicate clearly, and turn ambiguous productivity problems into concrete technical plans.
  • Are excited to apply AI to developer infrastructure in ways that make engineers faster without weakening quality or safety.

Skills

Bazel, Starlark, Buildkite, Python, FastAPI, Kubernetes, Terraform, Postgres, Kafka, CI/CD

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.