Software Engineer, Full-Stack — Developer Experience
Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.
About the job
In This Role, You Will
- Design, build, and operate CI infrastructure that gives engineers fast, reliable feedback on every change.
- Improve Bazel-based build and test workflows across a large, polyglot codebase, including dependency modeling, remote caching, remote execution, and build/test performance.
- Build systems that reduce unnecessary CI work through affected-target detection, test selection, caching, batching, and smarter scheduling.
- Partner closely with product and infrastructure teams to understand their workflows, pain points, and reliability needs, then turn those into practical platform improvements.
- Improve the observability and debuggability of build and CI failures, making it easier for engineers to distinguish product regressions, infrastructure failures, and flakes.
- Use modern AI tools to rethink how engineers interact with CI: failure explanation, fix suggestions, automatic retries, and agent-assisted debugging.
- Own the reliability of the systems you build, including participating in an on-call rotation for critical developer infrastructure.
Technologies Commonly Used
- Bazel and Starlark for large-scale build and test workflows
- Buildkite for CI orchestration
- Python and FastAPI for internal services
- Kubernetes for large-scale infrastructure
- Terraform for infrastructure as code
- Postgres, Kafka, and other systems used to power internal engineering platforms
You May Be A Strong Fit If You
- Have 5+ years of software engineering experience, including significant experience building infrastructure or tooling for developers.
- Have hands-on experience with Bazel, Buck, Pants, Gradle, or similar build systems, and understand the tradeoffs of hermetic builds, dependency graphs, caching, and remote execution.
- Have built or operated CI systems at scale, especially in environments where build time, queue time, test flakiness, and developer trust materially affect engineering velocity.
- Care deeply about developer experience and have empathy for the small sources of friction that slow teams down or create operational toil.
- Are comfortable debugging distributed systems and using metrics, logs, traces, and structured data to understand reliability and performance problems.
- Can work across teams, communicate clearly, and turn ambiguous productivity problems into concrete technical plans.
- Are excited to apply AI to developer infrastructure in ways that make engineers faster without weakening quality or safety.
Skills
Bazel, Starlark, Buildkite, Python, FastAPI, Kubernetes, Terraform, Postgres, Kafka, CI/CD
Similar jobs
DevOps / SRE jobsOwn Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.