Software Engineer, Developer Productivity
Builds and maintains foundational systems, tools, and processes to boost developer productivity and engineering velocity at OpenAI. Requires 5+ years engineering experience, including infrastructure tooling, with core tech like Kubernetes, Python, and Terraform. Onsite in SF HQ.
About the job
Responsibilities
- Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output.
- Use our latest AI tools to re-think how we can be the most productive team in the industry.
- Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements.
- Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations.
- Guide and advise product engineering teams on best practices for ensuring observable, scalable systems.
- Responsible for the reliability of the systems we build, including an on-call rotation to respond to critical incidents as needed.
Requirements
- 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers.
- Experience-driven empathy for the tools, frustrations, and processes that slow engineering teams down and lead to toil or burnout.
- Voracious and intrinsic desire to learn and fill in missing skills—and an equally strong talent for sharing learnings clearly and concisely with others.
- Comfortable with ambiguity and rapidly changing conditions. You view changes as an opportunity to add structure and order when necessary.
Technical Context
- Large-scale deployment of GPU nodes running in dozens of Kubernetes clusters across regions.
- Core technologies: Terraform, Buildkite, Postgres, Cosmos DB, Kafka, Python, FastAPI.
Skills
Kubernetes, Terraform, Buildkite, Postgres, Cosmos Db, Kafka, Python, FastAPI, GPU, AI Tools
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.