Software Engineer, Productivity - Inference Runtime
Builds and improves CI/CD, testing, validation, and release tooling for OpenAI's inference runtime teams to ensure reliable, performant model deployments across ChatGPT, API, and research workloads. Requires strong Python skills, developer productivity experience, and high ownership in ambiguous environments.
About the job
Responsibilities
- Improve systems that ensure inference engine releases are correct, performant, and regression-free by evolving tooling and infrastructure for deploy gate validation
- Bring rigor to release, validation, branching, and deployment processes across the inference stack
- Improve canary, async, and large-scale validation workflows for inference systems
- Harden CI, testing, and validation infrastructure so failures are actionable and trustworthy
- Reduce noisy or flaky failures caused by infrastructure instability, GPU scheduling, or test environment issues
- Build automation for failure triage, ownership detection, debugging, and escalation
- Partner closely with inference teams, research developer productivity, engine acceleration, and infrastructure teams to improve release quality and rollout safety
- Reduce developer friction in testing, debugging, and release workflows so engineers can move faster with confidence
Requirements
- Strong experience with CI/CD systems, testing infrastructure, release tooling, developer productivity, or large-scale build and validation systems
- Comfortable working in Python-heavy environments and debugging complex distributed systems
- C++ experience is helpful, especially for working near inference engine code, CI build issues, or performance-sensitive systems (not required)
- High ownership, developer empathy, pragmatic, collaborative
- Comfortable operating in ambiguous areas
Nice-to-Haves
- Excited to learn about large-scale inference systems
- Prior experience in inference environments (not required)
Skills
Python, CI/CD, Kubernetes, C++, Testing Infrastructure, Release Engineering, Observability, Distributed Systems, GPU, Automation
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.