Senior Software Engineer, Infrastructure
Builds and scales infrastructure for OpenAI's experimentation platform, including low-latency configuration delivery, high-throughput data ingestion, and analytics systems handling billions of evaluations. Requires expertise in large-scale distributed systems, performance optimization, and operational excellence.
About the job
In this role, you will:
- Design and operate low-latency configuration delivery systems powering progressive rollouts across OpenAI’s product suites.
- Build highly scalable data ingestion and analytics infrastructure to support experimentation, product analytics, and feature performance monitoring.
- Improve the performance, efficiency, and reliability of Statsig’s core infrastructure, ensuring systems remain fast and stable as OpenAI’s products scale globally.
- Optimize query performance and data availability for experimentation and analytics workflows used by teams across OpenAI.
- Lead large technical initiatives and shape the architecture of experimentation and rollout infrastructure used across the company.
You might thrive in this role if you:
- Have experience building large-scale distributed systems with strict performance and reliability requirements.
- Enjoy solving low-latency systems problems, such as real-time configuration delivery, high-throughput ingestion pipelines, or large-scale analytics systems.
- Have experience building and optimizing large-scale data platforms, event pipelines, etc.
- Care deeply about system reliability, observability, and operational excellence in production environments.
- Take ownership of complex technical problems end-to-end and enjoy building infrastructure that enables other teams to move faster.
Skills
Distributed Systems, Low-Latency Systems, Data Ingestion, Analytics Infrastructure, Observability, Scalability, Reliability, Event Pipelines, Query Optimization, Real-Time Configuration
Similar jobs
DevOps / SRE jobsLeads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Own and evolve a broad infrastructure platform spanning cloud, Kubernetes, deployment, reliability, security, and GPU-backed AI systems. The role requires 8+ years operating production distributed systems, strong incident and architecture experience, and practical cloud infrastructure expertise.
Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.