Software Engineer, Observability
Build observability infrastructure and AI-powered tools for OpenAI's large-scale production systems, including logging, metrics, and debugging UIs. Requires experience with distributed systems, Kubernetes, AWS, and observability tools.
About the job
What You’ll Do
- Own core observability infrastructure, including distributed logging, time series, and trace storage
- Build AI-native tools that help engineers detect, understand, and resolve issues autonomously.
- Contribute to UI experiences like dashboards, notebooking, or interactive debugging
- Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product
You Might Be a Fit If You:
- Have operated large-scale distributed systems in production (especially logging systems or some other time series databases)
- Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems.
- Have full-stack chops or product sensibilities—you're excited to build real tools people use.
- Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc).
Bonus: built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc).
Skills
Kubernetes, AWS, Prometheus, OpenTelemetry, Distributed Logging, Time Series Databases, Trace Storage, Distributed Systems, Full-Stack Development, Cloud Infrastructure
Similar jobs
DevOps / SRE jobsBuild and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.