Skip to content
OpenAIOpenAI

Software Engineer, Observability

Build observability infrastructure and AI-powered tools for OpenAI's large-scale production systems, including logging, metrics, and debugging UIs. Requires experience with distributed systems, Kubernetes, AWS, and observability tools.

About the job

What You’ll Do

  • Own core observability infrastructure, including distributed logging, time series, and trace storage
  • Build AI-native tools that help engineers detect, understand, and resolve issues autonomously.
  • Contribute to UI experiences like dashboards, notebooking, or interactive debugging
  • Collaborate closely with engineers, researchers, user ops, and other teams across the company to build the next generation observability product

You Might Be a Fit If You:

  • Have operated large-scale distributed systems in production (especially logging systems or some other time series databases)
  • Thrive in ambiguous environments and roll up your sleeves to solve unscoped problems.
  • Have full-stack chops or product sensibilities—you're excited to build real tools people use.
  • Have strong fundamentals in systems, networking, and cloud infra (Kubernetes, AWS, etc).

Bonus: built or contributed to observability systems (e.g. Prometheus, OpenTelemetry, etc).

Skills

Kubernetes, AWS, Prometheus, OpenTelemetry, Distributed Logging, Time Series Databases, Trace Storage, Distributed Systems, Full-Stack Development, Cloud Infrastructure

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.