Skip to content
OpenAIOpenAI

Software Engineer, Cloud Infrastructure

Builds and maintains cloud infrastructure abstractions for scalable, reliable product platforms like ChatGPT. Requires 5+ years in core infrastructure, Kubernetes at scale, and cloud abstractions; onsite in San Francisco with on-call duties.

About the job

Responsibilities

  • Design and build the development and production platforms that power our products, enabling reliability and security at scale
  • Ensure our infrastructure can scale to the next order of magnitude
  • Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think
  • Participate in on-call rotation to respond to critical incidents as needed

Requirements

  • 5+ years building core infrastructure
  • Experience operating orchestration systems such as Kubernetes at scale
  • Experience building abstractions over cloud platforms
  • Take pride in building and operating scalable, reliable, secure systems
  • Comfortable with ambiguity and rapid change

Skills

Kubernetes, Cloud Platforms, Infrastructure As Code, Scalable Systems, Reliable Systems, Secure Systems, Networking, Orchestration Systems

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.