Skip to content
OpenAIOpenAI

Software Engineer, Reliability

Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.

About the job

Responsibilities

  • Design and implement solutions to ensure the scalability of our infrastructure to meet rapidly increasing demands.
  • Build and maintain the load, chaos and synthetic testing software leveraged by development teams to make the systems they design and operate more reliable.
  • Build and maintain automation tools to streamline repetitive tasks and improve system reliability.
  • Build and maintain the platform for CPU/storage, GPU, and network lifecycle management to drive efficiency, accountability and support dynamic optimization of our resources.
  • Implement fault-tolerant and resilient design patterns to minimize service disruptions.
  • Develop and maintain service level objectives (SLOs) and service level indicators (SLIs) to measure and ensure system reliability.
  • Partner with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Participate in an on-call rotation to respond to critical incidents and ensure 24/7 system availability.

Requirements

  • Proven experience as an SWE focused on reliability or a similar role in a fast-paced, rapidly scaling company.
  • Strong proficiency in cloud infrastructure.
  • Proficiency in programming languages.
  • Experience with containerization technologies and container orchestration platforms like Kubernetes.
  • Knowledge of IaC tools such as Terraform or CloudFormation.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication and collaboration skills.
  • Experience with observability tools such as DataDog, Prometheus, Grafana and Splunk.
  • Experience with microservices architecture and service mesh technologies.
  • Knowledge of security best practices in cloud environments.

Nice-to-Haves

  • Track record of accelerating engineering reliability by empowering fellow engineers with excellent tooling and systems.
  • Experience utilizing Infrastructure as Code (IaC) principles to automate infrastructure provisioning and configuration management.
  • Experience collaborating with cross-functional teams to ensure reliability and scalability in design and development.

Skills

Kubernetes, Terraform, CloudFormation, Datadog, Prometheus, Grafana, Splunk, Infrastructure As Code, Microservices, Cloud Infrastructure

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.