Skip to content
OpenAIOpenAI

Software Engineer, Productivity - Model Performance

Builds and improves developer tools, CI/CD pipelines, and testing workflows to boost productivity for OpenAI's model performance engineering teams. Requires strong Python skills, experience with developer infrastructure, and ability to work in ambiguous environments.

About the job

Responsibilities

  • Improve development workflows for engineers working on model performance infrastructure
  • Design and improve CI/CD, release, validation, and testing pipelines
  • Build and maintain tools that improve reliability, iteration speed, and engineering confidence
  • Partner closely with engineers to identify friction in testing, debugging, deployment, and development workflows
  • Contribute to infrastructure efforts that support performance-critical training and inference systems
  • Help improve developer experience across Python-heavy codebases and performance-oriented infrastructure
  • Work in a high-context, ambiguous environment where ownership and good judgment matter

Requirements

  • Motivated by enabling other engineers and helping them do their best work
  • Strong experience with CI/CD, developer infrastructure, testing systems, tooling, or build/release workflows
  • Highly collaborative, empathetic, and comfortable partnering deeply with technical teams
  • Strong in Python and enjoy building reliable, scalable developer tools and infrastructure
  • Experience improving large-scale engineering workflows, especially around CI reliability, test infrastructure, and debugging velocity
  • Self-directed and comfortable operating with ambiguity
  • Excited to learn model performance domain

Nice-to-haves

  • Experience in the PyTorch ecosystem
  • Experience with C++ or Rust

Skills

Python, CI/CD, PyTorch, Triton, C++, Rust, Testing Frameworks, Developer Tooling, Infrastructure, Ci Pipelines

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.