Skip to content
OpenAIOpenAI

Software Engineer, Platform Systems

Designs and builds distributed failure detection, tracing, and observability systems for large-scale AI training jobs. Requires deep expertise in performance, distributed systems, hardware, networking, and low-level software engineering.

About the job

In This Role, You Will

  • Design and build distributed failure detection, tracing, and profiling systems for large-scale AI training jobs
  • Develop tooling to identify slow, faulty, or misbehaving nodes and provide actionable visibility into system behavior
  • Improve observability, reliability, and performance across OpenAI’s training platform
  • Debug and resolve issues in complex, high-throughput distributed systems
  • Collaborate with systems, infrastructure, and research teams to evolve platform capabilities
  • Extend and adapt failure detection systems or tracing systems to support new training paradigms and workloads

You Might Thrive in This Role If You

  • Care deeply about performance, stability, and observability in distributed systems
  • Enjoy finding and fixing issues in large-scale systems and automating operational workflows
  • Have experience writing low-level software where system details matter
  • Understand hardware, operating systems, networking, concurrency, and distributed systems
  • Have a background in high-performance computing or low-level systems engineering
  • Are excited to work on critical infrastructure that powers frontier AI research

Skills

Distributed Systems, Failure Detection, Tracing, Profiling, Observability, High-Performance Computing, Networking, Concurrency, Linux, C++

Anthropic

Anthropic

San Francisco, CA

DevOps / AgentOps Engineer, GTM Systems
$320k+/yrHybridDevOps / SRE

Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.

Anthropic

Anthropic

San Francisco, CA
Software Engineer, Infrastructure, Interpretability
$320k+/yrHybridDevOps / SRE

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Site Reliability Engineer (SRE)
$350k+/yrOn-siteDevOps / SRE

Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.