Skip to content
OpenAIOpenAI

Performance & Systems Engineer, Codex

Optimizes performance across Codex AI system's stack including LLM inference, cloud orchestration, and agent behavior to reduce latency and costs. Collaborates with researchers and engineers on high-impact improvements in a high-ownership role.

About the job

Responsibilities

  • Hunt down and address inefficiencies across the Codex system stack, from agent behavior to LLM inference to container orchestration, and beyond.
  • Build tooling to measure, profile, and optimize system performance at scale.
  • Collaborate with researchers and engineers to land high-ROI changes that improve latency and cost.

Requirements

  • Experience operating across both ML systems and cloud infrastructure.
  • Enjoy diving into messy, ambiguous problems and emerging with clear wins.
  • Think holistically about performance, balancing speed, cost, and user experience.

Skills

Llm Inference, Cloud Orchestration, Kubernetes, Performance Optimization, Ml Systems, Container Orchestration, Profiling Tools, System Monitoring, Latency Optimization, Cost Optimization

OpenAI

OpenAI

San Francisco, CA

Systems Integration Engineer, Build Systems | Consumer Devices
$293k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.

OpenAI

OpenAI

San Francisco, CA

Network Engineer
$293k+/yrHybridDevOps / SRE

Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.

Anthropic

Anthropic

San Francisco, CA

DevOps / AgentOps Engineer, GTM Systems
$320k+/yrHybridDevOps / SRE

Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.

Anthropic

Anthropic

San Francisco, CA
Software Engineer, Infrastructure, Interpretability
$320k+/yrHybridDevOps / SRE

Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Site Reliability Engineer (SRE)
$350k+/yrOn-siteDevOps / SRE

Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.