Skip to content
OpenAIOpenAI

Software Engineer, Caching Infrastructure

Design, build, and operate a multi-tenant caching platform powering OpenAI's inference, identity, and products. Requires 5+ years in distributed systems with deep Redis/Memcached expertise and Kubernetes experience.

About the job

In This Role, You Will:

  • Design, build, and operate OpenAI’s multi-tenant caching platform used across inference, identity, quota, and product experiences.
  • Define the long-term vision and roadmap for caching as a core infra capability, balancing performance, durability, and cost.
  • Collaborate with other infra teams (e.g., networking, observability, databases) and product teams to ensure our caching platform meets their needs.

You Might Thrive In This Role If You:

  • Have 5+ years of experience building and scaling distributed systems, with a strong focus on caching, load balancing, or storage systems.
  • Have deep expertise with Redis, Memcached or similar solutions, including clustering, durability configurations, client-side connection patterns, and performance tuning.
  • Have production experience with Kubernetes, service meshes (e.g., Envoy), and autoscaling systems.
  • Think rigorously about latency, reliability, throughput, and cost in designing platform capabilities.
  • Thrive in a fast-paced environment and enjoy balancing pragmatic engineering with long-term technical excellence.

Skills

Redis, Memcached, Kubernetes, Envoy, Distributed Systems, Caching, Load Balancing, Autoscaling, Networking, Storage Systems

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.