Skip to content
ExaExa

Software Engineer, Infrastructure

Builds and operates large-scale infrastructure including GPU clusters, Kubernetes orchestration, AWS batch jobs, and observability tooling to power AI search systems. Requires experience with massive-scale systems and focus on reliability and optimization.

About the job

Desired Experience

  • Experience designing and operating large-scale infrastructure - GPU clusters or large Kubernetes clusters or cloud batchjob systems
  • Obsessive mindset — always thinking about reliability, observability, and optimization across the entire stack

Example Projects

  • Build the Kubernetes orchestration on a $20m GPU cluster
  • Scale our AWS batchjob system to handle map reduce jobs over 10s of thousands of machines
  • Design GPU scheduling software so we max out our cluster utilization
  • Build observability into our production systems

Skills

Kubernetes, Rust, AWS, Ray, GPU, Observability, Mapreduce

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.