Software Engineer, Infrastructure
Builds and operates large-scale infrastructure including GPU clusters, Kubernetes orchestration, AWS batch jobs, and observability tooling to power AI search systems. Requires experience with massive-scale systems and focus on reliability and optimization.
About the job
Desired Experience
- Experience designing and operating large-scale infrastructure - GPU clusters or large Kubernetes clusters or cloud batchjob systems
- Obsessive mindset — always thinking about reliability, observability, and optimization across the entire stack
Example Projects
- Build the Kubernetes orchestration on a $20m GPU cluster
- Scale our AWS batchjob system to handle map reduce jobs over 10s of thousands of machines
- Design GPU scheduling software so we max out our cluster utilization
- Build observability into our production systems
Skills
Kubernetes, Rust, AWS, Ray, GPU, Observability, Mapreduce
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.