Software Engineer, Compute Infra
Designs, builds, and operates massive-scale compute clusters and custom container orchestration platforms for AI training and inference at exascale. Requires deep expertise in virtualization, containerization, systems programming in C++/Rust, and Linux kernel internals.
About the job
Responsibilities
- Build and manage massive-scale clusters to host, persist, train, and serve AI workloads with extreme reliability and performance.
- Design, develop, and extend an in-house container orchestration platform that achieves superior scalability, isolation, resource efficiency, and fault-tolerance compared to off-the-shelf solutions.
- Collaborate with research teams to architect and optimize compute clusters specifically for large-scale training runs, inference services, and real-time applications.
- Profile, debug, and resolve complex system-level performance bottlenecks, resource contention, scheduling issues, and reliability problems across the full stack.
- Own end-to-end infrastructure initiatives with first-principles design, rigorous testing, automation, and continuous optimization to support frontier AI compute demands.
Required Qualifications
- Deep expertise in virtualization technologies (KVM, Xen, QEMU) and advanced containerization/sandboxing (Kata, Firecracker, gVisor, Sysbox, or equivalent).
- Strong proficiency in systems programming languages such as C/C++ and Rust.
- Proven track record profiling, debugging, and optimizing complex system-level performance issues, with deep knowledge of Linux kernel internals, resource management, scheduling, memory management, and low-level engineering.
- Hands-on experience building or significantly enhancing distributed compute platforms, orchestration systems, or high-performance infrastructure at scale.
- Ability to thrive in a fast-paced, meritocratic environment with full ownership, high standards, and a focus on rigorous execution.
Preferred Qualifications
- Experience in Linux kernel development, hypervisor extensions, or low-level system programming for compute-intensive workloads.
- Proven track record operating or designing large-scale AI training/inference clusters (GPU/TPU scale).
- Experience with custom runtimes, isolation techniques, or bespoke platforms for specialized AI compute.
- Familiarity with performance tools, tracing, and debugging in production distributed environments.
Compensation and Benefits
Annual Salary Range: $180,000 - $440,000 USD
Base salary is just one part of total rewards, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
Skills
Kubernetes, Kvm, Xen, Qemu, Kata, Firecracker, Gvisor, Sysbox, C++, Rust, Linux Kernel
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.