# Staff Software Engineer, HPC

**Company:** [Zoox](https://hotfix.jobs/companies/zoox)
**Location:** Foster City, CA, Boston, MA, Seattle, WA
**Role:** DevOps / SRE
**Salary:** $230k – $295k/yr
**Experience:** 7+ years
**Skills:** Ray.Io, Ray Core, Ray Data, Slurm, Kubernetes, AWS, Python, Distributed Systems, High-Performance Computing, Job Scheduling, Autoscaling, Multi-Region Orchestration, Capacity Planning, Machine Learning
**Posted:** 2026-08-10

> Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.

## Job Description

## Responsibilities
- Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
- Work with customer and infrastructure teams to build a multiyear software engineering roadmap for the HPC platform.
- Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
- Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
- Design and improve job-scheduling algorithms and autoscaling policies to maximize reliability and resource availability.
- Design multi-region orchestration strategies optimized for data locality, reliability, and performance.
- Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
- Evaluate technologies and paradigms that improve computational and storage capabilities.
- Develop capacity-planning tools and forecasting models for growing compute needs.
- Mentor junior engineers and support their career development.

## Requirements
- Experience designing and operating large-scale distributed systems in production.
- Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
- Experience with Kubernetes for heterogeneous workloads.
- Experience with AWS or similar cloud infrastructure providers.
- Track record of shipping and operating reliable, scalable infrastructure.
- Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
- Proficiency with Python.

## Nice-to-haves
- Exposure to machine learning workloads, including training, inference, and data generation.
- Experience with Kubernetes or SLURM at scale, including clusters with more than 10,000 nodes.
- Experience with SLURM and advanced scheduling policies.
- Background in algorithmic optimization or operations research.
- Experience building developer tools and platforms used by large engineering organizations.

## Similar jobs

- [Sr. Staff Lead Site Reliability Engineer](https://hotfix.jobs/jobs/b123b52f-6dd7-417b-978e-717e23b19c0d) - Shield AI - San Mateo, CA - $220k – $330k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/75edbb52-9f2f-49c0-b1f8-8b59e87fffc3) - Skydio - Remote - $240k – $300k/yr
- [Staff Software Engineer, Developer Infrastructure](https://hotfix.jobs/jobs/9e5df28a-b02e-49c1-ae34-8ac4874bd491) - Coinbase - Remote - $218k – $257k/yr
- [Staff Infrastructure Engineer, Trading](https://hotfix.jobs/jobs/4be49748-b215-44fa-b91e-9975c38847a9) - Coinbase - Remote - $218k – $257k/yr
- [Staff Site Reliability Engineer - Site Experience](https://hotfix.jobs/jobs/d8fa80d2-f6f6-4a20-8485-60d2648f8daf) - Reddit - San Francisco, CA - $217k – $304k/yr

**Apply:** https://hotfix.jobs/jobs/03d819ae-4f6f-411e-b6bd-ee3bbc76b475
**Canonical:** https://hotfix.jobs/jobs/03d819ae-4f6f-411e-b6bd-ee3bbc76b475