Staff Software Engineer, HPC
Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.
About the job
Responsibilities
- Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
- Work with customer and infrastructure teams to build a multiyear software engineering roadmap for the HPC platform.
- Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
- Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
- Design and improve job-scheduling algorithms and autoscaling policies to maximize reliability and resource availability.
- Design multi-region orchestration strategies optimized for data locality, reliability, and performance.
- Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
- Evaluate technologies and paradigms that improve computational and storage capabilities.
- Develop capacity-planning tools and forecasting models for growing compute needs.
- Mentor junior engineers and support their career development.
Requirements
- Experience designing and operating large-scale distributed systems in production.
- Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
- Experience with Kubernetes for heterogeneous workloads.
- Experience with AWS or similar cloud infrastructure providers.
- Track record of shipping and operating reliable, scalable infrastructure.
- Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
- Proficiency with Python.
Nice-to-haves
- Exposure to machine learning workloads, including training, inference, and data generation.
- Experience with Kubernetes or SLURM at scale, including clusters with more than 10,000 nodes.
- Experience with SLURM and advanced scheduling policies.
- Background in algorithmic optimization or operations research.
- Experience building developer tools and platforms used by large engineering organizations.
Skills
Ray.Io, Ray Core, Ray Data, Slurm, Kubernetes, AWS, Python, Distributed Systems, High-Performance Computing, Job Scheduling, Autoscaling, Multi-Region Orchestration, Capacity Planning, Machine Learning
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.