Staff Software Engineer, HPC
Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.
$230k – $295k/yr
Hybrid7+ YOEDevOps / SRE
About the job
Responsibilities
- Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
- Work with customer and infrastructure teams to build a multiyear software engineering roadmap for the HPC platform.
- Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
- Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
- Design and improve job-scheduling algorithms and autoscaling policies to maximize reliability and resource availability.
- Design multi-region orchestration strategies optimized for data locality, reliability, and performance.
- Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
- Evaluate technologies and paradigms that improve computational and storage capabilities.
- Develop capacity-planning tools and forecasting models for growing compute needs.
- Mentor junior engineers and support their career development.
Requirements
- Experience designing and operating large-scale distributed systems in production.
- Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
- Experience with Kubernetes for heterogeneous workloads.
- Experience with AWS or similar cloud infrastructure providers.
- Track record of shipping and operating reliable, scalable infrastructure.
- Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
- Proficiency with Python.
Nice-to-haves
- Exposure to machine learning workloads, including training, inference, and data generation.
- Experience with Kubernetes or SLURM at scale, including clusters with more than 10,000 nodes.
- Experience with SLURM and advanced scheduling policies.
- Background in algorithmic optimization or operations research.
- Experience building developer tools and platforms used by large engineering organizations.
Skills
Ray.IoRay CoreRay DataSlurmKubernetesAWSPythonDistributed SystemsHigh-Performance ComputingJob SchedulingAutoscalingMulti-Region OrchestrationCapacity PlanningMachine Learning