Skip to content
ZooxZooxFoster City, CA

Staff Software Engineer, HPC

Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.

$230k – $295k/yr
Hybrid7+ YOEDevOps / SRE

About the job

Responsibilities

  • Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
  • Work with customer and infrastructure teams to build a multiyear software engineering roadmap for the HPC platform.
  • Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
  • Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
  • Design and improve job-scheduling algorithms and autoscaling policies to maximize reliability and resource availability.
  • Design multi-region orchestration strategies optimized for data locality, reliability, and performance.
  • Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
  • Evaluate technologies and paradigms that improve computational and storage capabilities.
  • Develop capacity-planning tools and forecasting models for growing compute needs.
  • Mentor junior engineers and support their career development.

Requirements

  • Experience designing and operating large-scale distributed systems in production.
  • Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
  • Experience with Kubernetes for heterogeneous workloads.
  • Experience with AWS or similar cloud infrastructure providers.
  • Track record of shipping and operating reliable, scalable infrastructure.
  • Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
  • Proficiency with Python.

Nice-to-haves

  • Exposure to machine learning workloads, including training, inference, and data generation.
  • Experience with Kubernetes or SLURM at scale, including clusters with more than 10,000 nodes.
  • Experience with SLURM and advanced scheduling policies.
  • Background in algorithmic optimization or operations research.
  • Experience building developer tools and platforms used by large engineering organizations.

Skills

Ray.IoRay CoreRay DataSlurmKubernetesAWSPythonDistributed SystemsHigh-Performance ComputingJob SchedulingAutoscalingMulti-Region OrchestrationCapacity PlanningMachine Learning