Skip to content
ZooxZoox

Staff Software Engineer, HPC

Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.

About the job

Responsibilities

  • Design and implement core services and abstractions for distributed compute infrastructure supporting thousands of concurrent jobs.
  • Work with customer and infrastructure teams to build a multiyear software engineering roadmap for the HPC platform.
  • Lead multi-quarter, cross-team initiatives that drive organization-wide improvements.
  • Create production-grade APIs, SDKs, and tools for running large-scale distributed workloads.
  • Design and improve job-scheduling algorithms and autoscaling policies to maximize reliability and resource availability.
  • Design multi-region orchestration strategies optimized for data locality, reliability, and performance.
  • Identify and resolve systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
  • Evaluate technologies and paradigms that improve computational and storage capabilities.
  • Develop capacity-planning tools and forecasting models for growing compute needs.
  • Mentor junior engineers and support their career development.

Requirements

  • Experience designing and operating large-scale distributed systems in production.
  • Experience with Ray.io, particularly Ray Core and Ray Data, or equivalent technologies.
  • Experience with Kubernetes for heterogeneous workloads.
  • Experience with AWS or similar cloud infrastructure providers.
  • Track record of shipping and operating reliable, scalable infrastructure.
  • Ability to prioritize development work and build cross-functional consensus around technical tradeoffs.
  • Proficiency with Python.

Nice-to-haves

  • Exposure to machine learning workloads, including training, inference, and data generation.
  • Experience with Kubernetes or SLURM at scale, including clusters with more than 10,000 nodes.
  • Experience with SLURM and advanced scheduling policies.
  • Background in algorithmic optimization or operations research.
  • Experience building developer tools and platforms used by large engineering organizations.

Skills

Ray.Io, Ray Core, Ray Data, Slurm, Kubernetes, AWS, Python, Distributed Systems, High-Performance Computing, Job Scheduling, Autoscaling, Multi-Region Orchestration, Capacity Planning, Machine Learning

Shield AI

Shield AI

San Mateo, CA
Sr. Staff Lead Site Reliability Engineer
$220k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Reddit

Reddit

San Francisco, CA

Staff Site Reliability Engineer - Site Experience
$217k+/yrOn-site8+ YOEDevOps / SRE

Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.