Senior Software Engineer, Test Infrastructure
Lead the architecture and core development of Zoox's test-orchestration platform for Build Integrity Tests across manufacturing and robotaxi fleet. Own distributed task scheduling/execution, build resilient scalable systems, design extensible test frameworks and APIs, and technically lead a small team of engineers and contractors.
About the job
Responsibilities
- Own the architecture and core of the test-orchestration platform, including the distributed task scheduling, execution, and result-collection engine that runs tests across manufacturing and the fleet.
- Personally build the most complex, reliability-critical components.
- Lead a small engineering team of two engineers plus 1–3 contractors: setting technical direction, reviewing code, and growing the team's impact.
- Make the platform resilient at scale: design for safe mid-run cancellation, no result loss under failure, and predictable behavior across hundreds of test runs.
- Build clean APIs and frameworks that let other teams plug their own test suites into the platform without forking it.
- Partner across the company with multiple teams to ship test tooling that production and validation campaigns depend on.
Requirements
- Strong production Python, including hands-on experience building distributed task/queue systems (e.g., Celery or equivalent) with a message broker and result backend (Redis / RabbitMQ) - task lifecycles, retries, cancellation, and worker reliability.
- Demonstrated ability to design and maintain test frameworks for large codebases (deep pytest experience: fixtures, parametrization, scalable suites).
- Design REST and library/framework APIs that other engineers build on.
- Solid software-engineering and distributed-systems fundamentals: systematic debugging across services, networking (HTTP, TCP/IP, container networking), CI/CD pipelines, Docker, and relational databases (PostgreSQL).
- Experience as a technical lead or architect - owning a system end to end, setting direction, and reviewing/mentoring other engineers' work.
Nice-to-Haves
- Experience with Bazel (or another large-monorepo build system with hermetic/reproducible builds).
- Experience deploying and operating services on AWS/Kubernetes (EKS, Helm, ArgoCD).
- Experience with observability tooling (OpenTelemetry, Grafana/Prometheus).
- Background in manufacturing, hardware, or vehicle test systems (CAN/DoIP, Automotive Ethernet, HIL).
- Experience working with cross-functional production teams.
Skills
Python, Celery, Redis, RabbitMQ, Pytest, REST APIs, Docker, Postgres, Kubernetes, AWS, Bazel, OpenTelemetry
Similar jobs
Engineering Management jobsLeads and scales an engineering team building a fault-tolerant managed AI platform for LLM workloads, including task queues, model management, scheduling, and agentic execution infrastructure. Requires 5+ years leading engineering teams plus depth in distributed systems, cloud-native platforms, and AI infrastructure.
Leads a six-person engineering team building trusted, transparent experiences for financial institutions, while setting technical strategy, shaping the roadmap, and ensuring reliable delivery. Requires 8+ years of industry experience, prior Staff-level engineering experience, and engineering management expertise.
Leads software engineering strategy, people management, hiring, and delivery of large-scale customer-facing applications. Requires 10+ years of software development experience, 5+ years managing engineering teams, and expertise in distributed systems and cross-functional execution.
Leads the engineering team creating Webflow’s code-native platform and integrating coding agents into visual and code-based workflows. The player-coach role requires technical depth in developer tooling or infrastructure, experience leading engineering teams, and strong hiring and coaching skills.
Leads the Growth Expansion engineering team, partnering with Product and Design to deliver revenue-driving product features and upgrade paths. The role requires 7+ years of software engineering experience, engineering people-management experience, strong communication, and technical decision-making skills.