Software Engineer, ML Platform
Build and operate scalable ML training frameworks, distributed infrastructure, and model lifecycle tools to power foundational models, reinforcement learning, and other ML use cases across Zoox's autonomous driving teams. Requires 2+ years ML infrastructure experience with PyTorch, DeepSpeed, JAX, Ray and AWS.
About the job
Responsibilities
- Build the Zoox Training framework leveraged by all ML teams within Zoox. This framework needs to be highly scalable, reliable, and efficient.
- Design, implement, and operate a robust and efficient ML platform to enable the training, validation, serving, and monitoring of ML models.
- Collaborate closely with cross-functional teams, including ML researchers, software engineers, and data engineers, to define requirements and align on architectural decisions.
Requirements
- 2+ years of ML infrastructure experience.
- Experience with training frameworks like PyTorch, DeepSpeed, JAX, Ray, etc.
- Experience working with cloud providers like AWS.
Nice-to-Haves
- Experience with building large-scale, cost-efficient distributed model training and ML compute infrastructure.
- Experience with building model lifecycle management tools and experimentation.
Skills
PyTorch, Deepspeed, JAX, Ray, AWS, ML Infrastructure, Distributed Training, Model Lifecycle Management
Similar jobs
ML Engineering jobsBuild and deploy production machine-learning models for fraud detection, identity verification, and financial risk products. The role suits new PhD graduates or early-career researchers with strong quantitative foundations, Python experience, and interest in owning the full ML lifecycle.
Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.
Build replayable enterprise environments, evaluation systems, graders, and post-training workflows for AI agents. The role spans machine-learning research and production engineering and requires 1–7 years of software or ML systems experience.
Research Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.
Research Engineer focused on making conversational AI agents safe, reliable, and controllable in production. The role develops evaluations, safeguards, post-training methods, and monitoring systems, requiring 2+ years of AI/ML or safety experience and strong Python and production engineering skills.