Staff AI Runtime Engineer
Designs, develops, and optimizes core runtime infrastructure for distributed AI training and inference using PyTorch-based stack. Requires 8+ years in systems engineering, deep learning runtimes, Python/C++, and multi-node GPU workloads.
About the job
What You'll Do
Lead Runtime Design & Development:
- Own the core runtime architecture supporting AI training and inference at scale.
- Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
- Optimize distributed training reliability, orchestration, and job-level fault tolerance.
Drive Performance at Scale:
- Profile and enhance low-level system performance across training and inference pipelines.
- Improve packaging, deployment, and integration of customer models in production environments.
- Ensure consistent throughput, latency, and reliability metrics across multi-node, multi-GPU setups.
Build Internal Tooling & Frameworks:
- Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
- Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
- Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.
Collaborate & Mentor:
- Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
- Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team’s capabilities.
What You’ll Need to Be Successful
- 8+ years of experience in systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.
- Experience in delivering PaaS services.
- Proven experience optimizing and scaling deep learning runtimes (e.g. PyTorch, TensorFlow, JAX) for large-scale training and/or inference.
- Strong programming skills in Python and C++ (Go or Rust is a plus).
- Familiarity with distributed training frameworks, low-level performance tuning, and resource orchestration.
- Experience working with multi-GPU, multi-node, or cloud-native AI workloads.
- Solid understanding of containerized workloads, job scheduling, and failure recovery in production environments.
Nice to Have
- Contributions to PyTorch internals or open-source DL infrastructure projects.
- Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
- Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
- Background in systems research, compilers, or runtime architecture for HPC or ML.
- Startup previous experience
Skills
PyTorch, TensorFlow, JAX, Python, C++, Kubernetes, Ray, Torchelastic, Distributed Training, Multi-Gpu
Similar jobs
ML Engineering jobsBuild and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.
Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.