Staff Software Engineer - AI Research Infrastructure
Build and operate the large-scale training and inference infrastructure that powers Databricks AI Research, enabling researchers to run experiments across thousands of GPUs. Partner with ML scientists and platform teams to deliver reliable, high-performance orchestration and tooling.
About the job
Responsibilities
- Design and implement infrastructure that supports large‑scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems)
- Build powerful abstractions for job submission, scheduling, and monitoring so researchers can go from idea to large‑scale experiment in minutes
- Create tooling that improves research developer productivity, such as experiment management systems, CI/testing infrastructure for research code, and workflows that reduce iteration time
- Influence the long‑term roadmap for research computation, shaping how Databricks AI Research trains, evaluates, and ships models
- Serve as a technical mentor and force multiplier for other engineers working on compute, infra, and AI systems
Requirements
- BS/MS or PhD in Computer Science or related field
- 5+ years of software engineering experience, including substantial time working on large‑scale distributed systems or infrastructure
- Deep experience building and operating distributed systems, data pipelines, or large‑scale backend services, ideally involving GPUs, clusters, or major cloud providers
- Proficient in one or more systems programming languages (e.g., C++, Rust, Go, Java, Scala) and able to design, implement, and debug complex services
- Experience building or significantly contributing to cluster schedulers, resource managers, or large‑scale job orchestration systems (e.g., Kubernetes, Slurm, Ray, custom internal systems)
- Understanding of modern ML training and inference workflows (e.g., distributed training, model parallelism, fine‑tuning, evaluation)
- Ability to move fast and be pragmatic while caring about operational excellence; experience driving complex systems from prototype to stable, well‑owned services
- Strong communication skills with both researchers and engineers
Skills
Kubernetes, Slurm, Ray, C++, Rust, Go, Java, Scala, Gpu Clusters, Distributed Systems
Similar jobs
DevOps / SRE jobsOwn reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.