Staff Software Engineer - AI Research Infrastructure
Builds and operates research infrastructure for large-scale AI model training and inference across GPU fleets. Partners with scientists and engineers to create scheduling, orchestration, and dev tooling for efficient experimentation. Requires 5+ years in distributed systems and systems programming.
About the job
Responsibilities
- Design and implement infrastructure that supports large-scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud-based systems)
- Enable researchers to go from idea to large-scale experiment in minutes by building powerful abstractions for job submission, scheduling, and monitoring
- Create tooling that improves research developer productivity, such as experiment management systems, CI/testing infrastructure for research code, and workflows that reduce iteration time
- Influence the long-term roadmap for research computation, shaping how Databricks AI Research train, evaluate, and ship models to customers
- Serve as a technical mentor and force multiplier for other engineers working on compute, infra, and AI systems
Requirements
- BS/MS or PhD in Computer Science or related field
- 5+ years of software engineering experience, including substantial time working on large-scale distributed systems or infrastructure
- Deep experience with building and operating distributed systems, data pipelines, or large-scale backend services, ideally involving GPUs, clusters, or major cloud providers
- Proficient in one or more systems programming languages (C++, Rust, Go, Java, Scala) and can design, implement, and debug complex services
- Built or significantly contributed to cluster schedulers, resource managers, or large-scale job orchestration systems (Kubernetes, Slurm, Ray, custom internal systems)
- Understand modern ML training and inference workflows (e.g., distributed training, model parallelism, fine-tuning, evaluation)
- Can move fast and be pragmatic while caring about operational excellence; driven complex systems from prototype to stable services
- Communicate clearly with researchers and engineers
Skills
Kubernetes, Slurm, Ray, C++, Rust, Go, Java, Scala, Distributed Systems, Gpus, Cloud Providers, Ml Training, Model Parallelism, Hpc Clusters
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.