Skip to content
DatabricksDatabricks

Staff Software Engineer - AI Research Infrastructure

Builds and operates research infrastructure for large-scale AI model training and inference across GPU fleets. Partners with scientists and engineers to create scheduling, orchestration, and dev tooling for efficient experimentation. Requires 5+ years in distributed systems and systems programming.

About the job

Responsibilities

  • Design and implement infrastructure that supports large-scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud-based systems)
  • Enable researchers to go from idea to large-scale experiment in minutes by building powerful abstractions for job submission, scheduling, and monitoring
  • Create tooling that improves research developer productivity, such as experiment management systems, CI/testing infrastructure for research code, and workflows that reduce iteration time
  • Influence the long-term roadmap for research computation, shaping how Databricks AI Research train, evaluate, and ship models to customers
  • Serve as a technical mentor and force multiplier for other engineers working on compute, infra, and AI systems

Requirements

  • BS/MS or PhD in Computer Science or related field
  • 5+ years of software engineering experience, including substantial time working on large-scale distributed systems or infrastructure
  • Deep experience with building and operating distributed systems, data pipelines, or large-scale backend services, ideally involving GPUs, clusters, or major cloud providers
  • Proficient in one or more systems programming languages (C++, Rust, Go, Java, Scala) and can design, implement, and debug complex services
  • Built or significantly contributed to cluster schedulers, resource managers, or large-scale job orchestration systems (Kubernetes, Slurm, Ray, custom internal systems)
  • Understand modern ML training and inference workflows (e.g., distributed training, model parallelism, fine-tuning, evaluation)
  • Can move fast and be pragmatic while caring about operational excellence; driven complex systems from prototype to stable services
  • Communicate clearly with researchers and engineers

Skills

Kubernetes, Slurm, Ray, C++, Rust, Go, Java, Scala, Distributed Systems, Gpus, Cloud Providers, Ml Training, Model Parallelism, Hpc Clusters

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

VGS

VGS

United States
Senior Staff Infrastructure Engineer
$185k+/yrRemote10+ YOEDevOps / SRE

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.