Skip to content
DeepgramDeepgram

Systems Architect AI/ML Infrastructure

Designs end-to-end infrastructure architecture for AI/ML workloads, including multi-cloud strategies, GPU orchestration, storage for massive datasets, and capacity planning to support production inference and research training at scale. Requires 7+ years in systems architecture with deep expertise in Kubernetes, AWS, and GPU infrastructure.

About the job

What You'll Do

  • Define and drive the end-to-end infrastructure architecture for Deepgram's AI/ML workloads across production inference and research training
  • Design multi-cloud and hybrid infrastructure strategies that balance performance, reliability, cost, and vendor flexibility
  • Architect compute orchestration systems that efficiently schedule and manage GPU and CPU workloads across heterogeneous infrastructure
  • Design storage architectures that handle the massive datasets required for speech and audio ML -- from high-throughput training data pipelines to low-latency model serving
  • Lead capacity planning across all infrastructure dimensions, modeling growth and ensuring Deepgram can scale ahead of demand
  • Drive cost optimization and FinOps practices, identifying opportunities to reduce infrastructure spend without compromising performance or reliability
  • Design burstable, elastic training infrastructure that can scale up for large training runs and scale down to minimize idle cost
  • Architect research compute infrastructure that gives ML teams the resources they need while maintaining operational efficiency
  • Establish architectural standards, design review processes, and technical documentation practices for infrastructure decisions
  • Collaborate with engineering leadership to align infrastructure strategy with product roadmap and business objectives
  • Evaluate emerging hardware, cloud services, and infrastructure technologies for potential adoption

Requirements

  • 7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure
  • Proven experience designing multi-cloud architectures spanning AWS and at least one other major cloud provider or on-premises environment
  • Deep expertise in storage system design -- block, object, and file storage, including performance tuning for large-scale data workloads
  • Strong experience with compute orchestration using Kubernetes, and an understanding of how to schedule diverse workloads efficiently
  • Hands-on experience with GPU infrastructure -- procurement considerations, cluster design, driver and runtime management
  • Track record of capacity planning and infrastructure scaling for high-growth environments
  • Ability to communicate complex architectural decisions clearly to both technical and non-technical stakeholders
  • Strong understanding of networking fundamentals as they relate to infrastructure architecture

Nice-to-Haves

  • Direct experience architecting infrastructure for ML training workloads -- distributed training, large dataset management, experiment infrastructure
  • Background in cost optimization and FinOps practices for large-scale cloud and bare metal infrastructure
  • Experience operating and managing bare metal infrastructure in colocation facilities
  • Expertise in network architecture design, including high-bandwidth GPU interconnects and global traffic routing
  • Experience with infrastructure modeling and simulation for capacity planning
  • Familiarity with Slurm, Ray, or other HPC/ML job scheduling systems
  • Understanding of power, cooling, and physical infrastructure considerations for GPU-dense deployments

Skills

Kubernetes, AWS, GPU, Storage Systems, Multi-Cloud, Finops, Slurm, Ray, Bare Metal, Networking

Voltus

Voltus

United States

Senior Software Engineer, Infrastructure
$160k+/yrRemote6+ YOEDevOps / SRE

Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.

VSCO

VSCO

San Francisco, CA

Senior Software Engineer, Infrastructure
$165k+/yrHybrid5+ YOEDevOps / SRE

Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.

Imply

Imply

United States

Senior Software Engineer
$155k+/yrRemote6+ YOEDevOps / SRE

Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.

Commure

Commure

Mountain View, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid6+ YOEDevOps / SRE

Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.