Skip to content
DeepgramDeepgram

Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)

Builds and operates hybrid AI/ML infrastructure using Kubernetes, AWS, Terraform, and Slurm for GPU workloads. Requires 5+ years SRE/DevOps experience, expert Kubernetes, and bare metal management.

About the job

Opportunity

Experienced Site Reliability Engineer to build and operate hybrid infrastructure for AI/ML research and product development. Architect, build, and run platform spanning AWS and bare metal data centers using Kubernetes, AWS, and Terraform, orchestrating GPU workloads with Slurm.

What You’ll Do

  • Architect and maintain core computing platform using Kubernetes on AWS and on-premise.
  • Develop and manage infrastructure using Infrastructure-as-Code with Terraform.
  • Design, build, and optimize AI/ML job scheduling with Slurm integrated with Kubernetes for GPU resources.
  • Provision, manage, and maintain on-premise bare metal server infrastructure for GPU computing.
  • Implement platform's networking (CNI, service mesh) and storage (CSI, S3) solutions.
  • Develop observability stack (monitoring, logging, tracing) and automation for operations.
  • Collaborate with AI researchers and ML engineers to build tools and workflows.

You’ll Love This Role If You

  • Passionate about building platforms for developers and researchers.
  • Enjoy automated solutions for complex infrastructure in cloud and data centers.
  • Thrive on optimizing hybrid infrastructure for performance, cost, and reliability.
  • Excited about platform engineering and AI intersection.

It’s Important To Us That You Have

  • 5+ years in Platform Engineering, DevOps, or SRE.
  • Hands-on experience with Terraform in production.
  • Expert-level Kubernetes architecture and operations.
  • Experience with Slurm for GPU-intensive AI workloads.
  • Managing bare metal infrastructure (PXE boot, MAAS).
  • Strong scripting (Python, Go, Bash).

It Would Be Great if You Had

  • Experience with CI/CD (GitLab CI, Jenkins, ArgoCD).
  • FinOps and cloud cost optimization.
  • Kubernetes networking (Calico, Cilium) and storage (Ceph, Rook).
  • Multi-region or hybrid cloud experience.

Skills

Kubernetes, Terraform, AWS, Slurm, Python, Go, Bash, Cni, Csi, S3

Cloudflare

Cloudflare

Austin, TX
Systems Engineer - Database Platform
$150k+/yrHybridDevOps / SRE

Build and operate a highly available, multi-region PostgreSQL platform, developing automation, monitoring, disaster recovery, and performance tooling. Requires experience with large-scale PostgreSQL clusters, infrastructure as code, scripting, containers, and observability.

Fluidstack

Fluidstack

New York, NY
Infrastructure Deployment Engineer
$150k+/yrOn-site5+ YOEDevOps / SRE

Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.

Trexquant

Trexquant

New York, NY

Python Engineer - Trade Operations
$150k+/yrOn-site3+ YOEDevOps / SRE

The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

SimplePractice

SimplePractice

United States

DevOps Engineer, Data & AI Platform
$144k+/yrOn-site3+ YOEDevOps / SRE

The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.