Skip to content

Senior Site Reliability Engineer, AI Infrastructure

The Senior Site Reliability Engineer will build and operate secure, observable AI/ML infrastructure across cloud platforms. The role requires at least five years of production SRE or infrastructure experience, strong Terraform and observability expertise, and hands-on incident response and automation skills.

About the job

Responsibilities

  • Own service level objectives, error budgets, and reliability targets for infrastructure underpinning cloud-based platforms.
  • Ensure observability across platform components and serving endpoints, including metrics, logs, traces, alert quality, and telemetry completeness.
  • Design, build, and maintain infrastructure as code, operational automation, and change-control workflows for AI/ML platforms.
  • Implement and maintain platform security controls, including network segmentation, secrets management, encryption, and data protection safeguards aligned with compliance requirements.
  • Lead incident response and blameless postmortems.
  • Validate backup, restore, and disaster recovery processes; conduct game days and resiliency testing.
  • Improve platform resiliency, cost efficiency, capacity planning, and operational standards.
  • Mentor engineers, influence design reviews, and collaborate across engineering teams.

Requirements

  • 5+ years of experience in SRE, platform engineering, or infrastructure roles supporting production cloud environments and mission-critical applications.
  • Strong proficiency with observability, including metrics, logging, distributed tracing, SLI/SLO frameworks, incident response, blameless postmortems, and on-call operations.
  • Strong proficiency with Infrastructure as Code, Terraform, GitOps, and CI/CD for infrastructure and platform changes.
  • Working proficiency with cloud platform administration, including compute, networking, storage, and managed data or AI/ML platform services in production.
  • Working proficiency with platform security, including network segmentation, secrets management, encryption at rest and in transit, and key management.
  • Strong programming skills for automation, operational tooling, and infrastructure management.
  • Strong communication and documentation skills, including writing runbooks, leading postmortems, influencing operational standards, and translating technical complexity for diverse audiences.

Nice-to-haves

  • Disaster recovery planning, multi-region patterns, and capacity or cost optimization (FinOps).
  • Container orchestration with Kubernetes, progressive delivery patterns such as blue/green and canary deployments, and data lineage tooling.
  • Experience with Databricks, Azure Machine Learning, or Kubernetes-hosted infrastructure.
  • Experience in healthcare, life sciences, or other highly regulated industries with data privacy requirements.

Skills

Terraform, GitOps, CI/CD, Kubernetes, Databricks, Azure Machine Learning, Distributed Tracing, Sli/Slo, Incident Response, Infrastructure As Code, Cloud Computing, Secrets Management, Encryption, Disaster Recovery, Finops

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Kindred

Kindred

United States
Senior Infrastructure Engineer
$170k+/yrRemote5+ YOEDevOps / SRE

Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.