Senior Site Reliability Engineer, AI Infrastructure
The Senior Site Reliability Engineer will build and operate secure, observable AI/ML infrastructure across cloud platforms. The role requires at least five years of production SRE or infrastructure experience, strong Terraform and observability expertise, and hands-on incident response and automation skills.
About the job
Responsibilities
- Own service level objectives, error budgets, and reliability targets for infrastructure underpinning cloud-based platforms.
- Ensure observability across platform components and serving endpoints, including metrics, logs, traces, alert quality, and telemetry completeness.
- Design, build, and maintain infrastructure as code, operational automation, and change-control workflows for AI/ML platforms.
- Implement and maintain platform security controls, including network segmentation, secrets management, encryption, and data protection safeguards aligned with compliance requirements.
- Lead incident response and blameless postmortems.
- Validate backup, restore, and disaster recovery processes; conduct game days and resiliency testing.
- Improve platform resiliency, cost efficiency, capacity planning, and operational standards.
- Mentor engineers, influence design reviews, and collaborate across engineering teams.
Requirements
- 5+ years of experience in SRE, platform engineering, or infrastructure roles supporting production cloud environments and mission-critical applications.
- Strong proficiency with observability, including metrics, logging, distributed tracing, SLI/SLO frameworks, incident response, blameless postmortems, and on-call operations.
- Strong proficiency with Infrastructure as Code, Terraform, GitOps, and CI/CD for infrastructure and platform changes.
- Working proficiency with cloud platform administration, including compute, networking, storage, and managed data or AI/ML platform services in production.
- Working proficiency with platform security, including network segmentation, secrets management, encryption at rest and in transit, and key management.
- Strong programming skills for automation, operational tooling, and infrastructure management.
- Strong communication and documentation skills, including writing runbooks, leading postmortems, influencing operational standards, and translating technical complexity for diverse audiences.
Nice-to-haves
- Disaster recovery planning, multi-region patterns, and capacity or cost optimization (FinOps).
- Container orchestration with Kubernetes, progressive delivery patterns such as blue/green and canary deployments, and data lineage tooling.
- Experience with Databricks, Azure Machine Learning, or Kubernetes-hosted infrastructure.
- Experience in healthcare, life sciences, or other highly regulated industries with data privacy requirements.
Skills
Terraform, GitOps, CI/CD, Kubernetes, Databricks, Azure Machine Learning, Distributed Tracing, Sli/Slo, Incident Response, Infrastructure As Code, Cloud Computing, Secrets Management, Encryption, Disaster Recovery, Finops
Similar jobs
DevOps / SRE jobsThe Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.