Skip to content
DISQODISQOLos Angeles, CA

Senior Site Reliability Engineer

Senior Site Reliability Engineer building agentic AI platforms, MCP servers, and automation tools to achieve zero KTLO. Drives SecDevOps culture with focus on security, reliability, observability while partnering on product roadmap and incident response. Requires 6+ years SRE/DevOps experience plus deep expertise in AWS, EKS, Terraform, Kubernetes.

170k – 190k/yr
Hybrid6+ YOEDevOps / SRE

About the role

What you will do

  • Partner with a team of high-performing engineers and developers who are focused on delivering best in class software products
  • Function as a Change Agent to introduce and evangelize our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability
  • Building Zero trust security designs and adapting DevSecOps mindset when building new services
  • Cultivate an automation-first attitude and work to champion code-centric solutions throughout our department to improve velocity and deliverability
  • Participate in the design process, representing, solving and planning for security, reliability and cost effectiveness prism in the product roadmap
  • Collaborate and communicate with cross-functional colleagues belonging to the same job family to drive SecDevOps centric cross-org initiatives, tooling and standards
  • Participate in incident response in collaboration with application owners and the platform team
  • Build tools that empower capacity planning and demand forecasting, software performance analysis, and system tuning
  • Design and build agentic AI platforms and AI agents — including home-grown MCP (Model Context Protocol) servers — to drive our organization toward a 0 KTLO north star, shifting engineering effort away from routine operational maintenance and toward high-value, forward-looking initiatives
  • Develop tooling that enable our product teams to through self service
  • Think implementing long-term solutions that are complete mechanisms by building tools, driving adoption and inspecting results for tuning

What you bring to the role

  • At least 6+ years of experience as SRE, DevOps or equivalent engineering roles
  • Hands-on experience building AI SRE/DevOps Agents and using AI agentic frameworks
  • Hands-on experience building internal MCPs for AI Agentic use
  • Strong hands-on experience working with foundational AWS services
  • Expert level experience working with AWS EKS
  • Strong hands-on hands-on experience with ArgoCD
  • Strong hands-on experience Terraform
  • Experience with at least one language for automation - Bash, Python, Golang, etc.
  • Experience deploying Serverless application using SAM
  • Experience building and managing CI/CD pipelines
  • Strong hands-on experience in Linux architecture, microservices and container orchestration (Docker, Kubernetes, etc.)
  • Experience with monitoring tools (New Relic, Prometheus, Grafana, Loki etc.)
  • Experience working on infrastructure projects in an Agile environment
  • Great communication, collaboration and presentation skills
  • Ability to team up with people from different disciplines and drive for a win-win
  • Ability to adapt and build security orchestration and automation at scale

Skills

AWSEKSArgo CDTerraformPythonGosamCI/CDLinuxDockerKubernetesnew relicPrometheusGrafanaloki

Similar roles

DevOps / SRE jobs
Crusoe

Senior Performance Engineer

CrusoeSan Francisco, CA

Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.

170k – 205k/yr
On-site5+ YOEDevOps / SRE
Lightning AI

Senior Network Engineer

Lightning AINew York, NY +2

Senior Network Engineer responsible for designing, deploying, and optimizing large-scale NVIDIA InfiniBand fabrics and UFM for AI/ML GPU clusters. Requires 10+ years data center networking experience with deep expertise in InfiniBand, spine-leaf architectures, automation, and HPC environments.

170k – 210k/yr
On-site10+ YOEDevOps / SRE
Illumio

Sr. Site Reliability Engineer

IllumioSunnyvale, CA

Senior Site Reliability Engineer responsible for monitoring, incident response, and optimizing the reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure and SaaS services. Requires 5+ years SRE experience with strong cloud platform expertise.

170k – 196k/yr
On-site5+ YOEDevOps / SRE
Render

Software Engineer, Compute Infrastructure

RenderSan Francisco, CA

Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.

170k – 290k/yr
Remote7+ YOEDevOps / SRE
Skydio

Senior Software Engineer, Developer Productivity Cloud Infrastructure

SkydioSan Mateo, CA

Senior engineer focused on developer productivity and cloud infrastructure. Designs scalable internal tools, re-architects build systems, and improves CI/CD workflows using Terraform, Go/Python/C++.

170k – 240k/yr
Hybrid5+ YOEDevOps / SRE