Skip to content

Site Reliability Engineer (Mid/Senior/Staff)

Owns and operates Kubernetes infrastructure at scale, builds CI/CD pipelines, automates incident resolution with AI, defines SLOs, and drives reliability through chaos engineering. Requires 5+ years managing production systems, Kubernetes expertise, and monitoring tools like Prometheus/Grafana.

180k – 250kSan Francisco, CADevOps / SREHybrid5+ YOE

About the role

Key Responsibilities

  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
  • Build dashboards, alerting, and anomaly detection across our systems
  • Define and enforce SLOs and build out incident response processes
  • Manage and improve our networking, load balancing, and service mesh configurations
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering

Requirements

  • 5+ years experience in managing critical production systems and software development workflows
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
  • Proficiency in Python and either Go or Bash for tooling and automation
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Nice to have

  • Experience with managing GPU and AI/ML workloads
  • Experience with kernel-based monitoring and routing (eBPF, XDP)
  • Experience with security tooling (Falco, Coroot, SIEM)
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
  • Experience with distributed storage systems (Ceph, Longhorn, etc.)

Compensation

$180,000-250,000 plus equity + benefits (Range is based across 3 levels)

Skills

KubernetesTerraformAnsiblePrometheusGrafanaCI/CDGitOpsFluxcdArgo CDPythonGoLinux NetworkingEbpfCiliumDatadog

Similar roles

DevOps / SRE jobs

Staff Software Engineer, Cloud FinOps

Staff-level engineer driving company-wide cloud cost optimization and FinOps initiatives across engineering teams. Requires 5+ years infrastructure experience and 2+ years FinOps/cloud cost management.

180k – 240kUnited StatesDevOps / SRERemote5+ YOEAWSJava

Staff Engineer, AI Productivity

Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.

180k – 400kUnited StatesDevOps / SRERemote7+ YOEGoDevin

Staff Infrastructure Engineer

Build infrastructure, observability, and developer tooling for a realtime AI platform serving 911 centers. Requires 6+ years infrastructure/platform/backend experience and comfort across the full stack.

180k – 240kSeattle, WADevOps / SREOn-site6+ YOELoggingClickHouse

Staff Software Engineer, AI Developer Tools

Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.

180k – 245kDenver, CO +3DevOps / SREHybrid8+ YOERAGLLMs

Staff Infrastructure Engineer

Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.

180k – 235kUnited StatesDevOps / SRERemote5+ YOEGoAWS