Director of Site Reliability Engineering
Lead and develop a distributed SRE team, setting vision and operating model for reliability practices. Own core infrastructure services (Kubernetes, CI/CD, observability) and drive service ownership frameworks across engineering teams.
About the job
Responsibilities
- Lead, coach, and develop a distributed SRE team, setting vision, charter, operating model, priorities, and success measures
- Define and roll out a Service Ownership & Maturity Framework across engineering
- Own and improve core engineering infrastructure services including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation
- Help engineering teams improve service ownership through standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices
- Make reliability, operational maturity, infrastructure health, and developer productivity measurable through metrics and operational intelligence
- Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability
- Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team
- Build paved paths and self-service infrastructure that reduce toil and lower cognitive load
- Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT on infrastructure, access management, cloud operations, and controls
- Evaluate AI-assisted and agentic workflows for infrastructure operations, service ownership, developer workflows, or toil reduction
Requirements
- 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles
- 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers
- Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk
- 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments
- 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices
- Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability
- Pragmatic approach to tooling decisions (build vs. buy vs. adapt vs. simplify vs. retire)
- Ability to operate effectively in a small or mid-sized engineering organization
- Clear executive communication skills and ability to partner directly with a CTO and senior engineering leaders
Nice-to-Haves
- Experience leading SRE, infrastructure, or platform work in a lean, high-agency organization
- Experience supporting globally distributed teams or 24/7 operational coverage
- Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil
- Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls
- Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability technical ecosystems
- Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline
- Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity
Compensation & Benefits
- Base salary range: $210,000 - $310,000
- Lumen-denominated grants
- Competitive health, dental & vision coverage (most plans 100% covered for employee + dependents)
- Flexible time off + 15 company holidays including company-wide holiday break
- Generous paid parental leave + paid pregnancy disability leave
- Gym reimbursement ($80/month)
- Life & ADD (up to $50K), Short & Long term disability
- 401K with 4% match
- Health & Dependent Care FSA Accounts
- Commuter benefits ($250/month employer contribution)
- Health Savings Account (HSA) with monthly employer contribution
- Family building benefits through Kindbody
- Wellbeing benefits (One Medical, Rightway, Headspace)
- L&D budget of $1,500/year
- Daily lunch and snacks in office
- Company retreats
Skills
SRE, Kubernetes, AWS, GCP, CI/CD, Infrastructure As Code, Observability, Slos/Slis, Incident Response, Postmortems, Cloud Infrastructure, Distributed Systems, Infrastructure Automation, Secrets Management
Similar jobs
DevOps / SRE jobsBuild and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
This principal-level role owns operational excellence for a hyperscale AI data center network fleet, leading readiness, high-risk changes, audits, and incident resolution across sites. It requires extensive mission-critical network operations experience, routing and optical networking expertise, and 50–75% travel.
Leads Snowflake’s cloud infrastructure performance strategy by evaluating new hardware, building benchmark and validation systems, and translating performance data into pricing, capacity, and rollout decisions. Requires 12+ years in performance, systems, or infrastructure engineering and deep cloud hardware expertise.
As a Principal Operations Engineer, Mechanical, you will be the senior technical authority for mechanical and cooling infrastructure across hyperscale AI data centers. You will lead site assessments, drive operational readiness, review designs, and ensure precision execution of critical systems.
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.