Skip to content
OktaOkta

Manager, Site Reliability Engineering

Leads the Site Reliability Engineering team for Auth0, setting technical direction, improving platform resilience, and guiding incident response at scale. Requires 8+ years of industry experience, 3+ years of team leadership, and deep expertise in cloud-native infrastructure, automation, and SRE practices.

About the job

Responsibilities

  • Lead the SRE team's technical direction and translate organizational vision into actionable roadmaps.
  • Drive complex, cross-functional initiatives across product and platform teams.
  • Participate hands-on in 24/7 on-call rotations, troubleshooting and remediating incidents on critical systems.
  • Build infrastructure resilience through monitoring, alerting, and automation improvements.
  • Establish reliability practices and standards centered on observability, resilience, and software engineering rigor.
  • Mentor and develop SRE talent through pair programming, design discussions, and code reviews.
  • Represent reliability in architectural reviews and strategic planning.

Requirements

  • 8+ years of total industry experience.
  • 3+ years of hands-on team leadership in SRE or software engineering roles.
  • Experience in cloud-native environments and architectures, including containers, Kubernetes, microservices, and databases.
  • Expertise with AWS, Azure, and Terraform.
  • Strong programming skills in Go or Python.
  • Experience building and maintaining production-grade tools, automation, and infrastructure solutions.
  • Knowledge of SRE principles, blameless incident response, systematic problem-solving, and software engineering approaches to operational challenges.
  • Excellent verbal and written communication skills.
  • Experience leading high-performing teams in globally distributed, remote-first environments.
  • Strategic vision, technical depth, leadership ability, and a commitment to mentoring senior engineers.
  • Ability to submit documentation establishing U.S. Person status upon hire.

Nice-to-haves

  • Experience leading reliability initiatives that improved uptime and reduced incident response times at scale.
  • Contributions to open-source infrastructure or observability tooling.
  • Experience designing incident response programs and runbook automation.

Compensation

  • Annual base salary: $182,000–$250,800 USD.
  • Equity, bonus, health, dental and vision insurance, 401(k), flexible spending account, PTO, and parental leave may be available according to applicable plans and policies.

Skills

AWS, Azure, Terraform, Kubernetes, Containers, Microservices, Go, Python, Databases, Observability, Incident Response, Infrastructure Automation

Formation Bio

Formation Bio

New York, NY
Engineering Manager, Infrastructure
$186k+/yrOn-site7+ YOEEngineering Management

Leads the infrastructure and SRE organization, owning technical direction, reliability, security, operational practices, and team development. Requires 7+ years of combined infrastructure/SRE and management experience, including 3+ years managing relevant engineering teams.

Pinterest

Pinterest

United States

Manager II, Config Deployments
$177k+/yrRemote7+ YOEEngineering Management

Leads the Config Deployments engineering team responsible for high-scale configuration distribution, feature flags, and service coordination infrastructure. The role requires 7+ years of engineering experience, infrastructure expertise, and engineering management experience.

Snowflake

Snowflake

Chicago, IL

Senior Manager, Technical Delivery
$177k+/yrOn-site10+ YOEEngineering Management

Leads Snowflake Professional Services delivery, managing Solutions Architects and Consultants while overseeing complex data migrations and AI or generative AI engagements. Requires 10+ years in customer-facing technical roles, 5+ years of people management, and substantial professional-services sales experience.

Chime

Chime

San Francisco, CA

Tech Lead Manager, Human Agent Tooling
$187k+/yrOn-site5+ YOEEngineering Management

Leads and contributes to backend platform development for Chime’s Human Agent Tooling team, guiding 3–5 engineers while owning architecture, delivery, reliability, and technical growth. Requires 5+ years of scaled production software experience, Ruby on Rails or comparable frameworks, and strong web application architecture expertise.

Crusoe

Crusoe

Shakopee, MN

Senior Manager, Data Center Facility Operations
$175k+/yrOn-site5+ YOEEngineering Management

Leads 24/7 critical facility operations for a 20 MW AI data center expanding to 40 MW, overseeing electrical, mechanical, safety, maintenance, staffing, and operational performance. Requires at least five years of data center operations management experience and expertise in infrastructure reliability and team scaling.