Skip to content

Engineering Manager, Infrastructure

Leads the infrastructure and SRE organization, owning technical direction, reliability, security, operational practices, and team development. Requires 7+ years of combined infrastructure/SRE and management experience, including 3+ years managing relevant engineering teams.

About the job

Responsibilities

  • Hire, coach, and develop infrastructure and site reliability engineers; set goals, provide feedback, and manage performance.
  • Set the technical direction and roadmap for infrastructure and SRE across reliability, security, cost, and developer experience.
  • Build operating practices and paved paths that support rapid, reliable, and secure software delivery.
  • Partner with engineering and data science teams to deploy secure, observable, reliable software across containerized applications, internal tools, ML pipelines, and AI model-training workloads.
  • Oversee infrastructure services across multiple cloud providers and environments, including compute, load balancers, databases, secrets management, and validated or GxP-relevant workloads.
  • Ensure quality and optimization of infrastructure as code, container images, and CI/CD pipelines.
  • Review requirements, design documents, operating procedures, and maintenance practices.
  • Drive infrastructure and SRE training, mentoring, documentation, and knowledge sharing.
  • Own support rotation coverage, escalation paths, incident retrospectives, and operational improvements.
  • Manage budgeting, resourcing, estimates, tracking, and reporting for infrastructure initiatives.

Requirements

  • 7+ years of combined hands-on infrastructure/SRE experience and people management.
  • 3+ years of direct management experience managing Site Reliability, Infrastructure, or DevOps engineers.
  • Track record of hiring, developing, and retaining strong engineers.
  • Strong operational and reliability judgment, including advanced diagnostics, incident response, and root-cause analysis.
  • Bias toward automation and ability to apply sound judgment to operational tradeoffs.
  • Point of view on AI-native engineering and the operational requirements of an agentic software development lifecycle.
  • Exceptional collaboration and communication skills across technical and non-technical stakeholders.
  • Experience managing or coordinating technical projects and programs, including budgeting and reporting.
  • Working knowledge of AWS and Snowflake.
  • Working knowledge of Docker, GitHub, Kubernetes, Python, Terraform/OpenTofu, and virtual networking.
  • Familiarity with multi-tenant COTS and FOSS software applications.

Nice-to-haves

  • Experience with Azure, Google Cloud, or Vercel.
  • Digital forensics experience.
  • Experience supporting regulated or validated workloads.
  • Pharmaceutical or biotech experience.

Compensation

  • Total compensation range: $185,500–$232,000.
  • Compensation may include equity, benefits, and perks.

Skills

AWS, Snowflake, Docker, GitHub, Kubernetes, Python, Terraform, Opentofu, Virtual Networking, CI/CD

Chime

Chime

San Francisco, CA

Tech Lead Manager, Human Agent Tooling
$187k+/yrOn-site5+ YOEEngineering Management

Leads and contributes to backend platform development for Chime’s Human Agent Tooling team, guiding 3–5 engineers while owning architecture, delivery, reliability, and technical growth. Requires 5+ years of scaled production software experience, Ruby on Rails or comparable frameworks, and strong web application architecture expertise.

Okta

Okta

New York, NY
Manager, Site Reliability Engineering
$182k+/yrOn-site8+ YOEEngineering Management

Leads the Site Reliability Engineering team for Auth0, setting technical direction, improving platform resilience, and guiding incident response at scale. Requires 8+ years of industry experience, 3+ years of team leadership, and deep expertise in cloud-native infrastructure, automation, and SRE practices.

Databricks

Databricks

Mountain View, CA

Engineering Manager, App Traffic
$190k+/yrOn-site9+ YOEEngineering Management

Leads and develops the App Traffic engineering team building reliable, scalable service-mesh and networking infrastructure across multiple clouds. Requires 9+ years of software engineering experience, including engineering leadership and distributed-systems or infrastructure expertise.

Snowflake

Snowflake

Houston, TX

Senior District Manager, Majors Expansion
$190k+/yrRemote15+ YOEEngineering Management

Leads and develops a Majors sales organization focused on expansion and new-logo acquisition, coaching Account Executives through complex enterprise sales cycles, negotiations, and technical migrations. Requires 15+ years selling software or cloud solutions, 10+ years managing sales teams, and financial-services experience.

Databricks

Databricks

New York, NY

Engineering Manager, CustomerLake Profile Agents
$190k+/yrOn-site8+ YOEEngineering Management

Leads the team building Profile Agents for Databricks’ CustomerLake agentic customer data platform, owning delivery, architecture, product direction, and team growth. The role requires engineering management experience, deep production data or distributed-systems expertise, and hands-on experience applying agents or LLMs to data engineering.