Skip to content
ForgeForge

Manager, Site Reliability Engineer

Leads Forge’s Site Reliability Engineering team, overseeing availability, incident response, observability, production operations, and reliability practices. Requires substantial infrastructure or software engineering experience, people leadership, and expertise with cloud platforms and operational tooling.

About the job

Responsibilities

  • Manage the Site Reliability Engineering team responsible for maintaining high availability for Forge systems.
  • Drive incident management practices, including response, mitigation, follow-up, and post-incident learning.
  • Build, improve, and manage observability infrastructure, including monitoring, alerting, dashboards, and operational metrics.
  • Improve monitoring coverage and alert quality to reduce noise and shorten time to detect.
  • Champion reliability practices including service ownership, operational readiness, disaster recovery, and production support standards.
  • Contribute to technical design, architecture, automation, infrastructure, and team delivery.
  • Troubleshoot production issues with engineering teams, identify recurring problems, and improve system reliability.
  • Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
  • Partner with Security, Compliance, and Risk teams to ensure reliability and infrastructure practices support a regulated business.

Requirements

  • 5+ years leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
  • 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
  • Bachelor’s degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
  • Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
  • Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
  • Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
  • Strong technical judgment, communication skills, and ability to influence engineering and non-engineering stakeholders.

Preferred Qualifications

  • Experience in FinTech, financial services, or another regulated industry.
  • Experience with AWS and/or Azure.
  • Familiarity with Kubernetes, container platforms, infrastructure as code, Terraform, Ansible, or similar automation tools.
  • Experience with observability platforms such as Datadog or CloudWatch.
  • Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure.
  • Experience supporting growth-stage companies while balancing speed, scale, reliability, and operational discipline.

Compensation

  • San Francisco/Bay Area, CA or New York, NY: $150,000–$220,000 annual salary plus annual bonus.

Skills

Site Reliability Engineering, DevOps, Cloud Infrastructure, Distributed Systems, Observability, Incident Response, CI/CD, Infrastructure Automation, AWS, Azure, Kubernetes, Terraform, Ansible, Datadog, CloudWatch

Crusoe

Crusoe

Denver, CO

Senior Manager, Commissioning
$160k+/yrOn-site10+ YOEEngineering Management

Leads commissioning programs across multiple data center projects, managing commissioning teams, third-party agents, and stakeholder coordination from pre-functional testing through turnover. Requires 10+ years of mission-critical commissioning experience, 5+ years of leadership, engineering knowledge, and a bachelor’s degree.

LegitScript

LegitScript

United States

Manager, Platform Engineering
$160k+/yrRemote8+ YOEEngineering Management

Leads a hands-on platform engineering team responsible for AWS infrastructure, Kubernetes deployment paths, developer self-service, CI/CD governance, reliability, and audit readiness. The role requires deep infrastructure experience, Terraform expertise, production Kubernetes operations, and people leadership.

ZoomInfo

ZoomInfo

United States

Manager, Data Platform
$140k+/yrRemote8+ YOEEngineering Management

Leads strategy, architecture, operations, and modernization of an enterprise data platform while managing and mentoring data platform engineers. Requires 8+ years in data engineering, platforms, or architecture and experience with production cloud data systems.

ZoomInfo

ZoomInfo

Bethesda, MD

Manager, Software Engineering - Integrations
$140k+/yrHybrid6+ YOEEngineering Management

Leads and develops a distributed engineering team responsible for scalable customer integrations and backend systems. The role combines people management, technical strategy, large-scale system design, and expertise in CRM integrations, APIs, authentication, databases, and cloud platforms.

Crusoe

Crusoe

Brighton, CO

Senior Manager, QMS
$135k+/yrOn-site8+ YOEEngineering Management

Leads the enterprise QMS across manufacturing facilities, overseeing ISO 9001 compliance, audits, CAPA/NCR governance, document control, quality metrics, risk management, and new-site readiness. Requires a bachelor’s degree and 7+ years of quality or manufacturing-quality experience.