# Senior Engineering Manager, Site Reliability

**Company:** [Upstart](https://hotfix.jobs/companies/upstart)
**Location:** Remote
**Role:** Engineering Management
**Salary:** $195k – $270k/yr
**Experience:** 7+ years
**Skills:** site reliability engineering, Incident Management, Observability, Distributed Systems, Kubernetes, AWS, Datadog, Grafana, Prometheus, OpenTelemetry, service level objectives, Cloud Infrastructure
**Posted:** 2026-07-17

> Lead the Site Reliability Engineering team to improve incident management, observability, operational readiness, and overall system reliability. Requires 7+ years software/SRE experience and 5+ years managing reliability functions, with deep expertise in distributed systems and production operations.

## Job Description

## How you’ll make an impact

### Team Leadership and Execution
- Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
- Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
- Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
- Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
- Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
- Set a high bar for technical quality, operating rigor, and executive communication
- Develop engineers and leaders who can independently own complex reliability initiatives

### Incident Management and Learning
- Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery
- Establish clear standards for managing high severity incidents and provide visible leadership during critical events
- Improve postmortem quality and ensure incident learnings result in durable engineering improvements
- Identify recurring failure patterns and drive systemic solutions across teams
- Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction

### Observability and Reliability Engineering
- Improve the quality, accessibility, and trustworthiness of signals used to understand production health
- Drive consistent practices across metrics, logs, traces, alerting, and service health
- Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
- Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil
- Define measurable reliability outcomes and use data to prioritize investments and communicate impact
- Partner with platform and product engineering teams to embed reliability into standard engineering workflows

### Operational Readiness and Resilience
- Establish scalable operational readiness standards for new services, major launches, and architectural changes
- Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
- Identify systemic reliability risks and partner with engineering teams to prioritize and address them
- Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
- Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
- Align stakeholders and dependencies before critical launches and engineering decisions

## Minimum Qualifications
- 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
- Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
- Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
- Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
- Experience leading high severity incident response and improving incident management practices at scale
- Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
- Track record of hiring, developing, and retaining high performing engineers and engineering leaders
- Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

## Preferred Qualifications
- Experience operating large scale, highly available distributed systems
- Experience implementing or evolving service-level objectives and error-budget practices
- Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
- Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
- Familiarity with Kubernetes, AWS, and modern cloud native architectures
- Experience supporting major platform or product launches

## Similar roles

- [Senior Engineering Manager, Servicing](https://hotfix.jobs/jobs/491fbb60-2372-49fa-b8ae-aff0d95891d5) - Upstart - Remote - $195k – $270k/yr
- [Senior Manager, Privacy Engineering](https://hotfix.jobs/jobs/ad589e46-13a7-4acf-a8e4-5b9cd4d3fe5d) - Upstart - Remote - $195k – $270k/yr
- [Senior Engineering Manager, Upstart Bank](https://hotfix.jobs/jobs/34703124-c4a3-4f6d-a2be-d59f39da5a7b) - Upstart - Remote - $195k – $257k/yr
- [Senior Engineering Manager, Upstart Bank](https://hotfix.jobs/jobs/84e496c1-c753-47e1-a60a-3b247620f9e7) - Upstart - Remote - $195k – $270k/yr
- [Senior Engineering Manager, Platform Delivery](https://hotfix.jobs/jobs/7ebec76f-a4ab-473a-a50d-6beecc7feb86) - Upstart - Remote - $195k – $270k/yr

**Apply:** https://hotfix.jobs/jobs/1c5076a9-895a-43d5-8562-eb8e0fc2df22
**Canonical:** https://hotfix.jobs/jobs/1c5076a9-895a-43d5-8562-eb8e0fc2df22