Lead the Site Reliability Engineering team to improve incident management, observability, operational readiness, and overall system reliability. Requires 7+ years software/SRE experience and 5+ years managing reliability functions, with deep expertise in distributed systems and production operations.
195k – 270k/yr
Remote7+ YOEEngineering Management
About the role
How you’ll make an impact
Team Leadership and Execution
Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
Set a high bar for technical quality, operating rigor, and executive communication
Develop engineers and leaders who can independently own complex reliability initiatives
Incident Management and Learning
Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery
Establish clear standards for managing high severity incidents and provide visible leadership during critical events
Improve postmortem quality and ensure incident learnings result in durable engineering improvements
Identify recurring failure patterns and drive systemic solutions across teams
Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction
Observability and Reliability Engineering
Improve the quality, accessibility, and trustworthiness of signals used to understand production health
Drive consistent practices across metrics, logs, traces, alerting, and service health
Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
Define measurable reliability outcomes and use data to prioritize investments and communicate impact
Partner with platform and product engineering teams to embed reliability into standard engineering workflows
Operational Readiness and Resilience
Establish scalable operational readiness standards for new services, major launches, and architectural changes
Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
Identify systemic reliability risks and partner with engineering teams to prioritize and address them
Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
Align stakeholders and dependencies before critical launches and engineering decisions
Minimum Qualifications
5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
Experience leading high severity incident response and improving incident management practices at scale
Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
Track record of hiring, developing, and retaining high performing engineers and engineering leaders
Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams
Preferred Qualifications
Experience operating large scale, highly available distributed systems
Experience implementing or evolving service-level objectives and error-budget practices
Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
Familiarity with Kubernetes, AWS, and modern cloud native architectures
Experience supporting major platform or product launches
Skills
site reliability engineeringIncident ManagementObservabilityDistributed SystemsKubernetesAWSDatadogGrafanaPrometheusOpenTelemetryservice level objectivesCloud Infrastructure
Lead a 10+ engineer team on Upstart's Get Back On Track (GBOT) servicing platform. Partner with Product, Risk, Operations and ML to build scalable, compliant systems for borrower repayment and debt management using workflow automation and event-driven architectures. Requires 10+ years experience including 5 years people management.
195k – 270k/yr
Remote10+ YOEEngineering Management
Senior Manager, Privacy Engineering
UpstartUnited States
Lead Upstart's Privacy Engineering team to build scalable privacy infrastructure, data governance, and privacy-by-design controls for their AI lending platform. Requires 8+ years engineering experience (3+ in people management) and expertise translating regulatory requirements into technical systems.
195k – 270k/yr
Remote8+ YOEEngineering Management
Senior Engineering Manager, Upstart Bank
UpstartUnited States
Lead engineering teams building critical infrastructure for Upstart Bank, a modern AI-powered banking platform. Translate regulatory and product needs into scalable, audit-ready systems for payments, funding, reporting and integrations in a regulated fintech environment.
195k – 257k/yr
Remote8+ YOEEngineering Management
Senior Engineering Manager, Upstart Bank
UpstartUnited States
Leads engineering teams building critical Upstart Bank platform components like payments, funding, and reporting. Translates regulatory and business needs into scalable, compliant infrastructure while managing cross-functional delivery and team growth. Requires 8+ years engineering and 3+ years management experience.
195k – 270k/yr
Remote8+ YOEEngineering Management
Senior Engineering Manager, Platform Delivery
UpstartUnited States
Leads engineering team building reliable CI/CD pipelines, self-service developer platforms, and tools to accelerate software delivery. Requires 5+ years management and 7+ years in software/DevOps/platform engineering with strong technical expertise in AWS and distributed systems.