Skip to content

Senior Software Engineer, Site Reliability

Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.

About the job

Responsibilities

  • Own complex production-support escalations and ticket triage, troubleshooting and resolving issues alongside reliability work.
  • Partner with Software Engineering to investigate production issues, identify root causes and reliability risks, and drive permanent fixes.
  • Apply and mature Site Reliability Engineering practices, including automation, continuous improvement, shared ownership, and toil reduction.
  • Lead incident response through triage, mitigation, recovery, root-cause analysis, and blameless post-incident reviews.
  • Build observability with meaningful metrics, logs, traces, dashboards, and actionable alerts.
  • Define and mature service-level indicators (SLIs) and service-level objectives (SLOs), including error budgets.
  • Develop synthetic monitoring for critical customer journeys.
  • Automate recurring operational toil through tooling, process improvements, or permanent fixes.
  • Use AI-assisted tools and source-code repositories for triage, troubleshooting, code analysis, automation, and investigation.
  • Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.

Requirements

  • Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
  • Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
  • Experience building monitoring, dashboards, alerts, and telemetry with tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar.
  • Experience with production incident management, root-cause analysis, blameless post-incident reviews, and corrective-action follow-through.
  • Strong programming and scripting skills for troubleshooting application code and building automation and operational tooling.
  • Strong SQL and relational-database skills for production troubleshooting and safe data correction; PostgreSQL preferred.
  • Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying reliability issues.
  • Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases.
  • Comfort navigating application stacks involving PHP, .NET, and Node.js; deep expertise in each is not required.
  • Experience using AI-assisted tools in day-to-day engineering workflows.
  • Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams.

Compensation and Benefits

  • Salary: $114,800–$150,000 annually.
  • Eligibility for a discretionary bonus.
  • Health, vision, and dental insurance, plus 24/7 healthcare access.
  • 20 PTO days, 3 flex days, 4 volunteer days, 12 paid holidays, and paid parental leave.
  • 401(k) match.
  • Company-provided equipment.

Skills

Site Reliability Engineering, Observability, Slis, SLOs, Error Budgets, Incident Management, Grafana, CloudWatch, SQL, Postgres, PHP, .Net, Node.js, Automation, Synthetic Monitoring

MongoDB

MongoDB

Palo Alto, CA
Senior Network Engineer
$118k+/yrHybrid6+ YOEDevOps / SRE

Senior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.

PrizePicks

PrizePicks

United States

Senior Site Reliability Engineer
$120k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.

Axle

Axle

Rockville, MD

Lead Scientific Imaging Systems Engineer
$120k+/yrOn-site7+ YOEDevOps / SRE

Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.

Upstart

Upstart

United States

Senior DevOps Engineer
$136k+/yrRemote3+ YOEDevOps / SRE

Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.