Senior Software Engineer, Site Reliability
Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.
About the job
Responsibilities
- Own complex production-support escalations and ticket triage, troubleshooting and resolving issues alongside reliability work.
- Partner with Software Engineering to investigate production issues, identify root causes and reliability risks, and drive permanent fixes.
- Apply and mature Site Reliability Engineering practices, including automation, continuous improvement, shared ownership, and toil reduction.
- Lead incident response through triage, mitigation, recovery, root-cause analysis, and blameless post-incident reviews.
- Build observability with meaningful metrics, logs, traces, dashboards, and actionable alerts.
- Define and mature service-level indicators (SLIs) and service-level objectives (SLOs), including error budgets.
- Develop synthetic monitoring for critical customer journeys.
- Automate recurring operational toil through tooling, process improvements, or permanent fixes.
- Use AI-assisted tools and source-code repositories for triage, troubleshooting, code analysis, automation, and investigation.
- Participate in a rotating on-call schedule, primarily during business hours, with limited after-hours and weekend support.
Requirements
- Hands-on Site Reliability Engineering experience applying software engineering practices to production reliability and helping establish or mature SRE practices.
- Strong knowledge of SLIs, SLOs, error budgets, observability, automation, and toil reduction.
- Experience building monitoring, dashboards, alerts, and telemetry with tools such as Honeycomb, New Relic, Grafana, CloudWatch, Kibana, or similar.
- Experience with production incident management, root-cause analysis, blameless post-incident reviews, and corrective-action follow-through.
- Strong programming and scripting skills for troubleshooting application code and building automation and operational tooling.
- Strong SQL and relational-database skills for production troubleshooting and safe data correction; PostgreSQL preferred.
- Strong code literacy and debugging skills, including navigating unfamiliar codebases, understanding application flow, reviewing code and change history, and identifying reliability issues.
- Experience troubleshooting cloud-hosted applications using source code, logs, APIs, telemetry, event streams, and databases.
- Comfort navigating application stacks involving PHP, .NET, and Node.js; deep expertise in each is not required.
- Experience using AI-assisted tools in day-to-day engineering workflows.
- Strong communication and collaboration skills across Software Engineering, Product, Support, DevOps, and other technical teams.
Compensation and Benefits
- Salary: $114,800–$150,000 annually.
- Eligibility for a discretionary bonus.
- Health, vision, and dental insurance, plus 24/7 healthcare access.
- 20 PTO days, 3 flex days, 4 volunteer days, 12 paid holidays, and paid parental leave.
- 401(k) match.
- Company-provided equipment.
Skills
Site Reliability Engineering, Observability, Slis, SLOs, Error Budgets, Incident Management, Grafana, CloudWatch, SQL, Postgres, PHP, .Net, Node.js, Automation, Synthetic Monitoring
Similar jobs
DevOps / SRE jobsSenior network engineer responsible for designing, operating, and securing MongoDB’s global network and VPN infrastructure. The role requires 6+ years of networking or systems engineering experience, strong enterprise networking expertise, automation skills, and the ability to lead complex infrastructure initiatives.
Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.
Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.