# Staff Site Reliability Engineer

**Company:** [Domino](https://hotfix.jobs/companies/domino-data-lab)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $200k – $230k/yr
**Experience:** 7+ years
**Skills:** Kubernetes, Linux, Python, Go, Observability, slo, sli, LLMs, AI Tools, cloud platforms
**Posted:** 2026-07-27

> Lead development of AI-assisted reliability tooling, own incident response, improve observability and SLOs for Domino's SaaS platform. Requires deep SRE or platform engineering experience, fluency in Kubernetes/Linux/cloud/observability, and strong Python/Go software engineering skills.

## Job Description

## Responsibilities
- Lead the development of internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil.
- Improve the observability coverage and signal quality for our most critical customer-facing systems.
- Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur.
- Guide the development of customer and user-facing observability tools within our products.
- Define and mature SLO/SLI frameworks for priority services.
- Scale cloud operations practices for Domino’s single-tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades.
- Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture.

## Requirements
- Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership.
- Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real-world production problems.
- A strong ability to perceive and close reliability gaps in technical products, tools and processes.
- Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on.
- Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done.
- A history of improving reliability through engineering and automation, not just putting out fires manually.
- Strong communication skills and real experience mentoring engineers or shaping technical decision-making on your team.
- Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal.

## Nice-to-Haves
- Experience with LLM-based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams.

## Similar roles

- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/f16271a9-7826-4f8b-b02d-beef4aeaffa1) - Bluesky Social - Remote - $200k – $270k/yr
- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/75f07283-34de-4989-b193-11c3c3676fd2) - Aurelian - Seattle, WA - $200k – $300k/yr
- [Senior / Staff Platform Engineer](https://hotfix.jobs/jobs/1e15e953-b6f5-46a6-afc4-36a6811e809e) - Radar Labs - New York, NY - $200k – $300k/yr
- [Staff Software Engineer, Infrastructure](https://hotfix.jobs/jobs/ecfe4e28-57ed-4e1d-a9fb-5917d180aeec) - F2 - San Francisco, CA - $200k – $300k/yr
- [Member of Technical Staff, DevOps](https://hotfix.jobs/jobs/c99fed69-8e72-438a-9f3d-3e18894667e7) - Vapi - San Francisco, CA - $200k – $270k/yr

**Apply:** https://hotfix.jobs/jobs/a3ec9454-be8f-4d77-bef0-6de1db57e3f6
**Canonical:** https://hotfix.jobs/jobs/a3ec9454-be8f-4d77-bef0-6de1db57e3f6