Skip to content
DominoDominoUnited States

Staff Site Reliability Engineer

Lead development of AI-assisted reliability tooling, own incident response, improve observability and SLOs for Domino's SaaS platform. Requires deep SRE or platform engineering experience, fluency in Kubernetes/Linux/cloud/observability, and strong Python/Go software engineering skills.

200k – 230k/yr
Remote7+ YOEDevOps / SRE

About the role

Responsibilities

  • Lead the development of internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil.
  • Improve the observability coverage and signal quality for our most critical customer-facing systems.
  • Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur.
  • Guide the development of customer and user-facing observability tools within our products.
  • Define and mature SLO/SLI frameworks for priority services.
  • Scale cloud operations practices for Domino’s single-tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades.
  • Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture.

Requirements

  • Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership.
  • Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real-world production problems.
  • A strong ability to perceive and close reliability gaps in technical products, tools and processes.
  • Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on.
  • Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done.
  • A history of improving reliability through engineering and automation, not just putting out fires manually.
  • Strong communication skills and real experience mentoring engineers or shaping technical decision-making on your team.
  • Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal.

Nice-to-Haves

  • Experience with LLM-based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams.

Skills

KubernetesLinuxPythonGoObservabilityslosliLLMsAI Toolscloud platforms

Similar roles

DevOps / SRE jobs
Bluesky Social

Staff Site Reliability Engineer

Bluesky SocialUnited States

Staff Site Reliability Engineer responsible for designing, implementing, and operating high-scale infrastructure on bare metal and cloud for Bluesky's AT Protocol federated social network. Requires 10+ years operating production systems, strong fundamentals in distributed systems, Go programming, Kubernetes, observability, and incident response.

200k – 270k/yr
Remote10+ YOEDevOps / SRE
Aurelian

Staff Infrastructure Engineer

AurelianSeattle, WA

Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.

200k – 300k/yr
On-site6+ YOEDevOps / SRE
Radar Labs

Senior / Staff Platform Engineer

Radar LabsNew York, NY

Build and operate Radar’s high-scale infrastructure, developer platform, and data systems to support 1B daily API calls. Generalist engineer focused on availability, self-serve capabilities, automation, and customer feedback.

200k – 300k/yr
On-site7+ YOEDevOps / SRE
F2

Staff Software Engineer, Infrastructure

F2San Francisco, CA

Hands-on Infrastructure Tech Lead building and scaling AWS cloud infrastructure from scratch for an AI-driven enterprise analytics platform. Owns architecture, IaC, security/compliance (SOC 2), and operational excellence.

200k – 300k/yr
Hybrid7+ YOEDevOps / SRE
Vapi

Member of Technical Staff, DevOps

VapiSan Francisco, CA

The Member of Technical Staff, DevOps will own progressive delivery, GitOps, and on-demand environment tooling to improve deployment safety and speed for engineering teams. This role requires a platform-as-a-product mindset and experience with infrastructure as code and CI/CD pipelines.

200k – 270k/yr
Hybrid5+ YOEDevOps / SRE