Lead development of AI-assisted reliability tooling, own incident response, improve observability and SLOs for Domino's SaaS platform. Requires deep SRE or platform engineering experience, fluency in Kubernetes/Linux/cloud/observability, and strong Python/Go software engineering skills.
200k – 230k/yr
Remote7+ YOEDevOps / SRE
About the role
Responsibilities
Lead the development of internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil.
Improve the observability coverage and signal quality for our most critical customer-facing systems.
Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur.
Guide the development of customer and user-facing observability tools within our products.
Define and mature SLO/SLI frameworks for priority services.
Scale cloud operations practices for Domino’s single-tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades.
Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture.
Requirements
Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership.
Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real-world production problems.
A strong ability to perceive and close reliability gaps in technical products, tools and processes.
Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on.
Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done.
A history of improving reliability through engineering and automation, not just putting out fires manually.
Strong communication skills and real experience mentoring engineers or shaping technical decision-making on your team.
Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal.
Nice-to-Haves
Experience with LLM-based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams.
Staff Site Reliability Engineer responsible for designing, implementing, and operating high-scale infrastructure on bare metal and cloud for Bluesky's AT Protocol federated social network. Requires 10+ years operating production systems, strong fundamentals in distributed systems, Go programming, Kubernetes, observability, and incident response.
200k – 270k/yr
Remote10+ YOEDevOps / SRE
Staff Infrastructure Engineer
AurelianSeattle, WA
Staff Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents that support 911 emergency call centers. Requires 6+ years in infrastructure/platform/backend roles with experience in reliability and scale.
200k – 300k/yr
On-site6+ YOEDevOps / SRE
Senior / Staff Platform Engineer
Radar LabsNew York, NY
Build and operate Radar’s high-scale infrastructure, developer platform, and data systems to support 1B daily API calls. Generalist engineer focused on availability, self-serve capabilities, automation, and customer feedback.
200k – 300k/yr
On-site7+ YOEDevOps / SRE
Staff Software Engineer, Infrastructure
F2San Francisco, CA
Hands-on Infrastructure Tech Lead building and scaling AWS cloud infrastructure from scratch for an AI-driven enterprise analytics platform. Owns architecture, IaC, security/compliance (SOC 2), and operational excellence.
200k – 300k/yr
Hybrid7+ YOEDevOps / SRE
Member of Technical Staff, DevOps
VapiSan Francisco, CA
The Member of Technical Staff, DevOps will own progressive delivery, GitOps, and on-demand environment tooling to improve deployment safety and speed for engineering teams. This role requires a platform-as-a-product mindset and experience with infrastructure as code and CI/CD pipelines.