# Senior Site Reliability Engineer

**Company:** [Tulip](https://hotfix.jobs/companies/tulip)
**Location:** Somerville, MA
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Observability, Prometheus, Grafana, loki, tempo, mimir, OpenTelemetry, promql, Kubernetes, Go, TypeScript, MongoDB, Postgres
**Posted:** 2026-07-23

> Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.

## Job Description

## About You
- Reason about systems at scale: edge cases, failure modes, and life cycles
- Excited about setting the technical agenda and coming up with novel, broad ideas
- Keep up with newest AI advancements in Observability & Monitoring
- Know what a good SLA looks like and can teach others
- Communicate as well as you code; value discussion and clear, frequent communication in teams

## Skills
- 5+ years of experience with open source Observability tools (Loki, Grafana, Tempo, Mimir stack)
- Hands-on experience instrumenting distributed systems using OpenTelemetry
- Managing metrics pipelines with Prometheus at scale
- Direct experience developing and distributing Claude Skills, Gemini Gems, or other generic AI processes; iterate on their efficacy
- Experience working with time-series data, ideally using PromQL

## Key Responsibilities
- Mentor and evangelize observability best practices, SLIs/SLOs, and reliability culture across engineering teams
- Contribute to and maintain triage & remediation processes as a player/coach
- Perform incident response and debug production issues across the entire stack
- Design, build, and maintain core infrastructure & tooling used by all engineering teams

## Tech Stack
- TS and Go services running on Kubernetes
- MongoDB and Postgres databases
- Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry observability tooling

## Key Collaborators
- Engineering
- Edge
- DevOps
- Hardware

## Benefits
- Direct impact on product and culture
- Company equity
- Competitive benefits: Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, FSA, Commuter Benefits, Parental Leave, 401(K)
- Flexible work schedule and unlimited vacation
- Virtual company events and happy hours
- Fitness subsidies

## Similar roles

- [Senior Software Engineer - Snowpark Container Service](https://hotfix.jobs/jobs/47626f55-409c-4469-9cea-906cfd5683e3) - Snowflake - Bellevue, WA - $200k – $288k/yr
- [Senior Software Engineer, Infrastructure & Systems](https://hotfix.jobs/jobs/23005e89-18ee-4619-858c-3e32bea46510) - Astronomer - New York, NY - $200k – $300k/yr
- [Cluster Operations Software Engineer](https://hotfix.jobs/jobs/0d57ab45-9014-486d-8274-bf6c1952d5af) - Cerebras Systems - Sunnyvale, CA
- [Senior Manager, Site Reliability Engineering - Infrastructure Platform](https://hotfix.jobs/jobs/318325bc-01ec-460b-8ebe-e0c2c64eab78) - Okta - Bellevue, WA - $176k – $264k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/fe0ac17b-00d1-4161-ac39-91193d074036) - Clickhouse - Remote - $133k – $232k/yr

**Apply:** https://hotfix.jobs/jobs/04fb7758-447e-4265-9660-723661492dbf
**Canonical:** https://hotfix.jobs/jobs/04fb7758-447e-4265-9660-723661492dbf