# Observability Tech Lead

**Company:** [Tulip](https://hotfix.jobs/companies/tulip)
**Location:** Somerville, MA
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Observability, loki, Grafana, tempo, mimir, OpenTelemetry, Prometheus, promql, Kubernetes, TypeScript, Go, MongoDB, Postgres, ai monitoring
**Posted:** 2026-07-23

> Tech Lead for Observability at Tulip, mentoring on best practices, SLIs/SLOs, and reliability while designing, building, and maintaining core observability infrastructure, tooling, and AI-enhanced monitoring for distributed systems and production incidents.

## Job Description

## About You
- You can reason about systems at scale: their edge cases, failure modes, and life cycles.
- You’re excited about setting the technical agenda and coming up with novel, broad ideas.
- You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring.
- You know what a good SLA looks like, and can teach others how to spot one.
- You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication.

## Skills
- 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack).
- Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale.
- Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy.
- Experience working with time-series data, ideally using promQL.

## Key Responsibilities
- Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.
- Contributing to and maintaining Tulip's triage & remediation processes as a player / coach.
- Perform incident response and debug production issues across the entire stack.
- Design, build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering teams.

## Tech Stack
- TS and Go Services running on Kubernetes.
- MongoDB and Postgres DBs.
- Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry Observability tooling.

## Key Collaborators
- Engineering.
- Edge.
- DevOps.
- Hardware.

## Benefits
- Direct impact on product and culture.
- Company equity.
- Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K).
- Flexible work schedule and unlimited vacation policy.
- Virtual company events and happy hours.
- Fitness subsidies.

## Similar roles

- [Senior Software Engineer - Snowpark Container Service](https://hotfix.jobs/jobs/47626f55-409c-4469-9cea-906cfd5683e3) - Snowflake - Bellevue, WA - $200k – $288k/yr
- [Senior Software Engineer, Infrastructure & Systems](https://hotfix.jobs/jobs/23005e89-18ee-4619-858c-3e32bea46510) - Astronomer - New York, NY - $200k – $300k/yr
- [Cluster Operations Software Engineer](https://hotfix.jobs/jobs/0d57ab45-9014-486d-8274-bf6c1952d5af) - Cerebras Systems - Sunnyvale, CA
- [Senior Manager, Site Reliability Engineering - Infrastructure Platform](https://hotfix.jobs/jobs/318325bc-01ec-460b-8ebe-e0c2c64eab78) - Okta - Bellevue, WA - $176k – $264k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/fe0ac17b-00d1-4161-ac39-91193d074036) - Clickhouse - Remote - $133k – $232k/yr

**Apply:** https://hotfix.jobs/jobs/2b7cd0e8-d169-4891-90d0-62329e1dd730
**Canonical:** https://hotfix.jobs/jobs/2b7cd0e8-d169-4891-90d0-62329e1dd730