Senior Site Reliability Engineer
Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.
About the job
About You
- Reason about systems at scale: edge cases, failure modes, and life cycles
- Excited about setting the technical agenda and coming up with novel, broad ideas
- Keep up with newest AI advancements in Observability & Monitoring
- Know what a good SLA looks like and can teach others
- Communicate as well as you code; value discussion and clear, frequent communication in teams
Skills
- 5+ years of experience with open source Observability tools (Loki, Grafana, Tempo, Mimir stack)
- Hands-on experience instrumenting distributed systems using OpenTelemetry
- Managing metrics pipelines with Prometheus at scale
- Direct experience developing and distributing Claude Skills, Gemini Gems, or other generic AI processes; iterate on their efficacy
- Experience working with time-series data, ideally using PromQL
Key Responsibilities
- Mentor and evangelize observability best practices, SLIs/SLOs, and reliability culture across engineering teams
- Contribute to and maintain triage & remediation processes as a player/coach
- Perform incident response and debug production issues across the entire stack
- Design, build, and maintain core infrastructure & tooling used by all engineering teams
Tech Stack
- TS and Go services running on Kubernetes
- MongoDB and Postgres databases
- Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry observability tooling
Key Collaborators
- Engineering
- Edge
- DevOps
- Hardware
Benefits
- Direct impact on product and culture
- Company equity
- Competitive benefits: Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, FSA, Commuter Benefits, Parental Leave, 401(K)
- Flexible work schedule and unlimited vacation
- Virtual company events and happy hours
- Fitness subsidies
Skills
Observability, Prometheus, Grafana, Loki, Tempo, Mimir, OpenTelemetry, Promql, Kubernetes, Go, TypeScript, MongoDB, Postgres
Similar jobs
DevOps / SRE jobsBuild and operate foundational developer infrastructure spanning CI, build, deployment, and test systems used by engineers across the company. The role requires 5+ years of production software experience, distributed systems expertise, and proficiency in Go or a similar language.
Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.
Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.
Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.