Skip to content
TulipTulipSomerville, MA

Senior Site Reliability Engineer

Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.

Salary not listed
Hybrid5+ YOEDevOps / SRE

About the role

About You

  • Reason about systems at scale: edge cases, failure modes, and life cycles
  • Excited about setting the technical agenda and coming up with novel, broad ideas
  • Keep up with newest AI advancements in Observability & Monitoring
  • Know what a good SLA looks like and can teach others
  • Communicate as well as you code; value discussion and clear, frequent communication in teams

Skills

  • 5+ years of experience with open source Observability tools (Loki, Grafana, Tempo, Mimir stack)
  • Hands-on experience instrumenting distributed systems using OpenTelemetry
  • Managing metrics pipelines with Prometheus at scale
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or other generic AI processes; iterate on their efficacy
  • Experience working with time-series data, ideally using PromQL

Key Responsibilities

  • Mentor and evangelize observability best practices, SLIs/SLOs, and reliability culture across engineering teams
  • Contribute to and maintain triage & remediation processes as a player/coach
  • Perform incident response and debug production issues across the entire stack
  • Design, build, and maintain core infrastructure & tooling used by all engineering teams

Tech Stack

  • TS and Go services running on Kubernetes
  • MongoDB and Postgres databases
  • Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry observability tooling

Key Collaborators

  • Engineering
  • Edge
  • DevOps
  • Hardware

Benefits

  • Direct impact on product and culture
  • Company equity
  • Competitive benefits: Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, FSA, Commuter Benefits, Parental Leave, 401(K)
  • Flexible work schedule and unlimited vacation
  • Virtual company events and happy hours
  • Fitness subsidies

Skills

ObservabilityPrometheusGrafanalokitempomimirOpenTelemetrypromqlKubernetesGoTypeScriptMongoDBPostgres

Similar roles

DevOps / SRE jobs
Snowflake

Senior Software Engineer - Snowpark Container Service

SnowflakeBellevue, WA +1

Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.

200k – 288k/yrHybrid7+ YOEDevOps / SRE
Astronomer

Senior Software Engineer, Infrastructure & Systems

AstronomerNew York, NY

Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.

200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cerebras Systems

Cluster Operations Software Engineer

Cerebras SystemsSunnyvale, CA

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

Salary not listedHybrid6+ YOEDevOps / SRE
Okta

Senior Manager, Site Reliability Engineering - Infrastructure Platform

OktaBellevue, WA

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

176k – 264k/yrHybrid6+ YOEDevOps / SRE
Clickhouse

Senior Cloud Software Engineer - Efficiency Engineering

ClickhouseUnited States

Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.

133k – 232k/yrRemote5+ YOEDevOps / SRE