Skip to content
TulipTulipSomerville, MA

Observability Tech Lead

Tech Lead for Observability at Tulip, mentoring on best practices, SLIs/SLOs, and reliability while designing, building, and maintaining core observability infrastructure, tooling, and AI-enhanced monitoring for distributed systems and production incidents.

Salary not listed
Hybrid5+ YOEDevOps / SRE

About the role

About You

  • You can reason about systems at scale: their edge cases, failure modes, and life cycles.
  • You’re excited about setting the technical agenda and coming up with novel, broad ideas.
  • You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring.
  • You know what a good SLA looks like, and can teach others how to spot one.
  • You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication.

Skills

  • 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack).
  • Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale.
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy.
  • Experience working with time-series data, ideally using promQL.

Key Responsibilities

  • Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.
  • Contributing to and maintaining Tulip's triage & remediation processes as a player / coach.
  • Perform incident response and debug production issues across the entire stack.
  • Design, build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering teams.

Tech Stack

  • TS and Go Services running on Kubernetes.
  • MongoDB and Postgres DBs.
  • Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry Observability tooling.

Key Collaborators

  • Engineering.
  • Edge.
  • DevOps.
  • Hardware.

Benefits

  • Direct impact on product and culture.
  • Company equity.
  • Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K).
  • Flexible work schedule and unlimited vacation policy.
  • Virtual company events and happy hours.
  • Fitness subsidies.

Skills

ObservabilitylokiGrafanatempomimirOpenTelemetryPrometheuspromqlKubernetesTypeScriptGoMongoDBPostgresai monitoring

Similar roles

DevOps / SRE jobs
Snowflake

Senior Software Engineer - Snowpark Container Service

SnowflakeBellevue, WA +1

Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.

200k – 288k/yrHybrid7+ YOEDevOps / SRE
Astronomer

Senior Software Engineer, Infrastructure & Systems

AstronomerNew York, NY

Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.

200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cerebras Systems

Cluster Operations Software Engineer

Cerebras SystemsSunnyvale, CA

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

Salary not listedHybrid6+ YOEDevOps / SRE
Okta

Senior Manager, Site Reliability Engineering - Infrastructure Platform

OktaBellevue, WA

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

176k – 264k/yrHybrid6+ YOEDevOps / SRE
Clickhouse

Senior Cloud Software Engineer - Efficiency Engineering

ClickhouseUnited States

Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.

133k – 232k/yrRemote5+ YOEDevOps / SRE