Tech Lead for Observability at Tulip, mentoring on best practices, SLIs/SLOs, and reliability while designing, building, and maintaining core observability infrastructure, tooling, and AI-enhanced monitoring for distributed systems and production incidents.
Salary not listed
Hybrid5+ YOEDevOps / SRE
About the role
About You
You can reason about systems at scale: their edge cases, failure modes, and life cycles.
You’re excited about setting the technical agenda and coming up with novel, broad ideas.
You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring.
You know what a good SLA looks like, and can teach others how to spot one.
You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication.
Skills
5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack).
Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale.
Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy.
Experience working with time-series data, ideally using promQL.
Key Responsibilities
Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams.
Contributing to and maintaining Tulip's triage & remediation processes as a player / coach.
Perform incident response and debug production issues across the entire stack.
Design, build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering teams.
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cluster Operations Software Engineer
Cerebras SystemsSunnyvale, CA
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listedHybrid6+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.