Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.
Salary not listed
Hybrid5+ YOEDevOps / SRE
About the role
About You
Reason about systems at scale: edge cases, failure modes, and life cycles
Excited about setting the technical agenda and coming up with novel, broad ideas
Keep up with newest AI advancements in Observability & Monitoring
Know what a good SLA looks like and can teach others
Communicate as well as you code; value discussion and clear, frequent communication in teams
Skills
5+ years of experience with open source Observability tools (Loki, Grafana, Tempo, Mimir stack)
Hands-on experience instrumenting distributed systems using OpenTelemetry
Managing metrics pipelines with Prometheus at scale
Direct experience developing and distributing Claude Skills, Gemini Gems, or other generic AI processes; iterate on their efficacy
Experience working with time-series data, ideally using PromQL
Key Responsibilities
Mentor and evangelize observability best practices, SLIs/SLOs, and reliability culture across engineering teams
Contribute to and maintain triage & remediation processes as a player/coach
Perform incident response and debug production issues across the entire stack
Design, build, and maintain core infrastructure & tooling used by all engineering teams
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cluster Operations Software Engineer
Cerebras SystemsSunnyvale, CA
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listedHybrid6+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.