Lead DevOps Engineer owning multi-cloud SaaS infrastructure at scale for Tulip's AI-native frontline operations platform. Design resilient cloud architecture, CI/CD automation, observability, and mentor engineers while partnering with application teams. Requires 5-7+ years infrastructure experience and leadership.
Salary not listed
Hybrid7+ YOEDevOps / SRE
About the role
What Skills Do I Need?
5-7+ years of hands-on DevOps or Infrastructure Engineering experience, with demonstrated ownership of production cloud environments at scale
Proficiency with modern cloud infrastructure tooling — experience with Kubernetes, Helm, Terraform, Ansible, and major cloud providers (AWS and/or Azure)
Proven experience mentoring and coaching engineers and a genuine interest in developing the people around you
Experience managing enterprise-grade data persistence layers, including NoSQL and SQL databases, key/value stores, and messaging systems (e.g., AMQP, MQTT)
Familiarity with observability and monitoring tooling (e.g., Prometheus, Mimir, Thanos, Grafana) and a strong understanding of SRE practices in a fast-growing SaaS environment
Comfort driving team rituals such as sprint planning, standups, and retrospectives
Exposure to modern programming or scripting languages used in infrastructure contexts (e.g., Go, TypeScript, Python, Bash)
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Key Responsibilities
Own the deployment, health, and continuous improvement of Tulip's multi-cloud, multi-region SaaS environments — including clusters spanning the US, Europe, and Asia
Design and evolve cloud architecture to ensure customer availability, stability, and performance as Tulip scales globally
Contribute to and help shape the infrastructure technical roadmap in partnership with engineering leadership
Own and continuously improve Tulip's CI/CD infrastructure, driving toward a fully automated, human-interaction-free software delivery lifecycle
Build automation tooling and internal systems that reduce operational toil and increase developer velocity
Define and maintain observability standards across Tulip's cloud environments, including metrics, alerting, logging, and distributed tracing
Proactively identify performance degradation and capacity risks before they impact customers; lead incident response and drive root cause analysis
Mentor and coach junior and mid-level engineers through code reviews, pairing sessions, and regular technical guidance
Serve as a close partner to application engineering teams throughout the SDLC, providing infrastructure guidance and support
Participate in the on-call rotation and help establish on-call best practices that scale as the team grows
Benefits
Company equity
Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K)
Flexible work schedule and unlimited vacation policy
Senior Software Engineer - Snowpark Container Service
SnowflakeBellevue, WA +1
Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.
200k – 288k/yrHybrid7+ YOEDevOps / SRE
Senior Software Engineer, Infrastructure & Systems
AstronomerNew York, NY
Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.
200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cluster Operations Software Engineer
Cerebras SystemsSunnyvale, CA
Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.
Salary not listedHybrid6+ YOEDevOps / SRE
Senior Manager, Site Reliability Engineering - Infrastructure Platform
OktaBellevue, WA
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.