Skip to content
TulipTulip

Lead DevOps Engineer

Lead DevOps Engineer owning multi-cloud SaaS infrastructure at scale for Tulip's AI-native frontline operations platform. Design resilient cloud architecture, CI/CD automation, observability, and mentor engineers while partnering with application teams. Requires 5-7+ years infrastructure experience and leadership.

About the job

What Skills Do I Need?

  • 5-7+ years of hands-on DevOps or Infrastructure Engineering experience, with demonstrated ownership of production cloud environments at scale
  • Proficiency with modern cloud infrastructure tooling — experience with Kubernetes, Helm, Terraform, Ansible, and major cloud providers (AWS and/or Azure)
  • Proven experience mentoring and coaching engineers and a genuine interest in developing the people around you
  • Experience managing enterprise-grade data persistence layers, including NoSQL and SQL databases, key/value stores, and messaging systems (e.g., AMQP, MQTT)
  • Familiarity with observability and monitoring tooling (e.g., Prometheus, Mimir, Thanos, Grafana) and a strong understanding of SRE practices in a fast-growing SaaS environment
  • Comfort driving team rituals such as sprint planning, standups, and retrospectives
  • Exposure to modern programming or scripting languages used in infrastructure contexts (e.g., Go, TypeScript, Python, Bash)
  • Bachelor's degree in Computer Science, Engineering, or equivalent practical experience

Key Responsibilities

  • Own the deployment, health, and continuous improvement of Tulip's multi-cloud, multi-region SaaS environments — including clusters spanning the US, Europe, and Asia
  • Design and evolve cloud architecture to ensure customer availability, stability, and performance as Tulip scales globally
  • Contribute to and help shape the infrastructure technical roadmap in partnership with engineering leadership
  • Own and continuously improve Tulip's CI/CD infrastructure, driving toward a fully automated, human-interaction-free software delivery lifecycle
  • Build automation tooling and internal systems that reduce operational toil and increase developer velocity
  • Define and maintain observability standards across Tulip's cloud environments, including metrics, alerting, logging, and distributed tracing
  • Proactively identify performance degradation and capacity risks before they impact customers; lead incident response and drive root cause analysis
  • Mentor and coach junior and mid-level engineers through code reviews, pairing sessions, and regular technical guidance
  • Serve as a close partner to application engineering teams throughout the SDLC, providing infrastructure guidance and support
  • Participate in the on-call rotation and help establish on-call best practices that scale as the team grows

Benefits

  • Company equity
  • Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K)
  • Flexible work schedule and unlimited vacation policy
  • Virtual company events and happy hours
  • Fitness subsidies

Skills

Kubernetes, Helm, Terraform, Ansible, AWS, Azure, Prometheus, Grafana, Go, Python, TypeScript, Bash, SQL, NoSQL

Applied Intuition

Applied Intuition

Sunnyvale, CA

Senior Software Engineer - Cloud Infrastructure
$190k+/yrOn-site5+ YOEDevOps / SRE

Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.

tastytrade

tastytrade

Chicago, IL

Senior Site Reliability Engineer - Linux Systems & Application Observability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer - SRE
$190k+/yrRemote5+ YOEDevOps / SRE

Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.

Zoox

Zoox

Foster City, CA

Senior Software Engineer - Pipeline Infrastructure & Integration
$219k+/yrHybrid7+ YOEDevOps / SRE

Senior engineer owning safety-critical software pipelines and infrastructure, from static and dynamic analysis through CI enforcement, dashboards, and reliability tooling. Requires an advanced technical degree, 7+ years working with large codebases, and expertise in Bazel, Python, backend infrastructure, and C++.

Bloomerang

Bloomerang

United States

Senior Software Engineer, Site Reliability
$115k+/yrRemote5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for production troubleshooting, incident response, observability, SLOs, automation, and permanent reliability improvements. Requires strong software engineering, SQL, debugging, cloud-application troubleshooting, and cross-functional collaboration skills.