Skip to content
LyftLyft

Software Engineer, Observability

Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.

About the job

Responsibilities

  • Maintain, improve, and develop tooling and systems that enhance platform reliability, scalability, and efficiency.
  • Assist engineering teams in defining service-level objectives (SLOs) and provide tooling to monitor and balance feature development speed with reliability.
  • Maintain and analyze metrics from operating systems, control planes, and applications to support fault detection and performance improvements.
  • Collaborate with cross-functional engineering teams to enhance observability and meet developer needs, including design and production readiness reviews, platform management, and capacity planning.
  • Maintain world-class documentation for infrastructure operations processes and insights.
  • Identify repeatable actions and automate repetitive tasks.
  • Participate in on-call rotations, respond to incidents, and help mitigate customer-impacting events.

Requirements

  • 3+ years of experience working on teams responsible for software development, automation, and systems engineering.
  • Bachelor's degree or equivalent experience in Computer Science or a relevant discipline.
  • Proficiency writing production-ready code in one or more high-level languages, such as Go or Python.
  • Experience operating infrastructure in public cloud environments, such as AWS, including familiarity with managed services.
  • Experience building and maintaining observability infrastructure for robust monitoring and analysis.
  • Familiarity with Kubernetes and managing multi-cluster environments in production.
  • Experience with modern observability stacks, including Prometheus, Grafana, Loki, and open-source tracing and alerting frameworks.

Benefits

  • Extended health and dental coverage, life insurance, and disability benefits
  • Mental health benefits
  • Family building, child care, and pet benefits
  • Lyft-funded Health Care Savings Account
  • RRSP plan with company match
  • Flexible paid time off for salaried team members; hourly team members receive 15 days paid time off, with an additional day for each year of service
  • 18 weeks of paid parental leave through a top-up plan
  • Subsidized commuter benefits and Lyft ride credits

Compensation

  • Expected base pay range: CAD $108,000–$135,000 in the Toronto area, excluding potential equity, bonus, and benefits.
  • Hybrid roles may work from anywhere for up to 4 weeks per year.

Skills

Go, Python, AWS, Kubernetes, Prometheus, Grafana, Loki, Distributed Tracing, Alerting, Service-Level Objectives, Cloud Infrastructure, Observability

Greenhouse

Greenhouse

British Columbia, Canada
Site Reliability Engineer
CA$88k+/yrRemote3+ YOEDevOps / SRE

Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.

Baseten

Baseten

San Francisco, CA
Software Engineer - Continuous Delivery
$165k+/yrHybridDevOps / SRE

Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.