# Software Engineer, Observability

**Company:** [Lyft](https://hotfix.jobs/companies/lyft)
**Location:** Toronto, Canada
**Role:** DevOps / SRE
**Salary:** CA$108k – CA$135k/yr
**Experience:** 3+ years
**Skills:** Go, Python, AWS, Kubernetes, Prometheus, Grafana, Loki, Distributed Tracing, Alerting, Service-Level Objectives, Cloud Infrastructure, Observability
**Posted:** 2026-08-14

> Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.

## Job Description

## Responsibilities

- Maintain, improve, and develop tooling and systems that enhance platform reliability, scalability, and efficiency.
- Assist engineering teams in defining service-level objectives (SLOs) and provide tooling to monitor and balance feature development speed with reliability.
- Maintain and analyze metrics from operating systems, control planes, and applications to support fault detection and performance improvements.
- Collaborate with cross-functional engineering teams to enhance observability and meet developer needs, including design and production readiness reviews, platform management, and capacity planning.
- Maintain world-class documentation for infrastructure operations processes and insights.
- Identify repeatable actions and automate repetitive tasks.
- Participate in on-call rotations, respond to incidents, and help mitigate customer-impacting events.

## Requirements

- 3+ years of experience working on teams responsible for software development, automation, and systems engineering.
- Bachelor's degree or equivalent experience in Computer Science or a relevant discipline.
- Proficiency writing production-ready code in one or more high-level languages, such as Go or Python.
- Experience operating infrastructure in public cloud environments, such as AWS, including familiarity with managed services.
- Experience building and maintaining observability infrastructure for robust monitoring and analysis.
- Familiarity with Kubernetes and managing multi-cluster environments in production.
- Experience with modern observability stacks, including Prometheus, Grafana, Loki, and open-source tracing and alerting frameworks.

## Benefits

- Extended health and dental coverage, life insurance, and disability benefits
- Mental health benefits
- Family building, child care, and pet benefits
- Lyft-funded Health Care Savings Account
- RRSP plan with company match
- Flexible paid time off for salaried team members; hourly team members receive 15 days paid time off, with an additional day for each year of service
- 18 weeks of paid parental leave through a top-up plan
- Subsidized commuter benefits and Lyft ride credits

## Compensation

- Expected base pay range: **CAD $108,000–$135,000** in the Toronto area, excluding potential equity, bonus, and benefits.
- Hybrid roles may work from anywhere for up to 4 weeks per year.

## Similar jobs

- [Site Reliability Engineer](https://hotfix.jobs/jobs/703f9a1f-45d8-458e-8362-002655fa42a0) - Greenhouse - Remote - CA$88k – CA$127k/yr
- [Software Engineer - Continuous Delivery](https://hotfix.jobs/jobs/0ceaa410-df08-45d6-80b3-68328c769297) - Baseten - San Francisco, CA - $165k – $330k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote

**Apply:** https://hotfix.jobs/jobs/3bea48bc-7046-4e2a-827e-c38c033a726e
**Canonical:** https://hotfix.jobs/jobs/3bea48bc-7046-4e2a-827e-c38c033a726e