Software Engineer, Observability
Build and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.
About the job
Responsibilities
- Maintain, improve, and develop tooling and systems that enhance platform reliability, scalability, and efficiency.
- Assist engineering teams in defining service-level objectives (SLOs) and provide tooling to monitor and balance feature development speed with reliability.
- Maintain and analyze metrics from operating systems, control planes, and applications to support fault detection and performance improvements.
- Collaborate with cross-functional engineering teams to enhance observability and meet developer needs, including design and production readiness reviews, platform management, and capacity planning.
- Maintain world-class documentation for infrastructure operations processes and insights.
- Identify repeatable actions and automate repetitive tasks.
- Participate in on-call rotations, respond to incidents, and help mitigate customer-impacting events.
Requirements
- 3+ years of experience working on teams responsible for software development, automation, and systems engineering.
- Bachelor's degree or equivalent experience in Computer Science or a relevant discipline.
- Proficiency writing production-ready code in one or more high-level languages, such as Go or Python.
- Experience operating infrastructure in public cloud environments, such as AWS, including familiarity with managed services.
- Experience building and maintaining observability infrastructure for robust monitoring and analysis.
- Familiarity with Kubernetes and managing multi-cluster environments in production.
- Experience with modern observability stacks, including Prometheus, Grafana, Loki, and open-source tracing and alerting frameworks.
Benefits
- Extended health and dental coverage, life insurance, and disability benefits
- Mental health benefits
- Family building, child care, and pet benefits
- Lyft-funded Health Care Savings Account
- RRSP plan with company match
- Flexible paid time off for salaried team members; hourly team members receive 15 days paid time off, with an additional day for each year of service
- 18 weeks of paid parental leave through a top-up plan
- Subsidized commuter benefits and Lyft ride credits
Compensation
- Expected base pay range: CAD $108,000–$135,000 in the Toronto area, excluding potential equity, bonus, and benefits.
- Hybrid roles may work from anywhere for up to 4 weeks per year.
Skills
Go, Python, AWS, Kubernetes, Prometheus, Grafana, Loki, Distributed Tracing, Alerting, Service-Level Objectives, Cloud Infrastructure, Observability
Similar jobs
DevOps / SRE jobsBuild and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.