Software Engineer, Async Platform
Build and operate scalable asynchronous platform infrastructure, developing automation and tooling while improving reliability, performance, and operational efficiency. The role requires 5+ years of software development, automation, and systems engineering experience, plus cloud and distributed-systems expertise.
About the job
Responsibilities
- Maintain and analyze metrics from operating systems, control planes, and applications to assist in fault detection and performance enhancement.
- Design, develop, and deploy tooling and systems that improve the reliability, scalability, and efficiency of the platform.
- Balance feature development speed and reliability with service-level objectives.
- Operate and improve infrastructure using industry best practices and tools.
- Participate in design and production-readiness reviews, platform management, and capacity-planning ceremonies with cross-functional teams.
- Document infrastructure operations processes and insights; identify repeatable actions and automate repetitive tasks.
- Participate in team on-call rotations, respond to incidents, and support other teams in mitigating customer-impacting events.
Requirements
- 5+ years of experience working on teams responsible for software development, automation, and systems engineering.
- Experience building large-scale infrastructure, distributed systems, or networks.
- Experience developing in Go, Python, or other object-oriented programming languages.
- Experience working with cloud-based environments such as AWS, GCP, or Azure.
- Commitment to reducing technical debt and maintaining clean, reliable code and configuration.
Nice-to-haves
- Knowledge of SQS, Kafka, and Kinesis.
Compensation and Benefits
- Expected base pay range in the Toronto area: CAD $108,000–$135,000.
- Extended health and dental coverage, life insurance, and disability benefits.
- Mental health, family-building, child-care, and pet benefits.
- Lyft-funded Health Care Savings Account.
- RRSP plan with company match.
- Flexible paid time off for salaried team members; hourly team members receive 15 days of paid time off, with additional days based on tenure.
- 18 weeks of paid parental leave top-up for eligible biological, adoptive, and foster parents.
- Subsidized commuter benefits and Lyft ride credits.
Skills
Go, Python, AWS, GCP, Microsoft Azure, SQS, Apache Kafka, Amazon Kinesis, Infrastructure As Code, Distributed Systems, Systems Engineering, Capacity Planning
Similar jobs
DevOps / SRE jobsBuild and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.
Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.