Site Reliability Engineer
Build and operate a self-service datastore platform, embedding provisioning, observability, reliability, disaster recovery, compliance, and cost controls into production infrastructure. The role requires at least three years of SRE or infrastructure experience with AWS, Kubernetes, production software delivery, and AI coding tools.
About the job
Responsibilities
- Extend self-service datastore platform provisioning automation, guardrails, and paved paths for product engineering teams.
- Build observability, alerting, and backup/disaster recovery into every datastore by default.
- Turn recurring reliability, scaling, and schema-change support work into reusable platform features or AI tooling.
- Codify compliance, capacity management, and cost management into the platform.
- Support the operational needs of a growing database fleet alongside product engineering teams.
Requirements
- 3+ years of experience as a Site Reliability Engineer or in a similar infrastructure-focused role.
- Experience directly managing or building tooling for production infrastructure in AWS and Kubernetes.
- Experience building and delivering software to production, including tooling, platforms, or product-facing services.
- Strong fluency with AI coding tools such as Claude.
- Experience with or interest in databases and datastores such as RDS, Redis, OpenSearch, or Kubernetes-hosted datastores.
Compensation
- National pay range: $87,700–$127,100 CAD annually.
- Some roles may also be eligible for stock options, bonuses, merit increases, or sales commissions.
Skills
AWS, Kubernetes, Amazon Rds, Redis, Opensearch, Datastores, Observability, Alerting, Disaster Recovery, Infrastructure Automation, Claude, Ai Coding Tools
Similar jobs
DevOps / SRE jobsBuild and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.