Site Reliability Engineer
Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.
About the job
Responsibilities
- Extend self-service datastore platform provisioning automation, guardrails, and paved paths for product engineering teams.
- Build observability, alerting, and backup/disaster recovery into datastore defaults without manual intervention.
- Turn recurring reliability, scaling, and schema-change support work into reusable platform features or AI tooling.
- Codify compliance, capacity management, and cost management into the platform.
- Support operational needs across the production database fleet alongside product engineering teams.
Requirements
- 3+ years of experience as a Site Reliability Engineer or in a similar infrastructure-focused role.
- Direct experience managing production infrastructure or building tooling to support it in AWS and Kubernetes.
- Experience building and delivering working software to production, including tooling, platforms, or product-facing services.
- Strong fluency with AI coding tools such as Claude.
- Legal eligibility to work in Canada; sponsorship is unavailable.
Nice-to-haves
- Experience with RDS, Redis, OpenSearch, or Kubernetes-hosted datastores.
- Experience with observability, alerting, backup, and disaster recovery.
- Experience automating compliance, capacity, and cost management.
Skills
AWS, Kubernetes, Rds, Redis, Opensearch, Datastores, Observability, Alerting, Disaster Recovery, Claude, Infrastructure Automation, Backup
Similar jobs
DevOps / SRE jobsBuild and operate observability tooling and infrastructure that improves platform reliability, scalability, and incident response. The role requires software development, public cloud and Kubernetes experience, and proficiency with modern monitoring and tracing technologies.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.