Senior Software Engineer - SRE
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.
About the job
Responsibilities
- Embed with product teams to improve operational maturity through on-call, monitoring, alerting, and runbook practices.
- Run game day exercises to improve incident investigation and remediation readiness.
- Implement reliability techniques in Haskell and TypeScript application code, including retries, error handling, logging, and circuit breaking.
- Establish meaningful SLOs tied to customer outcomes.
- Champion reliability practices through design-document and code reviews.
- Identify and close observability gaps affecting debugging, incident response, and business intelligence.
- Advocate for longer-term reliability improvements led by non-product engineering teams.
- Participate in product-team and general engineering on-call rotations and improve incident-learning processes.
Requirements
- Previous Site Reliability Engineering or DevOps experience.
- Demonstrated impact influencing an organization toward greater reliability.
- Significant PostgreSQL experience.
- Experience authoring and operating Temporal workflows.
- Experience with observability platforms such as Grafana or Honeycomb.
- Familiarity with OpenTelemetry.
Compensation
- US employees: $200,700–$250,900 USD base salary.
- Canadian employees: $189,700–$237,100 CAD base salary.
- Compensation also includes equity and benefits.
Skills
Site Reliability Engineering, DevOps, Haskell, TypeScript, Postgres, Temporal, Grafana, Honeycomb, OpenTelemetry, SLOs, Observability, Incident Response
Similar jobs
DevOps / SRE jobsBuild and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.