Senior Site Reliability Engineer
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.
About the job
Responsibilities
Platform & Reliability
- Design, build, and maintain core infrastructure for security SaaS offerings, ensuring high availability, performance, and scalability.
- Build and operate tooling for Snowflake data systems.
Automation
- Develop production-grade automation to eliminate toil and ensure consistency across environments.
- Automate infrastructure provisioning, application deployment, and incident response.
Security & Compliance
- Embed security-first practices into infrastructure and operational processes.
- Ensure systems and data platforms comply with industry standards.
Incident Response
- Participate in on-call rotations and respond to critical incidents.
- Lead root cause analysis and implement preventative measures.
Collaboration
- Partner with development, data science, and security teams on architecture, best practices, and new services.
Requirements
- Strong production-level coding skills for solving operational challenges.
- Deep experience with Terraform for infrastructure provisioning and management.
- Familiarity with CI/CD practices and Spinnaker.
- Expertise with container technologies and production Kubernetes clusters.
- Experience with database schema management and migrations using Flyway.
- Direct experience with large-scale data systems, specifically Snowflake.
- Excellent analytical and problem-solving skills with a proactive approach.
- Participation in on-call rotations.
Nice-to-Have
- Experience or strong interest in AI/ML applications for reliability, security, and operational efficiency, including AIOps and predictive analysis.
Compensation & Benefits
- Annual base salary: $147,000–$202,400 USD.
- Equity, bonus, health, dental and vision insurance, 401(k), flexible spending account, PTO, and parental leave may be available according to applicable plans and policies.
- In-person onboarding and travel to the Toronto, Canada, office are required during the first week of employment.
Skills
Terraform, Spinnaker, Kubernetes, Snowflake, Flyway, CI/CD, Infrastructure As Code, Containerization, Incident Response, AI/ML
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.