Senior Site Reliability Engineer
Senior SRE responsible for production infrastructure reliability, incident response, deployment automation, and scaling SaaS systems on Kubernetes and major cloud platforms.
About the job
What You’ll Do
- Responsible for ongoing reliability and robustness of Fivetran’s production infrastructure by monitoring availability, capacity, and throughput.
- Evolve systems by adding reliability into our product roadmap.
- Coordinate the re-prioritize or fix critical bugs for support or sales requirements as needed.
- Make recommendations to production infrastructure by interfacing with engineering to ensure 100% availability.
- Ensure scalable artifacts deployment to all environments by automation scripts.
- Constantly monitor infrastructure vulnerabilities and remedy them by working with the security team.
Technologies You’ll Use
Kubernetes, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal.
Skills We’re Looking For
- 5+ years of experience working with SaaS products at scale.
- Working knowledge of managed Kubernetes (EKS, AKS and GKE).
- Knowledge of Cloud Platforms and related tooling: AWS, Azure, GCP, Terraform, Ansible, Buildkite, Pulumi and ArgoCD.
- Experience in Python/Shell scripting. Bonus if you have Java, Go, etc.
- Experience with Linux operating systems internals and administration.
- Experience with cloud networking like VPNs, PrivateLinks, and Private Service Connect (GCP).
- Experience with databases such as PostgreSQL.
Optional Bonus Skills
- Java, GoLang Programming skills.
Skills
Kubernetes, Postgres, Terraform, Ansible, Python, AWS, GCP, Azure, Argo CD, Linux
Similar jobs
DevOps / SRE jobsLeads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.
Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.
Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.