Staff Site Reliability Engineer
Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a SaaS data platform. The role requires 7+ years of experience, strong managed Kubernetes and cloud-platform expertise, and proficiency with infrastructure-as-code and scripting.
About the job
Responsibilities
- Own the reliability and robustness of production infrastructure by monitoring availability, capacity, and throughput.
- Collaborate with engineering teams to integrate reliability best practices into product roadmaps.
- Support prioritization and resolution of critical bugs identified by support or sales.
- Implement automation for scalable deployments in collaboration with engineering teams.
- Ensure scalable artifact deployment across all environments through automation scripts.
- Proactively monitor infrastructure vulnerabilities and collaborate with the security team to address them promptly.
- Drive effective incident response, resolution, and issue avoidance.
Requirements
- 7+ years of experience working with SaaS platforms at scale.
- Expertise with managed Kubernetes, including EKS, AKS, and GKE.
- Knowledge of cloud platforms and related tooling, including AWS, Azure, GCP, Terraform, Ansible, Buildkite, Pulumi, and ArgoCD.
- Experience with Python, shell scripting, and Go; Java experience is a bonus.
- Experience with Linux operating systems, internals, and administration.
- Experience with cloud networking, including managed NAT gateways, VPNs, PrivateLinks, and Private Service Connect.
- Experience with databases such as PostgreSQL.
Technologies
- Managed Kubernetes
- GCP, AWS, and Azure
- Temporal Cloud
- ArgoCD
- Terraform
- Pulumi
- Grafana
- Buildkite
- PostgreSQL
- Python, Go, Java, and Bash/Shell Scripting
Benefits
- 100% employer-paid medical insurance*
- Generous paid time off, paid sick time, inclusive parental leave, holidays, a year-end Global Week of Rest, and volunteer days
- RSU stock grants*
- Professional development and training opportunities
- Company virtual happy hours, free food, and team-building activities
- Monthly cell phone stipend
- Mental health support platform with therapy, coaching, and self-guided mindfulness resources for employees and covered dependents
* Benefits may vary by country and worker type.
Skills
Kubernetes, AWS, Azure, GCP, Terraform, Ansible, Buildkite, Pulumi, Argo CD, Grafana, Python, Go, Linux, Postgres, Cloud Networking
Similar jobs
DevOps / SRE jobsBuild and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Own the reliability, scalability, deployment automation, and incident response of Fivetran’s production infrastructure. The role requires 7+ years of SaaS-scale experience plus deep expertise in Kubernetes, cloud platforms, infrastructure as code, Linux, networking, and programming.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.