Staff Site Reliability Engineer
Staff Site Reliability Engineer responsible for the performance, scalability, deployment robustness, vulnerability management, and incident response of a large-scale cloud data platform. The role requires deep Kubernetes and multi-cloud expertise, infrastructure automation, scripting, Linux administration, and cloud networking experience.
About the job
Responsibilities
- Own the performance and reliability of production infrastructure, deployment pipelines, and incident response.
- Monitor infrastructure availability, capacity, and throughput.
- Incorporate reliability improvements into the product roadmap.
- Reprioritize or fix critical issues based on support and sales requirements.
- Recommend production infrastructure improvements in collaboration with engineering.
- Automate artifact deployment across all environments.
- Monitor infrastructure vulnerabilities and remediate them with the security team.
Requirements
- Expertise with managed Kubernetes, including EKS, AKS, and GKE.
- Knowledge of AWS, Azure, Google Cloud, Terraform, Ansible, Buildkite, Pulumi, and ArgoCD.
- Experience with Python and Shell scripting and the Go programming language; Java is a bonus.
- Experience with Linux internals and administration.
- Experience with cloud networking, including site-to-site VPNs, PrivateLink, and Private Service Connect.
- Experience using AI tools for code review, automation, incident analysis, documentation, and reducing operational toil.
Nice to Have
- 5+ years of experience working with SaaS products at scale.
- Experience with PostgreSQL and other databases.
Benefits
- Employer-paid medical insurance.
- Paid time off, paid sick time, parental leave, holidays, a year-end Global Week of Rest, and volunteer days off.
- RSU stock grants.
- Professional development and training opportunities.
- Company events, free food, and team-building activities.
- Monthly cell phone stipend.
- Mental health support platform with therapy, coaching, and mindfulness resources.
Skills
AWS, Azure, GCP, Kubernetes, Terraform, Ansible, Buildkite, Pulumi, Argo CD, Python, Shell, Go, Java, Linux, Postgres
Similar jobs
DevOps / SRE jobsLeads the design, development, and operation of Stripe’s large-scale CI and developer productivity systems. Requires 10+ years of hands-on software development, distributed-systems expertise, technical leadership, and mentoring experience.
Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
The Staff Production Engineer will operate and improve reliable production infrastructure, with a strong focus on data center networking, capacity planning, troubleshooting, automation, and hardware operations. The role requires 5+ years of production engineering experience, networking expertise, and proficiency with infrastructure and CI/CD tools.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.