Senior Site Reliability Engineer
Owns the reliability, scalability, deployment automation, and incident response of production infrastructure for a large-scale data platform. The role requires expertise in managed Kubernetes, multi-cloud platforms, infrastructure as code, cloud networking, scripting, and Linux administration.
About the job
Technologies You’ll Use
- Cloud service providers: AWS, Azure, Google Cloud
- Kubernetes: EKS, AKS, GKE
- Continuous integration: GitHub Actions, Buildkite
- Continuous delivery: ArgoCD
- Databases: PostgreSQL and other major databases
- Languages: Go, Java
- Scripting: TypeScript, Python, Shell
- Infrastructure as code: Terraform, Pulumi
- RESTful APIs: FastAPI
- Cloud networking: Azure and AWS PrivateLink, GCP Private Service Connect, and site-to-site VPN tunnels
Responsibilities
- Own the performance and reliability of production infrastructure, deployment pipelines, and incident response.
- Monitor infrastructure availability, capacity, and throughput.
- Incorporate reliability improvements into the product roadmap.
- Coordinate prioritization and resolution of critical bugs affecting support or sales requirements.
- Recommend production infrastructure improvements in partnership with engineering teams.
- Automate deployment of scalable artifacts across all environments.
- Monitor infrastructure vulnerabilities and remediate them with the security team.
Requirements
- Expertise with managed Kubernetes, including EKS, AKS, and GKE.
- Knowledge of AWS, Azure, Google Cloud, Terraform, Ansible, Buildkite, Pulumi, and ArgoCD.
- Experience with Python and Shell scripting and the Go programming language; Java experience is a bonus.
- Experience with Linux operating-system internals and administration.
- Experience with cloud networking, including site-to-site VPNs, PrivateLink, and Private Service Connect.
- Experience using AI tools for code review, automation, incident analysis, documentation, and reducing operational toil.
Nice to Have
- 5+ years of experience working with SaaS products at scale.
- Experience with PostgreSQL.
Compensation and Benefits
- Employer-paid medical insurance.
- Paid time off, paid sick time, inclusive parental leave, holidays, a year-end Global Week of Rest, and volunteer days.
- RSU stock grants.
- Professional development and training opportunities.
- Company events, free food, and team-building activities.
- Monthly cell phone stipend.
- Mental health support platform with therapy, coaching, and self-guided mindfulness resources.
- Benefits may vary by country and worker type.
Skills
AWS, Azure, GCP, Kubernetes, Amazon Eks, Azure Aks, Google Gke, Terraform, Ansible, Buildkite, Pulumi, Argo CD, Python, Go, Linux
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.