Senior Site Reliability Engineer
The Senior Site Reliability Engineer will improve the reliability, resilience, monitoring, and incident response of Auth0’s large-scale production systems. The role requires 3+ years in SRE or cloud operations, experience with Go, shell scripting, and Terraform, and participation in rotational 24/7 on-call coverage.
About the job
Responsibilities
- Work with other teams to run, own, and improve incident response processes.
- Participate in regular on-call rotations to ensure 24/7 coverage of all critical systems.
- Use existing monitoring tools to identify problems and resolve or escalate them to service teams.
- Implement changes to improve infrastructure resilience, monitoring, and alerting.
Requirements
- Exceptional communication skills, including technical writing in English.
- Systematic problem-solving approach, strong ownership, and drive.
- Understanding of microservices, cloud infrastructure, databases, containers, web technologies, and networking.
- Familiarity with SLIs, SLOs, error budgets, and SLAs.
- Strong belief in automation and reducing operational toil.
- Ability to work effectively in a team and in a remote environment with self-directed tasks.
- Ability to independently handle rotational 24/7 on-call responsibilities.
- 3+ years as a Site Reliability Engineer or in a Cloud Operations/DevOps role.
- 2+ years using Go, shell scripting, and Terraform.
- 2+ years as a software developer in a SaaS environment.
- 3+ years supporting large-scale, mission-critical applications in production.
Nice to Have
- Knowledge of Datadog or another observability platform.
Benefits
- In-person onboarding experience.
- Opportunities for talent development, social impact, and community connection.
Skills
AWS, Azure, SQL, NoSQL, Docker, Kubernetes, WebSockets, Http, Ssl, Vpn, Datadog, Go, Shell Scripting, Terraform
Similar jobs
DevOps / SRE jobsSenior Site Reliability Engineer responsible for operating and improving reliable, scalable cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, Terraform, Go or Python, distributed systems, and reliability engineering expertise.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.