Site Reliability Engineer
Operates and improves large-scale cloud infrastructure, focusing on reliability, observability, incident response, automation, and disaster recovery. The role requires production cloud experience, scripting or programming skills, containers, infrastructure as code, CI/CD, and modern monitoring tools.
About the job
Responsibilities
- Build and maintain core infrastructure as code using tools such as Terraform and Ansible.
- Implement robust monitoring, logging, and alerting systems to ensure services meet and exceed 99.99% uptime.
- Resolve production incidents through systematic debugging and drive blameless postmortems to prevent recurrence.
- Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams.
- Collaborate with developers to embed reliability and scalability best practices into the application lifecycle.
- Contribute to capacity planning, disaster recovery drills, and security hardening processes.
- Participate in a fair and sustainable on-call rotation.
Requirements
- Experience operating production workloads on a major cloud provider such as AWS, Google Cloud, or Azure.
- Proficiency in at least one programming or scripting language, such as Go, Python, or Bash.
- Hands-on experience with containerization and orchestration technologies including Docker and Kubernetes.
- Knowledge of infrastructure-as-code principles and tools; Terraform experience is a plus.
- Familiarity with CI/CD concepts and pipeline tools such as GitLab CI and Jenkins.
- Understanding of modern observability stacks such as Prometheus, Grafana, and ELK.
Skills
Terraform, Ansible, AWS, GCP, Microsoft Azure, Go, Python, Bash, Docker, Kubernetes, Gitlab Ci, Jenkins, Prometheus, Grafana, Elk
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.