Senior Site Reliability Engineer
The Senior Site Reliability Engineer will improve the reliability, availability, scalability, and performance of ClickHouse Cloud by designing distributed systems, managing observability and incident response, and driving automation and chaos initiatives. The role requires 8+ years of SRE experience and hands-on Go or Python expertise.
About the job
Responsibilities
- Design and implement scalable, secure, highly available, and fault-tolerant distributed systems for ClickHouse Cloud.
- Establish and manage service-level objectives (SLOs) and service-level agreements (SLAs).
- Ensure infrastructure components have monitoring and alerting for timely incident detection and resolution.
- Improve incident response and outage postmortem processes, including blameless postmortems and customer communications.
- Continuously improve service reliability and performance.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and escalation best practices to minimize downtime.
- Develop software platforms and tools that improve operational and engineering efficiency.
Requirements
- Bachelor's or master's degree in Computer Science or a related field.
- At least 8 years of experience in Site Reliability Engineering or a related field.
- Hands-on experience with Go and/or Python.
- Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Understanding of distributed databases and SQL; ClickHouse experience is a major plus.
- Experience with Kubernetes or Docker Swarm.
- Experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
- Strong production debugging and problem-solving skills.
- Excellent communication and interpersonal skills.
Benefits
- Healthcare contributions.
- Company stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to attend company-wide global gatherings.
Skills
Go, Python, AWS, Microsoft Azure, GCP, SQL, ClickHouse, Kubernetes, Docker Swarm, Ansible, Terraform, Puppet, Chaos Engineering, Incident Management, Distributed Systems
Similar jobs
DevOps / SRE jobsBuild and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.
Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.
Designs and operates secure development infrastructure and CI/CD pipelines for autonomous defence systems. Requires at least five years of DevOps or related experience, plus UK defence or regulated national-security experience and knowledge of Secure by Design and assurance practices.
Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.
Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.