# Senior Site Reliability Engineer

**Company:** [Clickhouse](https://hotfix.jobs/companies/clickhouse)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** Go, Python, AWS, Microsoft Azure, GCP, SQL, ClickHouse, Kubernetes, Docker Swarm, Ansible, Terraform, Puppet, Chaos Engineering, Incident Management, Distributed Systems
**Posted:** 2026-03-13

> The Senior Site Reliability Engineer will improve the reliability, availability, scalability, and performance of ClickHouse Cloud by designing distributed systems, managing observability and incident response, and driving automation and chaos initiatives. The role requires 8+ years of SRE experience and hands-on Go or Python expertise.

## Job Description

## Responsibilities
- Design and implement scalable, secure, highly available, and fault-tolerant distributed systems for ClickHouse Cloud.
- Establish and manage service-level objectives (SLOs) and service-level agreements (SLAs).
- Ensure infrastructure components have monitoring and alerting for timely incident detection and resolution.
- Improve incident response and outage postmortem processes, including blameless postmortems and customer communications.
- Continuously improve service reliability and performance.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and escalation best practices to minimize downtime.
- Develop software platforms and tools that improve operational and engineering efficiency.

## Requirements
- Bachelor's or master's degree in Computer Science or a related field.
- At least 8 years of experience in Site Reliability Engineering or a related field.
- Hands-on experience with Go and/or Python.
- Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Understanding of distributed databases and SQL; ClickHouse experience is a major plus.
- Experience with Kubernetes or Docker Swarm.
- Experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
- Strong production debugging and problem-solving skills.
- Excellent communication and interpersonal skills.

## Benefits
- Healthcare contributions.
- Company stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to attend company-wide global gatherings.

## Similar jobs

- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/e14b1da4-1208-4648-a785-30451d299dd7) - Mozilla - Remote - CA$95k – CA$139k/yr
- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/d332c3cb-7ad5-4a0d-b2c2-af9f684bd506) - Mozilla - Remote
- [Senior DevSecOps Engineer](https://hotfix.jobs/jobs/9fad0d81-f196-425c-b1d6-cf7adc34ac02) - Shield AI - London, United Kingdom
- [Senior Production Engineer](https://hotfix.jobs/jobs/98eb4848-f8e8-4b98-95bd-c5f2b852c00f) - Clear Street - London, United Kingdom
- [Software Engineer, GPU Infrastructure](https://hotfix.jobs/jobs/0140c9f5-05ed-4ac4-8f3c-89978a1ea960) - Cohere

**Apply:** https://hotfix.jobs/jobs/4df7d29d-d2cf-466e-a169-de736a3728e4
**Canonical:** https://hotfix.jobs/jobs/4df7d29d-d2cf-466e-a169-de736a3728e4