# Database Reliability Engineer - Core Team

**Company:** [Clickhouse](https://hotfix.jobs/companies/clickhouse)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Post-Mortems, Production Debugging
**Posted:** 2026-04-02

> Build and lead reliability practices for ClickHouse Core, improving database performance, observability, incident response, and scalability. The role requires at least five years of reliability, QA, or customer-facing engineering experience plus production SQL database operations and cloud expertise.

## Job Description

## Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate recurring customer problems, identify root causes, submit bug fixes and issue reports, and recommend improvements.
- Enhance incident response and post-mortem processes for ClickHouse Core outages, including blameless postmortems and customer communication.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and establish escalation practices that minimize customer impact.
- Guide Control Plane, Dataplane, Security, Support, and Operations teams in deploying ClickHouse effectively.

## Requirements
- Bachelor’s or master’s degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong understanding of distributed database internals and SQL.
- Shell or Python scripting experience.
- Ability to read and understand C++ code.
- Knowledge of AWS, Azure, or Google Cloud Platform.
- Strong production debugging and problem-solving skills.
- Excellent communication skills.

## Nice-to-haves
- Experience with ClickHouse.
- Experience leading incident response, escalation management, or post-mortem analysis.
- Experience with chaos engineering initiatives.

## Compensation and Benefits
- Equity through company stock options.
- Healthcare contributions.
- Flexible time off; generous entitlement outside the United States.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to participate in company-wide offsites.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Site Reliability Engineer, Infrastructure Platforms](https://hotfix.jobs/jobs/e28d2ba0-df16-41d6-9c60-e694ee353996) - GitLab - Remote
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/e7b74501-38ae-4cd1-8b26-91e042b41460) - Writer - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/3933adb1-d8bc-4482-a616-d09b8e055938
**Canonical:** https://hotfix.jobs/jobs/3933adb1-d8bc-4482-a616-d09b8e055938