# Database Reliability Engineer - Core Team

**Company:** [Clickhouse](https://hotfix.jobs/companies/clickhouse)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Postmortems
**Posted:** 2026-04-02

> The Database Reliability Engineer will improve the reliability, scalability, performance, and incident response practices for ClickHouse Core. The role requires at least five years of reliability, QA, or customer-facing engineering experience, production database operations expertise, and strong scripting and debugging skills.

## Job Description

## Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate recurring customer problems, identify root causes, submit bug fixes and issue reports, and recommend improvements.
- Enhance incident response and postmortem processes for ClickHouse Core outages, including running blameless postmortems.
- Coordinate with Support and Cloud teams to communicate outage impacts to customers.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and establish best practices for escalation and incident resolution.
- Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams.

## Requirements
- Bachelor’s or master’s degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong understanding of distributed database internals and SQL.
- Scripting experience with Shell or Python.
- Ability to read and understand C++ code.
- Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
- Strong production debugging and problem-solving skills.
- Excellent communication skills.

## Nice to have
- Experience with ClickHouse.
- Experience with chaos engineering, incident response, escalation management, and blameless postmortems.

## Compensation and benefits
- Equity through company stock options.
- Healthcare contributions.
- Flexible time off in the United States and generous entitlement in other countries.
- $500 home office setup allowance for remote employees.
- Global company gatherings.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Release Engineer - Data Plane Internal Tooling and Productivity](https://hotfix.jobs/jobs/ee184847-0b76-42f4-8c1f-157f199626d3) - Clickhouse - Remote
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Capacity Ops Engineer](https://hotfix.jobs/jobs/f1904714-7dd3-4ee3-9e7a-e4fcf52083bd) - Baseten - San Francisco, CA - $170k – $230k/yr
- [IT Security and Automation Engineer](https://hotfix.jobs/jobs/604b87b5-13a2-4bba-88b2-f7d0fbbad141) - Teleport - Remote - $149k – $258k/yr

**Apply:** https://hotfix.jobs/jobs/e54279d5-7918-4419-ab7c-d9144389b8a3
**Canonical:** https://hotfix.jobs/jobs/e54279d5-7918-4419-ab7c-d9144389b8a3