# Database Reliability Engineer - Core Team

**Company:** [Clickhouse](https://hotfix.jobs/companies/clickhouse)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** SQL, ClickHouse, Distributed Databases, Shell, Python, C++, AWS, Azure, GCP, Chaos Engineering, Incident Response, Postmortems, Production Debugging
**Posted:** 2026-04-02

> Build and lead reliability engineering practices for ClickHouse Core, improving production performance, observability, incident response, and database operations. The role requires at least five years of reliability, QA, or customer-facing engineering experience, plus production SQL database and cloud expertise.

## Job Description

## Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate recurring customer problems, determine root causes, submit bug fixes and issue reports, and recommend improvements.
- Enhance incident response and post-mortem processes for ClickHouse Core outages, including blameless postmortems and customer communications.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and establish escalation best practices to resolve reliability and performance issues while minimizing customer impact.
- Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams.

## Requirements
- Bachelor's or master's degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong understanding of distributed database internals and SQL.
- Scripting experience with Shell or Python.
- Ability to read and understand C++ code.
- Knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
- Strong production debugging and problem-solving skills.
- Excellent communication skills and the ability to work effectively in a global, fast-paced team.
- High ownership, responsibility, and accountability.

## Nice to Have
- Experience with ClickHouse.
- Experience with chaos engineering and post-mortem analysis.

## Benefits
- Healthcare contributions.
- Company stock options.
- Flexible time off in the United States and generous leave allowances in other countries.
- USD $500 home-office setup allowance for remote employees.
- Company-wide global gatherings.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [AI Infrastructure Systems Engineer](https://hotfix.jobs/jobs/34b2c44e-c259-48df-a98e-0e98e0da2aab) - Together AI - San Francisco, CA
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [ClickHouse Operations Engineer](https://hotfix.jobs/jobs/b2da36d0-2f09-45b4-88c7-996eb12809a8) - PostHog - Remote

**Apply:** https://hotfix.jobs/jobs/991286c8-8504-43ac-9e57-e79b88cec642
**Canonical:** https://hotfix.jobs/jobs/991286c8-8504-43ac-9e57-e79b88cec642