Database Reliability Engineer - Core Team
The Database Reliability Engineer will improve ClickHouse Core’s reliability, scalability, performance, and production operations through observability, incident response, chaos initiatives, and database debugging. The role requires at least five years of reliability, QA, or customer-facing engineering experience and production SQL database expertise.
About the job
Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate recurring customer problems, identify root causes, submit bug fixes and issue reports, and recommend improvements.
- Enhance incident response and post-mortem processes for ClickHouse Core outages, including blameless postmortems and customer communications.
- Plan, enable, and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes, escalation coordination, and best practices to minimize customer impact.
- Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams.
Requirements
- Bachelor’s or master’s degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong understanding of distributed database internals and SQL.
- Scripting experience with Shell or Python.
- Ability to read and understand C++ code.
- Knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud.
- Strong problem-solving and production debugging skills.
- Excellent communication skills and the ability to work effectively in a global, fast-paced team.
- High ownership, responsibility, and accountability.
Nice-to-have
- Experience with ClickHouse.
Compensation and Benefits
- Healthcare contributions.
- Company stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home-office setup benefit for remote employees.
- Opportunities to attend company-wide global gatherings.
Skills
ClickHouse, SQL, Distributed Systems, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Postmortems
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.