Database Reliability Engineer - Core Team
Build and lead reliability practices for ClickHouse Core, improving production performance, observability, incident response, and scalability. The role requires at least five years in reliability, QA, or customer-facing engineering plus production database operations experience.
About the job
Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate common customer issues, identify root causes, submit bug fixes and issue reports, and recommend improvements.
- Improve incident response and postmortem processes for ClickHouse Core outages, including blameless postmortems and customer communications.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and establish escalation best practices to minimize customer impact.
Requirements
- Bachelor’s or master’s degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong problem-solving and production debugging skills.
- Excellent communication, ownership, accountability, and collaboration skills.
Nice-to-haves
- Knowledge of distributed database internals and SQL, particularly ClickHouse.
- Shell or Python scripting experience.
- Ability to read and understand C++ code.
- Familiarity with AWS, Azure, or Google Cloud Platform.
Compensation and Benefits
- Employer healthcare contributions.
- Company stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home office setup allowance for remote employees.
- Opportunities to attend company-wide offsites.
Skills
ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Postmortems
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.