Database Reliability Engineer - Core Team
Build and lead reliability practices for ClickHouse Core, improving database performance, observability, incident response, and scalability. The role requires at least five years of reliability, QA, or customer-facing engineering experience plus production SQL database operations and cloud expertise.
About the job
Responsibilities
- Improve the reliability, availability, scalability, and performance of ClickHouse Core.
- Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
- Investigate recurring customer problems, identify root causes, submit bug fixes and issue reports, and recommend improvements.
- Enhance incident response and post-mortem processes for ClickHouse Core outages, including blameless postmortems and customer communication.
- Plan and drive chaos engineering initiatives across engineering teams.
- Manage on-call processes and establish escalation practices that minimize customer impact.
- Guide Control Plane, Dataplane, Security, Support, and Operations teams in deploying ClickHouse effectively.
Requirements
- Bachelor’s or master’s degree in Computer Science or a related field.
- At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
- Experience operating ClickHouse or other SQL databases in production.
- Strong understanding of distributed database internals and SQL.
- Shell or Python scripting experience.
- Ability to read and understand C++ code.
- Knowledge of AWS, Azure, or Google Cloud Platform.
- Strong production debugging and problem-solving skills.
- Excellent communication skills.
Nice-to-haves
- Experience with ClickHouse.
- Experience leading incident response, escalation management, or post-mortem analysis.
- Experience with chaos engineering initiatives.
Compensation and Benefits
- Equity through company stock options.
- Healthcare contributions.
- Flexible time off; generous entitlement outside the United States.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to participate in company-wide offsites.
Skills
ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Post-Mortems, Production Debugging
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.