Skip to content
ClickhouseClickhouse

Database Reliability Engineer - Core Team

The Database Reliability Engineer will improve the reliability, scalability, performance, and incident response practices for ClickHouse Core. The role requires at least five years of reliability, QA, or customer-facing engineering experience, production database operations expertise, and strong scripting and debugging skills.

About the job

Responsibilities

  • Improve the reliability, availability, scalability, and performance of ClickHouse Core.
  • Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
  • Investigate recurring customer problems, identify root causes, submit bug fixes and issue reports, and recommend improvements.
  • Enhance incident response and postmortem processes for ClickHouse Core outages, including running blameless postmortems.
  • Coordinate with Support and Cloud teams to communicate outage impacts to customers.
  • Plan and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes and establish best practices for escalation and incident resolution.
  • Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams.

Requirements

  • Bachelor’s or master’s degree in Computer Science or a related field.
  • At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
  • Experience operating ClickHouse or other SQL databases in production.
  • Strong understanding of distributed database internals and SQL.
  • Scripting experience with Shell or Python.
  • Ability to read and understand C++ code.
  • Knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Strong production debugging and problem-solving skills.
  • Excellent communication skills.

Nice to have

  • Experience with ClickHouse.
  • Experience with chaos engineering, incident response, escalation management, and blameless postmortems.

Compensation and benefits

  • Equity through company stock options.
  • Healthcare contributions.
  • Flexible time off in the United States and generous entitlement in other countries.
  • $500 home office setup allowance for remote employees.
  • Global company gatherings.

Skills

ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Postmortems

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.