Skip to content
ClickhouseClickhouse

Database Reliability Engineer - Core Team

Build and lead reliability practices for ClickHouse Core, improving production performance, observability, incident response, and scalability. The role requires at least five years in reliability, QA, or customer-facing engineering plus production database operations experience.

About the job

Responsibilities

  • Improve the reliability, availability, scalability, and performance of ClickHouse Core.
  • Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
  • Investigate common customer issues, identify root causes, submit bug fixes and issue reports, and recommend improvements.
  • Improve incident response and postmortem processes for ClickHouse Core outages, including blameless postmortems and customer communications.
  • Plan and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes and establish escalation best practices to minimize customer impact.

Requirements

  • Bachelor’s or master’s degree in Computer Science or a related field.
  • At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
  • Experience operating ClickHouse or other SQL databases in production.
  • Strong problem-solving and production debugging skills.
  • Excellent communication, ownership, accountability, and collaboration skills.

Nice-to-haves

  • Knowledge of distributed database internals and SQL, particularly ClickHouse.
  • Shell or Python scripting experience.
  • Ability to read and understand C++ code.
  • Familiarity with AWS, Azure, or Google Cloud Platform.

Compensation and Benefits

  • Employer healthcare contributions.
  • Company stock options.
  • Flexible time off, with country-specific entitlements.
  • USD $500 home office setup allowance for remote employees.
  • Opportunities to attend company-wide offsites.

Skills

ClickHouse, SQL, Distributed Databases, Shell, Python, C++, AWS, Microsoft Azure, GCP, Chaos Engineering, Incident Response, Postmortems

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Together AI

Together AI

San Francisco, CA
AI Infrastructure Systems Engineer
No salary listedHybrid3+ YOEDevOps / SRE

Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.