Skip to content
ClickhouseClickhouse

Database Reliability Engineer - Core Team

Build and lead reliability engineering practices for ClickHouse Core, improving production performance, observability, incident response, and database operations. The role requires at least five years of reliability, QA, or customer-facing engineering experience, plus production SQL database and cloud expertise.

About the job

Responsibilities

  • Improve the reliability, availability, scalability, and performance of ClickHouse Core.
  • Create and improve metrics and alerts to identify and prevent production issues before they affect customers.
  • Investigate recurring customer problems, determine root causes, submit bug fixes and issue reports, and recommend improvements.
  • Enhance incident response and post-mortem processes for ClickHouse Core outages, including blameless postmortems and customer communications.
  • Plan and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes and establish escalation best practices to resolve reliability and performance issues while minimizing customer impact.
  • Collaborate with Control Plane, Dataplane, Security, Support, and Operations teams.

Requirements

  • Bachelor's or master's degree in Computer Science or a related field.
  • At least 5 years of experience in reliability engineering, QA, or customer-facing engineering.
  • Experience operating ClickHouse or other SQL databases in production.
  • Strong understanding of distributed database internals and SQL.
  • Scripting experience with Shell or Python.
  • Ability to read and understand C++ code.
  • Knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Strong production debugging and problem-solving skills.
  • Excellent communication skills and the ability to work effectively in a global, fast-paced team.
  • High ownership, responsibility, and accountability.

Nice to Have

  • Experience with ClickHouse.
  • Experience with chaos engineering and post-mortem analysis.

Benefits

  • Healthcare contributions.
  • Company stock options.
  • Flexible time off in the United States and generous leave allowances in other countries.
  • USD $500 home-office setup allowance for remote employees.
  • Company-wide global gatherings.

Skills

SQL, ClickHouse, Distributed Databases, Shell, Python, C++, AWS, Azure, GCP, Chaos Engineering, Incident Response, Postmortems, Production Debugging

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Together AI

Together AI

San Francisco, CA
AI Infrastructure Systems Engineer
No salary listedHybrid3+ YOEDevOps / SRE

Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

PostHog

PostHog

Remote

ClickHouse Operations Engineer
No salary listedRemoteDevOps / SRE

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.