Skip to content
PostHogPostHog

ClickHouse Operations Engineer

Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.

About the job

What you'll be doing

ClickHouse is the core piece of infrastructure at PostHog. Every product and customer relies on it to ingest, store, and query data.

We need someone to automate, manage, and maintain ClickHouse as we grow towards capturing trillions of events per year and having one of the world’s largest clusters.

This includes ClickHouse operations and scaling infrastructure, as well as node and instance-level performance optimization. Ensure the right hardware deployed at the right time for each workload on ClickHouse.

Build systems and automations for provisioning and scaling of large ClickHouse clusters, handling over 100 PB's of data. Investigate and experiment using the latest hardware that cloud providers have to offer. Use Terraform, Ansible, and Kubernetes to automate dynamic provisioning of instances and work on bleeding edge ClickHouse implementation, like open format backed tables, and query performance tooling.

You’ll fit right in if:

  • OLAP Database Experience. Focused on ClickHouse, but strong experience with other OLAP Databases is great. Experience with internals of ClickHouse and other OLAP Databases, not high level users.
  • Automating Dynamic Provisioning Instances. Strong experience with Terraform, Ansible and Kubernetes.
  • Experience with Scale and Complexity! Building and operating high-scale complex data storage solutions.
  • The Stack we need. Python, Terraform, Ansible, Kubernetes, AWS, and Zookeeper (or alternative).

Skills

ClickHouse, Terraform, Ansible, Kubernetes, Python, AWS, Zookeeper, Olap Databases

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.