Release Engineer - Data Plane Internal Tooling and Productivity
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
About the job
Responsibilities
- Plan and execute rolling upgrades across tens of thousands of ClickHouse clusters, ensuring safety, correctness, and minimal customer impact.
- Own the full release pipeline, including pre-upgrade validation, staged rollouts, post-upgrade monitoring, and incident response.
- Investigate and resolve production issues during a regular on-call rotation, including snowflake clusters and edge cases that automation cannot yet handle.
- Build and improve internal tooling and automation for reliable, repeatable large-scale database operations.
- Partner with core database and cloud infrastructure teams to identify operational pain points and develop solutions.
- Support and educate engineering teams using internal tools.
Requirements
- 5+ years of experience operating stateful distributed systems in production, such as databases, message queues, or storage systems.
- Hands-on experience running upgrades or maintenance operations on live production data stores at scale.
- Strong production debugging skills and comfort investigating unfamiliar systems under pressure.
- Experience with cloud infrastructure, including AWS, Azure, or Google Cloud.
- Experience with Kubernetes.
- Software development experience in Go, or strong experience in another language with willingness to learn Go.
Nice to Have
- Experience with ClickHouse as a user, operator, or contributor.
Compensation and Benefits
- Equity through company stock options.
- Healthcare contributions.
- Flexible time off.
- USD $500 home office setup allowance for remote employees.
- Opportunities to attend company-wide offsites.
Skills
Go, AWS, Azure, GCP, Kubernetes, ClickHouse, Distributed Systems, Database Operations, Production Debugging, Release Engineering
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.