Cloud Database Infrastructure Engineer
Build and operate scalable, highly available cloud-native database infrastructure for ClickHouse’s serverless platform. The role requires 5+ years of distributed-systems experience, production programming in Go, C++, or Java, cloud infrastructure expertise, and operational on-call experience.
About the job
Responsibilities
- Build a cloud-native database platform on public cloud infrastructure.
- Develop and maintain infrastructure that supports ClickHouse as a fully functional serverless database solution.
- Improve and operate the in-house Kubernetes operator for seamless infrastructure management.
- Enhance the metrics pipeline and build systems that generate statistics and recommendations.
- Collaborate with the ClickHouse core development team and data plane teams on infrastructure use cases and internal improvements.
- Architect and build robust, scalable, highly available distributed infrastructure.
- Participate in production on-call rotations, debug production issues, and solve operational problems.
Requirements
- 5+ years of relevant software development experience building and operating scalable, fault-tolerant, distributed systems.
- Production experience with Go, C++, or Java.
- Experience with a public cloud provider such as AWS, Google Cloud, or Azure and infrastructure-as-a-service offerings such as EC2.
- Experience with data storage, ingestion, and transformation tools such as Spark or Kafka.
- Experience working with distributed systems.
- Strong problem-solving and communication skills, with the ability to work effectively across engineering teams.
- Willingness to participate in PagerDuty on-call rotations and debug production systems.
Benefits
- Employer healthcare contributions.
- Company stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home office setup allowance for remote employees.
- Opportunities to participate in company-wide global gatherings.
Skills
Kubernetes, Go, C++, Java, AWS, GCP, Azure, EC2, Spark, Apache Kafka, Distributed Systems, Pagerduty
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.