Senior Cloud Performance Engineer
Build tools and practices for benchmarking, optimizing, and testing the resilience of large-scale distributed cloud database systems. The role requires 6+ years of software development experience, strong distributed-systems expertise, cloud infrastructure knowledge, and production debugging skills.
About the job
Responsibilities
- Benchmark system and database performance, analyze results, size capacity, and optimize systems.
- Troubleshoot and debug applications, server errors, and logs; triage issues accordingly.
- Recommend configuration tuning and optimizations for performance bottlenecks.
- Partner with core development, cloud, and security teams to improve ClickHouse Cloud performance.
- Plan and drive chaos engineering initiatives across engineering teams.
- Develop, deploy, and manage tools for systematically running chaos experiments and measuring their impact.
- Study software resilience, operational, and delivery challenges.
- Extend backend systems to enable chaos engineering techniques.
- Observe running systems and prioritize innovative ways to test their resilience.
Requirements
- 6+ years of relevant software development industry experience building and operating scalable, fault-tolerant, distributed systems.
- Software development experience in Go, C, C++, Java, or similar languages.
- Experience with concurrency, multithreading, and distributed system architectures.
- Experience developing cloud infrastructure services, preferably with Kubernetes.
- Experience leading and delivering large-scope technical projects collaboratively.
- Expertise with a public cloud provider such as AWS, Google Cloud, or Azure and its infrastructure-as-a-service offerings.
- Strong production debugging and problem-solving skills.
- Excellent communication and cross-functional collaboration skills.
- Passion for efficiency, availability, scalability, and data governance.
- Strong ownership, responsibility, and accountability.
Benefits
- Flexible work environment for remote employees.
- Employer healthcare contributions.
- Company stock options.
- Flexible or country-specific time off.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to participate in company-wide offsites.
Skills
Go, C, C++, Java, Kubernetes, AWS, GCP, Microsoft Azure, Distributed Systems, Chaos Engineering, Database Benchmarking, Multithreading, Capacity Planning, Production Debugging
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.