Senior Cloud Performance Engineer
Leads performance benchmarking, optimization, capacity planning, and chaos engineering for large-scale distributed cloud database systems. Requires 6+ years of software development experience, strong cloud infrastructure expertise, and proficiency in systems programming and production debugging.
About the job
Responsibilities
- Benchmark system and database performance, analyze results, and optimize capacity sizing.
- Troubleshoot and debug application and server errors and logs; triage issues accordingly.
- Recommend configuration tuning and optimizations for performance bottlenecks.
- Collaborate with core development, cloud, security, and partner engineering teams to improve ClickHouse Cloud performance.
- Plan, enable, and drive chaos initiatives across engineering teams.
- Develop, deploy, and manage tools for systematically running chaos experiments and measuring their impact.
- Study large-scale distributed systems and software resilience, operational, and delivery challenges.
- Extend backend systems to enable chaos engineering techniques.
- Observe running systems and prioritize innovative ways to test their resilience.
Requirements
- 6+ years of relevant software development experience building and operating scalable, fault-tolerant, distributed systems.
- Software development experience in Go, C, C++, Java, or similar languages.
- Experience with concurrency, multithreading, and distributed-system architecture deployment.
- Experience developing cloud infrastructure services, preferably with Kubernetes.
- Experience leading and delivering large-scope technical projects with experienced engineers.
- Expertise with a public cloud provider such as AWS, Google Cloud, or Azure and infrastructure-as-a-service offerings such as EC2.
- Strong production debugging and problem-solving skills.
- Excellent communication and cross-functional collaboration skills.
- Passion for efficiency, availability, scalability, and data governance.
- High ownership, responsibility, and accountability.
Benefits
- Healthcare contributions.
- Company equity through stock options.
- Flexible time off, with country-specific entitlements.
- USD $500 home-office setup allowance for remote employees.
- Opportunities to participate in company-wide offsites.
Skills
Go, C, C++, Java, Kubernetes, AWS, GCP, Microsoft Azure, Amazon Ec2, Chaos Engineering, Distributed Systems, Multithreading, Database Benchmarking, Performance Analysis
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.