Senior Cloud Performance Engineer
The Senior Cloud Performance Engineer benchmarks and optimizes distributed database and cloud infrastructure performance while developing chaos engineering tools and initiatives. The role requires 6+ years of experience with scalable distributed systems, programming in Go, C/C++, or Java, Kubernetes, and a major public cloud provider.
About the job
Responsibilities
- Benchmark system and database performance; perform performance analysis, capacity sizing, and optimization.
- Troubleshoot and debug application and server errors and logs, and triage issues.
- Recommend configuration tuning and optimizations for performance bottlenecks.
- Collaborate with ClickHouse core development, cloud, security, and partner engineering teams to improve ClickHouse Cloud performance.
- Plan, enable, and drive chaos engineering initiatives across engineering teams.
- Develop, deploy, and manage tools for systematically running chaos experiments and measuring their impact.
- Analyze large-scale distributed systems, software resilience, operational challenges, and delivery processes.
- Extend backend systems to support chaos engineering techniques.
- Observe running systems and prioritize innovative ways to test system resilience.
Requirements
- 6+ years of relevant software development industry experience building and operating scalable, fault-tolerant, distributed systems.
- Software development experience in Go, C/C++, Java, or similar languages.
- Experience with concurrency, multithreading, and distributed system architectures.
- Experience developing cloud infrastructure services, preferably with Kubernetes.
- Experience leading and shipping large-scope technical projects with multiple experienced engineers.
- Expertise with a public cloud provider such as AWS, Google Cloud, or Azure and its infrastructure-as-a-service offerings, such as EC2.
- Excellent communication, collaboration, problem-solving, production debugging, ownership, and accountability skills.
- Passion for efficiency, availability, scalability, and data governance.
Compensation and Benefits
- Equity in the company through stock options.
- Flexible time off; entitlement varies by country.
- USD $500 home-office setup allowance for remote employees.
- Healthcare contributions.
- Opportunities to attend company-wide global gatherings.
Skills
Go, C++, C, Java, Kubernetes, AWS, GCP, Microsoft Azure, Amazon Ec2, Chaos Engineering, Distributed Systems, Multithreading, Performance Analysis, Database Benchmarking
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.