Senior Cloud Performance Engineer
Senior cloud performance engineer responsible for benchmarking and optimizing distributed database and cloud infrastructure systems, while building chaos-engineering tools to improve resilience and scalability. Requires 6+ years of software development experience and expertise in distributed systems and public-cloud infrastructure.
About the job
Responsibilities
- Benchmark system and database performance; analyze results, size capacity, and optimize systems.
- Troubleshoot and debug application and server errors and logs, triaging issues appropriately.
- Recommend configuration tuning and optimizations for performance bottlenecks.
- Collaborate with core development, cloud, security, and partner engineering teams to improve ClickHouse Cloud performance.
- Plan, enable, and drive chaos initiatives across engineering teams.
- Develop, deploy, and manage tools for systematically running chaos experiments and measuring their impact.
- Extend backend systems to enable chaos engineering techniques.
- Study software resilience, operational, and delivery challenges in large-scale distributed systems.
- Observe running systems and prioritize innovative ways to test system resilience.
Requirements
- 6+ years of relevant software development industry experience building and operating scalable, fault-tolerant, distributed systems.
- Software development experience in Go, C/C++, Java, or similar languages.
- Experience with concurrency, multithreading, and deploying distributed system architectures.
- Experience developing cloud infrastructure services, preferably with Kubernetes.
- Experience leading and delivering large-scope technical projects with multiple engineers.
- Expertise with a public cloud provider such as AWS, Google Cloud, or Azure and infrastructure-as-a-service offerings such as Amazon EC2.
- Strong production debugging and problem-solving skills.
- Excellent communication and cross-functional collaboration skills.
- Passion for efficiency, availability, scalability, and data governance.
- High ownership, responsibility, and accountability.
Compensation and Benefits
- Equity through company stock options.
- Flexible time off; entitlement varies by country.
- USD $500 home-office setup allowance for remote employees.
- Healthcare contributions.
- Flexible work environment and company-wide gatherings.
Skills
Go, C++, C, Java, Kubernetes, AWS, GCP, Microsoft Azure, Amazon Ec2, Distributed Systems, Chaos Engineering, Multithreading, Database Benchmarking, Performance Analysis
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.