Skip to content
ClickhouseClickhouse

Senior Cloud Performance Engineer

Leads performance benchmarking, optimization, capacity planning, and chaos engineering for large-scale distributed cloud database systems. Requires 6+ years of software development experience, strong cloud infrastructure expertise, and proficiency in systems programming and production debugging.

About the job

Responsibilities

  • Benchmark system and database performance, analyze results, and optimize capacity sizing.
  • Troubleshoot and debug application and server errors and logs; triage issues accordingly.
  • Recommend configuration tuning and optimizations for performance bottlenecks.
  • Collaborate with core development, cloud, security, and partner engineering teams to improve ClickHouse Cloud performance.
  • Plan, enable, and drive chaos initiatives across engineering teams.
  • Develop, deploy, and manage tools for systematically running chaos experiments and measuring their impact.
  • Study large-scale distributed systems and software resilience, operational, and delivery challenges.
  • Extend backend systems to enable chaos engineering techniques.
  • Observe running systems and prioritize innovative ways to test their resilience.

Requirements

  • 6+ years of relevant software development experience building and operating scalable, fault-tolerant, distributed systems.
  • Software development experience in Go, C, C++, Java, or similar languages.
  • Experience with concurrency, multithreading, and distributed-system architecture deployment.
  • Experience developing cloud infrastructure services, preferably with Kubernetes.
  • Experience leading and delivering large-scope technical projects with experienced engineers.
  • Expertise with a public cloud provider such as AWS, Google Cloud, or Azure and infrastructure-as-a-service offerings such as EC2.
  • Strong production debugging and problem-solving skills.
  • Excellent communication and cross-functional collaboration skills.
  • Passion for efficiency, availability, scalability, and data governance.
  • High ownership, responsibility, and accountability.

Benefits

  • Healthcare contributions.
  • Company equity through stock options.
  • Flexible time off, with country-specific entitlements.
  • USD $500 home-office setup allowance for remote employees.
  • Opportunities to participate in company-wide offsites.

Skills

Go, C, C++, Java, Kubernetes, AWS, GCP, Microsoft Azure, Amazon Ec2, Chaos Engineering, Distributed Systems, Multithreading, Database Benchmarking, Performance Analysis

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.

Clickhouse

Clickhouse

Remote

Senior Cloud Software Engineer - Efficiency Engineering
No salary listedRemote5+ YOEDevOps / SRE

Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.