Skip to content
ClickhouseClickhouse

Senior Site Reliability Engineer

The Senior Site Reliability Engineer will improve the reliability, availability, scalability, and performance of ClickHouse Cloud by designing distributed systems, managing observability and incident response, and driving automation and chaos initiatives. The role requires 8+ years of SRE experience and hands-on Go or Python expertise.

About the job

Responsibilities

  • Design and implement scalable, secure, highly available, and fault-tolerant distributed systems for ClickHouse Cloud.
  • Establish and manage service-level objectives (SLOs) and service-level agreements (SLAs).
  • Ensure infrastructure components have monitoring and alerting for timely incident detection and resolution.
  • Improve incident response and outage postmortem processes, including blameless postmortems and customer communications.
  • Continuously improve service reliability and performance.
  • Plan and drive chaos engineering initiatives across engineering teams.
  • Manage on-call processes and escalation best practices to minimize downtime.
  • Develop software platforms and tools that improve operational and engineering efficiency.

Requirements

  • Bachelor's or master's degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Understanding of distributed databases and SQL; ClickHouse experience is a major plus.
  • Experience with Kubernetes or Docker Swarm.
  • Experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Strong production debugging and problem-solving skills.
  • Excellent communication and interpersonal skills.

Benefits

  • Healthcare contributions.
  • Company stock options.
  • Flexible time off, with country-specific entitlements.
  • USD $500 home-office setup allowance for remote employees.
  • Opportunities to attend company-wide global gatherings.

Skills

Go, Python, AWS, Microsoft Azure, GCP, SQL, ClickHouse, Kubernetes, Docker Swarm, Ansible, Terraform, Puppet, Chaos Engineering, Incident Management, Distributed Systems

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.

Shield AI

Shield AI

London, United Kingdom

Senior DevSecOps Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Designs and operates secure development infrastructure and CI/CD pipelines for autonomous defence systems. Requires at least five years of DevOps or related experience, plus UK defence or regulated national-security experience and knowledge of Secure by Design and assurance practices.

Clear Street

Clear Street

London, United Kingdom

Senior Production Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

Cohere

Cohere

United States
Software Engineer, GPU Infrastructure
No salary listedHybrid7+ YOEDevOps / SRE

Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.