# Senior Production Engineer

**Company:** [Clear Street](https://hotfix.jobs/companies/clear-street)
**Location:** London, United Kingdom
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Python, Kubernetes, AWS, Terraform, Argo CD, GitHub Actions, Datadog, Kafka, Redis, Snowflake, Postgres, Go, Java, gRPC, Protobuf
**Posted:** 2026-09-08

> Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

## Job Description

## Responsibilities
- Own the health, resilience, and recovery of production systems.
- Design and build monitoring and observability platforms that reduce alert fatigue and accelerate root-cause analysis.
- Develop automation and self-healing capabilities for diagnosis and recovery workflows.
- Analyze incidents, identify systemic trends, and prevent recurring failures.
- Build reusable runbooks, diagnostic tooling, and recovery playbooks.
- Create reliable, standardized operational workflows for engineering teams.
- Partner with Platform Engineering on CI/CD pipelines, deployment safety, and infrastructure resilience.
- Champion Infrastructure as Code, GitOps, and SRE practices.
- Measure production health through SLIs, SLOs, and SLAs and use reliability data to guide priorities.
- Explore AI-assisted diagnostics and developer tooling.

## Requirements
- Strong hands-on Python skills for automation and tooling.
- Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
- Demonstrated experience building automation and diagnostic tooling that improves recovery times or reduces operational toil.
- Deep familiarity with cloud-native technologies, Kubernetes, containers, and distributed systems.
- Experience with observability platforms such as Datadog.
- Exposure to Infrastructure as Code with Terraform and GitOps deployment workflows using ArgoCD, GitHub Actions, or similar tools.
- Familiarity with Java, Go, Kafka, Redis, Snowflake, and PostgreSQL.
- Strong analytical, problem-solving, communication, and cross-functional collaboration skills.
- Product mindset for internal operational tooling, including usability, adoption, and documentation.
- Self-starter mentality and continuous-learning mindset.

## Nice-to-have
- Fintech or financial industry experience.

## Technology
- Kubernetes, AWS, Terraform, ArgoCD, GitHub Actions
- Kafka, Redis, PostgreSQL, Snowflake
- Datadog, Python, Go, Java
- gRPC, Protobuf, internal platform APIs, and developer tooling

## Similar jobs

- [Senior DevSecOps Engineer](https://hotfix.jobs/jobs/9fad0d81-f196-425c-b1d6-cf7adc34ac02) - Shield AI - London, United Kingdom
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Engineer - Platform](https://hotfix.jobs/jobs/df84a55a-dda0-45b3-afb8-9680e622a064) - Hudl - Remote - £66k – £111k/yr
- [Senior Software Engineer, DevOps](https://hotfix.jobs/jobs/bfede940-874a-457b-a7d5-afaa6be82cff) - Muck Rack - Remote - €95k – €110k/yr
- [Senior Database Administrator - Core Infrastructure](https://hotfix.jobs/jobs/84ebc795-fb55-4998-94cf-dd630cc51e8a) - Kraken - Remote

**Apply:** https://hotfix.jobs/jobs/98eb4848-f8e8-4b98-95bd-c5f2b852c00f
**Canonical:** https://hotfix.jobs/jobs/98eb4848-f8e8-4b98-95bd-c5f2b852c00f