Skip to content

Senior Production Engineer

Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

About the job

Responsibilities

  • Own the health, resilience, and recovery of production systems.
  • Design and build monitoring and observability platforms that reduce alert fatigue and accelerate root-cause analysis.
  • Develop automation and self-healing capabilities for diagnosis and recovery workflows.
  • Analyze incidents, identify systemic trends, and prevent recurring failures.
  • Build reusable runbooks, diagnostic tooling, and recovery playbooks.
  • Create reliable, standardized operational workflows for engineering teams.
  • Partner with Platform Engineering on CI/CD pipelines, deployment safety, and infrastructure resilience.
  • Champion Infrastructure as Code, GitOps, and SRE practices.
  • Measure production health through SLIs, SLOs, and SLAs and use reliability data to guide priorities.
  • Explore AI-assisted diagnostics and developer tooling.

Requirements

  • Strong hands-on Python skills for automation and tooling.
  • Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
  • Demonstrated experience building automation and diagnostic tooling that improves recovery times or reduces operational toil.
  • Deep familiarity with cloud-native technologies, Kubernetes, containers, and distributed systems.
  • Experience with observability platforms such as Datadog.
  • Exposure to Infrastructure as Code with Terraform and GitOps deployment workflows using ArgoCD, GitHub Actions, or similar tools.
  • Familiarity with Java, Go, Kafka, Redis, Snowflake, and PostgreSQL.
  • Strong analytical, problem-solving, communication, and cross-functional collaboration skills.
  • Product mindset for internal operational tooling, including usability, adoption, and documentation.
  • Self-starter mentality and continuous-learning mindset.

Nice-to-have

  • Fintech or financial industry experience.

Technology

  • Kubernetes, AWS, Terraform, ArgoCD, GitHub Actions
  • Kafka, Redis, PostgreSQL, Snowflake
  • Datadog, Python, Go, Java
  • gRPC, Protobuf, internal platform APIs, and developer tooling

Skills

Python, Kubernetes, AWS, Terraform, Argo CD, GitHub Actions, Datadog, Kafka, Redis, Snowflake, Postgres, Go, Java, gRPC, Protobuf

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Hudl

Hudl

London, United Kingdom

Senior Engineer - Platform
£66k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.

Muck Rack

Muck Rack

Bulgaria
Senior Software Engineer, DevOps
€95k+/yrRemote5+ YOEDevOps / SRE

Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.

Kraken

Kraken

United Arab Emirates
Senior Database Administrator - Core Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.

Invisible Tech

Invisible Tech

London, United Kingdom

Senior DevOps Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Own and evolve a platform domain supporting reliable, secure, and cost-effective multi-tenant infrastructure. The role requires strong Kubernetes, cloud, Terraform, Helm, GitOps, Python, security, and agentic coding expertise, along with excellent technical judgment and communication.