Skip to content
EllipticElliptic

Lead DevOps Engineer

Leads and mentors a DevOps team while remaining hands-on in building secure, multi-region Kubernetes platforms, cloud infrastructure, GitOps delivery, and reliability practices. The role also supports AI platform infrastructure and requires strong expertise in Kubernetes, Terraform, security, observability, and team leadership.

About the job

Responsibilities

  • Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging, migration to EKS, and enabling reliable production delivery.
  • Lead by doing: engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure while setting technical standards.
  • Architect multi-cluster, multi-region environments using Istio, Linkerd, Cluster API, and Kyverno.
  • Build progressive delivery frameworks with Flux and Flagger for GitOps-driven canary and automated releases.
  • Implement Kubernetes-native cloud provisioning with Crossplane and ACK.
  • Define and enforce Zero Trust architecture with Vault, Boundary, service identity, and mTLS-secured service meshes.
  • Engineer policy-driven automation and compliance using OPA, Kyverno, and secure supply-chain configurations.
  • Establish Infrastructure-as-Code and GitOps standards with automated testing for infrastructure changes.
  • Prototype agentic infrastructure components, including deployment and observability platforms in service meshes.
  • Design AI gateways and registries for traffic and event routing between microservices and autonomous agents via CNCF Gateway APIs.
  • Champion DevSecOps maturity through SAST, DAST, chaos engineering, and error-budget monitoring.
  • Collaborate with Security, Data, and AI teams on DevOps and AI platform architectures and regulatory compliance.
  • Stay current with CNCF and AI ecosystem innovations, including eBPF observability and agent-aware orchestration.
  • Lead and mentor DevOps engineers while remaining hands-on with technical work.

Requirements

  • Experience leading or mentoring engineering teams, setting direction, and contributing hands-on.
  • Strong Kubernetes expertise, including cluster lifecycle management, API extensions, Operators, Helm, and the CNCF ecosystem.
  • Experience with multi-cluster, multi-region Kubernetes platforms and service meshes such as Istio, Consul, or Linkerd.
  • Experience writing modular Terraform Infrastructure-as-Code on AWS or GCP, integrating GitOps and automated testing such as Terratest or InSpec.
  • Experience implementing GitOps pipelines with ArgoCD or FluxCD for progressive delivery, drift correction, and multi-environment releases.
  • Experience building containerized, serverless, or event-driven systems with observability using Datadog, Splunk, or OpenTelemetry.
  • Experience with Vault-based secret management, least-privilege access, and compliance automation.
  • Experience designing CI/CD workflows with SAST, DAST, policy enforcement, and performance telemetry.
  • Experience improving reliability and resilience through SLOs, error budgets, and chaos engineering.
  • Passion for secure, scalable, and reliable systems and for helping others build them.
  • Strong customer and product mindset, ownership, thoughtful decision-making, collaboration, and commitment to inclusive team building.

Nice to Have

  • Leadership of platform modernization or reliability initiatives in scale-up or regulated environments.
  • Operator development, CRD automation, eBPF, or Cilium experience.
  • Policy-as-code expertise with OPA or Kyverno in secure supply-chain or CSPM frameworks.
  • Experience with AI-driven internal developer platforms or predictive observability using AI or LLMs.
  • Experience designing agentic or autonomous infrastructure with AI-agent observability.
  • Familiarity with MCP and A2A orchestration patterns in Kubernetes service-mesh environments.
  • Understanding of Agent Gateways and Registries connecting microservices and AI agents.
  • Experience with secure containers, sandboxing, or confidential computing for regulated workloads.
  • Experience with Spark, Databricks, or Data Mesh.
  • Programming experience in Go, Python, or TypeScript.
  • Contributions to open-source or CNCF community projects.

Compensation and Benefits

  • Hybrid working, with the option to work from almost anywhere for up to 90 days per year.
  • £500 remote-working budget for home-office setup.
  • $1,000 Learning & Development budget.
  • 25 days of annual leave plus bank holidays.
  • Additional day off for your birthday.
  • Enhanced parental leave, including 16 weeks of fully paid leave for eligible employees.
  • Private health insurance through Vitality.
  • Spill Mental Health Support.
  • Life assurance covering four times salary for beneficiaries.
  • Cycle to Work Scheme.

Skills

Kubernetes, Amazon Eks, Terraform, AWS, GCP, GitOps, Fluxcd, Argo CD, Istio, Linkerd, Helm, Kyverno, Open Policy Agent, Hashicorp Vault, OpenTelemetry

Shield AI

Shield AI

London, United Kingdom

Senior DevSecOps Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Designs and operates secure development infrastructure and CI/CD pipelines for autonomous defence systems. Requires at least five years of DevOps or related experience, plus UK defence or regulated national-security experience and knowledge of Secure by Design and assurance practices.

Clear Street

Clear Street

London, United Kingdom

Senior Production Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Hudl

Hudl

London, United Kingdom

Senior Engineer - Platform
£66k+/yrRemote5+ YOEDevOps / SRE

Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.

Muck Rack

Muck Rack

Bulgaria
Senior Software Engineer, DevOps
€95k+/yrRemote5+ YOEDevOps / SRE

Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.