# Infrastructure Engineer

**Company:** [Writer](https://hotfix.jobs/companies/writer)
**Location:** London, United Kingdom
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Elk, Claude Code, Droid, Codex
**Posted:** 2026-09-04

> Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.

## Job Description

## Responsibilities
- Build resilient, scalable, fault-tolerant infrastructure for a high-traffic enterprise AI platform.
- Work across SRE, DevOps, infrastructure, and platform initiatives, including on-call operations, release pipelines, multi-region infrastructure, and internal platform capabilities.
- Automate operational tasks and infrastructure management using Python or Go.
- Design and operate cloud infrastructure across AWS, GCP, and Azure, with Kubernetes, Helm, Terraform, and related cloud tooling.
- Use AI-assisted workflows to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review pull requests.
- Lead incident response, postmortems, and root-cause analyses; incorporate findings into architecture and preventative measures.
- Own reliability, performance, and efficiency for core services, including SLOs, error budgets, and on-call operations.
- Shape longer-term observability, cost, and reliability investments while addressing immediate production issues.
- Partner with product, security, and engineering teams on reliable, performant, scalable system design.

## Requirements
- 5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, highly available production systems.
- Production experience running containerized workloads and real clusters.
- Experience with Helm and Terraform or Pulumi on at least one major cloud provider; AWS experience preferred.
- Proficiency in Python or Go for automation and tooling.
- Daily experience using agentic or AI-assisted development and operations tooling, and experience building or adopting AI-assisted workflows.
- Strong first-principles reasoning and ability to identify systemic reliability weaknesses and evaluate tradeoffs.
- Ability to make reversible production changes, define rollback plans, and manage blast radius.
- Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent tools.
- Strong communication, collaboration, problem-solving, autonomy, and ownership skills.
- At least one end-to-end 0-to-1 infrastructure build with measurable outcomes.

## Nice to have
- Software engineering background with experience designing and shipping production services, libraries, or internal frameworks.
- Ability to work across infrastructure automation and feature engineering using Python, Go, or a comparable language.

## Compensation and benefits
- Competitive compensation and company stock options.
- Generous paid time off and company holidays.
- Medical and dental insurance.
- 16 weeks of paid parental leave for all parents.
- Fertility and family planning support.
- Early-detection cancer testing.
- Competitive pension scheme and company contribution.
- Wellness, learning and development, and work-life stipends.
- Company-wide and team off-sites.

## Similar jobs

- [Software Engineer: Resiliency - Deploy at Scale](https://hotfix.jobs/jobs/21bb6f4f-daea-46ab-bac8-fff083d17451) - Cloudflare - London, United Kingdom
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Site Reliability Engineer, Infrastructure Platforms](https://hotfix.jobs/jobs/e28d2ba0-df16-41d6-9c60-e694ee353996) - GitLab - Remote
- [Member of Technical Staff](https://hotfix.jobs/jobs/d6c912e7-2f16-4a86-8738-980dd0b47cd0) - Perplexity - Remote - $220k – $405k/yr
- [Infrastructure Software Engineer, Apps Platform](https://hotfix.jobs/jobs/5c50edd8-ecfb-4db4-86e8-8698ff8827cd) - Scale AI - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/e7b74501-38ae-4cd1-8b26-91e042b41460
**Canonical:** https://hotfix.jobs/jobs/e7b74501-38ae-4cd1-8b26-91e042b41460