Skip to content
WriterWriter

Infrastructure Engineer

Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.

About the job

Responsibilities

  • Build resilient, scalable, fault-tolerant infrastructure for a high-traffic enterprise AI platform.
  • Work across SRE, DevOps, infrastructure, and platform initiatives, including on-call operations, release pipelines, multi-region infrastructure, and internal platform capabilities.
  • Automate operational tasks and infrastructure management using Python or Go.
  • Design and operate cloud infrastructure across AWS, GCP, and Azure, with Kubernetes, Helm, Terraform, and related cloud tooling.
  • Use AI-assisted workflows to investigate incidents, draft infrastructure changes, write runbooks, scaffold tooling, and review pull requests.
  • Lead incident response, postmortems, and root-cause analyses; incorporate findings into architecture and preventative measures.
  • Own reliability, performance, and efficiency for core services, including SLOs, error budgets, and on-call operations.
  • Shape longer-term observability, cost, and reliability investments while addressing immediate production issues.
  • Partner with product, security, and engineering teams on reliable, performant, scalable system design.

Requirements

  • 5+ years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, highly available production systems.
  • Production experience running containerized workloads and real clusters.
  • Experience with Helm and Terraform or Pulumi on at least one major cloud provider; AWS experience preferred.
  • Proficiency in Python or Go for automation and tooling.
  • Daily experience using agentic or AI-assisted development and operations tooling, and experience building or adopting AI-assisted workflows.
  • Strong first-principles reasoning and ability to identify systemic reliability weaknesses and evaluate tradeoffs.
  • Ability to make reversible production changes, define rollback plans, and manage blast radius.
  • Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent tools.
  • Strong communication, collaboration, problem-solving, autonomy, and ownership skills.
  • At least one end-to-end 0-to-1 infrastructure build with measurable outcomes.

Nice to have

  • Software engineering background with experience designing and shipping production services, libraries, or internal frameworks.
  • Ability to work across infrastructure automation and feature engineering using Python, Go, or a comparable language.

Compensation and benefits

  • Competitive compensation and company stock options.
  • Generous paid time off and company holidays.
  • Medical and dental insurance.
  • 16 weeks of paid parental leave for all parents.
  • Fertility and family planning support.
  • Early-detection cancer testing.
  • Competitive pension scheme and company contribution.
  • Wellness, learning and development, and work-life stipends.
  • Company-wide and team off-sites.

Skills

AWS, GCP, Azure, Kubernetes, Helm, Terraform, Pulumi, Python, Go, Prometheus, Grafana, Elk, Claude Code, Droid, Codex

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

GitLab

GitLab

United Kingdom

Site Reliability Engineer, Infrastructure Platforms
No salary listedRemote5+ YOEDevOps / SRE

Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Scale AI

Scale AI

London, United Kingdom

Infrastructure Software Engineer, Apps Platform
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.