Skip to content
IntercomIntercom

Senior Engineer, Infrastructure Platform

Build and operate scalable infrastructure platforms that improve reliability, developer velocity, security, and operational efficiency across R&D. The role requires senior-level experience with cloud-based distributed systems, automation, incident response, and AI-assisted engineering.

About the job

Responsibilities

  • Automate manual, repetitive, and unscalable infrastructure work through software tooling.
  • Plan, design, and execute architectural changes to core infrastructure and build elasticity into systems.
  • Own reliability of core services, respond to operational events, investigate root causes, and implement preventative solutions.
  • Trace incidents beyond the infrastructure layer into application code and resolve customer-facing issues.
  • Abstract shared concerns such as security, compute, availability, and costs for product engineers.
  • Create agent skills, documentation, and reliable defaults that enable engineers to work more effectively.
  • Lead complex infrastructure initiatives through design, implementation, testing, and operational runbooks.
  • Use AI and coding agents to automate infrastructure work, investigations, coding, and operational tasks.
  • Review code, share expertise, unblock peers, and support continuous learning.

Requirements

  • Multiple years of experience designing, building, and operating high-scale distributed systems and cloud platforms, ideally AWS.
  • Accountability for system availability, performance, and costs.
  • Deep knowledge of modern programming languages.
  • Strong understanding of reliability, performance, failure modes, metrics, and telemetry.
  • Experience treating infrastructure as a software engineering discipline.
  • Understanding of product developers’ pain points and a track record of reducing developer friction.
  • Proven ability to independently drive complex, cross-team technical initiatives from conception through production.
  • Strong written and verbal communication, including technical documentation, blameless incident reviews, and trade-off discussions.
  • Experience using AI to accelerate engineering work and building safe, reliable tooling around coding agents.

Nice to Have

  • Experience with infrastructure as code.

Compensation and Benefits

  • Competitive salary and equity in a fast-growing startup.
  • Regular compensation reviews.
  • Lunch provided every weekday, snacks, and a fully stocked kitchen.
  • Unlimited access to Claude Code and other AI tools.
  • Pension scheme with matching up to 4%.
  • Life assurance and comprehensive health and dental insurance for employees and dependents.
  • Flexible paid time off.
  • Paid maternity leave and six weeks of paternity leave.
  • Cycle-to-Work Scheme and secure bike storage.
  • MacBooks are standard, with Windows available for certain roles.
  • Hybrid working policy with employees expected to be in the office at least three days per week.

Skills

AWS, Cloud Infrastructure, Distributed Systems, Infrastructure As Code, Programming Languages, Reliability Engineering, Observability, Telemetry, Coding Agents, Technical Documentation

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Muck Rack

Muck Rack

Bulgaria
Senior Software Engineer, DevOps
€95k+/yrRemote5+ YOEDevOps / SRE

Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.

Kraken

Kraken

United Arab Emirates
Senior Database Administrator - Core Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Crusoe

Crusoe

Dublin, Ireland

Senior Network Production Operations Engineer
No salary listedOn-site10+ YOEDevOps / SRE

Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.