Skip to content
Shield AIShield AI

Staff Engineer, Digital Infrastructure

Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.

About the job

Responsibilities

  • Work with a small team of engineers and administrators.
  • Integrate the AI training factory and tooling onto new hardware.
  • Provide software updates and configuration management across isolated environments.
  • Ensure proper operation of the AI training factory and tooling.
  • Coordinate with autonomy teams to build and meet requirements.
  • Debug and troubleshoot deployed, distributed problems.
  • Develop and deploy automation tooling.
  • Support onsite operations at the London office.

Requirements

  • Typically 7+ years of related experience with a bachelor's degree, 6+ years with a master's degree, 4+ years with a PhD, or equivalent experience.
  • 3+ years of experience with cloud computing solutions and architecture, focused on Kubernetes and container orchestration.
  • Experience supporting C++ or Python applications.
  • Strong understanding of Linux/Unix systems, including networking and networked application design and implementation.
  • Bachelor's or master's degree in computer science, a similar discipline, or equivalent practical experience.
  • Experience debugging and troubleshooting deployed, distributed systems.
  • Experience developing and deploying automation tooling such as Ansible, Chef, or Puppet.
  • Strong teamwork, ownership, reliability, and communication skills.

Skills

Kubernetes, Cloud Computing, Container Orchestration, C++, Python, Linux, Unix, Computer Networking, Ansible, Chef, Puppet, Configuration Management, Distributed Systems

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, Observability & Profiling
£325k+/yrHybrid10+ YOEDevOps / SRE

Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Anthropic

Anthropic

London, United Kingdom

Staff Software Engineer, AI Reliability Engineering
£325k+/yrHybrid7+ YOEDevOps / SRE

Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.