Staff Engineer, Digital Infrastructure
Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.
About the job
Responsibilities
- Work with a small team of engineers and administrators.
- Integrate the AI training factory and tooling onto new hardware.
- Provide software updates and configuration management across isolated environments.
- Ensure proper operation of the AI training factory and tooling.
- Coordinate with autonomy teams to build and meet requirements.
- Debug and troubleshoot deployed, distributed problems.
- Develop and deploy automation tooling.
- Support onsite operations at the London office.
Requirements
- Typically 7+ years of related experience with a bachelor's degree, 6+ years with a master's degree, 4+ years with a PhD, or equivalent experience.
- 3+ years of experience with cloud computing solutions and architecture, focused on Kubernetes and container orchestration.
- Experience supporting C++ or Python applications.
- Strong understanding of Linux/Unix systems, including networking and networked application design and implementation.
- Bachelor's or master's degree in computer science, a similar discipline, or equivalent practical experience.
- Experience debugging and troubleshooting deployed, distributed systems.
- Experience developing and deploying automation tooling such as Ansible, Chef, or Puppet.
- Strong teamwork, ownership, reliability, and communication skills.
Skills
Kubernetes, Cloud Computing, Container Orchestration, C++, Python, Linux, Unix, Computer Networking, Ansible, Chef, Puppet, Configuration Management, Distributed Systems
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.