# AI Platform Support Engineer

**Company:** [Lightning AI](https://hotfix.jobs/companies/lightning-ai)
**Location:** London, United Kingdom
**Role:** Support Engineering
**Salary:** £75k – £95k/yr
**Skills:** Kubernetes, Linux, Distributed Systems, Cloud Infrastructure, Prometheus, Grafana, OpenTelemetry, PyTorch, CUDA, Nccl, Gpu Infrastructure, Ray, Kubeflow, Slurm, Python
**Posted:** 2026-08-17

> Supports technical customers running production ML workloads by diagnosing distributed systems, Kubernetes, GPU, networking, storage, and inference issues. The role also drives reliability improvements, tooling, observability, and operational guidance across platform engineering teams.

## Job Description

## Responsibilities

### Customer and Incident Support
- Partner directly with customer engineering teams running training and inference workloads in production.
- Diagnose and resolve complex distributed-systems and ML-infrastructure issues.
- Act as a technical advisor during high-impact incidents and platform degradation events.
- Translate infrastructure-level issues into actionable guidance for ML engineers.
- Build credibility with customers through strong technical reasoning and clear communication.

### ML Infrastructure and Distributed Workloads
- Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems.
- Troubleshoot PyTorch, CUDA, NCCL, and inference-serving issues.
- Analyze logs, metrics, traces, and system behavior to isolate root causes.
- Debug containerized workloads running across Kubernetes and bare-metal GPU environments.
- Support customers scaling workloads across multi-node GPU systems.
- Diagnose performance bottlenecks involving compute, memory, networking, or storage.

### Reliability and Platform Operations
- Identify recurring patterns across customer issues and drive long-term reliability improvements.
- Contribute to post-incident reviews and operational improvements.
- Build internal tooling, automation, documentation, and runbooks.
- Partner with infrastructure, networking, and platform engineering teams.
- Improve observability, operational visibility, and troubleshooting workflows.
- Improve the customer experience through better processes and technical guidance.

## Requirements

- Strong software-engineering and systems-troubleshooting background.
- Experience with Kubernetes and containerized environments.
- Linux systems knowledge, including networking, storage, process management, and performance tuning.
- Experience with cloud infrastructure and distributed systems.
- Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
- Hands-on experience operating machine-learning workloads in production or research environments.
- Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL.
- Familiarity with GPU infrastructure and orchestration.
- Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure.
- Understanding of the operational challenges involved in running ML systems at scale.
- Strong communication skills and ability to work directly with highly technical customers and engineering teams.
- Comfort operating in fast-moving, highly ambiguous environments.
- Enjoyment of solving complex technical problems collaboratively.

## Nice-to-Haves

- Experience with large-scale model training or distributed inference systems.
- Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms.
- Experience with InfiniBand, RDMA, or high-performance networking.
- Experience operating bare-metal infrastructure.
- Familiarity with storage systems commonly used in ML environments.
- Experience at an AI-infrastructure, cloud, MLOps, or developer-tooling company.
- Contributions to platform engineering, developer infrastructure, or operational-tooling projects.
- Experience writing automation, tooling, or scripts in Python or similar languages.

## Compensation and Benefits

- Anticipated annual base salary: **£75,000–£95,000 GBP**.
- Discretionary bonus and meaningful equity component.
- Comprehensive health coverage, including medical, dental, and vision coverage for employees and eligible dependents.
- Retirement savings and pension contributions in the U.K.
- Unlimited paid time off, company holidays, and floating holidays.
- Two-week company-wide winter break.
- Paid parental and family leave.
- Annual learning and development allowance.
- Wellness and work-from-home stipends.
- Four weeks of paid sabbatical leave after four years of service.
- Flexible schedules and a hybrid work model for office-based teams.
- Complimentary meals at office hubs.
- Benefits may vary by location, team, and role.

## Similar jobs

- [Customer Support Engineer](https://hotfix.jobs/jobs/9854c864-e33d-4340-a5ee-2e1914eaf136) - Tailscale - Remote - CA$84k – CA$133k/yr
- [Technical Support Engineer](https://hotfix.jobs/jobs/281ff086-630f-4a0d-a9d5-1e94322ad10b) - OPSWAT - London, United Kingdom - £60k – £85k/yr
- [Product Support Engineer](https://hotfix.jobs/jobs/d6b9db9e-339c-45ec-96aa-9dce94b8898d) - Nominal - New York, NY - $94k – $140k/yr
- [Technical Support Engineer](https://hotfix.jobs/jobs/a4e5f7f6-852e-4ce2-94d4-f6d9280a8791) - ConductorOne - Remote - $120k – $140k/yr
- [Premium Support Engineer](https://hotfix.jobs/jobs/31235aa2-5b50-4e1f-a59a-92c6ac386212) - Replit - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/602943a3-4085-45ab-8120-bb00a2c2d5d3
**Canonical:** https://hotfix.jobs/jobs/602943a3-4085-45ab-8120-bb00a2c2d5d3