# AI Infrastructure Engineer, Serving Platform

**Company:** [Scale AI](https://hotfix.jobs/companies/scale-ai)
**Location:** London, United Kingdom
**Role:** ML Engineering
**Experience:** 4+ years
**Skills:** Python, Go, Rust, C++, Llm Serving, Load Balancing, Kubernetes, Docker, AWS, GCP, Terraform, vLLM, Sglang, Tensorrt-Llm, Monitoring
**Posted:** 2026-08-04

> Build scalable, fault-tolerant platforms for serving large language models across research and production environments. The role requires 4+ years of backend systems experience, strong programming skills, and familiarity with LLM serving, containers, cloud infrastructure, and infrastructure as code.

## Job Description

## Responsibilities
- Build and maintain fault-tolerant, high-performance systems for serving large language models and other models at scale.
- Build an internal platform for large language model capability discovery.
- Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
- Conduct architecture and design reviews to uphold best practices in system design and scalability.
- Develop monitoring and observability solutions to ensure system health and performance.
- Lead projects end-to-end, from requirements gathering through implementation, in a cross-functional environment.

## Requirements
- 4+ years of experience building large-scale, high-performance backend systems.
- Strong programming skills in one or more of Python, Go, Rust, or C++.
- Experience with large language model serving and routing fundamentals, including rate limiting, token streaming, load balancing, and budgets.
- Experience with large language model capabilities and concepts such as reasoning, tool calling, and prompt templates.
- Experience with containers and orchestration tools.
- Familiarity with cloud infrastructure and infrastructure as code.
- Proven ability to solve complex problems and work independently in fast-moving environments.

## Nice to Haves
- Experience with modern large language model serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.

## Similar jobs

- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Applied AI Engineer, Digital Natives](https://hotfix.jobs/jobs/74eb3e31-bab1-479e-90e1-2e172658343a) - OpenAI - London, United Kingdom
- [Agent Engineer](https://hotfix.jobs/jobs/627cc378-6ac6-467e-bbf0-2c07738ae0a5) - Elliptic - London, United Kingdom
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [AI Engineer - Assistant Experience](https://hotfix.jobs/jobs/00d11204-96cf-4c85-a12f-9058f0581c15) - Build - New York, NY - $120k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/72401dcb-35d1-4b38-ab71-121a028b6312
**Canonical:** https://hotfix.jobs/jobs/72401dcb-35d1-4b38-ab71-121a028b6312