Skip to content
Scale AIScale AI

AI Infrastructure Engineer, Serving Platform

Build scalable, fault-tolerant platforms for serving large language models across research and production environments. The role requires 4+ years of backend systems experience, strong programming skills, and familiarity with LLM serving, containers, cloud infrastructure, and infrastructure as code.

About the job

Responsibilities

  • Build and maintain fault-tolerant, high-performance systems for serving large language models and other models at scale.
  • Build an internal platform for large language model capability discovery.
  • Collaborate with researchers and engineers to integrate and optimize models for production and research use cases.
  • Conduct architecture and design reviews to uphold best practices in system design and scalability.
  • Develop monitoring and observability solutions to ensure system health and performance.
  • Lead projects end-to-end, from requirements gathering through implementation, in a cross-functional environment.

Requirements

  • 4+ years of experience building large-scale, high-performance backend systems.
  • Strong programming skills in one or more of Python, Go, Rust, or C++.
  • Experience with large language model serving and routing fundamentals, including rate limiting, token streaming, load balancing, and budgets.
  • Experience with large language model capabilities and concepts such as reasoning, tool calling, and prompt templates.
  • Experience with containers and orchestration tools.
  • Familiarity with cloud infrastructure and infrastructure as code.
  • Proven ability to solve complex problems and work independently in fast-moving environments.

Nice to Haves

  • Experience with modern large language model serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.

Skills

Python, Go, Rust, C++, Llm Serving, Load Balancing, Kubernetes, Docker, AWS, GCP, Terraform, vLLM, Sglang, Tensorrt-Llm, Monitoring

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

OpenAI

OpenAI

London, United Kingdom

Applied AI Engineer, Digital Natives
No salary listedHybridML Engineering

Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.

Elliptic

Elliptic

London, United Kingdom

Agent Engineer
No salary listedHybrid5+ YOEML Engineering

Build full-stack AI agent fleets, APIs, workflows, and internal services that automate complex business processes. The role requires at least five years of engineering experience, hands-on LLM framework experience, production AWS expertise, Kubernetes, and strong API and database skills.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Build

Build

New York, NY
AI Engineer - Assistant Experience
$120k+/yrOn-siteML Engineering

Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.