# Software Engineer, ML Platform
**Company:** [Luma AI](https://hotfix.jobs/companies/lumalabs-ai)
**Location:** Palo Alto, CA
**Salary:** $188K-$395K
**Experience:** 5+ years
**Skills:** Python, Kubernetes, Docker, Linux, Redis, AWS, PyTorch, CUDA, S3, Rdma
**Posted:** 2025-08-20
> Builds foundational ML platform infrastructure including model serving pipelines, GPU scheduling systems, and CI/CD for large-scale multimodal AI models. Requires 5+ years in distributed systems with expertise in Python, Kubernetes, and AWS.
## Job Description
## What You'll Do
- Architect end-to-end model serving pipelines and integrate new model architectures from our research team into our core, high-throughput inference engine.
- Build robust and sophisticated scheduling systems to manage jobs based on cluster availability and user priority, ensuring we optimally leverage thousands of expensive GPU resources.
- Design and implement dynamic, traffic-based systems for hotswapping models on our GPU workers to maximize fleet efficiency and meet product SLOs.
- Own the end-to-end CI/CD pipelines, including creating a resilient artifact store to manage all model checkpoints across multiple versions and providers.
- Develop and maintain user-friendly APIs and interaction patterns that empower our product and research teams to ship groundbreaking features at high velocity.
- Manage and optimize our complex inference workloads at scale, operating across multiple clusters and hardware providers.

## Who You Are
We are looking for a world-class builder who has a proven history of creating and managing large-scale, high-performance systems. You are a non-negotiable fit if you have:
- 5+ years of professional engineering experience with deep, hands-on proficiency in **Python** and complex distributed systems architecture.
- Extensive, practical experience building and managing systems at scale, specifically with queues, scheduling, traffic-control, and fleet management.
- Deep expertise in our core infrastructure stack: **Linux**, **Docker**, and **Kubernetes**.
- Strong experience with **Redis**, **S3-compatible storage**, and public cloud platforms (**AWS**).

## What Sets You Apart (Bonus Points)
- Experience with high-performance, large-scale ML systems (managing >100 GPUs).
- Deep familiarity with **PyTorch** and **CUDA**.
- Experience with modern networking stacks, including **RDMA** (RoCE, Infiniband, NVLink).
- Familiarity with **FFmpeg** and multimedia processing pipelines.

## Compensation
The base pay range for this role is **$187,500 – $395,000** per year.
**Apply:** https://hotfix.jobs/jobs/software-engineer-ml-platform-at-lumalabs-ai-a9f2a8c2-33a8-4d0d-91c4-60f6a5135e7d
**Canonical:** https://hotfix.jobs/jobs/software-engineer-ml-platform-at-lumalabs-ai-a9f2a8c2-33a8-4d0d-91c4-60f6a5135e7d