Skip to content
Luma AILuma AI

Software Engineer, Inference

Develops and optimizes inference engines for multimodal AI models, integrating new architectures, building scheduling systems, and managing large-scale GPU deployments. Requires strong Python, model serving frameworks like PyTorch/vLLM, and Kubernetes expertise.

About the job

Role & Responsibilities

  • Ship new model architectures by integrating them into our inference engine
  • Collaborate closely across research, engineering and infrastructure to streamline and optimize model efficiency and deployments
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows
  • Automate, test and maintain our inference services to ensure maximum uptime and reliability
  • Optimize deployment workflows to scale across thousands of machines
  • Manage and optimize our inference workloads across different clusters & hardware providers
  • Build sophisticated scheduling systems to optimally leverage our expensive GPU resources while meeting internal SLOs
  • Build and maintain CI/CD pipelines for processing/optimizing model checkpoints, platform components, and SDKs for internal teams to integrate into our products/internal tooling

Background

Must have:

  • Strong Python and system architecture skills
  • Experience with model deployment using PyTorch, Huggingface, vLLM, SGLang, tensorRT-LLM, or similar
  • Experience with queues, scheduling, traffic-control, fleet management at scale
  • Experience with Linux, Docker, and Kubernetes

Bonus points:

  • Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink)
  • Experience with high performance large scale ML systems (>100 GPUs)
  • Experience with FFmpeg and multimedia processing

Tech stack

Must have:

  • Python
  • Redis
  • S3-compatible Storage
  • Model serving (one of: PyTorch, vLLM, SGLang, Huggingface)
  • Understanding of large-scale orchestration, deployment, scheduling (via Kubernetes or similar)

Nice to have:

  • CUDA
  • FFmpeg

Compensation

The base pay range for this role is $187,500 – $395,000 per year.

Skills

Python, PyTorch, Huggingface, vLLM, Sglang, Tensorrt-Llm, Kubernetes, Docker, Linux, Redis

Earnin

Earnin

Mountain View, CA

AI Builder
$189k+/yrHybrid3+ YOEML Engineering

Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.

Nuro

Nuro

Mountain View, CA
Software Engineer, Applied AI Infrastructure
$194k+/yrOn-site5+ YOEML Engineering

Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

Mirage

Mirage

New York, NY

Software Engineer, Agents
$175k+/yrOn-site5+ YOEML Engineering

Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.

Mirage

Mirage

New York, NY

Research Engineer, Agentic Systems
$175k+/yrOn-siteML Engineering

Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.