Skip to content
AnyscaleAnyscale

Distributed LLM Inference Engineer

Build and optimize distributed LLM inference systems at scale using Ray, integrating with engines like vLLM to deliver high-throughput, low-latency batch and online inference solutions.

About the job

Responsibilities

  • Iterate quickly with product teams to ship end-to-end solutions for batch and online inference at high scale for Ray users and Anyscale customers
  • Work across the stack integrating Ray Data and LLM engines to provide optimizations for low-cost, large-scale ML inference
  • Integrate with open-source software like vLLM, work with the community to adopt techniques in Anyscale solutions, and contribute improvements to open source
  • Follow state-of-the-art developments in open source and research, implementing and extending best practices

Requirements

  • Familiarity with running ML inference at large scale with high throughput and low latency
  • Familiarity with deep learning and deep learning frameworks (e.g., PyTorch)
  • Solid understanding of distributed systems and ML inference challenges

Nice-to-Haves

  • ML Systems knowledge
  • Experience using Ray
  • Work with community on LLM engines like vLLM, TensorRT-LLM
  • Contributions to deep learning frameworks (PyTorch, TensorFlow)
  • Contributions to deep learning compilers (Triton, TVM, MLIR)
  • Prior experience working on GPUs / CUDA

Compensation & Benefits

  • Market-based compensation approach
  • Equity (stock options)
  • Healthcare plans with 99% premiums covered for employees and dependents
  • 401k Retirement Plan
  • Education & Wellbeing Stipend
  • Paid Parental Leave
  • Fertility Benefits
  • Paid Time Off
  • Commute reimbursement
  • 100% of in-office meals covered

Skills

PyTorch, Ray, vLLM, Tensorrt-Llm, Distributed Systems, Ml Inference, CUDA, Triton, Tvm, Mlir, TensorFlow

Clay

Clay

New York, NY

Software Engineer, Applied AI
$170k+/yrHybridML Engineering

Build and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.

Clay

Clay

San Francisco, CA

Machine Learning Engineer
$170k+/yrHybrid5+ YOEML Engineering

Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.

Mirage

Mirage

New York, NY

Software Engineer, Agents
$175k+/yrOn-site5+ YOEML Engineering

Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.

Mirage

Mirage

New York, NY

Research Engineer, Agentic Systems
$175k+/yrOn-siteML Engineering

Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.

Taste Labs

Taste Labs

San Francisco, CA

AI Engineer, RL
$175k+/yrOn-siteML Engineering

Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.