Distributed LLM Inference Engineer
Build and optimize distributed LLM inference systems at scale using Ray, integrating with engines like vLLM to deliver high-throughput, low-latency batch and online inference solutions.
About the job
Responsibilities
- Iterate quickly with product teams to ship end-to-end solutions for batch and online inference at high scale for Ray users and Anyscale customers
- Work across the stack integrating Ray Data and LLM engines to provide optimizations for low-cost, large-scale ML inference
- Integrate with open-source software like vLLM, work with the community to adopt techniques in Anyscale solutions, and contribute improvements to open source
- Follow state-of-the-art developments in open source and research, implementing and extending best practices
Requirements
- Familiarity with running ML inference at large scale with high throughput and low latency
- Familiarity with deep learning and deep learning frameworks (e.g., PyTorch)
- Solid understanding of distributed systems and ML inference challenges
Nice-to-Haves
- ML Systems knowledge
- Experience using Ray
- Work with community on LLM engines like vLLM, TensorRT-LLM
- Contributions to deep learning frameworks (PyTorch, TensorFlow)
- Contributions to deep learning compilers (Triton, TVM, MLIR)
- Prior experience working on GPUs / CUDA
Compensation & Benefits
- Market-based compensation approach
- Equity (stock options)
- Healthcare plans with 99% premiums covered for employees and dependents
- 401k Retirement Plan
- Education & Wellbeing Stipend
- Paid Parental Leave
- Fertility Benefits
- Paid Time Off
- Commute reimbursement
- 100% of in-office meals covered
Skills
PyTorch, Ray, vLLM, Tensorrt-Llm, Distributed Systems, Ml Inference, CUDA, Triton, Tvm, Mlir, TensorFlow
Similar jobs
ML Engineering jobsBuild and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.
Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.
Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.