# Research Engineer - Inference

**Company:** [ElevenLabs](https://hotfix.jobs/companies/elevenlabs)
**Location:** Remote
**Role:** ML Engineering
**Skills:** Gpu Programming, Inference Optimization, CUDA, Triton, TensorRT, vLLM, Sglang, Quantization, Knowledge Distillation, Kv Cache Optimization, Custom Kernels, Model Serving
**Posted:** 2026-08-28

> Deploy and optimize frontier AI models for fast, reliable, real-time production serving at scale. The role requires production ML serving experience, GPU programming and inference optimization expertise, and the ability to diagnose bottlenecks across the serving stack.

## Job Description

## Responsibilities
- Deploy state-of-the-art AI models to production, owning the path from research checkpoints to serving infrastructure.
- Optimize inference performance across latency, throughput, and cost using quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
- Build and tune high-performance serving systems for real-time, streaming workloads.
- Create tooling and infrastructure that enables researchers to ship models to production quickly and safely.
- Profile, diagnose, and eliminate bottlenecks across model architecture, kernels, serving infrastructure, and orchestration.

## Requirements
- Experience deploying and serving machine-learning models in production, ideally for latency-sensitive or real-time applications.
- Strong engineering skills in GPU programming and inference optimization.
- Ability to independently profile and measure serving-stack performance.
- Demonstrated ability to solve difficult engineering problems through projects, designs, or open-source contributions.

## Compensation and Benefits
- Annual professional development stipend.
- Annual stipend for social travel with colleagues.
- Monthly coworking stipend for employees outside major hubs.
- Annual company offsite.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Software Engineer, AI for Chip Design](https://hotfix.jobs/jobs/abfb017d-1ce2-4d24-901d-16084eb7b3bc) - OpenAI - San Francisco, CA - $266k – $468k/yr
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Machine Learning Engineer, Ranking & Retrieval](https://hotfix.jobs/jobs/adb5411d-8bfc-48b8-a8e7-15cbf5ea21b2) - ClickUp - Remote - $200k – $250k/yr
- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/ebb7c992-845d-4ab6-bf47-ce54b2d96084
**Canonical:** https://hotfix.jobs/jobs/ebb7c992-845d-4ab6-bf47-ce54b2d96084