# Inference Performance Engineer

**Company:** [Adaption Labs](https://hotfix.jobs/companies/adaption-labs)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** Python, C++, Rust, CUDA, Nccl, vLLM, Sglang, Tensorrt-Llm, Quantization, Kv Cache, Continuous Batching, Speculative Decoding, Gpu Kernels, Mixed Precision, Profiling
**Posted:** 2026-08-21

> Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.

## Job Description

## Responsibilities
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads using real production traffic.
- Tune routing between infrastructure and external providers based on cost, capacity, and performance.
- Work with serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
- Build profiling and measurement systems to identify where time, memory, and compute are being spent.

## Requirements
- 5+ years of experience in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

## Benefits
- Flexible work, including in-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Annual travel stipend to explore a country never visited.
- Weekly meal allowance for take-out or grocery delivery.
- Comprehensive medical benefits and generous paid time off.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Software Engineer, AI for Chip Design](https://hotfix.jobs/jobs/abfb017d-1ce2-4d24-901d-16084eb7b3bc) - OpenAI - San Francisco, CA - $266k – $468k/yr
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Machine Learning Engineer, Ranking & Retrieval](https://hotfix.jobs/jobs/adb5411d-8bfc-48b8-a8e7-15cbf5ea21b2) - ClickUp - Remote - $200k – $250k/yr
- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/7718c69a-6d49-445e-8c7c-840bd1b3f346
**Canonical:** https://hotfix.jobs/jobs/7718c69a-6d49-445e-8c7c-840bd1b3f346