Inference Performance Engineer
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.
About the job
Responsibilities
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads using real production traffic.
- Tune routing between infrastructure and external providers based on cost, capacity, and performance.
- Work with serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
- Build profiling and measurement systems to identify where time, memory, and compute are being spent.
Requirements
- 5+ years of experience in ML systems, inference infrastructure, or performance engineering, with measurable improvements in cost or latency.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
Benefits
- Flexible work, including in-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Annual travel stipend to explore a country never visited.
- Weekly meal allowance for take-out or grocery delivery.
- Comprehensive medical benefits and generous paid time off.
Skills
Python, C++, Rust, CUDA, Nccl, vLLM, Sglang, Tensorrt-Llm, Quantization, Kv Cache, Continuous Batching, Speculative Decoding, Gpu Kernels, Mixed Precision, Profiling
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.