Member of Technical Staff — Model Optimization and Inference
Optimize inference for real-time multimodal AI avatars. Specialize in LLM and diffusion model serving, KV cache strategies, quantization, and low-latency frameworks like vLLM and TensorRT-LLM.
About the job
What You'll Do
- Own end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
- Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
- Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads
- Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
- Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
- Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
- Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
- Work closely with research and infrastructure to ensure new models ship with optimized serving from day one
What We're Looking For
- Deep expertise in LLM inference optimization — you've worked on KV caching, memory layout, attention kernels, or batching strategies in a production or research context
- Proficiency with inference serving frameworks — vLLM, SGLang, TensorRT-LLM, or similar — including the ability to go beyond default configurations and adapt them to non-standard use cases
- Experience optimizing diffusion model inference (latency reduction, step distillation, caching, or kernel-level work)
- Strong Python and PyTorch skills; comfort reading and writing CUDA or Triton kernels is a significant plus
- A systematic approach to profiling and optimization — you measure first, then optimize
- Familiarity with speculative decoding or other inference-time acceleration techniques
Bonus Points
- Hands-on experience with post-training quantization (GPTQ, AWQ, or similar) and understanding of quality/performance tradeoffs
- Familiarity with multimodal or streaming inference architectures
- Experience deploying real-time AI systems with hard latency SLAs
- Prior work at an AI lab, inference startup, or on a high-traffic model serving platform
- Contributions to open-source inference frameworks
Compensation
- $250,000 – $350,000 base salary, plus meaningful equity
- Health: HSA plan with ~$2,000 in company contributions
- PTO: 15 days + public holidays, and we close for a full week over the holidays
- Lunch, beverages, and snacks provided every workday
- Commuter benefits
- 401K: In the works
Skills
Python, PyTorch, vLLM, Sglang, Tensorrt-Llm, CUDA, Triton, Kv Cache Optimization, Quantization, Gptq, Awq, Diffusion Models, Speculative Decoding
Similar jobs
ML Engineering jobsSenior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Leads the development and production deployment of large-scale ASR and TTS systems for conversational intelligence products. The role requires 5+ years of industry experience, deep speech-model expertise, and strong software engineering and ML operations capabilities.
Leads Discord’s Safety ML team, setting technical direction and overseeing production machine learning systems for content understanding, account integrity, and platform abuse. Requires substantial machine learning and engineering management experience, hands-on technical depth, and experience delivering ML systems at scale.
Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.