Member of Technical Staff
Member of Technical Staff conducting hands-on LLM inference research at Modal. Own end-to-end bets on techniques like speculative decoding, quantization, KV-cache management, and disaggregation to improve cost per token and tail latency on production workloads. Requires strong LLM serving stack expertise and a track record shipping research or systems.
About the job
What you'll do
- Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.
- Train custom speculators against real production traffic and feed what you learn back into target models -- acceptance length is the metric that decides the win.
- Work directly with customers alongside our Forward Deployed Engineers to deploy and tune models, and bring what you learn back into the research.
- Carry and expand collaborations with outside research labs, for example: our work with ZLab on DFlash, a speculator design built on KV injection and blockwise parallel drafting; our work with SGLang on specdec and multimodal inference performance; our work on Flash Attention 4 kernels.
- Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine.
- Help shape the research agenda.
Requirements
- A research-leaning or systems background in LLM inference, with work you can point to.
- Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling.
- A record of shipping research or systems that other people build on, whether in a lab or in industry.
- The drive to independently take a research bet from idea to result, working in the open with the rest of the team.
- Ability to work in-person, in our NYC or San Francisco office.
Skills
Llm Inference, Speculative Decoding, Disaggregated Prefill/Decode, Quantization, Fp8, Int4, Kv-Cache, Memory Management, Autoscaling, Flash Attention, Sglang, Kernels, Schedulers
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.