Member of Technical Staff, Model Efficiency
Engineers on this team optimize LLM inference for lower latency and higher throughput by identifying bottlenecks, developing optimizations across the execution stack, and collaborating with modeling teams. Requires 5+ years high-performance coding in C++/Python and LLM inference experience.
About the job
Responsibilities
- Work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations.
- Collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference.
- Build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.
Requirements
- 5+ years of experience writing high-performance, production-quality code.
- Strong programming skills in C++ or Python (Rust/Go also welcome).
- Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.).
- Ability to diagnose and resolve performance bottlenecks across the model execution stack.
- A strong bias for action — ship fast, measure impact, and iterate.
Nice-to-Haves
- Experience with GPU programming, CUDA, or low-level systems optimization.
- Language modeling with transformers (MoE, speculative decoding, KV-cache optimizations).
- Scaling performance-critical distributed systems (e.g., computation, search, storage).
Skills
C++, Python, Rust, Go, LLMs, vLLM, Sglang, CUDA, GPU, Transformers, Moe, Kv-Cache
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.