Staff Research Engineer, Model Efficiency
Develops and deploys techniques to enhance LLM inference efficiency, focusing on architecture optimization, decoding algorithms, and GPU acceleration. Requires PhD in ML, expertise in LLM optimization, strong software skills, and top-tier publications.
About the job
Responsibilities
- Develop, prototype, and deploy techniques that improve LLM inference efficiency in production.
- Explore breakthroughs across model execution stack, including model architecture and MoE routing optimization.
- Implement decoding and inference-time algorithm improvements.
- Perform software/hardware co-design for GPU acceleration.
- Optimize performance without compromising model quality.
Requirements
- PhD in Machine Learning or related field.
- Deep understanding of LLM architecture and optimization under resource constraints.
- Significant experience with model efficiency techniques.
- Strong software engineering skills.
- Experience in fast-paced, high-ambiguity startup environment.
- Publications at top-tier conferences (ICLR, ACL, NeurIPS).
- Passion for mentoring others.
Nice-to-Haves
- Appetite for working in startups (encouraged even if not perfect fit).
Skills
LLMs, Model Architecture, Moe, Inference Optimization, Gpu Acceleration, PyTorch, JAX, Machine Learning, Software Engineering, Decoding Algorithms
Similar jobs
ML Engineering jobsBuild and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.
Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.