Skip to content
CohereCohere

Staff Research Engineer, Model Efficiency

Develops and deploys techniques to enhance LLM inference efficiency, focusing on architecture optimization, decoding algorithms, and GPU acceleration. Requires PhD in ML, expertise in LLM optimization, strong software skills, and top-tier publications.

About the job

Responsibilities

  • Develop, prototype, and deploy techniques that improve LLM inference efficiency in production.
  • Explore breakthroughs across model execution stack, including model architecture and MoE routing optimization.
  • Implement decoding and inference-time algorithm improvements.
  • Perform software/hardware co-design for GPU acceleration.
  • Optimize performance without compromising model quality.

Requirements

  • PhD in Machine Learning or related field.
  • Deep understanding of LLM architecture and optimization under resource constraints.
  • Significant experience with model efficiency techniques.
  • Strong software engineering skills.
  • Experience in fast-paced, high-ambiguity startup environment.
  • Publications at top-tier conferences (ICLR, ACL, NeurIPS).
  • Passion for mentoring others.

Nice-to-Haves

  • Appetite for working in startups (encouraged even if not perfect fit).

Skills

LLMs, Model Architecture, Moe, Inference Optimization, Gpu Acceleration, PyTorch, JAX, Machine Learning, Software Engineering, Decoding Algorithms

Anthropic

Anthropic

San Francisco, CA

Staff+ Software Engineer, ML Inference Path
$320k+/yrHybrid7+ YOEML Engineering

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

Shield AI

Shield AI

San Diego, CA

Staff Engineer, Perception Software
$200k+/yrOn-site7+ YOEML Engineering

Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Talkiatry

Talkiatry

United States

Staff AI Enablement Engineer
$190k+/yrRemote8+ YOEML Engineering

Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.

Nuro

Nuro

Mountain View, CA

Senior/Staff Engineer, Machine Learning - Online Mapping
$194k+/yrOn-site7+ YOEML Engineering

Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.