Skip to content

Latest ML Engineering jobs at Cohere

Search
Location
8 jobs

Job results

Cohere

Senior Research Engineer - Safety Tooling and Data

CohereNew York, NY

Senior Research Engineer building data synthesis, analysis, and management tooling for AI safety model training and evaluation. Requires strong software engineering, statistics, and ML framework expertise.

230k – 380k/yrRemote5+ YOEML Engineering
Cohere

Staff Software Engineer, Inference Infrastructure

CohereSan Francisco, CA +1

Build and operate high-performance inference infrastructure for large language models, deploying optimized NLP models to production with low latency and high throughput using Kubernetes and cloud platforms. Requires 5+ years experience in scalable distributed systems and GPU workloads.

Salary not listedHybrid5+ YOEML Engineering
Cohere

Member of Technical Staff, Senior/Staff MLE

CohereSan Francisco, CA +1

Design and deliver custom LLM solutions for enterprise customers, train frontier models using Cohere's stack, and influence foundation model capabilities. Requires strong ML fundamentals, Python fluency, and customer-facing technical leadership.

Salary not listedHybridML Engineering
Cohere

Member of Technical Staff, MLE

CohereSan Francisco, CA

Design and deliver custom LLM solutions for enterprise customers, train frontier models using Cohere's stack, and contribute to foundation model improvements. Requires strong ML fundamentals, Python fluency, and experience with LLMs and large-scale data.

Salary not listedHybridML Engineering
Cohere

Applied AI Engineer – Agentic Workflows

CohereSan Francisco, CA +5

Designs, builds, and deploys production-grade AI agents using LLMs for enterprise workflows. Collaborates with customers to solve business problems, ensures reliability, and mentors teams on agentic architectures.

Salary not listedHybridML Engineering
Cohere

Audio Inference Engineer, Model Efficiency

CohereNew York, NY +1

Develops high-performance audio inference systems, optimizing latency, throughput, and quality for real-time streaming workloads. Requires expertise in C++, Python, and deep learning models for audio/speech, with collaboration across training and serving teams.

Salary not listedRemoteML Engineering
Cohere

Staff Research Engineer, Model Efficiency

CohereNew York, NY +1

Develops and deploys techniques to enhance LLM inference efficiency, focusing on architecture optimization, decoding algorithms, and GPU acceleration. Requires PhD in ML, expertise in LLM optimization, strong software skills, and top-tier publications.

Salary not listedRemoteML Engineering
Cohere

Member of Technical Staff, Model Efficiency

CohereNew York, NY +1

Engineers on this team optimize LLM inference for lower latency and higher throughput by identifying bottlenecks, developing optimizations across the execution stack, and collaborating with modeling teams. Requires 5+ years high-performance coding in C++/Python and LLM inference experience.

Salary not listedRemote5+ YOEML Engineering