Skip to content
CantinaCantina

Machine Learning Enginer, Core Evaluations

Designs evaluation pipelines, metrics, and user studies for speech generation and recognition models (ASR/TTS). Trains evaluation models, builds dashboards, and collaborates with ML, data, and product teams to improve performance on large-scale systems.

About the job

Responsibilities

  • Design model evaluation pipelines for models in development and production.
  • Design user studies for subjective model evaluations.
  • Convert requirements into measurable metrics.
  • Design and develop automated evaluation dashboards to monitor model performances and compare results.
  • Train new models to capture new and different evaluation metrics.
  • Communicate with model team to design better models based on evaluation results.
  • Communicate with data team to decide data needed to improve model performance.
  • Communicate with product manager to ensure product requirements are correctly measured.
  • Help grow the evaluation team as founding member and lead it in the future.

Requirements

  • Strong experience designing metrics that capture model performance.
  • Strong experience designing user studies on Mechanical Turk or similar platforms.
  • Strong experience with model training and fine-tuning for evaluation.
  • Strong statistical knowledge to compare evaluation results and make decisions.
  • Very strong engineering and programming skills.
  • Experience training ASR and TTS models.
  • Experience at ML teams working on large-scale problems (>3B models with >1m hours of data).

Skills

Asr, Tts, Model Evaluation, User Studies, Mechanical Turk, Model Training, Fine-Tuning, Statistics, Python, Dashboards

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.