Skip to content
ProtegeProtege

Forward Deployed Machine Learning Engineer

Builds benchmarks, model evaluations, and backend infrastructure for an AI data platform’s Benchmarks and Evaluations vertical. The role requires 4+ years of engineering experience, hands-on model evaluation, and prior ownership of backend and infrastructure systems.

About the job

Responsibilities

Build the Evaluation Foundation

  • Partner with the GM and early customers to define strong evaluations across domains.
  • Design and build benchmarks with researchers.
  • Establish standards for processing different modalities.

Own Infrastructure

  • Build backend infrastructure for the vertical, including data pipelines, execution environments, storage, and orchestration.
  • Create sandboxed environments for agentic evaluations involving tools, code execution, and multi-step tasks.

Move from Iteration to Product

  • Identify repeatable evaluation patterns, infrastructure gaps, and product opportunities from live engagements.
  • Partner with the research team on domain-specific data and research questions.
  • Own the engineering component of customer engagements end to end.

Requirements

  • 4+ years of engineering experience.
  • Hands-on machine learning experience evaluating models.
  • Previous ownership of backend and infrastructure systems.
  • High tolerance for ambiguity and a bias toward action.
  • Ability to work with urgency in a fast-paced environment.
  • Strong written communication.

Nice-to-Haves

  • Experience building benchmarks, evaluations, or human-data pipelines for large language models.
  • Experience at a frontier lab, evaluation-focused team, or research organization.
  • Founding or early-engineer experience at a fast-moving startup.
  • Familiarity with agentic systems, reinforcement-learning environments, code-execution sandboxes, and TEE/TREs.

Compensation and Benefits

  • Compensation and benefits information was not provided in the posting.

Skills

Machine Learning, Model Evaluation, Backend Infrastructure, Data Pipelines, Benchmarking, LLMs, Agentic Systems, Reinforcement Learning, Code Execution, Orchestration

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.