Skip to content
FirecrawlFirecrawl

Research Engineer – Evals

Build and own evaluation systems, metrics, pipelines, and datasets to rigorously measure the quality of Firecrawl's LLM-ready web data outputs at massive scale, driving model and product improvements.

About the job

What You'll Do

  • Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases
  • Build the pipelines and harnesses that measure quality rigorously and at scale
  • Generate and curate the datasets that make evaluation trustworthy
  • Own the feedback loop from output quality back to model and product decisions
  • Turn "did that work?" into an answer the whole team can act on

What We're Looking For

  • Engineering depth to build real evaluation systems, not just run existing ones
  • Care deeply about what "good" means and how to measure it rigorously
  • Comfortable owning ambiguous problems where the metric itself has to be invented
  • Move fast and close the loop - ship, measure, and iterate

Requirements

  • 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work
  • Must already be authorized to work in the US or our eligible remote-hire regions

Nice-to-Haves

  • Experience designing metrics, building evaluation pipelines, generating datasets for LLM/web data quality
  • Background in measuring complex AI system outputs at scale

Compensation & Benefits

Salary: $210,000–$275,000/year (U.S.-based in San Francisco, CA; adjusted for other locations based on cost of living)
Equity: Up to 0.05%
Benefits (US-based): Full medical, dental, vision (100% for employees); life & disability insurance; 401(k); generous PTO (15+ days); 12 weeks parental leave; wellness stipend; learning stipend; pet insurance.
SF-specific: HQ perks, e-bike loaner.
Other: Sabbatical after 4 years; team offsites.

Skills

Evaluation Systems, Metrics Design, Data Pipelines, Dataset Curation, Ml Engineering, Research Engineering, Llm Evaluation, Quality Measurement

Firecrawl

Firecrawl

San Francisco, CA

Machine Learning Engineer
$210k+/yrHybrid3+ YOEML Engineering

Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Machine Learning Performance Engineer - Offboard Training & Inference
$215k+/yrOn-siteML Engineering

Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

Atomicmachines

Atomicmachines

Emeryville, CA

MLOps Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.

Cinder

Cinder

New York, NY

AI/ML Engineer
$220k+/yrHybrid5+ YOEML Engineering

Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.