Research Engineer – Evals
Build and own evaluation systems, metrics, pipelines, and datasets to rigorously measure the quality of Firecrawl's LLM-ready web data outputs at massive scale, driving model and product improvements.
About the job
What You'll Do
- Design the metrics that define what "good output" actually means across millions of sites, formats, and edge cases
- Build the pipelines and harnesses that measure quality rigorously and at scale
- Generate and curate the datasets that make evaluation trustworthy
- Own the feedback loop from output quality back to model and product decisions
- Turn "did that work?" into an answer the whole team can act on
What We're Looking For
- Engineering depth to build real evaluation systems, not just run existing ones
- Care deeply about what "good" means and how to measure it rigorously
- Comfortable owning ambiguous problems where the metric itself has to be invented
- Move fast and close the loop - ship, measure, and iterate
Requirements
- 4+ years in ML, research engineering, or data-heavy backend, with real evaluation work
- Must already be authorized to work in the US or our eligible remote-hire regions
Nice-to-Haves
- Experience designing metrics, building evaluation pipelines, generating datasets for LLM/web data quality
- Background in measuring complex AI system outputs at scale
Compensation & Benefits
Salary: $210,000–$275,000/year (U.S.-based in San Francisco, CA; adjusted for other locations based on cost of living)
Equity: Up to 0.05%
Benefits (US-based): Full medical, dental, vision (100% for employees); life & disability insurance; 401(k); generous PTO (15+ days); 12 weeks parental leave; wellness stipend; learning stipend; pet insurance.
SF-specific: HQ perks, e-bike loaner.
Other: Sabbatical after 4 years; team offsites.
Skills
Evaluation Systems, Metrics Design, Data Pipelines, Dataset Curation, Ml Engineering, Research Engineering, Llm Evaluation, Quality Measurement
Similar jobs
ML Engineering jobsBuild and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.