Skip to content
MercorMercorSan Francisco, CA

Research Engineer, Environments

Build and ship high-fidelity RL environments, verifiers, and evaluation pipelines from enterprise workflow data to train and assess frontier AI models. Requires prior experience shipping environments or agentic evals plus strong full-stack engineering skills with a bias for action and detail.

180k – 500k/yr
On-siteML Engineering

About the role

What You'll Do

  • Ship models for workflow extraction, classification, and grading.
  • Engineer autonomous task refinement processes which distill data taste into pipelines.
  • Deliver data to customers and deploy into real engagements.
  • Help define the future of agentic transformation for enterprises around the world.
  • Deeply learn about the intricacies of enterprises through building evaluations for all aspects of work.
  • Build end-to-end environments for labs & enterprises by platformizing sandbox app clones, load real data into the sandboxes, build prompts from real workflows, and write verifiers leveraging enterprise expertise & golden outputs.
  • Systematize the production of environments to scale throughput while maintaining high-quality worlds and verifiers.

What We're Looking For

  • Prior experience shipping environments – you’ve contributed to an OSS framework, built environments at previous companies, or worked on agentic evaluations.
  • Strong full-stack engineering skills – you’ll be responsible for everything from infrastructure to app code to analytics.
  • Bias to action – this team is focused on shipping evals, not just philosophizing about them.
  • Curiosity – being biased towards understanding and digging deep into model behavior and actually looking at the data.
  • Sweat the details that make a simulation indistinguishable from the real thing and have systems-level thinking skills that allow you to scale up quality.

Nice to Have

  • Experience with Temporal, Modal, or similar orchestration/compute services.
  • Experience with synthetic data generation for frontier models.
  • Past work auditing and scrutinizing industry-standard evaluations.

Benefits

  • Semi-annual performance bonus structure.
  • Generous equity grant vested over 4 years.
  • Up to $15k Relocation bonus.
  • $10K housing bonus (if you live within 0.5 miles of our office).
  • $1.5K monthly stipend for meals.
  • Free Equinox membership.
  • $200 monthly laundry reimbursement.
  • $200 monthly personal wellness reimbursement.
  • Health, Dental, Vision insurance.

Skills

full stack engineeringworkflow extractionclassification modelsgrading modelsautonomous task refinementrl environmentsagentic evaluationssandbox app clonesPrompt EngineeringverifiersTemporalmodalsynthetic data generation

Similar roles

ML Engineering jobs
Scale AI

Frontier Agents Engineer

Scale AISan Francisco, CA +2

Build and deploy production AI agents and frontier systems for enterprise customers, combining LLMs with retrieval, memory, multi-agent architectures, and traditional ML. Design evaluations, run experiments, ensure reliability/safety, and translate customer problems into scalable AI solutions. Requires 4+ years applied AI/ML experience and strong Python skills.

180k – 225k/yrOn-site4+ YOEML Engineering
Applied Intuition

Perception Engineer

Applied IntuitionSunnyvale, CA

Perception Engineer owning outcomes for autonomous mining vehicles. Responsible for sensor selection, model adaptation to new sites/domains, diagnosing failures, data strategies, and translating customer needs into technical KPIs and solutions. Requires strong systems understanding of perception/full stack and real-world deployment experience.

180k – 255k/yrOn-site5+ YOEML Engineering
Sprinter Health

Applied Scientist, AI

Sprinter HealthSan Francisco, CA

Build, evaluate, and productionize ML/AI models (including LLMs and NLP) that solve ambiguous healthcare, product, and operational problems at Sprinter Health. Requires strong experimentation, error analysis, stakeholder collaboration with clinicians, and focus on real-world impact, bias, and evaluation.

180k – 260k/yrHybridML Engineering
Zoox

Software Development Engineer in Test, Machine Learning

ZooxFoster City, CA

Build and productionize agentic LLM-powered triage systems and ML/DL pipelines to automate failure analysis for autonomous robots. Requires Master's/PhD in STEM + 4+ years production ML/NLP experience with PyTorch, RAG, Databricks, and AWS.

180k – 225k/yrHybrid4+ YOEML Engineering
Notion

Software Engineer, AI Platform

NotionSan Francisco, CA +1

Build and scale the shared AI platform foundations at Notion, enabling fast and safe shipping of AI products. Requires experience with LLM/ML platforms, strong ownership, and comfort across backend, infrastructure, and product code.

180k – 201k/yrHybrid5+ YOEML Engineering