Skip to content
EmaEma

AI Resident

AI Resident who owns a hard ML/agent problem end-to-end: from proposal and building to evaluation, shipping in production, and rigorous write-up. Requires strong ML fundamentals, Python/PyTorch engineering, and depth in at least one area like post-training, reward modeling, agents, or eval.

About the job

Responsibilities

  • Own one hard problem end to end: write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of results (including what didn't work).
  • Work in the production codebase with a senior mentor and real production data.
  • Projects scoped collaboratively in the loop of production traces → data → training/evaluation → better agents.
  • Potential project areas:
    • Harness and inference-time work (context engineering, tool/skill design, orchestration, inference compute allocation).
    • Self-improvement loops.
    • Post-training for agents (SFT on trajectories, preference optimization, RL on agent tasks, reward design, process vs outcome supervision, distillation).
    • Environments and rewards (enterprise workflows into training/eval environments, verifiable rewards, defenses against reward hacking).
    • Data engines (mining production steps, failure mining, labeling with judges, synthetic augmentation).
    • Evaluation (behavior-level benchmarks, calibrated LLM judges, reliability for stochastic agents).
    • Efficiency (routing, ensembles, caching, small-model specialization; quality per dollar as metric).

Requirements

  • Demonstrated depth in ML or agent systems (strong undergrads, grad students, or self-taught welcome; no specific degree required).
  • Solid ML fundamentals and strong engineering skills: Python, PyTorch, and ability to ship in a large production codebase.
  • Real depth in at least one of: post-training (SFT/DPO/GRPO-family RL), reward modeling or LLM judges, agent and tool-use systems, retrieval and memory, eval design.
  • Statistical literacy (sizing experiments, understanding sample reliability).
  • Habit of honest measurement and rigorous experimentation.

Nice-to-Haves

  • Hands-on post-training with open models (TRL, veRL, OpenRLHF, or custom loops); experience debugging reward-hacked runs.
  • Built or trained in interactive agent environments (SWE, web, or tool-use gyms).
  • Large-scale trace analysis, data curation, or synthetic data work.
  • Serving and efficiency experience (vLLM/SGLang, distillation, quantization).
  • Multi-node GPU training or strong infra fluency.
  • Publications, open-source contributions, or writing that demonstrates thinking.
  • Security instincts (prompt injection, data governance, fencing for self-improving agents).

Compensation

  • $4,000 per month.
  • Compensation determined by location, level, knowledge, skills, and experience. May include variable compensation, equity, and benefits.

Skills

Python, PyTorch, Post-Training, Sft, Dpo, Rl, Reward Modeling, Llm Judges, Agent Systems, Tool Use, Retrieval, Memory Systems, Eval Design, vLLM, Sglang

Mercor

Mercor

San Francisco, CA
Mercor Research Fellowship - APEX
$40k+/yrRemoteAI Research

Research fellows propose, build, validate, and publish benchmarks or evaluation methodologies for measuring frontier AI performance on economically valuable professional and scientific work. The fellowship requires a specific research pitch, relevant technical or adjacent-field background, and a commitment of at least 20 hours per week.

Mercor

Mercor

San Francisco, CA

Research Engineer – Benchmarking
$130k+/yrOn-siteAI Research

Research Engineer focused on designing benchmarks, evaluation systems, rubrics, and failure-analysis workflows for frontier language models. The role requires strong applied AI research and coding experience, with expertise in model evaluation, data quality, and backend systems.

AI Digest

AI Digest

Remote

Research Scientist - Member of Technical Staff
$150k+/yrRemoteAI Research

Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.

AI Digest

AI Digest

Remote

Engineer - Member of Technical Staff
$150k+/yrRemoteAI Research

Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.

Counsel Health

Counsel Health

New York, NY
Research Scientist
$165k+/yrHybrid5+ YOEAI Research

Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.