# AI Resident

**Company:** [Ema](https://hotfix.jobs/companies/ema)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $48k – $48k/yr
**Skills:** Python, PyTorch, Post-Training, Sft, Dpo, Rl, Reward Modeling, Llm Judges, Agent Systems, Tool Use, Retrieval, Memory Systems, Eval Design, vLLM, Sglang
**Posted:** 2026-07-31

> AI Resident who owns a hard ML/agent problem end-to-end: from proposal and building to evaluation, shipping in production, and rigorous write-up. Requires strong ML fundamentals, Python/PyTorch engineering, and depth in at least one area like post-training, reward modeling, agents, or eval.

## Job Description

## Responsibilities
- Own one hard problem end to end: write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of results (including what didn't work).
- Work in the production codebase with a senior mentor and real production data.
- Projects scoped collaboratively in the loop of production traces → data → training/evaluation → better agents.
- Potential project areas:
  - Harness and inference-time work (context engineering, tool/skill design, orchestration, inference compute allocation).
  - Self-improvement loops.
  - Post-training for agents (SFT on trajectories, preference optimization, RL on agent tasks, reward design, process vs outcome supervision, distillation).
  - Environments and rewards (enterprise workflows into training/eval environments, verifiable rewards, defenses against reward hacking).
  - Data engines (mining production steps, failure mining, labeling with judges, synthetic augmentation).
  - Evaluation (behavior-level benchmarks, calibrated LLM judges, reliability for stochastic agents).
  - Efficiency (routing, ensembles, caching, small-model specialization; quality per dollar as metric).

## Requirements
- Demonstrated depth in ML or agent systems (strong undergrads, grad students, or self-taught welcome; no specific degree required).
- Solid ML fundamentals and strong engineering skills: Python, PyTorch, and ability to ship in a large production codebase.
- Real depth in at least one of: post-training (SFT/DPO/GRPO-family RL), reward modeling or LLM judges, agent and tool-use systems, retrieval and memory, eval design.
- Statistical literacy (sizing experiments, understanding sample reliability).
- Habit of honest measurement and rigorous experimentation.

## Nice-to-Haves
- Hands-on post-training with open models (TRL, veRL, OpenRLHF, or custom loops); experience debugging reward-hacked runs.
- Built or trained in interactive agent environments (SWE, web, or tool-use gyms).
- Large-scale trace analysis, data curation, or synthetic data work.
- Serving and efficiency experience (vLLM/SGLang, distillation, quantization).
- Multi-node GPU training or strong infra fluency.
- Publications, open-source contributions, or writing that demonstrates thinking.
- Security instincts (prompt injection, data governance, fencing for self-improving agents).

## Compensation
- $4,000 per month.
- Compensation determined by location, level, knowledge, skills, and experience. May include variable compensation, equity, and benefits.

## Similar jobs

- [Mercor Research Fellowship - APEX](https://hotfix.jobs/jobs/87b84beb-0193-4421-adbe-4b113a41781e) - Mercor - Remote - $40k – $80k/yr
- [Research Engineer – Benchmarking](https://hotfix.jobs/jobs/92ac70e0-232e-4223-b628-2416fc91ef52) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Research Scientist - Member of Technical Staff](https://hotfix.jobs/jobs/9807d0a2-a252-48f0-81fa-2175608c94b8) - AI Digest - Remote - $150k – $350k/yr
- [Engineer - Member of Technical Staff](https://hotfix.jobs/jobs/cba82dd2-84d4-48a7-8d75-58783156ab95) - AI Digest - Remote - $150k – $350k/yr
- [Research Scientist](https://hotfix.jobs/jobs/f8da4cc2-a216-4a79-b247-5d2bb83ac27e) - Counsel Health - New York, NY - $165k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/fa1bc613-b113-4dbe-b90e-2e2c23647c03
**Canonical:** https://hotfix.jobs/jobs/fa1bc613-b113-4dbe-b90e-2e2c23647c03