# Founding Research Engineer

**Company:** [Ambral](https://hotfix.jobs/companies/ambral)
**Location:** New York, NY, San Francisco, CA
**Role:** ML Engineering
**Salary:** $215k – $330k/yr
**Experience:** 4+ years
**Skills:** Reinforcement Learning, llm post-training, Machine Learning, evaluation infrastructure, agent harnesses, Python, large-scale data processing, reward design, environment design, context engineering, Observability, Model Training
**Posted:** 2026-08-04

> Build research and production infrastructure for replayable enterprise environments, agent evaluation, and model improvement. The role requires 4+ years building production software or ML systems, including 2+ years in reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or related work.

## Job Description

## Responsibilities

- Build an environment factory that converts recorded enterprise data and task definitions into runnable environments.
- Design graders that turn ambiguous business objectives into verifiable rewards.
- Develop methods for mining useful tasks, trajectories, and evaluation cases from historical workflows.
- Create eval sets that are representative, reproducible, and resistant to overfitting.
- Find combinations of models, tools, context, and policies that maximize performance while reducing inference cost.
- Train and evaluate agents operating over long horizons, incomplete information, and large tool spaces.
- Build replay and observability systems that make agent behavior explainable and measurable.
- Scale from individual environments to thousands of concurrent training and evaluation runs.
- Own the research and infrastructure needed to create a scalable model-improvement system.
- Deploy research into real enterprise workflows and work directly with the CTO.

## Requirements

- 4+ years of experience building production software or machine-learning systems.
- At least 2 years working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems.
- Understanding of how environment design, reward design, context, tooling, and policy behavior interact.
- Ability to turn ambiguous business objectives into reliably evaluable tasks and signals.
- Ability to diagnose whether model limitations originate in the model, context, tools, harness, or training.
- Ability to move between research questions and production implementation.
- Strong software engineering skills and experience building systems that process large, messy datasets at scale.
- Care for reproducibility, observability, and understanding model behavior.
- Credentials are not required; demonstrated work and problem-solving ability are valued.

## Benefits

- Significant equity and ownership
- Equinox membership
- Free meals, coffee, and snacks
- Health insurance
- Unlimited PTO

## Similar roles

- [Machine Learning Performance Engineer - Offboard Training & Inference](https://hotfix.jobs/jobs/b2f397d3-cd2d-4eeb-bc02-1dfe0c90a5a7) - Applied Intuition - Sunnyvale, CA - $215k – $285k/yr
- [Machine Learning Engineer](https://hotfix.jobs/jobs/97713038-58cb-46c5-a9c7-0e1554e1ba4f) - Liftoff - Remote - $215k – $275k/yr
- [AI Research Engineer](https://hotfix.jobs/jobs/eb72cd41-2cc0-499f-b1aa-f5721d12433c) - Hex - San Francisco, CA - $214k – $285k/yr
- [Software Engineer, Enterprise AI](https://hotfix.jobs/jobs/98385fe7-6f4c-4386-ab6e-d4033a6d2440) - Scale AI - New York, NY - $216k – $270k/yr
- [ML Research Engineer, ML Systems](https://hotfix.jobs/jobs/c814c5a3-7e79-4c03-8f4b-89eb8352da30) - Scale AI - San Francisco, CA - $218k – $273k/yr

**Apply:** https://hotfix.jobs/jobs/92328d5a-8c83-43a9-afb7-e29b5480350f
**Canonical:** https://hotfix.jobs/jobs/92328d5a-8c83-43a9-afb7-e29b5480350f