# Member of Technical Staff

**Company:** [Ambral](https://hotfix.jobs/companies/ambral)
**Location:** New York, NY, San Francisco, CA
**Role:** ML Engineering
**Salary:** $165k – $325k/yr
**Experience:** 1+ years
**Skills:** Python, Reinforcement Learning, Llm Post-Training, Evaluation Infrastructure, Agent Harnesses, Machine Learning, Production Software, Large-Scale Data Processing, Observability, Reproducibility
**Posted:** 2026-09-01

> Build replayable enterprise environments, evaluation systems, graders, and post-training workflows for AI agents. The role spans machine-learning research and production engineering and requires 1–7 years of software or ML systems experience.

## Job Description

## Responsibilities
- Build an environment factory that converts recorded enterprise data and task definitions into runnable environments.
- Design graders that turn ambiguous business objectives into verifiable rewards.
- Develop methods for mining useful tasks, trajectories, and evaluation cases from historical workflows.
- Create representative, reproducible eval sets resistant to overfitting.
- Find combinations of models, tools, context, and policies that maximize performance while reducing inference cost.
- Train and evaluate agents operating over long horizons, incomplete information, and large tool spaces.
- Build replay and observability systems that make agent behavior explainable and measurable.
- Scale from individual environments to thousands of concurrent training and evaluation runs.
- Contribute to research direction and production systems, working directly with the CTO and deploying into enterprise workflows.

## Requirements
- 1–7 years of experience building production software or machine-learning systems.
- Experience building strong software and systems that process large, messy datasets at scale.
- Ability to turn ambiguous business objectives into reliably evaluable tasks and signals.
- Understanding of how environment design, reward design, context, tooling, and policy behavior interact.
- Ability to diagnose whether model limitations originate in the model, context, tools, harness, or training.
- Ability to move between research questions and production implementation.
- Commitment to reproducibility, observability, and understanding model behavior.
- Demonstrated work and thoughtful problem-solving are valued over credentials or conventional career paths.

## Nice to Have
- Experience with reinforcement-learning environments.
- Experience with LLM post-training.
- Experience with evaluation infrastructure.
- Experience with agent harnesses or closely related systems.

## Compensation and Benefits
- Salary: $165,000–$325,000 per year.
- Significant equity and ownership.
- Equinox membership.
- Free meals, coffee, and snacks.
- Health insurance.
- Unlimited paid time off.

## Similar jobs

- [Software Engineer, ML Inference Platform](https://hotfix.jobs/jobs/74081e27-c2cd-4485-b56d-22edd2ae9e1d) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Software Engineer, ML Infrastructure Platform](https://hotfix.jobs/jobs/5ccc0259-2d85-4c10-8ed4-76d8ac33c4ea) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/533760a9-ed27-4f88-8f58-42c6ec72014b) - Fireworks AI - San Mateo, CA - $160k – $180k/yr
- [Applied Scientist II](https://hotfix.jobs/jobs/c6c7879a-d2d0-4c0c-a46c-43a9930a8457) - Garner Health - New York, NY - $158k – $190k/yr
- [Applied ML Scientist, New Grad](https://hotfix.jobs/jobs/adfb42c1-5eb0-4d5a-bd6a-582c98bca4a9) - SentiLink - Remote - $180k – $220k/yr

**Apply:** https://hotfix.jobs/jobs/30cfebf1-9d5b-4d60-999b-8205184b07e0
**Canonical:** https://hotfix.jobs/jobs/30cfebf1-9d5b-4d60-999b-8205184b07e0