Research Engineer - Environments, Data and Post-Training
Develops post-training pipelines, RLVR experiments, synthetic data generation, and large-scale LLM evaluation systems to enhance frontier language model performance in tool use, agentic behavior, and reasoning. Requires strong ML experience, coding skills, and research background.
About the job
Responsibilities
- Work on post-training and RLVR pipelines to understand how datasets, rewards, and training strategies impact model performance.
- Design and run reward-shaping experiments and algorithmic improvements (e.g., GRPO, DAPO) to improve LLM tool-use, agentic behavior, and real-world reasoning.
- Quantify data usability, quality, and performance uplift on key benchmarks.
- Build and maintain data generation and augmentation pipelines that scale with training needs.
- Create and refine rubrics, evaluators, and scoring frameworks that guide training and evaluation decisions.
- Build and operate LLM evaluation systems, benchmarks, and metrics at scale.
- Collaborate closely with AI researchers, applied AI teams, and experts producing training data.
- Operate in a fast-paced, experimental research environment with rapid iteration cycles and high ownership.
Requirements
- Strong applied research background, with a focus on post-training and/or model evaluation.
- Strong coding proficiency and hands-on experience working with machine learning models.
- Strong understanding of data structures, algorithms, backend systems, and core engineering fundamentals.
- Familiarity with APIs, SQL/NoSQL databases, and cloud platforms.
- Ability to reason deeply about model behavior, experimental results, and data quality.
- Excitement to work in person in San Francisco, five days a week (with optional remote Saturdays), and thrive in a high-intensity, high-ownership environment.
Nice To Have
- Real-world post-training team experience in industry (highest priority).
- Publications at top-tier conferences (NeurIPS, ICML, ACL).
- Experience training models or evaluating model performance.
- Experience in synthetic data generation, LLM evaluations, or RL-style workflows.
- Work samples, artifacts, or code repositories demonstrating relevant skills.
Benefits
- Generous equity grant vested over 4 years
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- Health insurance
Skills
PyTorch, Machine Learning, LLMs, RLHF, Rlvr, Synthetic Data, Post-Training, Evaluation Frameworks, SQL, NoSQL, APIs, Cloud Platforms, Data Structures, Algorithms, Backend Systems
Similar jobs
ML Engineering jobsBuild and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.
Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.
Build modular AI operations and evaluation systems that power complex real estate workflows. The role focuses on improving output quality, defining correctness with domain experts, and reducing human review while maintaining high standards.