Skip to content

Applied Research - Evals & Data

Customer-facing applied research role focused on building AI agents, evaluation systems, and post-training workflows for frontier models. The role combines reinforcement learning, distributed infrastructure, applied data, and close collaboration with customers and research teams.

About the job

Responsibilities

  • Design and iterate on AI agents for workflow automation, reasoning-intensive tasks, and decision-making.
  • Use applied deployment data to refine policies, improve reasoning, and enhance reliability and safety.
  • Develop distributed systems, evaluation pipelines, coordination frameworks, and data workflows for feedback, model traces, and reward signals.
  • Work directly with customers to understand workflows, data sources, bottlenecks, and technical requirements.
  • Prototype and deploy agents, evaluation harnesses, data pipelines, and verifiers for real-world use cases.
  • Translate customer insights and evaluation results into product roadmaps and research direction.
  • Design and implement reinforcement learning and post-training methods, including RLHF, RLVR, and GRPO.
  • Integrate data collection and analytics into post-training to identify regressions, emergent skills, and alignment opportunities.
  • Prototype multi-agent and memory-augmented systems.
  • Extend agent frameworks and architect distributed training and inference pipelines.
  • Develop observability and monitoring using metrics and tracing tools.

Requirements

  • Strong machine learning engineering background with experience in post-training, reinforcement learning, or large-scale model alignment.
  • Experience with applied data workflows and evaluation frameworks for large models or agents, such as SWE-Bench, HELM, EvalFlow, or internal evaluation pipelines.
  • Deep expertise in distributed training and inference frameworks, such as vLLM, SGLang, Ray, or Accelerate.
  • Experience deploying containerized systems at scale with Docker, Kubernetes, and Terraform.
  • Track record of research contributions through publications, open-source contributions, or benchmarks.
  • Strong interest in reasoning, measurement, and practical agentic AI systems.

Compensation and Benefits

  • Cash compensation: $150,000–$300,000 plus equity incentives.
  • Flexible work arrangement; remote or San Francisco options.
  • Visa sponsorship and relocation support.
  • Professional development budget.
  • Team off-sites and conference attendance.

Skills

Reinforcement Learning, Post-Training, RLHF, Rlvr, Grpo, Machine Learning, AI Agents, Evaluation Frameworks, Distributed Training, vLLM, Sglang, Ray, Docker, Kubernetes, Terraform

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

Fab2

Fab2

Austin, TX
Software Engineer, AI Platform
$140k+/yrOn-siteML Engineering

Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.