Skip to content

Applied Research - RL & Agents

Develops reinforcement learning, post-training, and agent systems that advance model reasoning and support real-world workflows. The role combines applied research with scalable training infrastructure, evaluations, and production deployment.

About the job

Responsibilities

  • Design and iterate on AI agents for workflow automation, reasoning-intensive tasks, and large-scale decision-making.
  • Develop reliable, scalable systems and frameworks for agent operation.
  • Translate ambiguous application objectives into technical requirements that guide product and research priorities.
  • Prototype and deploy agents, evaluations, and harnesses for real-world tasks.
  • Shape verifiers, environment hubs, training services, and research platform offerings.
  • Build reference implementations, examples, and recipes for extending the stack.
  • Design environments, evaluations, and verifiers with research teams, infrastructure-heavy customers, and open-source contributors.
  • Design and implement reinforcement learning and post-training methods, including RLHF, RLVR, and GRPO.
  • Build evaluations and harnesses for reasoning, robustness, and agentic behavior.
  • Prototype multi-agent and memory-augmented systems.
  • Experiment with post-training recipes to improve downstream performance.
  • Extend and integrate agent frameworks.
  • Architect and maintain distributed training and inference pipelines for scalability and cost efficiency.
  • Develop observability and monitoring using metrics and tracing for production reliability.

Requirements

  • Strong machine learning engineering background with experience in post-training, reinforcement learning, or large-scale model alignment.
  • Experience with agent frameworks and tooling such as DSPy, LangGraph, MCP, or Stagehand.
  • Familiarity with distributed training and inference frameworks such as vLLM, SGLang, Accelerate, Ray, or Torch.
  • Research contributions through publications, open-source contributions, or benchmarks in machine learning or reinforcement learning.
  • Strong technical writing skills for documentation, blogs, or papers.
  • Ability to collaborate with external partners and the open-source community.

Nice-to-haves

  • Web programming experience with React, TypeScript, or Next.js.
  • Experience running LLM evaluations or synthetic data generation.
  • Experience deploying containerized systems at scale with Docker, Kubernetes, or Terraform.

Compensation and Benefits

  • Cash compensation range of $150,000–$300,000 plus equity incentives.
  • Flexible work based in San Francisco or hybrid-remote.
  • Visa sponsorship and relocation support.
  • Professional development budget.
  • Team off-sites and conference attendance.

Skills

Reinforcement Learning, RLHF, Rlvr, Grpo, Machine Learning, Agent Frameworks, Dspy, LangGraph, Mcp, vLLM, Sglang, Ray, PyTorch, Kubernetes, Terraform

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

Fab2

Fab2

Austin, TX
Software Engineer, AI Platform
$140k+/yrOn-siteML Engineering

Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.