# Applied Research - RL & Agents

**Company:** [Prime Intellect](https://hotfix.jobs/companies/prime-intellect)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $150k – $300k/yr
**Skills:** Reinforcement Learning, RLHF, Rlvr, Grpo, Machine Learning, Agent Frameworks, Dspy, LangGraph, Mcp, vLLM, Sglang, Ray, PyTorch, Kubernetes, Terraform
**Posted:** 2026-07-08

> Develops reinforcement learning, post-training, and agent systems that advance model reasoning and support real-world workflows. The role combines applied research with scalable training infrastructure, evaluations, and production deployment.

## Job Description

## Responsibilities
- Design and iterate on AI agents for workflow automation, reasoning-intensive tasks, and large-scale decision-making.
- Develop reliable, scalable systems and frameworks for agent operation.
- Translate ambiguous application objectives into technical requirements that guide product and research priorities.
- Prototype and deploy agents, evaluations, and harnesses for real-world tasks.
- Shape verifiers, environment hubs, training services, and research platform offerings.
- Build reference implementations, examples, and recipes for extending the stack.
- Design environments, evaluations, and verifiers with research teams, infrastructure-heavy customers, and open-source contributors.
- Design and implement reinforcement learning and post-training methods, including RLHF, RLVR, and GRPO.
- Build evaluations and harnesses for reasoning, robustness, and agentic behavior.
- Prototype multi-agent and memory-augmented systems.
- Experiment with post-training recipes to improve downstream performance.
- Extend and integrate agent frameworks.
- Architect and maintain distributed training and inference pipelines for scalability and cost efficiency.
- Develop observability and monitoring using metrics and tracing for production reliability.

## Requirements
- Strong machine learning engineering background with experience in post-training, reinforcement learning, or large-scale model alignment.
- Experience with agent frameworks and tooling such as DSPy, LangGraph, MCP, or Stagehand.
- Familiarity with distributed training and inference frameworks such as vLLM, SGLang, Accelerate, Ray, or Torch.
- Research contributions through publications, open-source contributions, or benchmarks in machine learning or reinforcement learning.
- Strong technical writing skills for documentation, blogs, or papers.
- Ability to collaborate with external partners and the open-source community.

## Nice-to-haves
- Web programming experience with React, TypeScript, or Next.js.
- Experience running LLM evaluations or synthetic data generation.
- Experience deploying containerized systems at scale with Docker, Kubernetes, or Terraform.

## Compensation and Benefits
- Cash compensation range of $150,000–$300,000 plus equity incentives.
- Flexible work based in San Francisco or hybrid-remote.
- Visa sponsorship and relocation support.
- Professional development budget.
- Team off-sites and conference attendance.

## Similar jobs

- [AI Engineer, Enablement](https://hotfix.jobs/jobs/6ae315a1-a605-45dc-b1ee-88fa8e9dee24) - LangChain - New York, NY - $150k – $195k/yr
- [Member of Technical Staff — Frontier Data](https://hotfix.jobs/jobs/20efe17f-aa7c-401b-8add-7086d4571fba) - Roboflow - Remote - $150k – $300k/yr
- [Algorithm Engineer](https://hotfix.jobs/jobs/3bef67e8-d4e7-4866-a9b8-6c7863fe2e96) - Beacon Biosignals - Remote - $150k – $170k/yr
- [Software Engineer - Prediction and Planning ML](https://hotfix.jobs/jobs/21b9c778-e1ae-4695-b26d-fec68ea8a8cc) - Applied Intuition - Sunnyvale, CA - $151k – $258k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/3328296f-ae54-45c9-94af-5489d15345e2
**Canonical:** https://hotfix.jobs/jobs/3328296f-ae54-45c9-94af-5489d15345e2