Skip to content
Scale AIScale AI

Staff Machine Learning Research Engineer, Agent Post-training - Enterprise GenAI

Develops next-gen Agent RL training platform for enterprise GenAI, integrating cutting-edge research to train state-of-the-art models for complex use cases. Requires 5+ years LLM production experience, RLHF expertise, recent top publications, and advanced CS degree.

About the job

Responsibilities

  • Train state-of-the-art models (internal and community-developed) for enterprise customers.
  • Research and integrate cutting-edge algorithms into the training stack.
  • Design solutions for complex multi-agent systems to learn from process and outcome-based rewards.

Requirements

  • 5+ years of LLM training in production environments.
  • Experience with post-training methods (RLHF/RLVR) and algorithms (PPO/GRPO).
  • Publications in top conferences (NeurIPS, ICLR, ICML) within last 2 years.
  • PhD or Master's in Computer Science or related field.

Skills

Llm Training, RLHF, Rlvr, Ppo, Grpo, Multi-Agent Systems, Reinforcement Learning, PyTorch, Kubernetes, Machine Learning

Shield AI

Shield AI

Washington, DC
Senior Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site10+ YOEAI Research

Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.

Shield AI

Shield AI

San Mateo, CA

Senior Staff Software Engineer, Autonomy Capabilities
$281k+/yrOn-site10+ YOEAI Research

Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.

Upstart

Upstart

United States

Staff Machine Learning Model Risk Specialist
$140k+/yrRemote7+ YOEAI Research

Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.

Anthropic

Anthropic

San Francisco, CA

Staff+ Researcher, Cybersecurity Products
$405k+/yrHybrid7+ YOEAI Research

Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.

Order.co

Order.co

United States

Staff Applied AI Scientist
No salary listedRemote10+ YOEAI Research

Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.