Skip to content
LabelboxLabelbox

Forward Deployed Engineer, RL Environments

Builds and maintains sandboxed, reproducible RL environments for AI agent training, including terminal, browser, and tool-augmented setups. Requires 2+ years Python/systems engineering, containerization, and RL concepts understanding.

About the job

What You’ll Do

  • Design, build, and maintain sandboxed RL environments for agentic AI training—including terminal emulators, browser automation harnesses, computer-use simulators, and tool-augmented workspaces (e.g., environments built on frameworks like TerminalBench, OSWorld, and Tau-bench)
  • Develop reproducible, containerized execution environments (Docker, VMs, lightweight sandboxes) that support deterministic task rollouts and reward signal collection
  • Integrate with and extend open-source agentic tooling and custom CLI/API harnesses to enable multi-step agent interaction
  • Build instrumentation and observability layers—structured logging, trajectory capture, state snapshotting—so training runs and human annotation sessions produce clean, auditable data
  • Collaborate with data operations to design task curricula and evaluation protocols that stress-test model capabilities across environment types
  • Own environment deployment and reliability: CI/CD pipelines, automated testing of environment configurations, and monitoring for drift or breakage across versions
  • Rapidly prototype new environment types as client and internal requirements evolve, moving from spec to working system in days, not weeks

What We’re Looking For

Required

  • 2+ years of professional software engineering experience, with strong fundamentals in Python and at least one systems-level language (Go, Rust, C++)
  • Demonstrated experience with containerization and sandboxing (Docker, Podman, Firecracker, or similar) in production or near-production contexts
  • Familiarity with RL concepts: MDPs, reward shaping, episode structure, observation/action spaces. You don’t need to have trained models, but you need to understand what an environment must provide to an RL training loop
  • Experience building or maintaining developer tooling, CLI tools, or infrastructure automation
  • Comfort working with browser automation frameworks or terminal interaction tooling
  • Strong debugging instincts—you can trace failures across process boundaries, container layers, and network calls
  • Ability to read and implement from academic papers and open-source benchmark repositories without extensive hand-holding

Preferred

  • Direct experience building or contributing to RL environments (Gymnasium/Gym, PettingZoo, or custom environment implementations)
  • Experience with agentic AI evaluation frameworks (SWE-bench, WebArena, OSWorld, TerminalBench, or similar)
  • Familiarity with GCP or AWS infrastructure (Compute Engine, ECS/EKS, Cloud Build)
  • Prior work at an AI data company, ML platform company, or AI research lab
  • Contributions to open-source projects in the RL, agents, or dev-tools space

Compensation

Annual base salary range $140,000—$200,000 USD

Skills

Python, Docker, Go, Rust, C++, Reinforcement Learning, Gymnasium, Pettingzoo, GCP, AWS

Ambral

Ambral

New York, NY
Member of Technical Staff
$140k+/yrOn-siteML Engineering

Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.

Lyft

Lyft

New York, NY
Machine Learning Engineer
$141k+/yrHybrid2+ YOEML Engineering

Design, deploy, and improve real-time machine learning systems for Lyft’s ride fulfillment and marketplace products. The role requires 2+ years of ML experience, production programming skills, and expertise with deep learning and recommendation systems.

Pinterest

Pinterest

San Francisco, CA

Machine Learning Engineer II, Responsible AI
$139k+/yrHybrid2+ YOEML Engineering

Develop and deploy responsible AI and machine learning fairness solutions across Pinterest’s large-scale, user-facing products, including generative AI, search, and recommendations. The role requires production ML experience, expertise in fairness interventions and modern architectures, and a master’s or PhD in computer science or a related field.

Benchling

Benchling

San Francisco, CA

Software Engineer, Model Evaluation and Improvement
$136k+/yrOn-site2+ YOEML Engineering

Build datasets, evaluations, and scalable data systems that improve frontier AI models on challenging biological and scientific tasks. The role partners with scientists and AI labs and requires at least two years of experience applying biology and AI, plus hands-on LLM experience.

DataVisor

DataVisor

Mountain View, CA

Software Engineer, Artificial Intelligence
$130k+/yrOn-site2+ YOEML Engineering

Builds high-scale data pipelines, distributed systems, and AI agent workflows using LLMs for fraud intelligence platform. Requires 2+ years software engineering, Python proficiency, big data tools, AWS/K8s, and ML foundations.