Staff Software Engineer, RL Environments
Staff engineer responsible for designing and scaling the infrastructure, execution environments, verifiers, and tooling used to train and evaluate AI agents. Requires 8+ years of software engineering experience, strong Python and distributed-systems expertise, and familiarity with sandboxing, high-throughput systems, and LLM workflows.
About the job
Responsibilities
- Own the technical foundation for building, running, verifying, and delivering reinforcement learning environments at scale.
- Design platforms for sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and authoring tools.
- Instrument real applications and design task suites that expose specific capability gaps.
- Build robust graders and verifiable reward signals that withstand adversarial optimization and reward hacking.
- Set technical direction across multiple teams while remaining hands-on and writing complex production code.
Requirements
- 8+ years of software engineering experience, with strong fundamentals in distributed systems, system design, data structures, and algorithms.
- Strong Python skills and production software experience, plus familiarity with another part of the stack such as TypeScript/React, Go, or Rust.
- Deep experience with containerization and sandboxed execution, including Docker, virtual machines, gVisor, Firecracker, Kubernetes, or equivalent technologies.
- Experience building or operating high-throughput backend systems involving orchestration, job scheduling, queuing, and large-scale data pipelines.
- Hands-on experience with LLMs, including agent loops, tool calling, MCP, or evaluation harnesses.
- Ability to own ambiguous problems end to end and deliver shipped systems.
- Excellent written and verbal communication, with the ability to align engineers, researchers, and non-engineering partners.
Nice-to-Haves
- Experience building reinforcement learning environments, agentic benchmarks, or evaluation harnesses, including SWE-bench-style task suites, terminal or browser environments, or tool-use benchmarks.
- Familiarity with RLHF, RLAIF, RLVR, GRPO/PPO-family algorithms, rejection sampling, reward modeling, and their practical failure modes.
- Experience designing verifiable reward signals and defending against reward hacking.
- Experience with RL training or serving stacks such as verl, TRL, Ray, vLLM, or SGLang.
- Experience with high-scale sandbox or code-execution infrastructure, remote development environments, or CI systems.
- Experience with AWS, Google Cloud, Azure, Infrastructure as Code, and CI/CD.
- Strong observability practices, including tracing, structured logging, and metrics.
- Experience building internal tools and data-dense review or annotation interfaces.
- Experience translating research goals into production systems, working with sophisticated technical customers, and providing staff-level technical leadership.
Compensation
- Base salary range: $252,000–$315,000 USD.
- Eligible roles may include equity and benefits such as health, dental and vision coverage, retirement benefits, learning and development support, generous paid time off, and potentially a commuter stipend.
Skills
Python, Distributed Systems, Docker, Kubernetes, Gvisor, Firecracker, Virtual Machines, TypeScript, React, Go, Rust, Reinforcement Learning, LLMs, Ray, vLLM
Similar jobs
ML Engineering jobsLeads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Staff Machine Learning Engineer building and operating production ML systems for causal marketing measurement, optimization, and planning. The role requires deep statistical and machine learning expertise, production programming experience, cross-functional collaboration, and technical mentorship.
Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
Leads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.