Research Engineer
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
About the job
Responsibilities
- Build and scale post-training systems for supervised fine-tuning, preference optimization, and reinforcement learning involving long-horizon tool use, transformation, and reconciliation over enterprise systems.
- Develop memory and context systems for long-running agents, including retention, retrieval, compaction, revision, and training models to use memory effectively.
- Build representation layers using ontologies and knowledge graphs derived from enterprise systems, plus pipelines to construct, validate, and maintain them.
- Design data-generation and curation pipelines for synthetic landscapes, transformation traces, tool-call trajectories, and curriculum infrastructure.
- Build sandboxed reinforcement-learning environments and execution-and-verification harnesses with automatically verifiable rewards.
- Develop offline evaluation infrastructure for long-horizon agent behavior, including trajectory scoring, task suites, reproducibility, and experiment tracking.
- Run experiments end to end: design, launch, debug, analyze, and distinguish real effects from noise or bugs.
- Optimize training and inference throughput across kernels, parallelism, memory, batching, and serving; address long-context workloads.
- Move training results into production through quantization, serving configuration, and rollback paths.
- Establish standards for reproducibility, experiment tracking, and research-result hygiene.
Requirements
- Significant experience training, fine-tuning, or post-training language models, with demonstrable results owned.
- Experience with reinforcement-learning approaches such as RLHF, RLAIF, RLVR, GRPO, or agentic RL.
- Experience with memory and context for long-running agents through architecture, retrieval, or training.
- Strong software-engineering fundamentals and reproducible experiment code.
- Fluency in Python and PyTorch or JAX.
- Ability to debug distributed training systems and interpret loss behavior.
- Experience with GPU infrastructure at scale and training-performance optimization.
- Ability to design, run, and interpret rigorous empirical experiments.
- Clear written communication.
Nice-to-haves
- Experience building RL environments, execution sandboxes, or verifiable-reward task suites.
- Long-context modeling experience, including context extension, efficient attention, positional methods, or evaluation.
- Experience with knowledge graphs, ontologies, or semantic layers over structured enterprise data.
- Production experience with agent memory systems.
- Experience with code models, repository-scale context, program synthesis, automated repair, or transpilation.
- Contributions to open-source ML systems such as vLLM, SGLang, PyTorch, Triton, DeepSpeed, Ray, Megatron, or TRL.
- Publications, technical reports, or open-source releases.
- An advanced degree in computer science, machine learning, mathematics, physics, or a related quantitative field.
Compensation
- Salary range: $200,000-$300,000.
Skills
Python, PyTorch, JAX, Reinforcement Learning, RLHF, Rlaif, Rlvr, Grpo, Gpu Infrastructure, Distributed Training, Long-Context Modeling, Knowledge Graphs, Ontologies, vLLM, Triton
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.