Build research and production infrastructure for replayable enterprise environments, agent evaluation, and model improvement. The role requires 4+ years building production software or ML systems, including 2+ years in reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or related work.
215k – 330k/yr
On-site4+ YOEML Engineering
About the role
Responsibilities
Build an environment factory that converts recorded enterprise data and task definitions into runnable environments.
Design graders that turn ambiguous business objectives into verifiable rewards.
Develop methods for mining useful tasks, trajectories, and evaluation cases from historical workflows.
Create eval sets that are representative, reproducible, and resistant to overfitting.
Find combinations of models, tools, context, and policies that maximize performance while reducing inference cost.
Train and evaluate agents operating over long horizons, incomplete information, and large tool spaces.
Build replay and observability systems that make agent behavior explainable and measurable.
Scale from individual environments to thousands of concurrent training and evaluation runs.
Own the research and infrastructure needed to create a scalable model-improvement system.
Deploy research into real enterprise workflows and work directly with the CTO.
Requirements
4+ years of experience building production software or machine-learning systems.
At least 2 years working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems.
Understanding of how environment design, reward design, context, tooling, and policy behavior interact.
Ability to turn ambiguous business objectives into reliably evaluable tasks and signals.
Ability to diagnose whether model limitations originate in the model, context, tools, harness, or training.
Ability to move between research questions and production implementation.
Strong software engineering skills and experience building systems that process large, messy datasets at scale.
Care for reproducibility, observability, and understanding model behavior.
Credentials are not required; demonstrated work and problem-solving ability are valued.
Benefits
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
Health insurance
Unlimited PTO
Skills
Reinforcement Learningllm post-trainingMachine Learningevaluation infrastructureagent harnessesPythonlarge-scale data processingreward designenvironment designcontext engineeringObservabilityModel Training
Machine Learning Performance Engineer - Offboard Training & Inference
Applied IntuitionSunnyvale, CA
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
215k – 285k/yrOn-siteML Engineering
Machine Learning Engineer
LiftoffCalifornia
Machine Learning Engineer building statistical models, optimization systems, and experiments for mobile ad tech economics on the Revenue Engine team. Requires PhD in CS/ML/Economics and industry experience applying ML or economics at scale.
215k – 275k/yrRemoteML Engineering
AI Research Engineer
HexSan Francisco, CA +1
Builds and deploys production AI features like Notebook Agent for data science workflows, partnering with product teams on experiments, model fine-tuning, and infra. Requires senior AI/ML engineering experience with MLOps, Python/TS proficiency.
214k – 285k/yrHybrid5+ YOEML Engineering
Software Engineer, Enterprise AI
Scale AINew York, NY +1
Build and scale enterprise Generative AI platform, owning large product areas across backend, frontend, LLMs, and ML models. Requires 4+ years experience, proficiency in Python/JavaScript/SQL, Kubernetes, and cloud providers.
216k – 270k/yrOn-site4+ YOEML Engineering
ML Research Engineer, ML Systems
Scale AISan Francisco, CA +2
Builds and optimizes distributed frameworks for LLM training and inference on Scale's RLXF platform. Collaborates with ML teams to accelerate research, requiring expertise in PyTorch, CUDA, transformers, and large-scale distributed systems.