Research Engineer
Research Engineer designing post-training infrastructure and running controlled experiments to measure how datasets affect foundation-model behavior. Requires at least 2 years of ML or research engineering experience, strong Python, and hands-on experience with PyTorch, JAX, Ray, Slurm, and LLM post-training.
About the job
Responsibilities
- Design and build post-training infrastructure for supervised fine-tuning (SFT), reinforcement learning (RL), and evaluation workflows.
- Build reliable pipelines for data preparation, dataset versioning, sampling, training, checkpoint management, and evaluation.
- Develop experiment orchestration and tracking systems that make runs reproducible, comparable, and easy to debug.
- Create reusable abstractions for launching experiments across datasets, models, and training recipes.
- Integrate systems with partner-lab training stacks, model APIs, compute environments, and evaluation infrastructure.
- Design and run controlled training experiments to measure how data sources, structures, and selection strategies affect model capability, generalization, and alignment.
- Formulate hypotheses, build pipelines, run models, analyze results, identify confounders, and determine next steps.
- Debug distributed systems, data pipelines, training infrastructure, and model behavior.
Requirements
- At least 2 years of professional experience in machine learning engineering, research engineering, ML infrastructure, or a closely related field; 2–4+ years preferred.
- Strong Python and software engineering skills.
- Hands-on experience with PyTorch, JAX, Ray, and Slurm.
- Experience building production-quality ML training, evaluation, or data infrastructure.
- Hands-on experience with large language model fine-tuning, post-training, and evaluation.
- Ability to build reliable, reproducible systems for launching and comparing ML experiments.
- Understanding of experimental design and ability to extract actionable conclusions from noisy results.
- Ability to move quickly between infrastructure engineering and hands-on experimentation.
- Bias toward building, testing, and shipping.
Benefits
- Medical, vision, and dental insurance.
- 401(k) with employer match.
- Daily meals and UberEats stipend.
- Wellness stipend, including monthly Equinox membership coverage.
- Commuting costs covered.
Skills
Python, PyTorch, JAX, Ray, Slurm, Machine Learning, ML Infrastructure, Llm Fine-Tuning, Reinforcement Learning, Supervised Fine-Tuning, Experiment Tracking, Distributed Systems, Data Pipelines, Model Evaluation, Experimental Design
Similar jobs
ML Engineering jobsBuild and productionize scalable machine learning models and systems for underwriting and portfolio management. The role requires a bachelor's degree and at least two years of experience shipping ML systems, plus expertise in model development, deployment, data pipelines, and deep learning.
Research Engineer focused on building and deploying real-time audio and speech models for conversational voice agents. The role requires experience with speech or multimodal machine learning, production inference, Python, and PyTorch, with emphasis on taking research from prototype to measurable production impact.
Research Engineer focused on making conversational AI agents safe, reliable, and controllable in production. The role develops evaluations, safeguards, post-training methods, and monitoring systems, requiring 2+ years of AI/ML or safety experience and strong Python and production engineering skills.
Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.
Build and deploy production machine-learning models for fraud detection, identity verification, and financial risk products. The role suits new PhD graduates or early-career researchers with strong quantitative foundations, Python experience, and interest in owning the full ML lifecycle.