Research Scientist, Post-Training
Leads research on post-training data curation for foundation models, designing algorithms to generate/improve instruction and preference datasets, and unifying pre/post-training optimization. Requires 3+ years deep learning research, post-training experience with vision/language/multimodal models, and PyTorch proficiency.
About the job
What You'll Work On
- Post-training data curation: conduct research on algorithmically curating post-training data (e.g., generating/refining preference and instruction-following data, curating capability/domain-specific data, making post-training more effective/controllable/generalizable).
- Unifying pre-training and post-training data curation: pursue research on end-to-end data curation (curate pre-training data to improve post-trainability, jointly optimize pre/post-training data to maximize final model performance).
- Transform messy literature into practical improvements: source, vet, implement, and improve promising ideas from literature or your own creation.
- Conduct science driven by real-world needs: guided by customer needs and product improvements.
About You
Required:
- 3+ years of deep learning research experience
- Experience with post-training large vision, language, and multimodal models
- Post-training algorithm development, data curation, and/or synthetic data methods for:
- Preference-based tuning (e.g. DPO, RLVR, RRHF)
- Alternative supervision & self-supervision techniques (e.g. self-training, chain-of-thought distillation)
- SFT (e.g. instruction tuning, demonstration fine-tuning)
- Post-training tooling development and engineering experience
- Strong understanding of deep learning fundamentals
- Software engineering + deep learning framework (PyTorch or willingness to learn) skills for large-scale experiments and production prototypes
- Track record of success in deep learning research (papers, tools, artifacts)
Nice-to-haves:
- Experience with data management and distributed data processing (Spark, Snowflake, etc.)
- Experience building + shipping ML products
Compensation
- Base salary: $180,000 - $300,000
- Significant equity
- 100% covered health benefits (medical, vision, dental)
- 401(k) with 4% company match
- Unlimited PTO
- Annual $2,000 wellness stipend
- Annual $1,000 learning stipend
- Daily lunches/snacks
- Relocation assistance to Bay Area
Skills
PyTorch, Dpo, Rlvr, Rrhf, Sft, Instruction Tuning, Preference Tuning, Synthetic Data, Data Curation, Multimodal Models
Similar jobs
AI Research jobsDevelops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.
The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.
Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.