Skip to content

Research Scientist, Post-Training

Leads research on post-training data curation for foundation models, designing algorithms to generate/improve instruction and preference datasets, and unifying pre/post-training optimization. Requires 3+ years deep learning research, post-training experience with vision/language/multimodal models, and PyTorch proficiency.

About the job

What You'll Work On

  • Post-training data curation: conduct research on algorithmically curating post-training data (e.g., generating/refining preference and instruction-following data, curating capability/domain-specific data, making post-training more effective/controllable/generalizable).
  • Unifying pre-training and post-training data curation: pursue research on end-to-end data curation (curate pre-training data to improve post-trainability, jointly optimize pre/post-training data to maximize final model performance).
  • Transform messy literature into practical improvements: source, vet, implement, and improve promising ideas from literature or your own creation.
  • Conduct science driven by real-world needs: guided by customer needs and product improvements.

About You

Required:

  • 3+ years of deep learning research experience
  • Experience with post-training large vision, language, and multimodal models
  • Post-training algorithm development, data curation, and/or synthetic data methods for:
    • Preference-based tuning (e.g. DPO, RLVR, RRHF)
    • Alternative supervision & self-supervision techniques (e.g. self-training, chain-of-thought distillation)
    • SFT (e.g. instruction tuning, demonstration fine-tuning)
  • Post-training tooling development and engineering experience
  • Strong understanding of deep learning fundamentals
  • Software engineering + deep learning framework (PyTorch or willingness to learn) skills for large-scale experiments and production prototypes
  • Track record of success in deep learning research (papers, tools, artifacts)

Nice-to-haves:

  • Experience with data management and distributed data processing (Spark, Snowflake, etc.)
  • Experience building + shipping ML products

Compensation

  • Base salary: $180,000 - $300,000
  • Significant equity
  • 100% covered health benefits (medical, vision, dental)
  • 401(k) with 4% company match
  • Unlimited PTO
  • Annual $2,000 wellness stipend
  • Annual $1,000 learning stipend
  • Daily lunches/snacks
  • Relocation assistance to Bay Area

Skills

PyTorch, Dpo, Rlvr, Rrhf, Sft, Instruction Tuning, Preference Tuning, Synthetic Data, Data Curation, Multimodal Models

Hightouch

Hightouch

United States

Software Engineer, Applied AI Research
$180k+/yrRemote5+ YOEAI Research

Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.

Greptile

Greptile

San Francisco, CA

Research Engineer
$180k+/yrOn-siteAI Research

The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.

Scale AI

Scale AI

San Francisco, CA
Machine Learning Research Scientist, Evaluations
$181k+/yrOn-siteAI Research

Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.

Scale AI

Scale AI

San Francisco, CA
Machine Learning Research Scientist / Research Engineer, Post-Training
$181k+/yrOn-siteAI Research

Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.

Earnin

Earnin

Mountain View, CA

Software Engineer (Gen AI)
$181k+/yrHybrid3+ YOEAI Research

Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.