Build and evaluate post-trained machine learning models for subjective domains such as design and writing. The role owns experiments, reward models, judges, infrastructure, data and evaluation pipelines, and research publications while collaborating with AI labs.
200k – 350k/yr
On-siteML Engineering
About the role
Responsibilities
Own full-stack machine learning experiments, training runs, infrastructure, data pipelines, and evaluation pipelines.
Develop internal research, evaluation design, judges, and reward models for subjective domains such as design, writing, and visual style.
Train reward models, classifiers, judges, and verifiers.
Develop frontier evaluations and benchmarks for subjective domains.
Run post-training experiments on open-source models to test new data formats and post-training techniques.
Set up infrastructure to run experiments.
Collaborate with AI labs and creative experts on pilots and experiments around taste.
Own the end-to-end pipeline.
Publish blogs and whitepapers.
Requirements
Experience with post-training and evaluations or judges.
Experience with large language models; purely classical machine learning experience is not sufficient.
Research-oriented mindset with strong engineering execution.
Creativity, scrappiness, and comfort working in ambiguity.
Technical Product Engineer advising Digital Native Businesses on integrating Claude API into products. Guides customers from discovery to deployment with expertise in LLMs, prompt engineering, agents, and evaluations; requires 4+ years experience and strong Python/TypeScript skills.
200k – 320k/yrHybrid4+ YOEML Engineering
Machine Learning Engineer - Voice Conversion
CantinaUnited States
Build and productionize large-scale generative speech models for voice conversion and related capabilities. The role combines research, data, evaluation, distributed training, performance optimization, and responsible deployment of speech systems.
200k – 220k/yrRemoteML Engineering
Research Engineer, Large-Scale Training
Together AISan Francisco, CA
Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.
200k – 290k/yrOn-siteML Engineering
Software Engineer, Backend
SiftstackSan Francisco, CA +1
Build and operate dependable agentic AI systems that analyze large-scale hardware telemetry for aerospace and defense teams. Design tools, execution environments, distributed job systems on Kubernetes, and evaluation frameworks while owning the product end-to-end and speaking directly with customers.
200k – 250k/yrHybrid3+ YOEML Engineering
Research Engineer
ConsoleSan Francisco, CA
Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.