Research Engineer - Mid-Training
Trains frontier LLMs on semiconductor design/verification data (RTL, netlists, PDKs) for automated chip development. Develops synthetic data generation, model distillation, evals, and scales training across thousands of GPUs.
About the job
Responsibilities
- Train frontier models to become highly knowledgeable semiconductor design and verification experts for reinforcement learning and automated chip development.
- Develop methods for generating and curating synthetic design data, performing model distillation, and enabling continual learning at scale.
- Work with hardware engineers, RL researchers, and verification specialists to create evals that guide design data quality and model improvement.
- Collaborate with compute engineers to scale efficient training across thousands of GPUs and RL environments.
- Build high-performance tools to investigate how data and simulation shape model-driven design intelligence.
Requirements
- Experience training LLMs or foundation models on semiconductor design and verification corpora (e.g., RTL, netlists, PDKs, simulation logs).
- Modeling design scaling laws and optimizing compute budgets for chip-design-specific workloads.
- Generating large-scale synthetic design data (e.g., RTL variants, testbenches, verification traces).
- Building evals that correlate with downstream design metrics (e.g., timing closure, power, area, verification coverage).
Skills
LLMs, Foundation Models, Rtl, Netlists, Pdks, Simulation Logs, Scaling Laws, Synthetic Data, Testbenches, Verification Traces, Evals, Timing Closure, PyTorch, Gpu Training, Reinforcement Learning
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.