Researcher, Training
Researcher developing and scaling architectures and optimization techniques for flagship large language models. The role requires experience contributing to major LLM training runs, strong knowledge of inference and Transformer efficiency, and an empirical approach to experiments and debugging.
About the job
Responsibilities
- Design, prototype, and scale new architectures to improve model intelligence.
- Execute and analyze experiments autonomously and collaboratively.
- Study, debug, and optimize model performance and computational performance.
- Contribute to training and inference infrastructure.
Requirements
- Experience landing contributions to major LLM training runs.
- Ability to thoroughly evaluate and improve deep learning architectures independently.
- Motivation to safely deploy LLMs in the real world.
- Strong understanding of state-of-the-art Transformer modifications for efficiency.
- Deep understanding of LLM architectures and model inference.
- Hands-on empirical approach to research and development.
Compensation
- Annual salary: 170000–445000.
- Hybrid schedule with three days per week in the office.
- Relocation assistance for new employees.
- Option to work from home on Thursdays and Fridays.
- Office amenities include adjustable desks, phone booths, conference rooms, stocked kitchens, and spaces to unwind or collaborate.
Skills
LLMs, Deep Learning, Transformer Architectures, Model Inference, Model Training, Efficient Attention, Long-Context Modeling, Optimization, Scaling, Experiment Design, Model Evaluation, Python, Training Infrastructure, Inference Infrastructure
Similar jobs
AI Research jobsResearch Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.
The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.
Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.