Research Engineer, Pretraining
Research Engineer on a pretraining team, developing and scaling large language models through research, experimentation, infrastructure optimization, and model engineering. Requires advanced ML or computer science education, strong software engineering skills, and expertise in Python and deep learning frameworks.
About the job
Responsibilities
- Conduct research and implement solutions in model architecture, algorithms, data processing, and optimizer development.
- Independently lead small research projects and collaborate on larger initiatives.
- Design, run, and analyze scientific experiments to advance understanding of large language models.
- Optimize and scale training infrastructure for efficiency and reliability.
- Develop and improve developer tooling.
- Contribute across the stack, from low-level optimizations to high-level model design.
Requirements
- Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field.
- Strong software engineering skills and experience building complex systems.
- Expertise in Python and experience with deep learning frameworks, preferably PyTorch.
- Familiarity with large-scale machine learning, particularly language models.
- Ability to balance research goals with practical engineering constraints.
- Strong problem-solving, communication, and collaboration skills.
Nice-to-haves
- Experience with high-performance, large-scale ML systems.
- Familiarity with GPUs, Kubernetes, and operating-system internals.
- Experience with language modeling and Transformer architectures.
- Knowledge of reinforcement learning techniques.
- Background in large-scale ETL processes.
Sample Projects
- Optimizing throughput for novel attention mechanisms.
- Comparing compute efficiency of Transformer variants.
- Preparing large-scale datasets for efficient model consumption.
- Scaling distributed training jobs to thousands of GPUs.
- Designing fault-tolerance strategies for training infrastructure.
- Creating visualizations of model internals, such as attention patterns.
Compensation and Benefits
- Annual salary: £260,000–£630,000 GBP.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Collaborative office environment.
- Visa sponsorship may be available.
Work Arrangement
- Staff are expected to work from an office at least 25% of the time; some roles may require more office time.
Skills
Python, PyTorch, Machine Learning, LLMs, Transformer Architectures, GPU, Kubernetes, Reinforcement Learning, ETL, Distributed Training, Model Architecture, Data Processing, Optimizer Development, Linux Os Internals, Developer Tooling
Similar jobs
ML Engineering jobsResearch Engineer responsible for operating and improving large-scale production pretraining systems, from performance optimization and hardware debugging to experiments, observability, and launch incident response. Requires deep ML systems expertise and experience with LLM training, JAX, TPU, PyTorch, or distributed systems.
Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.
Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.