# Research Engineer, Pretraining Scaling

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** London, United Kingdom
**Role:** ML Engineering
**Salary:** £260k – £630k/yr
**Skills:** JAX, Tpu, PyTorch, LLMs, Distributed Systems, Machine Learning Frameworks, Observability, Logging, Monitoring Dashboards, Evaluation Infrastructure, Networking, Openlm, Llm Foundry, Mesh Transformer Jax
**Posted:** 2026-08-28

> Research Engineer responsible for operating and improving large-scale production pretraining systems, from performance optimization and hardware debugging to experiments, observability, and launch incident response. Requires deep ML systems expertise and experience with LLM training, JAX, TPU, PyTorch, or distributed systems.

## Job Description

## Responsibilities
- Own critical aspects of the production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
- Debug and resolve issues across hardware, networking, training dynamics, and evaluation infrastructure.
- Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.
- Respond to on-call incidents during model launches and coordinate solutions across teams.
- Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.
- Add capabilities to the training codebase, such as long-context support and novel architectures.
- Collaborate with teams across San Francisco and London, including Tokens, Architectures, and Systems.
- Document systems, debugging approaches, and lessons learned.

## Requirements
- Hands-on experience training large language models or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
- Interest and experience spanning research and engineering work.
- Comfort with on-call responsibilities, extended launch hours, production incidents, and changing priorities.
- Ability to debug complex, ambiguous problems across multiple layers of the stack.
- Clear communication and effective collaboration across time zones and during high-stress incidents.
- Passion for research engineering and responsible AI development.
- Bachelor's degree or equivalent combination of education, training, and experience in a relevant field.

## Nice-to-haves
- Experience training LLMs or working extensively with JAX, TPU, PyTorch, or other ML frameworks at scale.
- Contributions to open-source LLM frameworks such as OpenLM, LLM Foundry, or Mesh Transformer JAX.
- Published research on model training, scaling laws, or ML systems.
- Experience with production ML systems, observability tools, or evaluation infrastructure.
- Background in systems engineering, quantitative research, or roles requiring technical depth and operational excellence.

## Compensation
- Annual salary: £260,000–£630,000 GBP.

## Similar jobs

- [AI Engineer (Assistant)](https://hotfix.jobs/jobs/4c7ec3aa-a4bf-4dc9-987f-c541087f251b) - Build - New York, NY - $125k – $225k/yr
- [AI Engineer - Assistant Experience](https://hotfix.jobs/jobs/00d11204-96cf-4c85-a12f-9058f0581c15) - Build - New York, NY - $120k – $240k/yr
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Applied AI Engineer, Digital Natives](https://hotfix.jobs/jobs/74eb3e31-bab1-479e-90e1-2e172658343a) - OpenAI - London, United Kingdom
- [Agent Engineer](https://hotfix.jobs/jobs/627cc378-6ac6-467e-bbf0-2c07738ae0a5) - Elliptic - London, United Kingdom

**Apply:** https://hotfix.jobs/jobs/155ea3ee-e487-4cb8-8343-cab2d40cf642
**Canonical:** https://hotfix.jobs/jobs/155ea3ee-e487-4cb8-8343-cab2d40cf642