Skip to content
AnthropicAnthropic

Research Engineer, Pretraining Scaling

Research Engineer responsible for operating and improving large-scale production pretraining systems, from performance optimization and hardware debugging to experiments, observability, and launch incident response. Requires deep ML systems expertise and experience with LLM training, JAX, TPU, PyTorch, or distributed systems.

About the job

Responsibilities

  • Own critical aspects of the production pretraining pipeline, including model operations, performance optimization, observability, and reliability.
  • Debug and resolve issues across hardware, networking, training dynamics, and evaluation infrastructure.
  • Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.
  • Respond to on-call incidents during model launches and coordinate solutions across teams.
  • Build and maintain production logging, monitoring dashboards, and evaluation infrastructure.
  • Add capabilities to the training codebase, such as long-context support and novel architectures.
  • Collaborate with teams across San Francisco and London, including Tokens, Architectures, and Systems.
  • Document systems, debugging approaches, and lessons learned.

Requirements

  • Hands-on experience training large language models or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
  • Interest and experience spanning research and engineering work.
  • Comfort with on-call responsibilities, extended launch hours, production incidents, and changing priorities.
  • Ability to debug complex, ambiguous problems across multiple layers of the stack.
  • Clear communication and effective collaboration across time zones and during high-stress incidents.
  • Passion for research engineering and responsible AI development.
  • Bachelor's degree or equivalent combination of education, training, and experience in a relevant field.

Nice-to-haves

  • Experience training LLMs or working extensively with JAX, TPU, PyTorch, or other ML frameworks at scale.
  • Contributions to open-source LLM frameworks such as OpenLM, LLM Foundry, or Mesh Transformer JAX.
  • Published research on model training, scaling laws, or ML systems.
  • Experience with production ML systems, observability tools, or evaluation infrastructure.
  • Background in systems engineering, quantitative research, or roles requiring technical depth and operational excellence.

Compensation

  • Annual salary: £260,000–£630,000 GBP.

Skills

JAX, Tpu, PyTorch, LLMs, Distributed Systems, Machine Learning Frameworks, Observability, Logging, Monitoring Dashboards, Evaluation Infrastructure, Networking, Openlm, Llm Foundry, Mesh Transformer Jax

Build

Build

New York, NY
AI Engineer (Assistant)
$125k+/yrOn-siteML Engineering

Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.

Build

Build

New York, NY
AI Engineer - Assistant Experience
$120k+/yrOn-siteML Engineering

Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

OpenAI

OpenAI

London, United Kingdom

Applied AI Engineer, Digital Natives
No salary listedHybridML Engineering

Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.

Elliptic

Elliptic

London, United Kingdom

Agent Engineer
No salary listedHybrid5+ YOEML Engineering

Build full-stack AI agent fleets, APIs, workflows, and internal services that automate complex business processes. The role requires at least five years of engineering experience, hands-on LLM framework experience, production AWS expertise, Kubernetes, and strong API and database skills.