Skip to content
AnthropicAnthropic

ML/Research Engineer, Safeguards

Build ML systems to detect and mitigate AI misuse, including classifiers for anomalous behavior, multi-exchange harm monitoring, and agentic safety evaluations. Requires 4+ years ML experience, Python proficiency, and research-to-deployment skills.

About the job

Responsibilities

  • Develop classifiers to detect misuse and anomalous behavior at scale. This includes developing synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on
  • Build systems to monitor for harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts
  • Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks
  • Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse

You may be a good fit if you

  • Have 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry
  • Have proficiency in Python and experience building ML systems
  • Are comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems
  • Are worried about misuse risks of AI systems, and want to work to mitigate them
  • Have strong communication skills and ability to explain complex technical concepts to non-technical stakeholders

Strong candidates may also have experience with

  • Language modeling and transformers
  • Building classifiers, anomaly detection systems, or behavioral ML
  • Adversarial machine learning or red-teaming
  • Interpretability or probes
  • Reinforcement learning
  • High-performance, large-scale ML systems

Annual Salary: $350,000—$500,000 USD

Skills

Python, Machine Learning, Transformers, Classifiers, Anomaly Detection, Adversarial Machine Learning, Red-Teaming, Interpretability, Reinforcement Learning, Large-Scale Ml Systems

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

AI Infrastructure Engineer
$350k+/yrOn-site4+ YOEML Engineering

Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research, General Agents
$350k+/yrHybridML Engineering

Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research, RL Scaling
$350k+/yrHybridML Engineering

Researcher focused on scaling reinforcement learning for frontier models, with ownership spanning asynchronous RL algorithms, inference and distributed training systems, and large-scale empirical studies. Requires strong Python and deep learning experience, scalable systems debugging, and rigorous research judgment.

OpenAI

OpenAI

San Francisco, CA

Machine Learning Engineer, Multimodal Perception and Authentication
$342k+/yrHybridML Engineering

Develop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.