Latest ML Engineering jobs at Anthropic
Job results
Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Research Engineer developing machine-learning and reinforcement-learning systems for defensive cybersecurity, including agentic investigations, experiments, evaluations, and production training runs. Requires cybersecurity research experience, strong software engineering, and a bachelor’s degree or equivalent experience.
Research Engineer responsible for operating and improving large-scale production pretraining systems, from performance optimization and hardware debugging to experiments, observability, and launch incident response. Requires deep ML systems expertise and experience with LLM training, JAX, TPU, PyTorch, or distributed systems.
Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Technical Product Engineer advising Digital Native Businesses on integrating Claude API into products. Guides customers from discovery to deployment with expertise in LLMs, prompt engineering, agents, and evaluations; requires 4+ years experience and strong Python/TypeScript skills.
Develops and optimizes reinforcement learning systems and infrastructure for training large AI models like Claude, focusing on performance, reliability, and researcher productivity. Requires 4+ years software engineering experience.
Research Engineer on a pretraining team, developing and scaling large language models through research, experimentation, infrastructure optimization, and model engineering. Requires advanced ML or computer science education, strong software engineering skills, and expertise in Python and deep learning frameworks.
Conduct large-scale reinforcement learning experiments, develop long-horizon benchmarks, investigate scaling behavior, and translate validated research into production training recipes. The role requires strong empirical research skills, Python, distributed ML experience, and a bachelor's degree or equivalent experience.
Build and operate high-performance distributed inference infrastructure serving Claude across large-scale accelerator fleets. The role requires strong software engineering experience with production distributed systems, Kubernetes, cloud platforms, and machine learning infrastructure.
Staff Software Engineer driving reinforcement learning infrastructure for Claude's coding capabilities at Anthropic. Design APIs/frameworks, embed with research teams to build and hand off maintainable systems, improve research code reliability, and ensure production RL run health. Requires deep Python expertise, API design track record, and failure-mode intuition.
Build and own Python frameworks, APIs, and infrastructure for Anthropic's RL environments and agent runtimes. Embed with research teams to productionize their work, design for correctness in stateful distributed systems, and create self-service tooling for production debugging.
Build and own ML infrastructure, data pipelines, and tooling for Safeguards research at Anthropic. Focus on fast researcher iteration for training/evaluating lightweight detectors on model internals while ensuring correctness at scale. Requires strong Python, distributed systems, and production infrastructure experience.
Build and maintain large-scale distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and deployment pipelines across diverse AI accelerators while maximizing compute efficiency for production and research workloads. Requires significant distributed systems experience.
Build and own validation pipelines, CI/CD infrastructure, and platform integrations to launch frontier models and inference features reliably across AWS, GCP, and Azure. Requires strong large-scale distributed systems experience and track record improving release velocity.
Research Engineer advancing RL for silicon chip design at Anthropic. Design RL environments for RTL generation, verification, and physical optimization; requires deep ASIC/FPGA expertise from spec to tapeout.
Build and run evaluations to measure Claude's capabilities, safety, and performance. Design metrics, implement scalable distributed eval infrastructure and dashboards, debug training runs, and partner with researchers to characterize and improve AI systems.
Research Engineer developing novel evaluation frameworks and training strategies for AI systems in life sciences and biology. Requires experience training/evaluating LLMs, Python/ML proficiency, and data pipeline expertise; biology background preferred but not required.
Research Engineer advancing Claude's computer use capabilities through experiments, RL environments, evaluations, and infrastructure for perception and agentic tasks. Requires Python, ML training/evaluation experience, and a focus on safe AI.
Research engineer/scientist building and evaluating vision capabilities for Claude models. Requires 7+ years ML/computer vision experience and work across pretraining, RL, and agentic infrastructure.
Own end-to-end data strategy and RL environment creation for domain-specific knowledge work (finance, healthcare, legal). Combine applied research with hands-on data sourcing, vendor management, and model performance measurement.
Technical lead for the shared, accelerator-agnostic inference runtime serving Claude. Owns architecture, performance, and validation for GPU/TPU/Trainium platforms in a high-scale distributed systems environment.
Research Engineer advancing Claude's code generation capabilities through reinforcement learning. Design RL environments, build verifiers, run training experiments on frontier models, and improve training pipelines for real software engineering tasks.
Build and maintain distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and high-performance infrastructure across diverse AI accelerators.
Designs new information architectures for LLMs to interact with external data sources, implements finetuning/RL training, builds evaluation sets, and develops agentic search capabilities. Requires strong Python/ML skills and LLM experience.
Builds and optimizes RL training infrastructure, removes bottlenecks in the RL stack, and partners with researchers to accelerate model development at scale. Requires strong software engineering, ML infra experience, and comfort across the stack.
Fellows develop and optimize ML systems and performance infrastructure for AI models, focusing on scaling, efficiency, and reliability in a remote-friendly program.
Research Engineer on the Code RL team advancing AI models' ability to write efficient code for accelerators. Requires deep expertise in accelerators like CUDA/ROCm and ML frameworks like JAX/PyTorch, plus experience across kernels, model code, and distributed systems.
Build and optimize large-scale ML systems for safe, steerable AI, handling infrastructure, experiments, and dev tooling. Requires strong software engineering and interest in ML research.
Build ML systems to detect and mitigate AI misuse, including classifiers for anomalous behavior, multi-exchange harm monitoring, and agentic safety evaluations. Requires 4+ years ML experience, Python proficiency, and research-to-deployment skills.
Software Engineer on Claude Code team builds evaluation systems, tooling, and infrastructure to enhance AI coding capabilities. Collaborates with researchers in fast-paced environment; requires 5+ years experience building complex systems.
Software Engineer focused on AI reliability engineering, improving robustness of Claude's serving infrastructure across SDK to accelerators. Partners cross-team on SLOs, monitoring, high-availability systems, incident response, and safeguard models.
Designs and optimizes TPU kernels to address performance issues in ML research, training, and inference systems. Provides feedback on model impacts and solves large-scale systems problems, requiring deep accelerator expertise.