Skip to content
AnthropicAnthropic

Research Engineer, Discovery

Builds large-scale infrastructure for AI scientist training, evaluation, and deployment, resolving bottlenecks in distributed systems for scientific AGI. Requires 6+ years in infrastructure engineering with expertise in ML stacks, containers, and data pipelines.

About the job

Responsibilities

  • Design and implement large-scale infrastructure systems to support AI scientist training, evaluation, and deployment across distributed environments
  • Identify and resolve infrastructure bottlenecks impeding progress toward scientific capabilities
  • Develop robust and reliable evaluation frameworks for measuring progress towards scientific AGI
  • Build scalable and performant VM/sandboxing/container architectures to safely execute long-horizon AI tasks and scientific workflows
  • Collaborate to translate experimental requirements into production-ready infrastructure
  • Develop large scale data pipelines to handle advanced language model training requirements
  • Optimize large scale training and inference pipelines for stable and efficient reinforcement learning

You may be a good fit if you

  • Have 6+ years of highly-relevant experience in infrastructure engineering with demonstrated expertise in large-scale distributed systems
  • Are a strong communicator and enjoy working collaboratively
  • Possess deep knowledge of performance optimization techniques and system architectures for high-throughput ML workloads
  • Have experience with containerization technologies (Docker, Kubernetes) and orchestration at scale
  • Have proven track record of building large-scale data pipelines and distributed storage systems
  • Excel at diagnosing and resolving complex infrastructure challenges in production environments
  • Can work effectively across the full ML stack from data pipelines to performance optimization
  • Have experience collaborating with other researchers to scale experimental ideas
  • Thrive in fast-paced environments and can rapidly iterate from experimentation to production

Strong candidates may also have

  • Experience with language model training infrastructure and distributed ML frameworks (PyTorch, JAX, etc.)
  • Background in building infrastructure for AI research labs or large-scale ML organizations
  • Knowledge of GPU/TPU architectures and language model inference optimization
  • Experience with cloud platforms (AWS, GCP) at enterprise scale
  • Familiarity with VM and container orchestration
  • Experience with workflow orchestration tools and experiment management systems
  • History working with large scale reinforcement learning
  • Comfort with large scale data pipelines (Beam, Spark, Dask)

Annual Salary: $350,000 — $850,000 USD

Education requirements: At least a Bachelor's degree in a related field or equivalent experience.

Skills

Kubernetes, Docker, PyTorch, JAX, AWS, GCP, Apache Beam, Spark, Dask, Distributed Systems

Anthropic

Anthropic

San Francisco, CA
Applied AI, Research Engineer
$300k+/yrHybrid6+ YOEAI Research

Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Research Scientist - Humanoid Robotics
$250k+/yrOn-site7+ YOEAI Research

Leads the research agenda for humanoid robotics, developing foundation-model and reinforcement-learning methods for dexterous manipulation and deploying them on real robotic systems. Requires a PhD, strong robotics research publications, and senior-level technical leadership.

Decagon

Decagon

San Francisco, CA
Senior Research Engineer, Safety
$200k+/yrOn-site4+ YOEAI Research

Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.

Maybern

Maybern

New York, NY

Senior Software Engineer, AI
$180k+/yrHybrid5+ YOEAI Research

Build and expand customer-facing agentic AI products, MCP integrations, and automated reconciliation workflows for private fund management. The role requires senior-level software engineering, strong systems thinking, product judgment, and hands-on experience building and evaluating AI systems.

Vanta

Vanta

Remote

Senior Product Builder, Organizational Intelligence
$176k+/yrRemote5+ YOEAI Research

Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.