Build and operate research infrastructure like evaluation frameworks, RL training systems, experiment tracking, and visualization tools. Partner directly with ML researchers to identify bottlenecks, ensure high adoption, and accelerate research velocity. Requires strong software engineering skills and Python/Rust proficiency.
350k – 475k/yr
On-siteML Engineering
About the role
What You'll Do
Design, build, and operate research infrastructure including evaluation frameworks, RL training systems, experiment tracking platforms, visualization tools, and shared utilities.
Develop high-throughput, scalable pipelines for distributed evaluation, reward modeling, and multimodal assessment.
Build systems for reproducibility, traceability, and robust quality control across research experiments and model training runs. Implement monitoring and observability.
Partner directly with researchers to identify bottlenecks and unlock new capabilities. Own research tooling like a product manager, proactively seeking feedback and tracking adoption.
Collaborate with infrastructure, data, and product teams to integrate tools across the technical stack.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, machine learning, or similar.
Strong software engineering fundamentals with a track record of building reliable, maintainable systems.
Proficiency in at least one backend language (we use Python or Rust).
Comfort operating across the stack and owning projects end-to-end.
Experience in highly collaborative environments involving many different cross-functional partners and subject matter experts.
Preferred qualifications:
Track record building tooling for researchers that achieved high adoption without top down mandates.
Experience building or maintaining ML research infrastructure such as training frameworks, evaluation libraries, or experiment tracking systems.
Contributions to open-source ML tools or widely-used internal frameworks at research-focused organizations.
Record of publications or technical writing on ML systems, infrastructure, or tooling.
Background working closely with ML researchers to understand and solve their tooling needs.
Familiarity with distributed systems, modern ML frameworks (PyTorch, JAX), and data processing at scale.
Experience with research observability tools, distributed compute frameworks (Ray, Spark), or large-scale evaluation pipelines.
Machine Learning Infrastructure Engineer, Safeguards Research
AnthropicSan Francisco, CA +1
Build and own ML infrastructure, data pipelines, and tooling for Safeguards research at Anthropic. Focus on fast researcher iteration for training/evaluating lightweight detectors on model internals while ensuring correctness at scale. Requires strong Python, distributed systems, and production infrastructure experience.
350k – 500k/yr
Hybrid5+ YOEML Engineering
Research Infrastructure Engineer
Thinking Machines LabSan Francisco, CA
Build and operate research infrastructure like evaluation frameworks, RL training systems, and experiment tracking platforms. Partner directly with ML researchers to identify bottlenecks, ensure high adoption of tools, and accelerate research velocity.
350k – 475k/yr
On-siteML Engineering
Research Engineer, Safeguards Labs
AnthropicSan Francisco, CA +1
Research engineer on the Safeguards Labs team building and evaluating novel safety methods to detect misuse, strengthen model safeguards, and reduce real-world harm from Claude.
350k – 850k/yr
HybridML Engineering
Research Engineer, Knowledge Team
AnthropicSan Francisco, CA +2
Designs new information architectures for LLMs to interact with external data sources, implements finetuning/RL training, builds evaluation sets, and develops agentic search capabilities. Requires strong Python/ML skills and LLM experience.
350k – 850k/yr
HybridML Engineering
Research, Vision Expertise
Thinking Machines LabSan Francisco, CA
Conducts research on visual perception, multimodal learning, and large-scale AI model training. Designs architectures, builds datasets and evaluations, and collaborates on frontier models. Requires ML expertise, Python proficiency, and experimental rigor.