Skip to content

Research Engineer Intern, Evaluations

Designs evaluation frameworks and benchmarks to test AI agents' autonomy, reasoning, and reliability in data pipelines and warehouses. Requires experience in LLM benchmarking, reinforcement learning, Python, PyTorch/JAX, and data engineering tools.

About the job

What You’ll Do

  • Develop evaluation environments to test AI agents' ability to reason, plan, and act autonomously within mission-critical data pipelines.
  • Design benchmarks to assess model capabilities in failure detection, pipeline optimization, and agentic decision-making in data workflows.
  • Implement automated assessment frameworks for language model-based agents operating over data lakes and warehouses.
  • Work with synthetic and real-world datasets to create robust testing environments for AI-driven data automation.
  • Collaborate with research engineers to refine reward shaping strategies, guiding models toward more efficient and agentic behaviors in data-intensive tasks.

What We’re Looking For

  • Experience in language model research, with a focus on benchmarking LLMs in mission-critical domains.
  • Strong background in AI evaluation methodologies, reinforcement learning, and RLHF techniques.
  • Familiarity with benchmarking language models for structured and unstructured data tasks.
  • Proficiency in Python and experience with ML frameworks like PyTorch or JAX.
  • Hands-on experience with data lakes, warehouses, and data engineering tools (Snowflake, BigQuery, dbt, Spark, Kafka).
  • High agency—proactive, resourceful, and comfortable working in a fast-paced research environment with minimal supervision.
  • Attention to detail—ability to design rigorous, reproducible experiments and evaluations.

Bonus Points

  • Contributions to open-source AI benchmarks (e.g., SweBench, BIRD, SPIDER).
  • Contributions to open-source agentic frameworks.
  • Experience developing custom RL environments for AI evaluation.
  • Strong understanding of ETL, ELT, and data transformation pipelines.

Skills

Python, PyTorch, JAX, Reinforcement Learning, RLHF, Snowflake, BigQuery, dbt, Spark, Kafka

DataVisor

DataVisor

Mountain View, CA

Software Engineer, Artificial Intelligence
$130k+/yrOn-site2+ YOEML Engineering

Builds high-scale data pipelines, distributed systems, and AI agent workflows using LLMs for fraud intelligence platform. Requires 2+ years software engineering, Python proficiency, big data tools, AWS/K8s, and ML foundations.

Coinbase

Coinbase

San Francisco, CA

Machine Learning Engineer Intern
$60+/hrHybridML Engineering

Machine learning intern pursuing a Ph.D. who will build and deploy production-scale models and pipelines, conduct an end-to-end research project, and collaborate with engineering and product teams on blockchain and cryptocurrency applications.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff
$160k+/yrOn-siteML Engineering

Build, deploy, and optimize AI applications and machine learning models for customer use cases while contributing to an internal ML platform. This new graduate role requires a technical master’s degree, hands-on ML or LLM experience, and strong customer communication skills.

PathAI

PathAI

Boston, MA
Machine Learning Engineer II/III
$107k+/yrOn-site2+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI product development, collaborating with scientific, engineering, and product teams. Requires a master's degree with 2–4 years of experience or a PhD with 0–2 years, plus strong Python and ML development skills.

Earnin

Earnin

Mountain View, CA

Machine Learning Engineer
$187k+/yrHybrid2+ YOEML Engineering

Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.