Research Engineer Intern, Evaluations
Designs evaluation frameworks and benchmarks to test AI agents' autonomy, reasoning, and reliability in data pipelines and warehouses. Requires experience in LLM benchmarking, reinforcement learning, Python, PyTorch/JAX, and data engineering tools.
About the job
What You’ll Do
- Develop evaluation environments to test AI agents' ability to reason, plan, and act autonomously within mission-critical data pipelines.
- Design benchmarks to assess model capabilities in failure detection, pipeline optimization, and agentic decision-making in data workflows.
- Implement automated assessment frameworks for language model-based agents operating over data lakes and warehouses.
- Work with synthetic and real-world datasets to create robust testing environments for AI-driven data automation.
- Collaborate with research engineers to refine reward shaping strategies, guiding models toward more efficient and agentic behaviors in data-intensive tasks.
What We’re Looking For
- Experience in language model research, with a focus on benchmarking LLMs in mission-critical domains.
- Strong background in AI evaluation methodologies, reinforcement learning, and RLHF techniques.
- Familiarity with benchmarking language models for structured and unstructured data tasks.
- Proficiency in Python and experience with ML frameworks like PyTorch or JAX.
- Hands-on experience with data lakes, warehouses, and data engineering tools (Snowflake, BigQuery, dbt, Spark, Kafka).
- High agency—proactive, resourceful, and comfortable working in a fast-paced research environment with minimal supervision.
- Attention to detail—ability to design rigorous, reproducible experiments and evaluations.
Bonus Points
- Contributions to open-source AI benchmarks (e.g., SweBench, BIRD, SPIDER).
- Contributions to open-source agentic frameworks.
- Experience developing custom RL environments for AI evaluation.
- Strong understanding of ETL, ELT, and data transformation pipelines.
Skills
Python, PyTorch, JAX, Reinforcement Learning, RLHF, Snowflake, BigQuery, dbt, Spark, Kafka
Similar jobs
ML Engineering jobsBuilds high-scale data pipelines, distributed systems, and AI agent workflows using LLMs for fraud intelligence platform. Requires 2+ years software engineering, Python proficiency, big data tools, AWS/K8s, and ML foundations.
Machine learning intern pursuing a Ph.D. who will build and deploy production-scale models and pipelines, conduct an end-to-end research project, and collaborate with engineering and product teams on blockchain and cryptocurrency applications.
Build, deploy, and optimize AI applications and machine learning models for customer use cases while contributing to an internal ML platform. This new graduate role requires a technical master’s degree, hands-on ML or LLM experience, and strong customer communication skills.
Develop and deploy machine learning models for biomedical research and AI product development, collaborating with scientific, engineering, and product teams. Requires a master's degree with 2–4 years of experience or a PhD with 0–2 years, plus strong Python and ML development skills.
Machine learning engineer who trains, evaluates, and productionizes models and LLM-powered applications for financial products. Requires 2+ years of ML systems experience, strong Python and PyTorch skills, production data pipelines, model evaluation, and API development.