# AI Engineer, Intern

**Company:** [Postman](https://hotfix.jobs/companies/postman)
**Location:** Berkeley, CA, San Francisco, CA
**Role:** ML Engineering
**Skills:** Python, PyTorch, TypeScript, LLMs, llm agents, tool calling, Reinforcement Learning, model fine-tuning, quantization, Docker, CI/CD, cloud gpus, benchmarking, prompt injection, Statistical Analysis
**Posted:** 2026-08-09

> AI Engineer Intern working with senior engineers to build and evaluate agentic AI systems, benchmarks, model-training pipelines, and safety evaluations. Requires current study in a quantitative field, hands-on ML experience, strong Python fundamentals, and familiarity with PyTorch.

## Job Description

## Responsibilities

### Benchmarks and Evaluation
- Contribute to APIFlow-Bench, an open-source benchmark for API-development work.
- Design and review benchmark tasks and mock API environments.
- Extend the evaluation harness and task-generation pipeline in Python.
- Maintain a public multi-model leaderboard with statistical confidence intervals.
- Help build an action-level AI safety benchmark for simulated enterprise API environments.
- Design scenarios, threat models, and auditable evaluations covering prompt injection, data exfiltration, and permission overreach.

### Model Training and Efficiency
- Fine-tune open-weight models for tool calling and agentic tasks using supervised fine-tuning, distillation, and reinforcement learning.
- Run training on managed platforms and self-managed cloud GPUs.
- Design rigorous experiments, evaluate training runs on benchmarks, conduct ablation studies and error analysis, track experiments, and report results including cost.
- Evaluate ultra-low-bit quantized models for on-device use and analyze differences between quantized and full-precision models.

### Agent Systems and Engineering
- Help build Postman’s in-product AI agent using tool loops, multi-step execution, and checkpointing, primarily in TypeScript.
- Read open-source agent harnesses and turn findings into design specifications and prototypes.
- Document experiments, design decisions, and runbooks.
- Flag safety, fairness, and privacy concerns in model or agent behavior.

## Requirements
- Currently pursuing a BS, MS, or PhD in Computer Science, Data Science, or a related quantitative field.
- Hands-on experience training or evaluating machine-learning models through coursework, research, hackathons, or internships.
- Solid Python fundamentals, including data structures, functions, and basic testing.
- Working knowledge of at least one deep-learning framework, preferably PyTorch.
- Clear written and verbal communication and strong documentation habits.

## Nice-to-Haves
- Experience fine-tuning open-weight large language models using supervised fine-tuning, LoRA, reinforcement learning, or distillation.
- Experience building LLM agents, tool-calling systems, evaluation harnesses, or benchmarks.
- Experience shipping software end-to-end, including APIs, services, CLIs, Docker, CI/CD, and cloud systems.
- Interest or experience in AI safety and robustness, including red-teaming, prompt injection, agent security, fairness, or interpretability.
- Exposure to quantization, low-bit inference, or serving optimization.
- Publications, technical blog posts, ablation studies, or self-directed projects with quantified results.
- Fluency with AI coding tools such as Claude Code, Cursor, or Codex.

## Compensation and Benefits
- Pay-on-performance philosophy and flexible schedule.
- Full medical coverage, flexible PTO, wellness reimbursement, and monthly lunch stipend.
- Wellness programs, team-building events, and donation matching.
- This role is based in the San Francisco Bay Area and requires working in the office five days per week.

## Similar roles

- [Applied Scientist II](https://hotfix.jobs/jobs/c6c7879a-d2d0-4c0c-a46c-43a9930a8457) - Garner Health - New York, NY - $158k – $190k/yr
- [Applied ML Scientist, New Grad](https://hotfix.jobs/jobs/adfb42c1-5eb0-4d5a-bd6a-582c98bca4a9) - SentiLink - Remote - $180k – $220k/yr
- [Software Engineer 3](https://hotfix.jobs/jobs/ffa1abd6-cf7d-4412-95c2-7350792135b4) - MongoDB - Remote - $109k – $215k/yr
- [Software Engineer, ML Inference Platform](https://hotfix.jobs/jobs/74081e27-c2cd-4485-b56d-22edd2ae9e1d) - Nuro - Mountain View, CA - $160k – $241k/yr
- [Software Engineer, ML Infrastructure Platform](https://hotfix.jobs/jobs/5ccc0259-2d85-4c10-8ed4-76d8ac33c4ea) - Nuro - Mountain View, CA - $160k – $241k/yr

**Apply:** https://hotfix.jobs/jobs/aa12f053-f702-4f11-9c48-2b8fed35e3e6
**Canonical:** https://hotfix.jobs/jobs/aa12f053-f702-4f11-9c48-2b8fed35e3e6