Skip to content
PostmanPostmanBerkeley, CA

AI Engineer, Intern

AI Engineer Intern working with senior engineers to build and evaluate agentic AI systems, benchmarks, model-training pipelines, and safety evaluations. Requires current study in a quantitative field, hands-on ML experience, strong Python fundamentals, and familiarity with PyTorch.

Salary not listed
On-siteML Engineering

About the role

Responsibilities

Benchmarks and Evaluation

  • Contribute to APIFlow-Bench, an open-source benchmark for API-development work.
  • Design and review benchmark tasks and mock API environments.
  • Extend the evaluation harness and task-generation pipeline in Python.
  • Maintain a public multi-model leaderboard with statistical confidence intervals.
  • Help build an action-level AI safety benchmark for simulated enterprise API environments.
  • Design scenarios, threat models, and auditable evaluations covering prompt injection, data exfiltration, and permission overreach.

Model Training and Efficiency

  • Fine-tune open-weight models for tool calling and agentic tasks using supervised fine-tuning, distillation, and reinforcement learning.
  • Run training on managed platforms and self-managed cloud GPUs.
  • Design rigorous experiments, evaluate training runs on benchmarks, conduct ablation studies and error analysis, track experiments, and report results including cost.
  • Evaluate ultra-low-bit quantized models for on-device use and analyze differences between quantized and full-precision models.

Agent Systems and Engineering

  • Help build Postman’s in-product AI agent using tool loops, multi-step execution, and checkpointing, primarily in TypeScript.
  • Read open-source agent harnesses and turn findings into design specifications and prototypes.
  • Document experiments, design decisions, and runbooks.
  • Flag safety, fairness, and privacy concerns in model or agent behavior.

Requirements

  • Currently pursuing a BS, MS, or PhD in Computer Science, Data Science, or a related quantitative field.
  • Hands-on experience training or evaluating machine-learning models through coursework, research, hackathons, or internships.
  • Solid Python fundamentals, including data structures, functions, and basic testing.
  • Working knowledge of at least one deep-learning framework, preferably PyTorch.
  • Clear written and verbal communication and strong documentation habits.

Nice-to-Haves

  • Experience fine-tuning open-weight large language models using supervised fine-tuning, LoRA, reinforcement learning, or distillation.
  • Experience building LLM agents, tool-calling systems, evaluation harnesses, or benchmarks.
  • Experience shipping software end-to-end, including APIs, services, CLIs, Docker, CI/CD, and cloud systems.
  • Interest or experience in AI safety and robustness, including red-teaming, prompt injection, agent security, fairness, or interpretability.
  • Exposure to quantization, low-bit inference, or serving optimization.
  • Publications, technical blog posts, ablation studies, or self-directed projects with quantified results.
  • Fluency with AI coding tools such as Claude Code, Cursor, or Codex.

Compensation and Benefits

  • Pay-on-performance philosophy and flexible schedule.
  • Full medical coverage, flexible PTO, wellness reimbursement, and monthly lunch stipend.
  • Wellness programs, team-building events, and donation matching.
  • This role is based in the San Francisco Bay Area and requires working in the office five days per week.

Skills

PythonPyTorchTypeScriptLLMsllm agentstool callingReinforcement Learningmodel fine-tuningquantizationDockerCI/CDcloud gpusbenchmarkingprompt injectionStatistical Analysis

Similar roles

ML Engineering jobs
Garner Health

Applied Scientist II

Garner HealthNew York, NY

Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.

158k – 190k/yrHybrid2+ YOEML Engineering
SentiLink

Applied ML Scientist, New Grad

SentiLinkUnited States

Build and deploy production machine-learning models for fraud detection, identity verification, and financial risk products. The role suits new PhD graduates or early-career researchers with strong quantitative foundations, Python experience, and interest in owning the full ML lifecycle.

180k – 220k/yrRemoteML Engineering
MongoDB

Software Engineer 3

MongoDBUnited States

Build and maintain tooling, evaluation systems, quality gates, and infrastructure for MongoDB's agent skills and AI platform. Requires 2+ years building production software, developer tools, CLIs, test infrastructure, or CI/CD, with strong fundamentals in API design, testing, and reasoning about nondeterministic AI systems.

109k – 215k/yrRemote2+ YOEML Engineering
Nuro

Software Engineer, ML Inference Platform

NuroMountain View, CA

Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.

160k – 241k/yrOn-site1+ YOEML Engineering
Nuro

Software Engineer, ML Infrastructure Platform

NuroMountain View, CA

Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.

160k – 241k/yrOn-site1+ YOEML Engineering