Skip to content

Software Engineer - Benchmarking

Build reproducible systems for AI model benchmarking, including datasets, evaluation pipelines, containerized environments, scoreboards, and analysis tools. The role requires 4+ years of professional engineering experience, strong Python, dataset rigor, and Docker expertise.

About the job

Responsibilities

  • Prepare and maintain benchmark datasets, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance.
  • Validate that tasks are complete, consistent, and executable, and flag ambiguities that could compromise results.
  • Build and maintain evaluation pipelines across model APIs and terminal agents to ensure reproducible and comparable results.
  • Create lightweight, containerized evaluation environments and viewers for tasking and tool-use evaluation.
  • Support fine-tuning of small open-source large language models and compare baseline with post-training performance.
  • Develop published scoreboards and leaderboards covering model-, benchmark-, task-, domain-, and rubric-level performance.
  • Build analysis tools to identify recurring failure modes, compare models and agent scaffolds, and track capability improvements and regressions.
  • Build tooling for expert annotation and data-generation projects, including HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.
  • Collaborate with researchers and engineers to ensure evaluation data and outputs are accurate, consistent, and integrated into published results.

Requirements

  • 4+ years of professional experience building and maintaining complex systems.
  • Strong Python skills and the ability to write robust, maintainable code.
  • Experience preparing, cleaning, and maintaining datasets.
  • Experience with Docker and reproducible execution environments.
  • Ability to collaborate with researchers and scientists and translate methodology into working systems.

Nice-to-haves

  • Experience running AI evaluations or using frameworks such as Harbor, Terminal-Bench, or Inspect.
  • Experience fine-tuning or post-training open-source large language models.
  • Experience with agentic, multi-turn, long-context, or tool-use evaluation.
  • Experience validating LLM-as-judge or rubric-based grading setups.
  • Background or strong interest in a scientific or technical domain.
  • Experience building data-heavy dashboards, leaderboards, or visualizations.
  • Open-source contributions or published work related to benchmarks and measurement.

Tech Stack

  • Evaluations: Python, model APIs, agent and evaluation frameworks, custom evaluation tooling
  • Models: Major AI provider APIs, terminal agents, Hugging Face ecosystem, PyTorch
  • Environments: Docker
  • Publishing: React, Next.js, Tailwind
  • Workflow: GitHub, Slack, Notion, Linear

Compensation and Benefits

  • Base salary range: $160,000–$210,000, based on seniority, relevant experience, and location.
  • Equity.
  • Medical, dental, and vision coverage.
  • 401(k).
  • Monthly wellness and fitness stipend.
  • Paid time off and company holidays.
  • Annual company off-sites.
  • Parent-friendly policies, remote flexibility, and paid family leave.

Skills

Python, Docker, React, Next.js, Tailwind, PyTorch, Hugging Face, Model Apis, Llm Evaluation, Benchmarking, Data Pipelines, GitHub

Pindrop

Pindrop

United States

Research Scientist II
$160k+/yrRemote3+ YOEML Engineering

Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.

Roboflow

Roboflow

San Francisco, CA

Member of Technical Staff — Frontier Data
$150k+/yrRemoteML Engineering

Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.

Beacon Biosignals

Beacon Biosignals

Boston, MA
Algorithm Engineer
$150k+/yrRemote4+ YOEML Engineering

Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.