Skip to content

Senior / Staff AI Model Engineer

Own reliability and quality for an AI copilot in a trading platform. Design evaluation systems, benchmarks, quality gates, model improvement loops, and AI monitoring for correctness, safety, and performance in market analysis and trading workflows. Requires 8+ years production software experience and strong ML eval expertise.

About the job

Responsibilities

  • Own the reliability and quality bar for an AI copilot embedded in a trading platform used by sophisticated investors.
  • Design and build evaluation systems that measure correctness, safety, latency, and regression risk across market analysis, portfolio/risk reasoning, and trading workflows (including order placement).
  • Develop and maintain benchmarks: curated “golden sets,” scenario suites, stress/adversarial cases, and continuously refreshed market/regime-based test corpora.
  • Build automated quality gates and regression workflows that block releases when key metrics degrade.
  • Partner with engineering and product to define safe tool/action contracts (deterministic previews, confirmations, auditability) and ensure predictable assistant behavior.
  • Own model improvement loops tied to evals: data collection/labeling strategies, error taxonomy, prompt/tooling changes, and when appropriate, fine-tuning or preference optimization to measurably improve benchmark performance.
  • Design and operate monitoring + incident response for AI: telemetry, alerting, RCA, and “fix-forward” processes.
  • Develop a deep understanding of trading concepts (margin, shorting, portfolio margin, risk, execution) and how to express them accurately and understandably to users.

Requirements

  • At least 8 years of experience shipping production software; strong proficiency with any programming language.
  • Strong knowledge of computer science fundamentals, testing methodology, and systems design.
  • Experience building evaluation frameworks, test harnesses, and benchmark suites for complex systems (LLMs/agents/search/retrieval/ranking/recommenders).
  • Experience running model improvement cycles: dataset curation, labeling/QA, offline experimentation, and deploying changes with measurable impact on benchmarks.
  • Ability to define metrics, build measurement pipelines, and drive engineering/product decisions from data.
  • Comfort working across the stack: debugging model/tooling failures, instrumenting services, and partnering with frontend/product on UX patterns that improve safety and trust.
  • High degree of self-motivation and willingness to jump into unfamiliar areas to solve problems.

Nice-to-Haves

  • Experience with fine-tuning, preference optimization, distillation, or prompt/compiler-style techniques for improving tool-use reliability.
  • Experience creating domain-specific benchmarks and adversarial suites (e.g., “known-bad” scenarios) for high-stakes applications.
  • Deep experience with trading across asset classes, margin types, etc.
  • Experience with Rust and performance-sensitive services.
  • Experience designing incident response and SLOs for ML/AI systems.

Compensation

  • Base Salary Range: $200,000 - $350,000 (does not include bonuses or equity).
  • Competitive compensation packages, company equity, 401k matching, gender neutral parental leave, and full medical, dental and vision insurance.

Skills

Rust, TypeScript, LLM APIs, Model Serving, Evaluation Frameworks, Benchmark Suites, Fine-Tuning, Preference Optimization, Prompt Engineering, Python, Testing Methodology, System Design

Shield AI

Shield AI

San Diego, CA

Staff Engineer, Perception Software
$200k+/yrOn-site7+ YOEML Engineering

Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.

Nuro

Nuro

Mountain View, CA

Senior/Staff Engineer, Machine Learning - Online Mapping
$194k+/yrOn-site7+ YOEML Engineering

Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.

Nuro

Nuro

Mountain View, CA

Senior/Staff Software Engineer, ML Inference Platform
$194k+/yrOn-site5+ YOEML Engineering

Build and operate ML infrastructure for autonomy teams, including training and deployment pipelines, model observability, inference serving, and compiler platforms across hardware targets. Requires a degree, 3+ years of relevant experience, Python proficiency, and distributed-systems expertise.

Talkiatry

Talkiatry

United States

Staff AI Enablement Engineer
$190k+/yrRemote8+ YOEML Engineering

Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.

Airbnb

Airbnb

United States

Staff Machine Learning Engineer, Relevance and Personalization
$212k+/yrRemote9+ YOEML Engineering

Staff machine learning engineer leading scalable ranking, search, recommendation, and personalization systems. The role requires 9+ years of applied machine learning experience, strong programming and data engineering skills, and expertise productionizing models and pipelines.