Skip to content
HarveyHarvey

Staff Software Engineer, Model Infrastructure

Lead design and development of Harvey's Model Infrastructure platform powering all AI requests, including unified model controller, intelligent routing, multi-provider integrations, observability, and capacity management for high reliability, low latency, and efficiency. Requires 7+ years building large-scale distributed systems with strong programming and leadership skills; AI/LLM infrastructure experience preferred.

About the job

What You’ll Do

  • Lead the design and implementation of Harvey's Model Infrastructure platform.
  • Build systems to ensure high availability, low latency, and operational excellence for AI inference.
  • Design and improve Harvey's Unified Model Controller (UMC) and Model Selector platform to automatically detect model degradations and intelligently route traffic based on reliability, latency, quality, compliance, and cost.
  • Develop systems for model provisioning, capacity management, failover, and traffic engineering across multiple AI providers.
  • Integrate new model providers and maintain provider APIs and SDKs, enabling Harvey to rapidly adopt emerging frontier models.
  • Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end telemetry.
  • Partner with Product Engineering to support model launches, experimentation, and proactive monitoring of production AI workloads.
  • Drive infrastructure efficiency through capacity planning, utilization optimization, and cost visibility.
  • Collaborate with AI Research to build the infrastructure foundation for future model evaluation, training, and deployment.
  • Lead cross-functional technical initiatives and mentor engineers across the organization.

What You’ll Build

  • Model Reliability & Operations: Model health monitoring, automated failover and recovery, capacity provisioning, operational tooling and incident automation, Unified Model Controller (UMC), policy-based model routing, Intelligent Model Selector, traffic management, reliability and latency optimization.
  • Provider Platform: Multi-provider architecture, API and SDK integrations (OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, and future providers), rapid adoption of new frontier models.
  • Observability & Cost Platform: Token usage analytics, cost attribution, latency and reliability dashboards, capacity forecasting, utilization optimization.
  • AI Platform Foundation: Infrastructure supporting model evaluation, model deployment and operations, future model training platform, agent infrastructure and CcaaS.

What You Have

  • 7+ years of software engineering experience building large-scale distributed systems.
  • Experience designing and operating highly available production services.
  • Strong programming skills in Go, Java, Python, Rust, or C++.
  • Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
  • Experience leading technical projects across multiple engineering teams.
  • Ability to balance long-term architecture with pragmatic execution.
  • Strong communication and collaboration skills.
  • Passion for building foundational platforms that enable other engineering teams.

Nice to Have

  • Experience with AI infrastructure, LLM serving, or machine learning platforms.
  • Experience with model routing, inference gateways, or policy-based serving systems.
  • Experience working with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs.
  • Experience with Kubernetes, cloud infrastructure, and service mesh technologies.
  • Experience with large-scale observability and SRE best practices.
  • Experience with data infrastructure technologies such as Kafka, Spark, Flink, Airflow, or Iceberg.
  • Familiarity with GPU infrastructure or model training platforms.

Compensation

$236,000 - $290,000 USD

Skills

Go, Java, Python, Rust, C++, Kubernetes, Kafka, Spark, Flink, Airflow, Iceberg, AWS, Distributed Systems, Observability, Llm Serving

Snowflake

Snowflake

Bellevue, WA

Staff Software Engineer - Snowflake Feature Store
$236k+/yrOn-site10+ YOEML Engineering

Leads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.

Shield AI

Shield AI

San Mateo, CA

Staff Software Engineer, Autonomy Capabilities
$234k+/yrOn-site7+ YOEML Engineering

Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Model Infrastructure
$231k+/yrHybrid7+ YOEML Engineering

Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.

Airbnb

Airbnb

United States

Senior Staff Machine Learning Engineer, Post Training
$248k+/yrRemote10+ YOEML Engineering

Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.

Artisan

Artisan

San Francisco, CA

Staff AI Engineer - Agent Architecture & Behavior
$250k+/yrOn-site7+ YOEML Engineering

Staff AI Engineer responsible for designing and shipping agent architecture, multi-agent coordination, reliable execution, memory, tooling, and evaluation systems. Requires substantial shipped agentic-system experience and strong software engineering skills in Python, TypeScript, or a comparable language.