# Staff Software Engineer, Model Infrastructure

**Company:** [Harvey](https://hotfix.jobs/companies/harvey)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $236k – $290k/yr
**Experience:** 7+ years
**Skills:** Go, Java, Python, Rust, C++, Kubernetes, Kafka, Spark, Flink, Airflow, iceberg, AWS, Distributed Systems, Observability, llm serving
**Posted:** 2026-07-22

> Lead design and development of Harvey's Model Infrastructure platform powering all AI requests, including unified model controller, intelligent routing, multi-provider integrations, observability, and capacity management for high reliability, low latency, and efficiency. Requires 7+ years building large-scale distributed systems with strong programming and leadership skills; AI/LLM infrastructure experience preferred.

## Job Description

## What You’ll Do
- Lead the design and implementation of Harvey's Model Infrastructure platform.
- Build systems to ensure high availability, low latency, and operational excellence for AI inference.
- Design and improve Harvey's Unified Model Controller (UMC) and Model Selector platform to automatically detect model degradations and intelligently route traffic based on reliability, latency, quality, compliance, and cost.
- Develop systems for model provisioning, capacity management, failover, and traffic engineering across multiple AI providers.
- Integrate new model providers and maintain provider APIs and SDKs, enabling Harvey to rapidly adopt emerging frontier models.
- Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end telemetry.
- Partner with Product Engineering to support model launches, experimentation, and proactive monitoring of production AI workloads.
- Drive infrastructure efficiency through capacity planning, utilization optimization, and cost visibility.
- Collaborate with AI Research to build the infrastructure foundation for future model evaluation, training, and deployment.
- Lead cross-functional technical initiatives and mentor engineers across the organization.

## What You’ll Build
- **Model Reliability & Operations**: Model health monitoring, automated failover and recovery, capacity provisioning, operational tooling and incident automation, Unified Model Controller (UMC), policy-based model routing, Intelligent Model Selector, traffic management, reliability and latency optimization.
- **Provider Platform**: Multi-provider architecture, API and SDK integrations (OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, and future providers), rapid adoption of new frontier models.
- **Observability & Cost Platform**: Token usage analytics, cost attribution, latency and reliability dashboards, capacity forecasting, utilization optimization.
- **AI Platform Foundation**: Infrastructure supporting model evaluation, model deployment and operations, future model training platform, agent infrastructure and CcaaS.

## What You Have
- 7+ years of software engineering experience building large-scale distributed systems.
- Experience designing and operating highly available production services.
- Strong programming skills in Go, Java, Python, Rust, or C++.
- Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
- Experience leading technical projects across multiple engineering teams.
- Ability to balance long-term architecture with pragmatic execution.
- Strong communication and collaboration skills.
- Passion for building foundational platforms that enable other engineering teams.

## Nice to Have
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience with model routing, inference gateways, or policy-based serving systems.
- Experience working with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs.
- Experience with Kubernetes, cloud infrastructure, and service mesh technologies.
- Experience with large-scale observability and SRE best practices.
- Experience with data infrastructure technologies such as Kafka, Spark, Flink, Airflow, or Iceberg.
- Familiarity with GPU infrastructure or model training platforms.

## Compensation
$236,000 - $290,000 USD

## Similar roles

- [Senior/Staff Software Engineer - Machine Learning Platform (Inference)](https://hotfix.jobs/jobs/c165a122-9da2-4d67-994b-5f193b19f971) - Snowflake - Menlo Park, CA - $236k – $339k/yr
- [Staff AI Engineer - Cortex Code Quality](https://hotfix.jobs/jobs/1274746b-863d-4937-84f2-c41ac854c1b9) - Snowflake - Menlo Park, CA - $236k – $339k/yr
- [Staff Software Engineer](https://hotfix.jobs/jobs/a1dd767e-3d27-40b5-bb17-9a5f3a5afe35) - Confluent - Remote - $236k – $277k/yr
- [Senior/Staff Software Engineer, Labeling Platform](https://hotfix.jobs/jobs/91ce91e6-1f34-4392-b863-d7b69689bea5) - Nuro - Mountain View, CA - $235k – $352k/yr
- [Senior Staff Software Engineer, AI Model LifeCycle](https://hotfix.jobs/jobs/75876f8d-3aba-4f82-9049-17169f5eec63) - Crusoe - San Francisco, CA - $238k – $318k/yr

**Apply:** https://hotfix.jobs/jobs/636a493c-2661-46b6-959e-de5f7eb37f60
**Canonical:** https://hotfix.jobs/jobs/636a493c-2661-46b6-959e-de5f7eb37f60