# Senior Engineering Manager, Model Infrastructure

**Company:** [Harvey](https://hotfix.jobs/companies/harvey)
**Location:** San Francisco, CA
**Role:** Engineering Management
**Salary:** $272k – $355k/yr
**Experience:** 8+ years
**Skills:** Kubernetes, Spark, Kafka, Flink, Airflow, iceberg, llm serving, Distributed Systems, Cloud Infrastructure, Observability, gpu infrastructure
**Posted:** 2026-07-22

> Lead the Model Infrastructure engineering team at Harvey to build reliable, scalable platforms for multi-provider AI model operations, routing, observability, and future training infrastructure. Requires 8+ years software engineering experience including multiple years managing teams on large-scale distributed systems.

## Job Description

## Responsibilities
- Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.
- Define the technical roadmap for model reliability, scalability, and operational excellence.
- Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.
- Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.
- Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.
- Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end model telemetry.
- Partner with Product Engineering to support new model launches, capacity planning, experimentation, and proactive production monitoring.
- Lead initiatives to improve inference efficiency, reduce infrastructure costs, and increase model utilization across providers.
- Build the infrastructure foundation for Harvey's future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities.
- Partner with executive leadership on long-term AI infrastructure strategy and vendor relationships.
- Recruit, mentor, and develop exceptional engineering talent while fostering a culture of technical excellence and operational ownership.

## Requirements
- 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.
- Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.
- Strong technical background that enables you to guide architectural decisions and mentor senior engineers.
- Experience operating highly available production services with strong reliability and operational excellence.
- Experience building platforms that require scalability, observability, automation, and cost optimization.
- Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.
- Excellent communication skills and the ability to influence technical strategy across organizations.
- A passion for building teams and developing engineering talent.

## Nice-to-Haves
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.
- Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure.
- Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms.
- Experience supporting GPU infrastructure, model training platforms, or ML infrastructure.
- Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg.
- Experience leading organizations through periods of rapid growth and technical transformation.

## Similar roles

- [Senior Engineering Manager, Production Engineering](https://hotfix.jobs/jobs/3af05ebc-49b4-4e6d-96eb-31ec60bcafd7) - Harvey - San Francisco, CA - $272k – $355k/yr
- [Senior Engineering Manager, Flink Control Plane](https://hotfix.jobs/jobs/5e52de10-d7b5-46e4-a67b-1cedda96eec3) - Confluent - Remote - $272k – $319k/yr
- [Production Engineering Lead, Compute](https://hotfix.jobs/jobs/e99e0d51-6bc7-41d6-b4af-bf447d38f4a6) - Fluidstack - San Francisco, CA - $269k – $335k/yr
- [Software Engineer Tech Lead](https://hotfix.jobs/jobs/174ca984-9a78-481c-b068-65c2bc510343) - Fluidstack - Austin, TX - $269k – $335k/yr
- [Software Engineer Team Lead](https://hotfix.jobs/jobs/dc2249ff-d71d-4baa-ba63-a76e4e57438c) - Fluidstack - Austin, TX - $269k – $335k/yr

**Apply:** https://hotfix.jobs/jobs/01cf9918-eb1c-4cb0-9687-5b869cdea637
**Canonical:** https://hotfix.jobs/jobs/01cf9918-eb1c-4cb0-9687-5b869cdea637