Skip to content
HarveyHarveySan Francisco, CA

Senior Engineering Manager, Model Infrastructure

Lead the Model Infrastructure engineering team at Harvey to build reliable, scalable platforms for multi-provider AI model operations, routing, observability, and future training infrastructure. Requires 8+ years software engineering experience including multiple years managing teams on large-scale distributed systems.

272k – 355k/yr
Hybrid8+ YOEEngineering Management

About the role

Responsibilities

  • Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.
  • Define the technical roadmap for model reliability, scalability, and operational excellence.
  • Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.
  • Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.
  • Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.
  • Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end model telemetry.
  • Partner with Product Engineering to support new model launches, capacity planning, experimentation, and proactive production monitoring.
  • Lead initiatives to improve inference efficiency, reduce infrastructure costs, and increase model utilization across providers.
  • Build the infrastructure foundation for Harvey's future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities.
  • Partner with executive leadership on long-term AI infrastructure strategy and vendor relationships.
  • Recruit, mentor, and develop exceptional engineering talent while fostering a culture of technical excellence and operational ownership.

Requirements

  • 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.
  • Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.
  • Strong technical background that enables you to guide architectural decisions and mentor senior engineers.
  • Experience operating highly available production services with strong reliability and operational excellence.
  • Experience building platforms that require scalability, observability, automation, and cost optimization.
  • Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.
  • Excellent communication skills and the ability to influence technical strategy across organizations.
  • A passion for building teams and developing engineering talent.

Nice-to-Haves

  • Experience with AI infrastructure, LLM serving, or machine learning platforms.
  • Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.
  • Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure.
  • Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms.
  • Experience supporting GPU infrastructure, model training platforms, or ML infrastructure.
  • Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg.
  • Experience leading organizations through periods of rapid growth and technical transformation.

Skills

KubernetesSparkKafkaFlinkAirflowicebergllm servingDistributed SystemsCloud InfrastructureObservabilitygpu infrastructure
Harvey

Senior Engineering Manager, Production Engineering

HarveySan Francisco, CA

Lead the Infrastructure Foundation & Production Quality Engineering team at Harvey, owning core compute, networking, Kubernetes, and workflow orchestration platforms. Drive reliability, scalability, security, and cost optimization for rapidly growing AI workloads while mentoring engineers and partnering with cross-functional leaders.

272k – 355k/yr
Hybrid7+ YOEEngineering Management
Confluent

Senior Engineering Manager, Flink Control Plane

ConfluentCalifornia

Lead a team of engineers to develop and execute the roadmap for the Flink Control Plane, focusing on architectural excellence, reliability, and scaling for Confluent's managed Flink offering. This role involves significant technical strategy and people leadership.

272k – 319k/yr
Remote10+ YOEEngineering Management
Fluidstack

Production Engineering Lead, Compute

FluidstackSan Francisco, CA

Lead the compute production engineering team at Fluidstack, owning availability SLOs, automation for node lifecycle (provisioning to remediation), and on-call models for tens of thousands of GPUs at nation-scale. Requires prior leadership of large-fleet SRE/production teams, proven availability improvements, and automated remediation experience.

269k – 335k/yr
On-site7+ YOEEngineering Management
Fluidstack

Software Engineer Tech Lead

FluidstackAustin, TX +3

Tech lead setting technical direction and building core pieces of an internal platform for AI compute infrastructure, including orchestration, integrations, and shared systems. Requires experience designing platforms others depend on, high-bar code review, and daily use of AI coding tools.

269k – 335k/yr
On-site7+ YOEEngineering Management
Fluidstack

Software Engineer Team Lead

FluidstackAustin, TX +3

Lead a small high-velocity engineering team building internal operational software and tools for AI compute infrastructure buildout. Stay hands-on coding while owning priorities, delivery, engineer growth, hiring, and cross-functional product decisions.

269k – 335k/yr
On-site5+ YOEEngineering Management