Senior Engineering Manager, Model Infrastructure
Lead the Model Infrastructure engineering team at Harvey to build reliable, scalable platforms for multi-provider AI model operations, routing, observability, and future training infrastructure. Requires 8+ years software engineering experience including multiple years managing teams on large-scale distributed systems.
$272k – $355k/yr
Hybrid8+ YOEEngineering Management
About the job
Responsibilities
- Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.
- Define the technical roadmap for model reliability, scalability, and operational excellence.
- Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.
- Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.
- Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.
- Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end model telemetry.
- Partner with Product Engineering to support new model launches, capacity planning, experimentation, and proactive production monitoring.
- Lead initiatives to improve inference efficiency, reduce infrastructure costs, and increase model utilization across providers.
- Build the infrastructure foundation for Harvey's future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities.
- Partner with executive leadership on long-term AI infrastructure strategy and vendor relationships.
- Recruit, mentor, and develop exceptional engineering talent while fostering a culture of technical excellence and operational ownership.
Requirements
- 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.
- Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.
- Strong technical background that enables you to guide architectural decisions and mentor senior engineers.
- Experience operating highly available production services with strong reliability and operational excellence.
- Experience building platforms that require scalability, observability, automation, and cost optimization.
- Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.
- Excellent communication skills and the ability to influence technical strategy across organizations.
- A passion for building teams and developing engineering talent.
Nice-to-Haves
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.
- Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure.
- Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms.
- Experience supporting GPU infrastructure, model training platforms, or ML infrastructure.
- Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg.
- Experience leading organizations through periods of rapid growth and technical transformation.
Skills
KubernetesSparkKafkaFlinkAirflowIcebergLlm ServingDistributed SystemsCloud InfrastructureObservabilityGpu Infrastructure