Senior Engineering Manager, Model Infrastructure
Lead the Model Infrastructure engineering team at Harvey to build reliable, scalable platforms for multi-provider AI model operations, routing, observability, and future training infrastructure. Requires 8+ years software engineering experience including multiple years managing teams on large-scale distributed systems.
About the job
Responsibilities
- Lead and grow a high-performing team of software engineers responsible for Harvey's Model Infrastructure platform.
- Define the technical roadmap for model reliability, scalability, and operational excellence.
- Build highly reliable systems for model provisioning, capacity management, failover, and incident response across multiple AI providers.
- Own Harvey's multi-provider model platform, including provider integrations, SDK upgrades, API migrations, and onboarding new model providers.
- Drive the evolution of our Unified Model Controller (UMC) and Model Selector platform to automatically detect degraded models and intelligently route traffic based on health, latency, quality, compliance, and cost.
- Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end model telemetry.
- Partner with Product Engineering to support new model launches, capacity planning, experimentation, and proactive production monitoring.
- Lead initiatives to improve inference efficiency, reduce infrastructure costs, and increase model utilization across providers.
- Build the infrastructure foundation for Harvey's future model training efforts, including data pipelines, model operations, training environments, and AI platform capabilities.
- Partner with executive leadership on long-term AI infrastructure strategy and vendor relationships.
- Recruit, mentor, and develop exceptional engineering talent while fostering a culture of technical excellence and operational ownership.
Requirements
- 8+ years of software engineering experience, including multiple years managing high-performing engineering teams.
- Experience leading teams responsible for large-scale distributed systems or cloud infrastructure.
- Strong technical background that enables you to guide architectural decisions and mentor senior engineers.
- Experience operating highly available production services with strong reliability and operational excellence.
- Experience building platforms that require scalability, observability, automation, and cost optimization.
- Strong cross-functional leadership skills with the ability to partner effectively across Engineering, Research, Product, and external vendors.
- Excellent communication skills and the ability to influence technical strategy across organizations.
- A passion for building teams and developing engineering talent.
Nice-to-Haves
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience working with multiple model providers such as OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source model ecosystems.
- Experience building inference platforms, model gateways, traffic routing systems, or policy-based serving infrastructure.
- Experience with Kubernetes, cloud infrastructure, distributed systems, and large-scale observability platforms.
- Experience supporting GPU infrastructure, model training platforms, or ML infrastructure.
- Familiarity with data platforms and technologies such as Spark, Kafka, Flink, Airflow, or Iceberg.
- Experience leading organizations through periods of rapid growth and technical transformation.
Skills
Kubernetes, Spark, Kafka, Flink, Airflow, Iceberg, Llm Serving, Distributed Systems, Cloud Infrastructure, Observability, Gpu Infrastructure
Similar jobs
Engineering Management jobsLeads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.
Leads the Machine Learning Engineering team, setting technical strategy, developing engineers, and shipping reliable ML and agentic AI systems for product intelligence, fraud detection, and workforce integrity. Requires 10+ years building production software and ML/AI systems plus substantial engineering management experience.
Leads and grows a team building shared test systems and tooling for production-representative, workload, performance, and failure-mode validation. Requires engineering management experience, distributed-systems expertise, and a record of delivering widely adopted internal platforms.
Leads machine learning engineering teams responsible for relevance, personalization, and video discovery systems serving Reddit’s core product. Requires 5+ years managing ML teams, hands-on experience with production ML systems, and strong recommender-systems expertise.
Leads a team building retrieval and machine-learned ranking systems that match patients with therapists across a healthcare marketplace. Requires substantial engineering management experience, production software or ML expertise, and hands-on ownership of search, ranking, recommendation, or personalization systems.