Staff Software Engineer, Model Infrastructure
Lead design and development of Harvey's Model Infrastructure platform powering all AI requests, including unified model controller, intelligent routing, multi-provider integrations, observability, and capacity management for high reliability, low latency, and efficiency. Requires 7+ years building large-scale distributed systems with strong programming and leadership skills; AI/LLM infrastructure experience preferred.
About the job
What You’ll Do
- Lead the design and implementation of Harvey's Model Infrastructure platform.
- Build systems to ensure high availability, low latency, and operational excellence for AI inference.
- Design and improve Harvey's Unified Model Controller (UMC) and Model Selector platform to automatically detect model degradations and intelligently route traffic based on reliability, latency, quality, compliance, and cost.
- Develop systems for model provisioning, capacity management, failover, and traffic engineering across multiple AI providers.
- Integrate new model providers and maintain provider APIs and SDKs, enabling Harvey to rapidly adopt emerging frontier models.
- Improve observability through health dashboards, alerting, token usage analytics, cost reporting, and end-to-end telemetry.
- Partner with Product Engineering to support model launches, experimentation, and proactive monitoring of production AI workloads.
- Drive infrastructure efficiency through capacity planning, utilization optimization, and cost visibility.
- Collaborate with AI Research to build the infrastructure foundation for future model evaluation, training, and deployment.
- Lead cross-functional technical initiatives and mentor engineers across the organization.
What You’ll Build
- Model Reliability & Operations: Model health monitoring, automated failover and recovery, capacity provisioning, operational tooling and incident automation, Unified Model Controller (UMC), policy-based model routing, Intelligent Model Selector, traffic management, reliability and latency optimization.
- Provider Platform: Multi-provider architecture, API and SDK integrations (OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, and future providers), rapid adoption of new frontier models.
- Observability & Cost Platform: Token usage analytics, cost attribution, latency and reliability dashboards, capacity forecasting, utilization optimization.
- AI Platform Foundation: Infrastructure supporting model evaluation, model deployment and operations, future model training platform, agent infrastructure and CcaaS.
What You Have
- 7+ years of software engineering experience building large-scale distributed systems.
- Experience designing and operating highly available production services.
- Strong programming skills in Go, Java, Python, Rust, or C++.
- Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
- Experience leading technical projects across multiple engineering teams.
- Ability to balance long-term architecture with pragmatic execution.
- Strong communication and collaboration skills.
- Passion for building foundational platforms that enable other engineering teams.
Nice to Have
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience with model routing, inference gateways, or policy-based serving systems.
- Experience working with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs.
- Experience with Kubernetes, cloud infrastructure, and service mesh technologies.
- Experience with large-scale observability and SRE best practices.
- Experience with data infrastructure technologies such as Kafka, Spark, Flink, Airflow, or Iceberg.
- Familiarity with GPU infrastructure or model training platforms.
Compensation
$236,000 - $290,000 USD
Skills
Go, Java, Python, Rust, C++, Kubernetes, Kafka, Spark, Flink, Airflow, Iceberg, AWS, Distributed Systems, Observability, Llm Serving
Similar jobs
ML Engineering jobsLeads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.
Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.
Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.
Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.
Staff AI Engineer responsible for designing and shipping agent architecture, multi-agent coordination, reliable execution, memory, tooling, and evaluation systems. Requires substantial shipped agentic-system experience and strong software engineering skills in Python, TypeScript, or a comparable language.