Senior Software Engineer, Model Infrastructure
Leads development of Harvey’s model infrastructure platform, including reliable AI inference, model routing, provider integrations, observability, and capacity management. Requires distributed-systems experience, strong programming skills, and technical leadership across engineering teams.
About the job
Responsibilities
- Lead the design and implementation of Harvey’s Model Infrastructure platform.
- Build highly available, low-latency, operationally excellent systems for AI inference.
- Design and improve the Unified Model Controller and Model Selector to detect model degradation and route traffic based on reliability, latency, quality, compliance, and cost.
- Develop model provisioning, capacity management, failover, and traffic-engineering systems across multiple AI providers.
- Integrate model providers and maintain provider APIs and SDKs.
- Build observability capabilities including health dashboards, alerting, token-usage analytics, cost reporting, and end-to-end telemetry.
- Partner with Product Engineering on model launches, experimentation, and production monitoring.
- Drive capacity planning, utilization optimization, and cost visibility.
- Collaborate with AI Research on infrastructure for model evaluation, training, and deployment.
- Lead cross-functional technical initiatives and mentor engineers.
Requirements
- 4+ years of software engineering experience building large-scale distributed systems.
- Experience designing and operating highly available production services.
- Strong programming skills in Go, Java, Python, Rust, or C++.
- Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
- Experience leading technical projects across multiple engineering teams.
- Ability to balance long-term architecture with pragmatic execution.
- Strong communication and collaboration skills.
- Passion for building foundational platforms for other engineering teams.
Nice to Have
- Experience with AI infrastructure, LLM serving, or machine learning platforms.
- Experience with model routing, inference gateways, or policy-based serving systems.
- Experience with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs.
- Experience with Kubernetes, cloud infrastructure, and service mesh technologies.
- Experience with large-scale observability and SRE practices.
- Experience with Kafka, Spark, Flink, Airflow, or Iceberg.
- Familiarity with GPU infrastructure or model training platforms.
Compensation
- $193,400–$290,000 USD
Skills
Go, Java, Python, Rust, C++, Distributed Systems, Cloud Infrastructure, Networking, Observability, Kubernetes, Service Mesh, Kafka, Spark, Flink, Airflow
Similar jobs
Backend Engineering jobsBuilds real-time video-streaming pipelines, simulation frameworks, and communication infrastructure connecting autonomous vehicles with remote operators. Requires a computer science degree, 4+ years of relevant experience, and proficiency in C/C++ or Go plus networking expertise.
Build networking and real-time systems for reliable remote vehicle operation over cellular networks. The role requires 5+ years of industry experience, strong C++ and Linux networking expertise, and familiarity with real-time protocols, streaming, and reliability techniques.
Leads architecture and development of resilient fleet connectivity software spanning cellular and Wi-Fi networks, telemetry, routing, and networking infrastructure for autonomous vehicles. Requires 8+ years of experience, technical leadership, strong systems-design skills, and expertise with modems and Linux networking.
Build and operate the backend runtime platform that schedules, executes, and persists security and compliance tests at scale. The role requires strong production systems experience, ownership of complex projects, debugging expertise, and effective technical collaboration and mentorship.
Senior software engineer building backend systems and generalized platform solutions for data-intensive, cross-product challenges. Requires 5+ years of backend experience, strong product intuition, data infrastructure expertise, and the ability to lead complex technical projects.