Skip to content
HarveyHarvey

Staff Software Engineer, Model Infrastructure

Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.

About the job

Responsibilities

  • Lead the design and implementation of the Model Infrastructure platform.
  • Build highly available, low-latency, operationally excellent systems for AI inference.
  • Design and improve the Unified Model Controller and Model Selector for degradation detection and policy-based traffic routing.
  • Develop model provisioning, capacity management, failover, and traffic-engineering systems across multiple AI providers.
  • Integrate model providers and maintain provider APIs and SDKs.
  • Improve observability with health dashboards, alerting, token-usage analytics, cost reporting, and end-to-end telemetry.
  • Support model launches, experimentation, and monitoring of production AI workloads.
  • Drive capacity planning, utilization optimization, infrastructure efficiency, and cost visibility.
  • Collaborate on infrastructure for model evaluation, training, deployment, and agent platforms.
  • Lead cross-functional technical initiatives and mentor engineers.

Requirements

  • 7+ years of software engineering experience building large-scale distributed systems.
  • Experience designing and operating highly available production services.
  • Strong programming skills in Go, Java, Python, Rust, or C++.
  • Deep understanding of distributed systems, cloud infrastructure, networking, and observability.
  • Experience leading technical projects across multiple engineering teams.
  • Ability to balance long-term architecture with pragmatic execution.
  • Strong communication and collaboration skills.

Nice to Have

  • Experience with AI infrastructure, LLM serving, or machine-learning platforms.
  • Experience with model routing, inference gateways, or policy-based serving systems.
  • Experience with OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, or open-source LLMs.
  • Experience with Kubernetes, cloud infrastructure, and service-mesh technologies.
  • Experience with large-scale observability and SRE best practices.
  • Experience with Kafka, Spark, Flink, Airflow, or Iceberg.
  • Familiarity with GPU infrastructure or model-training platforms.

Compensation

  • $231,000–$340,000 USD annually.

Skills

Go, Java, Python, Rust, C++, Distributed Systems, Cloud Infrastructure, Kubernetes, Model Routing, Llm Serving, Observability, Networking, Kafka, Spark, Flink

Shield AI

Shield AI

San Mateo, CA

Staff Software Engineer, Autonomy Capabilities
$234k+/yrOn-site7+ YOEML Engineering

Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.

Snowflake

Snowflake

Bellevue, WA

Staff Software Engineer - Snowflake Feature Store
$236k+/yrOn-site10+ YOEML Engineering

Leads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.

Shield AI

Shield AI

Washington, DC
Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site7+ YOEML Engineering

Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.

Ironclad

Ironclad

San Francisco, CA

Senior Staff Software Engineer, Agentic Search
$220k+/yrHybrid10+ YOEML Engineering

Leads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.

Idme

Idme

Mountain View, CA

Staff Software Engineer - AI Agent Evaluations
$218k+/yrOn-site8+ YOEML Engineering

Leads the engineering discipline for evaluating, testing, and monitoring production AI agents, while building scalable eval infrastructure and developer tooling. Requires 8+ years of production software experience, strong backend skills, and expertise with LLM evaluation and agentic systems.