Skip to content
HarveyHarvey

Senior Engineering Manager, Production Engineering

Lead the Infrastructure Foundation & Production Quality Engineering team at Harvey, owning core compute, networking, Kubernetes, and workflow orchestration platforms. Drive reliability, scalability, security, and cost optimization for rapidly growing AI workloads while mentoring engineers and partnering with cross-functional leaders.

About the job

What You'll Do

Leadership & Strategy

  • Lead, mentor, and grow a team of high-performing infrastructure engineers responsible for Harvey's production infrastructure foundation.
  • Foster a culture of operational excellence, engineering quality, customer ownership, and continuous improvement.
  • Partner with Engineering, Security, Product, and AI Infrastructure leaders to define long-term infrastructure strategy and execution priorities.
  • Drive technical direction for compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead cross-functional initiatives to improve reliability, scalability, security, operational efficiency, and infrastructure cost optimization.

Infrastructure Foundation & Production Operations

  • Own and operate Harvey's global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
  • Manage compute resources to maximize utilization, performance, and service availability while supporting rapidly growing AI workloads.
  • Lead capacity planning, demand forecasting, and fleet lifecycle management to ensure infrastructure scales efficiently with business growth.
  • Operate and continuously improve Harvey's Kubernetes platform, including cluster provisioning, upgrades, monitoring, reliability, performance, and operational automation.
  • Own Harvey's Temporal-based workflow orchestration platform, ensuring reliable, scalable, and observable execution of distributed application workflows.
  • Drive infrastructure cost optimization through capacity management, resource rightsizing, workload efficiency improvements, and utilization monitoring.
  • Build and maintain secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
  • Establish comprehensive observability, monitoring, alerting, incident response, and operational readiness practices across the infrastructure platform.

What You Have

  • 7+ years of software or infrastructure engineering experience, including 5+ years leading engineering teams.
  • Deep expertise operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
  • Experience building and operating large-scale distributed systems with strong reliability, scalability, and performance characteristics.
  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
  • Experience designing and operating observability platforms, including monitoring, logging, alerting, and incident response processes.
  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
  • Demonstrated success leading complex cross-functional technical initiatives and influencing engineering strategy across organizations.
  • Excellent communication skills with the ability to communicate technical concepts clearly to both engineering and executive audiences.
  • A systems-thinking mindset and passion for building simple, reliable, and scalable infrastructure platforms.

Nice to Have

  • Experience operating workflow orchestration platforms such as Temporal.
  • Experience supporting AI/ML or LLM infrastructure at scale.
  • Experience managing GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.
  • Experience with multi-cloud infrastructure or hybrid cloud environments.

Skills

Kubernetes, Terraform, Pulumi, AWS, Azure, GCP, Temporal, Infrastructure As Code, Observability, Capacity Planning

Checkr

Checkr

San Francisco, CA

Senior Engineering Manager, Mortgage
$269k+/yrHybrid8+ YOEEngineering Management

Leads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.

Checkr

Checkr

San Francisco, CA

Senior Engineering Manager, Machine Learning
$268k+/yrOn-site10+ YOEEngineering Management

Leads the Machine Learning Engineering team, setting technical strategy, developing engineers, and shipping reliable ML and agentic AI systems for product intelligence, fraud detection, and workforce integrity. Requires 10+ years building production software and ML/AI systems plus substantial engineering management experience.

Temporal

Temporal

United States

Senior Engineering Manager - Test Systems & Tooling
$268k+/yrRemote5+ YOEEngineering Management

Leads and grows a team building shared test systems and tooling for production-representative, workload, performance, and failure-mode validation. Requires engineering management experience, distributed-systems expertise, and a record of delivering widely adopted internal platforms.

Reddit

Reddit

United States

Senior Machine Learning Manager, Video Ranking
$266k+/yrRemote5+ YOEEngineering Management

Leads machine learning engineering teams responsible for relevance, personalization, and video discovery systems serving Reddit’s core product. Requires 5+ years managing ML teams, hands-on experience with production ML systems, and strong recommender-systems expertise.

Headway

Headway

San Francisco, CA
Senior Engineering Manager
$265k+/yrHybrid8+ YOEEngineering Management

Leads a team building retrieval and machine-learned ranking systems that match patients with therapists across a healthcare marketplace. Requires substantial engineering management experience, production software or ML expertise, and hands-on ownership of search, ranking, recommendation, or personalization systems.