Skip to content
HarveyHarvey

Staff Software Engineer, Production Engineering

Staff Production Engineer building and operating Harvey's core compute, networking, Kubernetes, and workflow orchestration infrastructure to support rapidly growing AI workloads. Requires 10+ years experience with large-scale cloud infrastructure, Kubernetes, IaC, observability, and security.

About the job

What You'll Do

Infrastructure Engineering & Technical Leadership

  • Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.
  • Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
  • Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.
  • Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.
  • Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.
  • Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.

Infrastructure Foundation & Production Operations

  • Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
  • Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.
  • Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.
  • Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.
  • Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.
  • Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
  • Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
  • Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.
  • Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.

What You Have

  • 10+ years of software, infrastructure, site reliability, or production engineering experience.
  • Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
  • Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
  • Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
  • Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
  • Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
  • Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
  • Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
  • A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority.
  • Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.
  • A systems-thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.

Nice to Have

  • Experience supporting AI/ML or LLM infrastructure at scale.
  • Experience operating GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.
  • Experience with multi-cloud infrastructure or hybrid cloud environments.
  • Experience building internal platforms or developer tooling that improves engineering velocity and production safety.

Skills

Kubernetes, Terraform, Pulumi, AWS, Azure, GCP, Infrastructure As Code, Observability, Distributed Systems, IAM, Network Security

Skydio

Skydio

San Mateo, CA
Staff Site Reliability Engineer
$240k+/yrRemote8+ YOEDevOps / SRE

Owns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.

Shield AI

Shield AI

San Mateo, CA
Sr. Staff Lead Site Reliability Engineer
$220k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.

Coinbase

Coinbase

United States

Staff Software Engineer, Developer Infrastructure
$218k+/yrRemote8+ YOEDevOps / SRE

Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.

Coinbase

Coinbase

United States

Staff Infrastructure Engineer, Trading
$218k+/yrRemote8+ YOEDevOps / SRE

Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.

Datadog

Datadog

Boston, MA
Staff Engineer - Cloud Networks
$244k+/yrHybrid7+ YOEDevOps / SRE

Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.