# Staff Software Engineer, Production Engineering

**Company:** [Harvey](https://hotfix.jobs/companies/harvey)
**Location:** New York, NY, San Francisco, CA
**Role:** DevOps / SRE
**Salary:** $231k – $340k/yr
**Experience:** 10+ years
**Skills:** Kubernetes, Terraform, Pulumi, AWS, Azure, GCP, Infrastructure As Code, Observability, Distributed Systems, IAM, Network Security
**Posted:** 2026-07-30

> Staff Production Engineer building and operating Harvey's core compute, networking, Kubernetes, and workflow orchestration infrastructure to support rapidly growing AI workloads. Requires 10+ years experience with large-scale cloud infrastructure, Kubernetes, IaC, observability, and security.

## Job Description

## What You'll Do

**Infrastructure Engineering & Technical Leadership**
- Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.
- Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
- Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.
- Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.
- Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.
- Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.

**Infrastructure Foundation & Production Operations**
- Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
- Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.
- Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.
- Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.
- Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.
- Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
- Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
- Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.
- Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.

## What You Have
- 10+ years of software, infrastructure, site reliability, or production engineering experience.
- Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
- Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
- Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
- Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
- Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
- Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
- Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
- A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority.
- Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.
- A systems-thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.

## Nice to Have
- Experience supporting AI/ML or LLM infrastructure at scale.
- Experience operating GPU fleets, high-performance compute infrastructure, or large-scale capacity planning.
- Experience with multi-cloud infrastructure or hybrid cloud environments.
- Experience building internal platforms or developer tooling that improves engineering velocity and production safety.

## Similar jobs

- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/75edbb52-9f2f-49c0-b1f8-8b59e87fffc3) - Skydio - Remote - $240k – $300k/yr
- [Sr. Staff Lead Site Reliability Engineer](https://hotfix.jobs/jobs/b123b52f-6dd7-417b-978e-717e23b19c0d) - Shield AI - San Mateo, CA - $220k – $330k/yr
- [Staff Software Engineer, Developer Infrastructure](https://hotfix.jobs/jobs/9e5df28a-b02e-49c1-ae34-8ac4874bd491) - Coinbase - Remote - $218k – $257k/yr
- [Staff Infrastructure Engineer, Trading](https://hotfix.jobs/jobs/4be49748-b215-44fa-b91e-9975c38847a9) - Coinbase - Remote - $218k – $257k/yr
- [Staff Engineer - Cloud Networks](https://hotfix.jobs/jobs/66aff9c2-57a3-4f6b-bc39-51a61a03c8ae) - Datadog - Boston, MA - $244k – $305k/yr

**Apply:** https://hotfix.jobs/jobs/d6a0bb77-5d16-40cd-a028-780e94d41257
**Canonical:** https://hotfix.jobs/jobs/d6a0bb77-5d16-40cd-a028-780e94d41257