Build and operate Harvey's core production infrastructure powering AI workloads, including Kubernetes, compute fleets, networking, and orchestration platforms. Drive reliability, scalability, security, and cost efficiency for rapidly growing LLM infrastructure while partnering across engineering teams.
161k – 242k/yr
Hybrid5+ YOEDevOps / SRE
About the role
What You'll Do
Infrastructure Engineering & Technical Leadership
Design, build, and operate the production infrastructure that powers Harvey’s products and AI workloads.
Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
Lead complex, cross-functional technical initiatives that improve reliability, scalability, security, operational efficiency, and infrastructure cost.
Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate product and business requirements into resilient infrastructure solutions.
Establish reusable patterns, tooling, and paved paths that help engineering teams ship and operate production services safely.
Raise the engineering bar through thoughtful design reviews, clear technical documentation, operational rigor, and mentorship.
Infrastructure Foundation & Production Operations
Build and operate Harvey’s global compute and network infrastructure, ensuring high availability, scalability, reliability, and performance.
Improve compute utilization, performance, and service availability while supporting rapidly growing AI workloads.
Develop capacity models, demand forecasts, and fleet lifecycle automation to help infrastructure scale efficiently with business growth.
Operate and continuously improve Harvey’s Kubernetes platform, including cluster provisioning, upgrades, networking, monitoring, reliability, performance, and operational automation.
Drive infrastructure cost efficiency through capacity management, resource rightsizing, workload optimization, and utilization monitoring.
Build secure infrastructure foundations, including identity and access management, network isolation, secrets management, auditing, and compliance controls.
Develop scalable Infrastructure-as-Code and automation frameworks using technologies such as Terraform and Pulumi.
Improve observability, monitoring, alerting, incident response, and operational readiness across the infrastructure platform.
Participate in the on-call rotation, lead incident response when needed, and turn production learnings into durable engineering improvements.
What You Have
5+ years of software, infrastructure, site reliability, or production engineering experience.
Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
Experience with infrastructure automation and Infrastructure-as-Code using tools such as Terraform or Pulumi.
Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance best practices.
A track record of driving complex, cross-functional technical initiatives and influencing engineering decisions without relying on formal authority.
Excellent communication skills and the ability to explain technical concepts clearly to engineering partners and other stakeholders.
A systems-thinking mindset and a passion for building simple, reliable, and scalable infrastructure platforms.
Nice to Have
Experience supporting AI/ML or LLM infrastructure at scale.
Senior engineer building self-service internal platform capabilities at Docker, focusing on multi-region networking, continuous deployment, EKS foundations, and AI-assisted operations to enable faster, safer provisioning for engineering teams.
161k – 261k/yrRemote6+ YOEDevOps / SRE
Senior Site Reliability Engineer
SquareSan Francisco, CA +1
Senior SRE improves platform reliability using AI-driven automation, leads incident response and oncall for critical services, and implements safe deployment practices. Requires 5+ years experience in backend/platform engineering, observability, and high-availability systems.
161k – 284k/yrOn-site5+ YOEDevOps / SRE
Senior Site Reliability Engineer
TalkiatryUnited States
Join as the first SRE to define reliability practices, SLOs, observability, and toil reduction across six product teams at a leading mental health platform. 7+ years software/infra engineering with hands-on SRE experience required; product teams retain on-call ownership.
160k – 185k/yrRemote7+ YOEDevOps / SRE
Senior Infrastructure Engineer
AurelianSeattle, WA
Senior Infrastructure Engineer building analytics, observability, and developer tooling for Aurelian's real-time AI agents used in 911 emergency response centers. Requires 4+ years in infrastructure/platform/backend roles with experience in reliability and scale.
160k – 220k/yrOn-site4+ YOEDevOps / SRE
Senior Microsoft Cloud Infrastructure Engineer
CrusoeSan Francisco, CA
Senior Cloud Infrastructure Engineer owning design, implementation, and management of Microsoft 365, Entra ID, Azure, and Azure Arc hybrid infrastructure. Requires 8+ years infrastructure experience with deep Azure/M365 expertise, IaC, Windows Server admin, and hands-on data center hardware work.