Latest DevOps / SRE jobs at Harvey
Job results
Build and operate Harvey's core production infrastructure powering AI workloads, including Kubernetes, compute fleets, networking, and orchestration platforms. Drive reliability, scalability, security, and cost efficiency for rapidly growing LLM infrastructure while partnering across engineering teams.
Staff Production Engineer building and operating Harvey's core compute, networking, Kubernetes, and workflow orchestration infrastructure to support rapidly growing AI workloads. Requires 10+ years experience with large-scale cloud infrastructure, Kubernetes, IaC, observability, and security.
Builds and scales developer platforms, CI/CD systems, testing infrastructure, and AI integrations to boost engineering velocity and reliability at an AI-native company. Requires 7+ years backend experience, leadership, and tools like Python, Kubernetes, and Terraform.
Designs and builds scalable multi-cloud infrastructure powering AI platform, focusing on Kubernetes orchestration, reliability, observability, and operational excellence. Requires 10+ years in infrastructure engineering with deep IaC and cloud expertise.
Build and operate reliable, scalable infrastructure for a legal AI platform, leading observability, incident response, automation, capacity planning, and security practices. The role requires 12+ years of SRE or comparable production experience and strong cloud, Kubernetes, programming, and infrastructure-as-code expertise.
Leads PostgreSQL database infrastructure for a global legal AI platform, owning migration governance, multi-region scaling, reliability, performance tuning, and self-service tooling. Requires 10+ years experience with expert PostgreSQL knowledge and staff-level impact.
Designs, builds, and scales core infrastructure for Harvey's AI platform, focusing on multi-cloud systems (Azure, GCP), Kubernetes, and observability. Requires 4+ years in infrastructure engineering with strong IaC and distributed systems expertise.
Senior SRE ensures reliability, scalability, and performance of legal AI platform by managing global infrastructure, leading incident response, automating operations, and optimizing costs. Requires 5+ years SRE experience, IaC expertise, cloud proficiency, and strong programming skills.
Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.