Skip to content
HarveyHarvey

Senior Software Engineer, Core Infrastructure

Designs, builds, and scales core infrastructure for Harvey's AI platform, focusing on multi-cloud systems (Azure, GCP), Kubernetes, and observability. Requires 4+ years in infrastructure engineering with strong IaC and distributed systems expertise.

About the job

What You’ll Do

  • Design and build scalable, fault-tolerant infrastructure systems that power Harvey's AI platform across multiple cloud regions
  • Own and evolve our multi-cloud infrastructure (Azure, GCP), including Kubernetes orchestration, networking, and container management
  • Lead technical initiatives around observability, incident response, and operational excellence
  • Architect and optimize our distributed systems for reliability, including load balancing, quota management, and failover mechanisms
  • Partner with Product Engineering and Security teams to ensure infrastructure accelerates product development
  • Drive infrastructure-as-code practices using tools like Terraform and Pulumi
  • Mentor junior engineers and raise the technical bar through code reviews, design reviews, and technical leadership

What You Have

  • 4+ years of experience in Infrastructure Engineering or Platform Engineering in a production environment
  • Long track record building and scaling complex, large-scale distributed systems
  • Deep proficiency with cloud infrastructure platforms (Azure preferred; GCP or AWS experience transfers well)
  • Strong fluency in Infrastructure as Code (IaC) tools — Terraform, Pulumi, or CloudFormation
  • Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale
  • Experience with observability tools (Datadog, Sentry) and incident response practices (PagerDuty, Incident.io)
  • Strong programming skills in Python, Go, or similar languages

Nice to Have

  • Experience building infrastructure for AI/ML workloads or high-throughput inference systems
  • Background with distributed rate limiting, load balancing, or quota management systems
  • Experience operating multi-tenant platforms with strict security and compliance requirements
  • Track record of leading complex cross-functional projects

Skills

Kubernetes, Azure, GCP, Terraform, Pulumi, Python, Go, Datadog, Sentry, Pagerduty

Skydio

Skydio

San Mateo, CA

Senior Software Engineer, Developer Productivity
$200k+/yrOn-site5+ YOEDevOps / SRE

Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.

Anyscale

Anyscale

San Francisco, CA

Senior Site Reliability Engineer, Platform Infrastructure
$200k+/yrHybrid5+ YOEDevOps / SRE

Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.

Onos Health

Onos Health

San Francisco, CA

Lead Infrastructure Engineer
$200k+/yrHybrid7+ YOEDevOps / SRE

Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.

Lightspark

Lightspark

Remote

Senior Production Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.