Skip to content
HarveyHarvey

Staff Software Engineer, Core Infrastructure

Designs and builds scalable multi-cloud infrastructure powering AI platform, focusing on Kubernetes orchestration, reliability, observability, and operational excellence. Requires 10+ years in infrastructure engineering with deep IaC and cloud expertise.

About the job

What You’ll Do

  • Design and build scalable, fault-tolerant infrastructure systems that power Harvey's AI platform across multiple cloud regions
  • Own and evolve our multi-cloud infrastructure (Azure, GCP), including Kubernetes orchestration, networking, and container management
  • Lead technical initiatives around observability, incident response, and operational excellence — building systems that enable rapid detection and resolution of issues
  • Architect and optimize our distributed systems for reliability, including load balancing, quota management, and failover mechanisms
  • Partner with Product Engineering and Security teams to ensure our infrastructure is an accelerant, not a constraint
  • Drive infrastructure-as-code practices using tools like Terraform and Pulumi to enable reproducible, auditable deployments
  • Mentor engineers and raise the technical bar across the organization through code reviews, design reviews, and technical leadership

Representative Projects

  • Design and implement a next-generation model proxy architecture that routes millions of daily inference requests while maintaining model API compatibility and enabling seamless model integration
  • Build distributed rate limiting and quota management systems using Redis-backed algorithms to handle bursty traffic patterns without degrading user experience
  • Architect multi-region deployment strategies that meet strict data residency requirements for global enterprise customers
  • Develop comprehensive observability infrastructure with granular SLA monitoring, burn rate alerts, and detailed token attribution for cost tracking
  • Lead the evolution of our CI/CD pipelines to improve developer velocity while maintaining production stability

What You Have

  • 10+ years of experience in Infrastructure Engineering or Platform Engineering in a production environment
  • Long track record building and scaling complex, large-scale distributed systems
  • Deep proficiency with cloud infrastructure platforms (Azure preferred; GCP or AWS experience transfers well)
  • Strong fluency in Infrastructure as Code (IaC) tools — Terraform, Pulumi, or CloudFormation
  • Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale
  • Experience with observability tools (Datadog, Sentry) and incident response practices (PagerDuty, Incident.io)
  • Strong programming skills in Python, Go, or similar languages
  • Excellent problem-solving skills, a "spidey sense" of where things could go wrong, and a commitment to operational excellence

Nice to Have

  • Experience building infrastructure for AI/ML workloads or high-throughput inference systems
  • Background with distributed rate limiting, load balancing, or quota management systems
  • Experience operating multi-tenant platforms with strict security and compliance requirements
  • Track record of leading complex cross-functional projects and delivering measurable impact

Compensation Range

$201,000 - $264,000 USD

Skills

Kubernetes, Terraform, Pulumi, Azure, GCP, Python, Go, Datadog, Sentry, Pagerduty, Redis, CI/CD, Distributed Systems, Observability, Infrastructure As Code

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Operations
$195k+/yrOn-site8+ YOEDevOps / SRE

Own reliability, incident response, observability, and automation for Crusoe Cloud’s global network infrastructure supporting large-scale GPU workloads. The role requires 8+ years of production network engineering experience, expertise in data center and lossless fabrics, Python automation skills, and strong operational leadership.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

Airbnb

Airbnb

United States

Staff Software Engineer, Service Tools
$212k+/yrRemote9+ YOEDevOps / SRE

Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.

Temporal

Temporal

United States

Staff Software Engineer, Traffic
$212k+/yrRemote8+ YOEDevOps / SRE

Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.

OpenSea

OpenSea

United States

Staff Platform Engineer
$190k+/yrRemote7+ YOEDevOps / SRE

Build and operate scalable platform services, infrastructure, and developer tooling that enable reliable product delivery. The role requires 7+ years of software engineering experience, JVM expertise, distributed-systems experience, and strong platform, cloud, CI/CD, and observability skills.