Senior Software Engineer, Core Infrastructure
Designs, builds, and scales core infrastructure for Harvey's AI platform, focusing on multi-cloud systems (Azure, GCP), Kubernetes, and observability. Requires 4+ years in infrastructure engineering with strong IaC and distributed systems expertise.
About the job
What You’ll Do
- Design and build scalable, fault-tolerant infrastructure systems that power Harvey's AI platform across multiple cloud regions
- Own and evolve our multi-cloud infrastructure (Azure, GCP), including Kubernetes orchestration, networking, and container management
- Lead technical initiatives around observability, incident response, and operational excellence
- Architect and optimize our distributed systems for reliability, including load balancing, quota management, and failover mechanisms
- Partner with Product Engineering and Security teams to ensure infrastructure accelerates product development
- Drive infrastructure-as-code practices using tools like Terraform and Pulumi
- Mentor junior engineers and raise the technical bar through code reviews, design reviews, and technical leadership
What You Have
- 4+ years of experience in Infrastructure Engineering or Platform Engineering in a production environment
- Long track record building and scaling complex, large-scale distributed systems
- Deep proficiency with cloud infrastructure platforms (Azure preferred; GCP or AWS experience transfers well)
- Strong fluency in Infrastructure as Code (IaC) tools — Terraform, Pulumi, or CloudFormation
- Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale
- Experience with observability tools (Datadog, Sentry) and incident response practices (PagerDuty, Incident.io)
- Strong programming skills in Python, Go, or similar languages
Nice to Have
- Experience building infrastructure for AI/ML workloads or high-throughput inference systems
- Background with distributed rate limiting, load balancing, or quota management systems
- Experience operating multi-tenant platforms with strict security and compliance requirements
- Track record of leading complex cross-functional projects
Skills
Kubernetes, Azure, GCP, Terraform, Pulumi, Python, Go, Datadog, Sentry, Pagerduty
Similar jobs
DevOps / SRE jobsBuild and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.