Build and operate resilient, large-scale platform services across AWS and Azure, managing Kubernetes clusters, infrastructure automation, deployment orchestration, and observability. The role requires 4+ years of software engineering experience with Kubernetes, Terraform, Go, and distributed systems.
139k – 204k/yr
Remote4+ YOEDevOps / SRE
About the role
Responsibilities
Design, build, and operate services and automations to manage Kubernetes clusters at scale.
Partner with product management and technical leadership to break down complex system requirements into iterative milestones.
Drive rigorous code reviews and maintainable coding patterns.
Ensure high testing standards across unit, integration, and component testing.
Manage and enhance cloud configurations across AWS and Azure using Terraform.
Standardize metrics, alerts, and distributed tracing across core data pipelines.
Identify technical debt, system bottlenecks, and single points of failure while balancing feature delivery with platform refactoring.
Mentor junior engineers, lead technical sprint planning, and share expertise across distributed engineering teams.
Requirements
4+ years of professional software engineering experience building and operating resilient backend services at scale using Kubernetes.
Experience with CAPI, EKS, and zero-downtime Kubernetes cluster upgrades, including node draining, API deprecations, and PodDisruptionBudgets.
Experience using AI-assisted development tools, or a strong desire to learn.
Hands-on experience implementing GitOps workflows with Argo CD and automated pipeline orchestration with Harness or an equivalent enterprise CI/CD platform.
Strong hands-on experience with Shell, Terraform, YAML, and Go.
Experience deploying and managing production workloads in cloud environments, ideally AWS or equivalent Azure services.
Understanding of container networking, including VPC/VNet, pod IPAM, and CNI plugins.
Proficiency with Terraform for infrastructure automation and maintaining environment parity.
Understanding of distributed datastores, caching layers, and asynchronous event streaming such as Kafka.
Strong computer science fundamentals, including data structures and self-healing cloud architectures.
Nice-to-haves
Experience managing high-throughput applications in Docker and Kubernetes environments.
Familiarity with canary and blue/green deployment strategies.
Experience with OPA/Gatekeeper or similar policy-as-code enforcement in Kubernetes.
Experience implementing OpenTelemetry or distributed tracing across decoupled microservice platforms.
Exposure to network topology, proxy layers, or mail transfer agent protocol constraints.
New York, New Jersey, Washington State, or California outside the San Francisco Bay area: $146,800–$183,600
San Francisco Bay area, California: $163,100–$203,900
May be eligible for equity and a corporate bonus plan.
Benefits generally include health insurance, a 401(k) retirement account, paid sick time, paid personal time off, paid parental leave, wellness leave, and other benefits.
Infrastructure Engineer building and operating scalable, reliable production systems for an enterprise AI platform. Owns end-to-end reliability, automates with Python/Go, integrates AI agents into workflows, leads incident response, and collaborates cross-functionally on high-availability infrastructure using Kubernetes, Terraform, and multi-cloud tooling. Requires 5+ years experience and daily use of AI tooling.
140k – 274k/yrHybrid5+ YOEDevOps / SRE
Software Engineer, Compute Infrastructure
GleanMountain View, CA
Build and operate Kubernetes-based compute and runtime infrastructure powering AI search, assistant, and agent workloads across multi-cloud environments. Own reliability, scalability, cost-efficiency, and on-call for production platform services.
140k – 220k/yrHybrid5+ YOEDevOps / SRE
Platform Engineer
HarperSan Francisco, CA
Build and own core infrastructure, observability, and tooling for a hyper-growth AI insurance platform running 200+ services and thousands of daily agentic AI decisions. Focus on scale, reliability, developer velocity, and AI eval systems.
140k – 280k/yrOn-site5+ YOEDevOps / SRE
Release Engineer
ZooxFoster City, CA
As a Release Engineer, you will orchestrate software releases for autonomous vehicle technology, ensuring secure and streamlined delivery from development to production. This role involves managing simulation tools and autonomy software releases, coordinating vehicle-level testing, and scaling automation systems.
140k – 190k/yrHybrid3+ YOEDevOps / SRE
Site Reliability Engineering
ZooxFoster City, CA
Site Reliability Engineer owns the lifecycle of services powering autonomous vehicles, designing fault-tolerant systems, building monitoring tools, leading incident response, and ensuring infrastructure resilience with large-scale data processing on CPUs/GPUs. Requires 5+ years SRE experience, cloud/IaC expertise, Kubernetes, and strong programming skills.