Build and operate secure, resilient, multi-region Kubernetes infrastructure across AWS environments. The role requires deep Kubernetes and Terraform expertise, GitOps and progressive delivery experience, and familiarity with deploying and securing AI gateway infrastructure.
140k – 260k/yr
Hybrid5+ YOEDevOps / SRE
About the role
Responsibilities
Kubernetes Platform & Infrastructure
Engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure.
Operate and improve multi-cluster, multi-region environments using Istio, Linkerd, Cluster API, and Kyverno.
Contribute to the DevOps and platform roadmap, including Kubernetes platform evolution, application packaging, migration to EKS, and enabling reliable production delivery.
AI Gateway
Contribute to the Kong AI Gateway, including Kubernetes operator setup, dataplane deployments, ACM/SSL integration, and observability through Datadog.
Delivery & Automation
Build and maintain progressive delivery pipelines using GitOps, canary releases, and automated releases with Flux and Flagger.
Manage CircleCI CI pipelines alongside GitOps tooling.
Maintain infrastructure-as-code and GitOps standards, including automated testing for infrastructure changes.
Security & Compliance
Support Zero Trust practices using Vault, service identity, and mTLS-secured meshes.
Build policy-driven automation using OPA or Kyverno.
Work with provisioning tools such as Crossplane or ACK for Kubernetes-native cloud integration.
Collaboration
Collaborate with Security, Data, and AI teams on DevOps and platform decisions.
Requirements
Strong hands-on Kubernetes experience, including cluster lifecycle management, API extensions, Operators, Helm, and the broader CNCF ecosystem.
Strong Terraform experience, including complex state management and module design across cloud environments.
Experience designing and operating multi-cluster, multi-region Kubernetes platforms with service meshes such as Istio, Consul, or Linkerd.
Experience with policy-based workload placement.
Experience implementing GitOps pipelines with ArgoCD or FluxCD for progressive delivery, drift correction, and multi-environment releases.
Working knowledge of LLM-based services and AI infrastructure, including deploying, securing, and operating AI gateways and services.
Ability to build secure, scalable, and reliable systems in a fast-paced environment.
Strong ownership, collaboration, adaptability, and data-informed decision-making.
Nice-to-Haves
Operator development, CRD automation, eBPF, or Cilium experience.
Policy-as-code using OPA or Kyverno within secure supply-chain frameworks.
Exposure to AI-driven internal developer platforms or predictive observability using AI/LLMs.
Datadog experience for observability, dashboards, and monitoring pipeline integrations.
Contributions to open-source or CNCF community projects.
Compensation & Benefits
$650 remote work budget for home-office setup.
$1,000 annual Learning & Development budget.
25 days of annual leave plus 8 US public holidays.
Birthday leave.
16 weeks of fully paid enhanced parental leave for eligible employees.
Comprehensive medical, dental, and vision coverage with generous premium contributions.
401(k) with company match.
Mental health support through Spill.
Hybrid work, with the option to work from almost anywhere for up to 90 days per year.
Skills
KubernetesTerraformAWSamazon eksHelmistiolinkerdGitOpsfluxcdCircleCIvaultkyvernoopaDatadogkong ai gateway
Build and operate production-critical GitOps deployment platforms, shared service tooling, and infrastructure automation in Go and TypeScript. The role requires 8+ years of software engineering experience plus expertise with Argo CD, Helm, Kubernetes, cloud platforms, and scalable APIs.
140k – 220k/yrRemote8+ YOEDevOps / SRE
Site Reliability Engineer
AxleFrederick, MD
Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.
140k – 155k/yrOn-site6+ YOEDevOps / SRE
DevOps Engineer
Pump.coSan Francisco, CA
Hands-on DevOps role owning AWS infrastructure, building developer tooling, and driving technical roadmap at an early-stage YC startup. Requires 6+ years infra/DevOps experience and strong AWS/K8s/Terraform skills.
140k – 200k/yrOn-site6+ YOEDevOps / SRE
Senior Network Systems Engineer
ForterraEast Palo Alto, CA +2
Deploys, operates, and troubleshoots network infrastructure including routers, switches, Linux appliances, and AWS resources for edge-deployed communications in DDIL environments. Requires 5+ years network engineering experience, Linux proficiency, IaC automation, and 50% domestic travel.
140k – 185k/yrHybrid5+ YOEDevOps / SRE
Senior Infrastructure Engineer
ScrunchNew York, NY +16
Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.