Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
165k – 330k/yr
HybridDevOps / SRE
About the role
Responsibilities
Design and build continuous deployment infrastructure that safely rolls out changes across dozens of Kubernetes clusters and global regions.
Develop systems for progressive delivery, including canary releases, staged rollouts, and automated rollback.
Improve engineering velocity by reducing friction in the release pipeline and automating manual operational workflows.
Work with product and infrastructure teams to ensure their services are deployable, observable, and resilient at scale.
Implement and evolve deployment methodologies such as GitOps, infrastructure-as-code, and progressive delivery patterns.
Build systems that automatically evaluate deployment health using metrics, logs, traces, and alerts to detect regressions and trigger safe rollbacks.
Explore systems that support agent-assisted or autonomous deployment workflows using modern AI tooling.
Requirements
Strong proficiency in Go and/or Python.
Experience with Kubernetes-based deployment systems at scale.
Experience building or operating continuous deployment platforms.
Familiarity with GitOps tooling such as Argo CD or Flux.
Commitment to safe production rollouts and minimizing blast radius.
Strong developer-tooling culture and focus on best practices.
Comfort working independently on ambiguous, high-impact technical challenges.
Infrastructure Engineer responsible for securing, scaling, and maintaining cloud architecture, Kubernetes clusters, ML pipelines, and CI/CD systems at a computer vision AI startup. Must have production Kubernetes, IaC, and AWS/GCP experience.
165k – 200k/yrHybridDevOps / SRE
DevOps Engineer
OctusNew York, NY
Design, implement, and maintain cloud infrastructure and CI/CD pipelines. Collaborate with developers, SRE, and Security to ensure system reliability, scalability, and security.
165k – 190k/yrOn-site5+ YOEDevOps / SRE
OS / K8s Systems Engineer
BasetenSan Francisco, CA +1
Build automation and systems to provision and orchestrate GPU hardware into scalable Kubernetes clusters. Requires deep Linux expertise, provisioning experience, and strong programming in Python/Go.
165k – 330k/yrHybridDevOps / SRE
Platform Engineer
ZooxFoster City, CA
Build and operate reliable infrastructure and testing services for autonomous-vehicle development. The role requires 5+ years supporting production services and SRE responsibilities, plus proficiency in Python or Golang and experience with automation, observability, CI/CD, and resilient infrastructure.
165k – 208k/yrHybrid5+ YOEDevOps / SRE
Site Reliability Engineer (SRE)
BasetenSan Francisco, CA +1
Site Reliability Engineer builds and maintains scalable infrastructure for ML model deployment, automates CI/CD pipelines, and ensures reliability using tools like Kubernetes and Terraform. Collaborates cross-functionally, owns projects end-to-end, and mentors juniors; bachelor's in CS or related field required.