Skip to content
Worth AIWorth AI

Senior DevOps Engineer, Infrastructure & Reliability

Build and operate reliable, secure cloud infrastructure across AWS and Kubernetes while automating delivery, observability, disaster recovery, and cost optimization. The role requires 8+ years in DevOps, SRE, or infrastructure engineering and strong hands-on experience with Terraform, Kubernetes, AWS, and CI/CD.

About the job

Responsibilities

  • Implement scalable infrastructure-as-code patterns with Terraform to standardize cloud provisioning and reduce configuration drift.
  • Own and evolve the Kubernetes platform, including EKS or self-managed Kubernetes, with secure, scalable, and resilient workloads.
  • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
  • Design and enforce secure networking, IAM, and secrets-management strategies across environments.
  • Improve observability through actionable metrics, logs, and tracing using DataDog.
  • Optimize cloud costs through rightsizing, autoscaling, and architectural improvements.
  • Implement disaster recovery, backup, and multi-region resilience strategies.
  • Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems.
  • Introduce infrastructure tooling and architectural changes, driving adoption through documentation, workshops, and hands-on support.
  • Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments.
  • Communicate technical trade-offs across engineering and product stakeholders.

Requirements

  • 8+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Experience designing and operating production Kubernetes environments at scale.
  • Deep hands-on expertise with AWS infrastructure and cloud networking.
  • Strong experience building and maintaining Terraform modules across large cloud environments.
  • Experience owning CI/CD systems and improving DORA metrics.
  • Experience leading incident response and driving effective postmortems.
  • Strong understanding of distributed systems, event-driven architectures, Kafka, and PostgreSQL performance.
  • Ability to modernize legacy infrastructure and eliminate manual operational toil.
  • Ability to take infrastructure projects from ambiguity through production independently.
  • Ability to build trust across teams while improving reliability.

Success Metrics

  • Maintain or exceed SLO/SLA targets while reducing incident frequency and duration.
  • Reduce incidents caused by misconfiguration, manual processes, or infrastructure drift.
  • Increase the percentage of infrastructure managed through code and automation.
  • Improve cloud cost efficiency without sacrificing reliability or performance.

Nice-to-Haves

  • Application coding experience.
  • Experience operating high-throughput Kafka clusters, including MSK or self-managed deployments.
  • Database performance tuning with PostgreSQL and Redis.
  • Autoscaling strategies for high-traffic systems.
  • Service mesh technologies.
  • Internal developer platforms.
  • Security practices such as zero-trust networking and policy-as-code.
  • Multi-region or globally distributed systems.
  • Reliability frameworks including SLOs, error budgets, and chaos testing.

Technology Stack

  • Cloud and infrastructure: AWS, EKS, RDS, MSK, S3, Lambda, IAM, VPC
  • Containerization and orchestration: Kubernetes, ArgoCD
  • Infrastructure as code: Terraform
  • CI/CD: GitHub Actions
  • Monitoring and observability: DataDog
  • Data and messaging: PostgreSQL, Kafka, Redis
  • Languages: Bash, Python, TypeScript, JavaScript

Compensation and Benefits

  • Medical, dental, and vision health care plan
  • 401(k) retirement plan
  • Life insurance
  • Flexible paid time off
  • 9 paid holidays
  • Family leave
  • Wellness resources
  • Remote work
  • Remote hires must travel to Orlando, Florida at least twice per year for town halls and collaboration, in addition to orientation in Orlando.

Skills

AWS, Kubernetes, Terraform, Amazon Eks, Argo CD, GitHub Actions, Datadog, Kafka, Postgres, Redis, Python, Bash, IAM, Vpc, TypeScript

Shield AI

Shield AI

San Diego, CA
Senior Platform Engineer
$141k+/yrHybrid7+ YOEDevOps / SRE

Designs and operates shared cloud and private-cloud platforms, infrastructure automation, Kubernetes capabilities, and developer self-service tools. Requires 7+ years in platform, cloud infrastructure, DevOps, or SRE, with strong Terraform, Ansible, Linux, Kubernetes, and public-cloud experience.

Shield AI

Shield AI

San Mateo, CA
Senior Network Engineer
$140k+/yrOn-site6+ YOEDevOps / SRE

Designs, deploys, and operates secure, resilient enterprise and cloud networks across data centers, on-premises environments, and AWS and Azure. Requires 6+ years of production network experience plus expertise in routing, switching, firewalls, automation, and hybrid connectivity.

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

Coinbase

Coinbase

United States

Senior Software Engineer, Core Infra Systems
$186k+/yrRemote5+ YOEDevOps / SRE

Senior software engineer building standardized, self-service cloud infrastructure across AWS, Google Cloud, and networking systems. Requires 5+ years of software engineering experience, production cloud infrastructure expertise, and proficiency in Go or Python.

Shield AI

Shield AI

Seattle, WA
Senior Site Infrastructure Engineer
$110k+/yrOn-site5+ YOEDevOps / SRE

Designs and supports physical IT infrastructure across offices, labs, manufacturing facilities, and data centers, including racks, cabling, power, cooling, documentation, and capacity planning. Requires 5+ years of physical infrastructure engineering experience and strong cross-functional project execution.