Skip to content
Worth AIWorth AIOrlando, FL

Senior DevOps Engineer, Infrastructure & Reliability

The Senior DevOps Engineer will build and operate reliable cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability across AWS environments. The role requires 5+ years in DevOps, SRE, or infrastructure engineering, with strong Terraform, Kubernetes, AWS, networking, and distributed-systems experience.

Salary not listed
Remote5+ YOEDevOps / SRE

About the role

Responsibilities

  • Implement scalable infrastructure-as-code patterns with Terraform to standardize cloud provisioning and reduce configuration drift.
  • Own and evolve the Kubernetes platform, including EKS or self-managed Kubernetes, ensuring workloads are secure, scalable, and resilient.
  • Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
  • Design and enforce secure networking, IAM, and secrets-management strategies across environments.
  • Improve observability by refining metrics, logs, and tracing with tools such as Datadog.
  • Optimize cloud costs through rightsizing, autoscaling strategies, and architectural improvements.
  • Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
  • Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
  • Introduce infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.
  • Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.
  • Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.

Technology Stack

  • Cloud and infrastructure: AWS, EKS, RDS, MSK, S3, Lambda, IAM, VPC
  • Containerization and orchestration: Kubernetes, ArgoCD
  • Infrastructure as code: Terraform
  • CI/CD: GitHub Actions
  • Monitoring and observability: Datadog
  • Data and messaging: PostgreSQL, Kafka, Redis
  • Languages: Bash, Python, TypeScript, JavaScript

Requirements

  • 5+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Experience designing and operating production Kubernetes environments at scale.
  • Deep hands-on expertise with AWS infrastructure and cloud networking.
  • Strong experience building and maintaining Terraform modules across large cloud environments.
  • Experience owning CI/CD systems and improving DORA metrics.
  • Experience leading incident-response processes and driving meaningful postmortem outcomes.
  • Strong understanding of distributed systems, event-driven architectures, Kafka, and PostgreSQL performance.
  • Ability to modernize legacy infrastructure and eliminate manual operational toil.
  • Ability to take infrastructure projects from ambiguous beginnings through production without daily direction.
  • Ability to build trust across teams while raising the reliability bar.

Nice-to-Haves

  • Experience coding applications.
  • Experience operating high-throughput Kafka clusters, including MSK or self-managed deployments.
  • Background in PostgreSQL and Redis performance tuning.
  • Experience implementing autoscaling strategies for high-traffic systems.
  • Familiarity with service mesh technologies.
  • Experience building internal developer platforms.
  • Background in security best practices, including zero-trust networking and policy as code.
  • Experience with multi-region or globally distributed systems.
  • Experience introducing reliability frameworks such as SLOs, error budgets, and chaos testing.

Success Metrics

  • Maintain or exceed defined SLO and SLA targets while reducing incident frequency and duration.
  • Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.
  • Increase the percentage of infrastructure managed through code and automation.
  • Improve cloud cost efficiency without sacrificing reliability or performance.

Benefits

  • Medical, dental, and vision coverage
  • 401(k) retirement plan
  • Life insurance
  • Flexible paid time off
  • 9 paid holidays
  • Family leave
  • Wellness resources
  • Remote work; hybrid work for Orlando associates
  • Free food and snacks in Orlando
  • Remote hires travel to Orlando, Florida at least twice per year for town halls and team collaboration, in addition to orientation in Orlando

Skills

TerraformKubernetesAWSamazon eksArgo CDGitHub ActionsDatadogKafkaPostgresRedisPythonTypeScriptJavaScriptIAMvpc

Similar roles

DevOps / SRE jobs
Alpaca

Senior Site Reliability Engineer

AlpacaUnited States

Operates and improves reliability for a trading-critical brokerage platform across cloud infrastructure, Kubernetes, observability, messaging, and PostgreSQL. Requires 4+ years of production operations experience, strong PostgreSQL fundamentals, incident response expertise, and proficiency in Go or Python.

Salary not listedRemote5+ YOEDevOps / SRE
Shield AI

Senior Engineer, Platform Infrastructure

Shield AISan Diego, CA +2

Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.

120k – 180k/yrOn-site5+ YOEDevOps / SRE
The Voleon Group

Storage and Datacenter Team Lead

The Voleon GroupBerkeley, CA +1

Leads a storage engineering team while architecting and operating highly available Linux-based storage, datacenter, and data-protection infrastructure. The role requires deep Ceph experience, PB-scale archiving and backup expertise, hands-on troubleshooting, and team leadership.

215k – 245k/yrRemote5+ YOEDevOps / SRE
Bestow

Senior Platform Engineer

BestowUnited States

Own platform initiatives that improve cloud scalability, reliability, automation, and developer productivity. The role requires 5+ years of cloud infrastructure experience plus hands-on expertise with infrastructure as code, Kubernetes, CI/CD, programming or scripting, and AI-assisted engineering.

145k – 171k/yrRemote5+ YOEDevOps / SRE
Discord

Senior Software Engineer, Enterprise Platform

DiscordUnited States

Build and operate Discord’s greenfield Enterprise Platform, turning identity, device management, infrastructure, and application delivery into reusable self-service software. The role requires production software engineering experience, IAM and Terraform expertise, and end-to-end ownership of complex platform projects.

196k – 221k/yrOn-site5+ YOEDevOps / SRE