# Senior DevOps Engineer, Infrastructure & Reliability

**Company:** [Worth AI](https://hotfix.jobs/companies/worth-ai)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Terraform, Kubernetes, AWS, amazon eks, Argo CD, GitHub Actions, Datadog, Kafka, Postgres, Redis, Python, TypeScript, JavaScript, IAM, vpc
**Posted:** 2026-08-04

> The Senior DevOps Engineer will build and operate reliable cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability across AWS environments. The role requires 5+ years in DevOps, SRE, or infrastructure engineering, with strong Terraform, Kubernetes, AWS, networking, and distributed-systems experience.

## Job Description

## Responsibilities
- Implement scalable infrastructure-as-code patterns with Terraform to standardize cloud provisioning and reduce configuration drift.
- Own and evolve the Kubernetes platform, including EKS or self-managed Kubernetes, ensuring workloads are secure, scalable, and resilient.
- Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
- Design and enforce secure networking, IAM, and secrets-management strategies across environments.
- Improve observability by refining metrics, logs, and tracing with tools such as Datadog.
- Optimize cloud costs through rightsizing, autoscaling strategies, and architectural improvements.
- Implement disaster recovery planning, backup strategies, and multi-region resilience initiatives.
- Refactor brittle or manually managed infrastructure into automated, testable, and reproducible systems.
- Introduce infrastructure tooling or architectural shifts and drive adoption through documentation, workshops, and hands-on support.
- Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments.
- Communicate technical trade-offs clearly across engineering and product stakeholders, balancing speed with safety.

## Technology Stack
- **Cloud and infrastructure:** AWS, EKS, RDS, MSK, S3, Lambda, IAM, VPC
- **Containerization and orchestration:** Kubernetes, ArgoCD
- **Infrastructure as code:** Terraform
- **CI/CD:** GitHub Actions
- **Monitoring and observability:** Datadog
- **Data and messaging:** PostgreSQL, Kafka, Redis
- **Languages:** Bash, Python, TypeScript, JavaScript

## Requirements
- 5+ years of experience in DevOps, SRE, or infrastructure engineering.
- Experience designing and operating production Kubernetes environments at scale.
- Deep hands-on expertise with AWS infrastructure and cloud networking.
- Strong experience building and maintaining Terraform modules across large cloud environments.
- Experience owning CI/CD systems and improving DORA metrics.
- Experience leading incident-response processes and driving meaningful postmortem outcomes.
- Strong understanding of distributed systems, event-driven architectures, Kafka, and PostgreSQL performance.
- Ability to modernize legacy infrastructure and eliminate manual operational toil.
- Ability to take infrastructure projects from ambiguous beginnings through production without daily direction.
- Ability to build trust across teams while raising the reliability bar.

## Nice-to-Haves
- Experience coding applications.
- Experience operating high-throughput Kafka clusters, including MSK or self-managed deployments.
- Background in PostgreSQL and Redis performance tuning.
- Experience implementing autoscaling strategies for high-traffic systems.
- Familiarity with service mesh technologies.
- Experience building internal developer platforms.
- Background in security best practices, including zero-trust networking and policy as code.
- Experience with multi-region or globally distributed systems.
- Experience introducing reliability frameworks such as SLOs, error budgets, and chaos testing.

## Success Metrics
- Maintain or exceed defined SLO and SLA targets while reducing incident frequency and duration.
- Reduce production incidents caused by misconfiguration, manual processes, or infrastructure drift.
- Increase the percentage of infrastructure managed through code and automation.
- Improve cloud cost efficiency without sacrificing reliability or performance.

## Benefits
- Medical, dental, and vision coverage
- 401(k) retirement plan
- Life insurance
- Flexible paid time off
- 9 paid holidays
- Family leave
- Wellness resources
- Remote work; hybrid work for Orlando associates
- Free food and snacks in Orlando
- Remote hires travel to Orlando, Florida at least twice per year for town halls and team collaboration, in addition to orientation in Orlando

## Similar roles

- [Senior DevOps Engineer](https://hotfix.jobs/jobs/9954f52d-c8b0-48dc-afd2-d34daed0f45c) - StackAI - San Francisco, CA - $130k – $210k/yr
- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/c0d74eb1-6ed1-49b2-9a56-1729408b2c86) - Alpaca - Remote
- [Senior Engineer, Platform Infrastructure](https://hotfix.jobs/jobs/cf15e8c3-ee2a-4f6a-abd6-8a54bce776fd) - Shield AI - San Diego, CA - $120k – $180k/yr
- [Storage and Datacenter Team Lead](https://hotfix.jobs/jobs/367f1d16-228b-4832-8a2a-6ee32afbf54f) - The Voleon Group - Remote - $215k – $245k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/12d9928a-3411-4a84-a041-e8c5e3456aa4) - Bestow - Remote - $145k – $171k/yr

**Apply:** https://hotfix.jobs/jobs/b7baf476-cfc4-4010-b099-ca7fc1001dea
**Canonical:** https://hotfix.jobs/jobs/b7baf476-cfc4-4010-b099-ca7fc1001dea