# Senior DevOps Engineer, Infrastructure & Reliability

**Company:** [Worth AI](https://hotfix.jobs/companies/worth-ai)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 8+ years
**Skills:** AWS, Kubernetes, Terraform, Amazon Eks, Argo CD, GitHub Actions, Datadog, Kafka, Postgres, Redis, Python, Bash, IAM, Vpc, TypeScript
**Posted:** 2026-08-25

> Build and operate reliable, secure cloud infrastructure across AWS and Kubernetes while automating delivery, observability, disaster recovery, and cost optimization. The role requires 8+ years in DevOps, SRE, or infrastructure engineering and strong hands-on experience with Terraform, Kubernetes, AWS, and CI/CD.

## Job Description

## Responsibilities
- Implement scalable infrastructure-as-code patterns with Terraform to standardize cloud provisioning and reduce configuration drift.
- Own and evolve the Kubernetes platform, including EKS or self-managed Kubernetes, with secure, scalable, and resilient workloads.
- Optimize CI/CD pipelines to improve deployment frequency, reduce lead time, and increase release confidence.
- Design and enforce secure networking, IAM, and secrets-management strategies across environments.
- Improve observability through actionable metrics, logs, and tracing using DataDog.
- Optimize cloud costs through rightsizing, autoscaling, and architectural improvements.
- Implement disaster recovery, backup, and multi-region resilience strategies.
- Refactor brittle or manually managed infrastructure into automated, testable, reproducible systems.
- Introduce infrastructure tooling and architectural changes, driving adoption through documentation, workshops, and hands-on support.
- Partner with engineering teams to reduce friction in CI/CD, deployments, and cloud environments.
- Communicate technical trade-offs across engineering and product stakeholders.

## Requirements
- 8+ years of experience in DevOps, SRE, or infrastructure engineering.
- Experience designing and operating production Kubernetes environments at scale.
- Deep hands-on expertise with AWS infrastructure and cloud networking.
- Strong experience building and maintaining Terraform modules across large cloud environments.
- Experience owning CI/CD systems and improving DORA metrics.
- Experience leading incident response and driving effective postmortems.
- Strong understanding of distributed systems, event-driven architectures, Kafka, and PostgreSQL performance.
- Ability to modernize legacy infrastructure and eliminate manual operational toil.
- Ability to take infrastructure projects from ambiguity through production independently.
- Ability to build trust across teams while improving reliability.

## Success Metrics
- Maintain or exceed SLO/SLA targets while reducing incident frequency and duration.
- Reduce incidents caused by misconfiguration, manual processes, or infrastructure drift.
- Increase the percentage of infrastructure managed through code and automation.
- Improve cloud cost efficiency without sacrificing reliability or performance.

## Nice-to-Haves
- Application coding experience.
- Experience operating high-throughput Kafka clusters, including MSK or self-managed deployments.
- Database performance tuning with PostgreSQL and Redis.
- Autoscaling strategies for high-traffic systems.
- Service mesh technologies.
- Internal developer platforms.
- Security practices such as zero-trust networking and policy-as-code.
- Multi-region or globally distributed systems.
- Reliability frameworks including SLOs, error budgets, and chaos testing.

## Technology Stack
- **Cloud and infrastructure:** AWS, EKS, RDS, MSK, S3, Lambda, IAM, VPC
- **Containerization and orchestration:** Kubernetes, ArgoCD
- **Infrastructure as code:** Terraform
- **CI/CD:** GitHub Actions
- **Monitoring and observability:** DataDog
- **Data and messaging:** PostgreSQL, Kafka, Redis
- **Languages:** Bash, Python, TypeScript, JavaScript

## Compensation and Benefits
- Medical, dental, and vision health care plan
- 401(k) retirement plan
- Life insurance
- Flexible paid time off
- 9 paid holidays
- Family leave
- Wellness resources
- Remote work
- Remote hires must travel to Orlando, Florida at least twice per year for town halls and collaboration, in addition to orientation in Orlando.

## Similar jobs

- [Senior Platform Engineer](https://hotfix.jobs/jobs/71245322-afd8-4019-8525-65088fed1493) - Shield AI - San Diego, CA - $141k – $212k/yr
- [Senior Network Engineer](https://hotfix.jobs/jobs/d4ecdaa5-c0ab-49f3-baed-0b13deaaf6c0) - Shield AI - San Mateo, CA - $140k – $211k/yr
- [Senior Platform Engineer](https://hotfix.jobs/jobs/3317571d-7a67-4059-b23e-af1a9033cf96) - Astra - Remote - $190k – $230k/yr
- [Senior Software Engineer, Core Infra Systems](https://hotfix.jobs/jobs/6261dc33-aecd-4734-b00d-46f60f4f38c9) - Coinbase - Remote - $186k – $219k/yr
- [Senior Site Infrastructure Engineer](https://hotfix.jobs/jobs/847213bd-91b8-44f9-b450-477e4aa3714e) - Shield AI - Seattle, WA - $110k – $210k/yr

**Apply:** https://hotfix.jobs/jobs/2295b0ce-fcdd-4983-b3ff-99a6ba994d38
**Canonical:** https://hotfix.jobs/jobs/2295b0ce-fcdd-4983-b3ff-99a6ba994d38