# Senior Staff Infrastructure Engineer

**Company:** [VGS](https://hotfix.jobs/companies/vgs)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $185k – $290k/yr
**Experience:** 10+ years
**Skills:** AWS, Terraform, Kubernetes, Amazon Eks, Docker, GitOps, Flux, Argo, GitHub Actions, Python, Go, Bash, Prometheus, Grafana, OpenTelemetry
**Posted:** 2026-09-02

> Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

## Job Description

## Responsibilities
- Design, build, and optimize multi-region, high-availability AWS infrastructure for mission-critical, high-throughput payments applications.
- Evolve hand-crafted environments into a standardized, globally scalable fleet managed entirely through code.
- Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and infrastructure as code.
- Build end-to-end telemetry to identify bottlenecks proactively.
- Own incident management and conduct blameless post-mortems to improve reliability.
- Partner with Product, Security, and Core Engineering teams to create standardized platform “Golden Paths.”
- Mentor engineering teams on effective, aligned development practices.
- Architect and operate high-performance, low-latency private connectivity for external customers.
- Partner with engineering teams on platform enablement and adoption.

## Requirements
- 10+ years of experience owning outcomes in complex, large-scale distributed systems within mission-critical environments.
- Advanced AWS experience, including Terraform for reproducible infrastructure.
- Hands-on experience with Kubernetes, including EKS, Docker, and GitOps workflows.
- Experience with Flux, Argo, and GitHub Actions.
- Strong coding skills in Python, Go, or Bash for infrastructure automation and operational tooling.
- Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
- Understanding of cloud security, API gateways, load balancing, and network isolation.

## Nice-to-haves
- Experience with tokenization, payment processing, or security products.
- BA/BS degree.
- Experience managing distributed data streaming platforms such as Kafka/MSK.
- Database performance tuning and query optimization skills.
- Familiarity with Java and Spring Framework services.
- Ability to think creatively in a fast-paced startup environment.

## Similar jobs

- [Staff Infrastructure Engineer](https://hotfix.jobs/jobs/63668c94-fd42-45fb-9699-5d39d126e387) - Komodo Health - Remote - $187k – $265k/yr
- [Staff DevSecOps Engineer](https://hotfix.jobs/jobs/593e5319-afa3-4408-bf94-a4885cd3c100) - Shield AI - San Mateo, CA - $182k – $274k/yr
- [Senior Staff Lead Site Reliability Engineer](https://hotfix.jobs/jobs/c5516e9e-9df1-44ef-b77a-caf084a94fc8) - Shield AI - San Diego, CA - $190k – $280k/yr
- [Senior/Staff Kubernetes Infrastructure Engineer](https://hotfix.jobs/jobs/5276ef82-8104-4269-8e3f-7f0e02d35c2d) - Fal - Remote - $180k – $250k/yr
- [Staff Site Reliability Engineer](https://hotfix.jobs/jobs/87068331-357d-4049-9cef-b6b8edc93f79) - Attentive - Remote - $180k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/f99a2e1a-7c43-4a02-ba1c-9d76880c0acd
**Canonical:** https://hotfix.jobs/jobs/f99a2e1a-7c43-4a02-ba1c-9d76880c0acd