# Staff Infrastructure Engineer

**Company:** [VGS](https://hotfix.jobs/companies/vgs)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $145k – $260k/yr
**Experience:** 8+ years
**Skills:** AWS, Terraform, Kubernetes, Amazon Eks, Docker, GitOps, Flux, Argo, GitHub Actions, Python, Go, Bash, Prometheus, Grafana, OpenTelemetry
**Posted:** 2026-09-02

> Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.

## Job Description

## Responsibilities
- Design, build, and optimize multi-region, highly available AWS infrastructure for mission-critical, high-throughput payments applications.
- Lead the evolution from manually configured environments to standardized, globally scalable infrastructure managed as code.
- Replace manual operational work with self-healing automation using GitOps, CI/CD pipelines, and infrastructure as code.
- Build end-to-end telemetry and observability to identify bottlenecks proactively.
- Own incident management and conduct blameless postmortems to improve reliability.
- Design Golden Paths that improve engineering velocity and mentor teams on effective engineering practices.
- Architect and operate high-performance, low-latency private connectivity for external customers.
- Partner with Product, Security, and Core Engineering teams on platform architecture, enablement, and adoption.

## Requirements
- 8+ years of experience owning outcomes in complex, large-scale distributed systems within mission-critical environments.
- Advanced AWS proficiency and experience using Terraform to build reproducible environments.
- Hands-on Kubernetes/EKS, Docker, and GitOps experience, including Flux, Argo, or GitHub Actions.
- Strong programming skills in Python, Go, or Bash for infrastructure automation and operational tooling.
- Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
- Understanding of cloud security, API gateways, load balancing, and network isolation.

## Nice to Have
- Experience with tokenization, payment processing, cryptology, or security products.
- BA/BS degree.
- Experience managing distributed data-streaming platforms such as Kafka/MSK.
- Database performance tuning and query optimization experience.
- Familiarity with Java and the Spring Framework.

## Similar jobs

- [Senior Staff Performance Engineer, Firefox](https://hotfix.jobs/jobs/64810199-b6d8-436c-b310-10d723f9bffc) - Mozilla - Remote - CA$149k – CA$220k/yr
- [Staff Engineer, Digital Factory Lead](https://hotfix.jobs/jobs/d5dafcaf-27b3-41d6-ac87-dc85218ea3d3) - Shield AI - Seattle, WA - $150k – $230k/yr
- [Staff Engineer, Platform & Infrastructure](https://hotfix.jobs/jobs/48b953ba-a6fb-4833-9ad0-140dfe4943dc) - Nango - Remote - $140k – $220k/yr
- [Staff Platform Engineer](https://hotfix.jobs/jobs/37cd4cd2-d007-4013-8153-5443ae70f1dd) - Nango - Remote - $140k – $220k/yr
- [Staff Cloud Engineer](https://hotfix.jobs/jobs/837c447b-ed05-4ff2-afbb-f70401e8c7e4) - Shield AI - San Diego, CA - $152k – $228k/yr

**Apply:** https://hotfix.jobs/jobs/b34c9349-da22-498a-9337-4a1f426a0b7f
**Canonical:** https://hotfix.jobs/jobs/b34c9349-da22-498a-9337-4a1f426a0b7f