Skip to content
VGSVGS

Staff Infrastructure Engineer

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.

About the job

Responsibilities

  • Design, build, and optimize multi-region, highly available AWS infrastructure for mission-critical, high-throughput payments applications.
  • Lead the evolution from manually configured environments to standardized, globally scalable infrastructure managed as code.
  • Replace manual operational work with self-healing automation using GitOps, CI/CD pipelines, and infrastructure as code.
  • Build end-to-end telemetry and observability to identify bottlenecks proactively.
  • Own incident management and conduct blameless postmortems to improve reliability.
  • Design Golden Paths that improve engineering velocity and mentor teams on effective engineering practices.
  • Architect and operate high-performance, low-latency private connectivity for external customers.
  • Partner with Product, Security, and Core Engineering teams on platform architecture, enablement, and adoption.

Requirements

  • 8+ years of experience owning outcomes in complex, large-scale distributed systems within mission-critical environments.
  • Advanced AWS proficiency and experience using Terraform to build reproducible environments.
  • Hands-on Kubernetes/EKS, Docker, and GitOps experience, including Flux, Argo, or GitHub Actions.
  • Strong programming skills in Python, Go, or Bash for infrastructure automation and operational tooling.
  • Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
  • Understanding of cloud security, API gateways, load balancing, and network isolation.

Nice to Have

  • Experience with tokenization, payment processing, cryptology, or security products.
  • BA/BS degree.
  • Experience managing distributed data-streaming platforms such as Kafka/MSK.
  • Database performance tuning and query optimization experience.
  • Familiarity with Java and the Spring Framework.

Skills

AWS, Terraform, Kubernetes, Amazon Eks, Docker, GitOps, Flux, Argo, GitHub Actions, Python, Go, Bash, Prometheus, Grafana, OpenTelemetry

Mozilla

Mozilla

Canada

Senior Staff Performance Engineer, Firefox
CA$149k+/yrRemote7+ YOEDevOps / SRE

Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.

Shield AI

Shield AI

Seattle, WA

Staff Engineer, Digital Factory Lead
$150k+/yrOn-site8+ YOEDevOps / SRE

Leads the design and deployment of AI-enabled manufacturing systems, MES, connected-factory infrastructure, and automation for aircraft production. Requires a bachelor’s degree and 8+ years of experience in digital manufacturing, industrial automation, or software-enabled operations.

Nango

Nango

United States
Staff Engineer, Platform & Infrastructure
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.

Nango

Nango

United States
Staff Platform Engineer
$140k+/yrRemote10+ YOEDevOps / SRE

Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.

Shield AI

Shield AI

San Diego, CA
Staff Cloud Engineer
$152k+/yrOn-site7+ YOEDevOps / SRE

Designs, automates, and operates AWS infrastructure, shared development environments, and container platforms. The role requires strong experience with Kubernetes, infrastructure as code, environment lifecycle automation, cloud security, compliance, and cost optimization.