Skip to content
VGSVGS

Senior Staff Infrastructure Engineer

Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.

About the job

Responsibilities

  • Design, build, and optimize multi-region, high-availability AWS infrastructure for mission-critical, high-throughput payments applications.
  • Evolve hand-crafted environments into a standardized, globally scalable fleet managed entirely through code.
  • Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and infrastructure as code.
  • Build end-to-end telemetry to identify bottlenecks proactively.
  • Own incident management and conduct blameless post-mortems to improve reliability.
  • Partner with Product, Security, and Core Engineering teams to create standardized platform “Golden Paths.”
  • Mentor engineering teams on effective, aligned development practices.
  • Architect and operate high-performance, low-latency private connectivity for external customers.
  • Partner with engineering teams on platform enablement and adoption.

Requirements

  • 10+ years of experience owning outcomes in complex, large-scale distributed systems within mission-critical environments.
  • Advanced AWS experience, including Terraform for reproducible infrastructure.
  • Hands-on experience with Kubernetes, including EKS, Docker, and GitOps workflows.
  • Experience with Flux, Argo, and GitHub Actions.
  • Strong coding skills in Python, Go, or Bash for infrastructure automation and operational tooling.
  • Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.
  • Understanding of cloud security, API gateways, load balancing, and network isolation.

Nice-to-haves

  • Experience with tokenization, payment processing, or security products.
  • BA/BS degree.
  • Experience managing distributed data streaming platforms such as Kafka/MSK.
  • Database performance tuning and query optimization skills.
  • Familiarity with Java and Spring Framework services.
  • Ability to think creatively in a fast-paced startup environment.

Skills

AWS, Terraform, Kubernetes, Amazon Eks, Docker, GitOps, Flux, Argo, GitHub Actions, Python, Go, Bash, Prometheus, Grafana, OpenTelemetry

Komodo Health

Komodo Health

United States

Staff Infrastructure Engineer
$187k+/yrRemote8+ YOEDevOps / SRE

Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.

Shield AI

Shield AI

San Mateo, CA
Staff DevSecOps Engineer
$182k+/yrOn-site7+ YOEDevOps / SRE

Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.

Shield AI

Shield AI

San Diego, CA

Senior Staff Lead Site Reliability Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Attentive

Attentive

United States

Staff Site Reliability Engineer
$180k+/yrRemote7+ YOEDevOps / SRE

Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.