Skip to content

Senior Software Engineer - Cloud Infrastructure

Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.

About the job

Responsibilities

  • Design, build, and operate multi-cluster Kubernetes infrastructure across compute, networking, storage, autoscaling, observability, and security.
  • Build multi-tenant platform primitives, including tenant-safe storage, shared secrets, RBAC, and workload identity.
  • Design and orchestrate secure, sandboxed execution environments for agentic and untrusted workloads.
  • Act as a solutions architect for internal platform customers, translating scaling and reliability needs into infrastructure designs and driving them to production.
  • Improve developer effectiveness by building tooling that removes infrastructure and workflow friction.
  • Debug and profile distributed-systems edge cases; own production incidents end-to-end, including incident command, postmortems, and remediation.
  • Mentor engineers and raise technical standards through design reviews and collaboration.

Requirements

  • 5+ years of experience building and operating large-scale infrastructure, platform, SRE, or DevOps systems.
  • Experience with Kubernetes or other container orchestration frameworks.
  • Deep experience with at least one major cloud provider: AWS, Google Cloud, Azure, or OCI.
  • Production-quality programming skills and expertise in at least one of Go, Python, Rust, or C++.
  • Experience with Infrastructure as Code and GitOps workflows, such as Terraform, OpenTofu, Pulumi, or Crossplane.
  • Strong cross-functional communication and collaboration skills.
  • Bachelor’s degree in Computer Science or a related field.

Nice-to-haves

  • Experience with serverless or scale-to-zero container platforms such as Knative, Cloud Run, or KEDA.
  • Expertise in sandboxing and workload isolation, including Linux namespaces, cgroups, seccomp, gVisor, Firecracker, or Kata.
  • Expertise in cluster and cloud networking, including CNI, Cilium, eBPF, NetworkPolicy, service mesh, and cross-cloud private connectivity.
  • Deep multi-cluster or multi-region Kubernetes experience supporting batch, data-processing, GPU, and agentic workloads at scale.
  • Experience with scheduling and autoscaling systems such as Karpenter, Kueue, or Volcano.
  • Platform security experience, including admission control, least-privilege IAM, workload identity, image provenance, and supply-chain hardening.
  • Incident command experience for customer-facing production systems.
  • Experience building enterprise AI infrastructure or delivering platform solutions to internal or enterprise customers.
  • Contributions to open-source infrastructure tooling.

Skills

Kubernetes, AWS, GCP, Azure, Go, Python, Rust, C++, Terraform, GitOps, Infrastructure As Code, Linux, Ebpf, RBAC, OIDC

Astra

Astra

United States

Senior Platform Engineer
$190k+/yrRemote5+ YOEDevOps / SRE

Build and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.

inKind

inKind

Austin, TX

Senior Platform Engineer
$190k+/yrRemote8+ YOEDevOps / SRE

Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.

Mercury

Mercury

San Francisco, CA
Senior Software Engineer - SRE
$190k+/yrRemote5+ YOEDevOps / SRE

Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.

Idme

Idme

McLean, VA
Senior Software Engineer – Platform & Data Infrastructure
$191k+/yrOn-site8+ YOEDevOps / SRE

Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.

Garner Health

Garner Health

United States

Senior Site Reliability Engineer
$191k+/yrRemote5+ YOEDevOps / SRE

Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.