Senior Software Engineer - Cloud Infrastructure
Build and operate multi-cloud, multi-cluster infrastructure and platform primitives for large-scale simulations and enterprise AI workloads. The role requires 5+ years in infrastructure, platform, SRE, or DevOps systems, strong Kubernetes and cloud expertise, production programming skills, and Infrastructure as Code experience.
About the job
Responsibilities
- Design, build, and operate multi-cluster Kubernetes infrastructure across compute, networking, storage, autoscaling, observability, and security.
- Build multi-tenant platform primitives, including tenant-safe storage, shared secrets, RBAC, and workload identity.
- Design and orchestrate secure, sandboxed execution environments for agentic and untrusted workloads.
- Act as a solutions architect for internal platform customers, translating scaling and reliability needs into infrastructure designs and driving them to production.
- Improve developer effectiveness by building tooling that removes infrastructure and workflow friction.
- Debug and profile distributed-systems edge cases; own production incidents end-to-end, including incident command, postmortems, and remediation.
- Mentor engineers and raise technical standards through design reviews and collaboration.
Requirements
- 5+ years of experience building and operating large-scale infrastructure, platform, SRE, or DevOps systems.
- Experience with Kubernetes or other container orchestration frameworks.
- Deep experience with at least one major cloud provider: AWS, Google Cloud, Azure, or OCI.
- Production-quality programming skills and expertise in at least one of Go, Python, Rust, or C++.
- Experience with Infrastructure as Code and GitOps workflows, such as Terraform, OpenTofu, Pulumi, or Crossplane.
- Strong cross-functional communication and collaboration skills.
- Bachelor’s degree in Computer Science or a related field.
Nice-to-haves
- Experience with serverless or scale-to-zero container platforms such as Knative, Cloud Run, or KEDA.
- Expertise in sandboxing and workload isolation, including Linux namespaces, cgroups, seccomp, gVisor, Firecracker, or Kata.
- Expertise in cluster and cloud networking, including CNI, Cilium, eBPF, NetworkPolicy, service mesh, and cross-cloud private connectivity.
- Deep multi-cluster or multi-region Kubernetes experience supporting batch, data-processing, GPU, and agentic workloads at scale.
- Experience with scheduling and autoscaling systems such as Karpenter, Kueue, or Volcano.
- Platform security experience, including admission control, least-privilege IAM, workload identity, image provenance, and supply-chain hardening.
- Incident command experience for customer-facing production systems.
- Experience building enterprise AI infrastructure or delivering platform solutions to internal or enterprise customers.
- Contributions to open-source infrastructure tooling.
Skills
Kubernetes, AWS, GCP, Azure, Go, Python, Rust, C++, Terraform, GitOps, Infrastructure As Code, Linux, Ebpf, RBAC, OIDC
Similar jobs
DevOps / SRE jobsBuild and operate core platform infrastructure, developer tooling, CI/CD, observability, and cloud reliability systems for a regulated payments platform. Requires 5+ years of infrastructure or backend experience, strong infrastructure-as-code skills, and production cloud expertise.
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.
Senior SRE who embeds with product teams to improve reliability, observability, performance, and incident preparedness. The role requires SRE or DevOps experience, strong PostgreSQL and Temporal expertise, and familiarity with observability platforms and OpenTelemetry.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.