Software Engineer, Compute Infrastructure
Build and own core compute infrastructure for Render's cloud platform, including Kubernetes clusters on hyperscalers and bare metal. Design, scale, debug, and optimize large-scale orchestration, scheduling, and distributed systems with deep Kubernetes and systems expertise.
About the job
What You'll Do
- Own Render's core compute infrastructure across multiple cloud providers, regions, and data centers. Shape how the compute platform evolves as the company rapidly scales.
- Design and build capabilities that give users greater performance and flexibility in how their services are built, deployed, perform, and stay available even when underlying resources go down.
- Investigate challenging cloud and compute issues across the stack, from the kernel and data plane to the Kubernetes cluster, control plane, and other orchestration mechanisms.
- Improve the performance and reliability of the infrastructure through systematic profiling, experimentation, and tuning.
- Partner with engineers across the company to build a platform that is stable, predictable, and secure.
- Participate in the on-call rotation. Help continuously improve how incidents are detected, responded to, and learned from.
What We're Looking For
- At least 7 years of experience building and operating large-scale platform or compute infrastructure.
- Deep expertise in operating, scaling, and enhancing Kubernetes clusters or similar resource/container orchestration systems.
- Experience developing in Go, Rust, or similar languages to develop custom infrastructure components, scheduling, controllers that apply business logic to resource management.
- Comfort going broad and deep in complex systems, making tradeoffs to improve performance and efficiency without sacrificing reliability.
- Strong experience designing, debugging, and operating distributed systems.
- Experience planning and executing rapid, high-risk upgrades and changes with minimal downtime to user services.
Nice-to-Haves
- Background in virtualization technologies like Firecracker, gVisor, Kata, or similar.
- Experience optimizing the performance of node, pod, and container spin-up times.
- Familiarity with eBPF, Linux kernel internals, resource management.
- Comfort securing and isolating workloads in multi-tenant execution environments.
Skills
Kubernetes, Go, Rust, Distributed Systems, Linux Kernel, Ebpf, Firecracker, Gvisor, Kata, Controllers, Operators
Similar jobs
DevOps / SRE jobsOwn foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.