Staff Software Engineer - Container Platform
Staff-level engineer owning design and operations of Snowflake's large-scale Kubernetes container platform across AWS, Azure, and GCP. Focus on reliability, automation, and developer experience for internal engineering teams.
About the job
What You'll Do
- Own the design and delivery of large, complex platform initiatives spanning cluster lifecycle management, multi-cloud automation, and internal developer tooling.
- Identify and drive cross-team technical improvements across the platform, from architecture through adoption.
- Make and defend architectural trade-offs grounded in reliability, scalability, and operational reality.
- Act as a technical anchor for the team, developing expertise in others, raising engineering standards, and being the person engineers turn to on hard problems.
- Treat internal engineers as your primary customers and measure success by their velocity and the reliability of their experience on the platform.
What We're Looking For
- Significant experience designing and operating large-scale distributed systems in production.
- Experience owning Kubernetes or similar orchestration systems in production at scale: cluster lifecycle, upgrades, and fleet management across a large heterogeneous fleet.
- Proficient in Go or a comparable systems language.
- Experience operating across multiple clouds (AWS, Azure, GCP).
Nice to Have
- Experience with GPU infrastructure or AI/ML training workloads at scale.
- Open source contributions to Kubernetes or related ecosystem projects.
Skills
Kubernetes, Go, AWS, Azure, GCP, Distributed Systems, Cluster Management, Automation, Multi-Cloud Platforms, Gpu Infrastructure
Similar jobs
DevOps / SRE jobsOwns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.