Staff Software Engineer, Managed Orchestration (Managed Kubernetes)
Staff Software Engineer designs, builds, and scales managed Kubernetes and AI training clusters, focusing on reliability, performance, and orchestration using Go, Terraform, and GCP. Oversees architecture, CI/CD pipelines, and critical infrastructure projects requiring 8+ years experience.
About the job
What You'll Be Working On
- Contribute to the development of scalable and robust software solutions, closely aligning with the strategic objectives outlined in the Crusoe Cloud roadmap
- Work collaboratively with tech leads and engineers to create a dynamic environment where creativity and technical excellence are encouraged, leading to the development of cutting-edge cloud solutions
- Continuously stay abreast of the latest trends and techniques in cloud software, incorporating these insights to keep Crusoe’s offerings innovative
- Support the development of your peers by sharing knowledge and providing guidance in technical discussions
What You'll Bring to the Team
- 8+ years of experience working in software engineering, with strong experience in Systems Engineering
- 2+ years of programming experience in GoLang
- Experience with Kubernetes and Linux Engineering and debugging
- Skilled in infrastructure as code and familiar with systems-level challenges
- Experience with Terraform and GCP (preferred)
- Understand Argo, CI/CD, and Automated Testing pipelines
- Build and manage Kubernetes operators and controllers
- Develop scalable systems to compete with GKE and EKS
- Oversee critical projects with broad impact on networking, quality control, and automation
- Design system architecture, taking ownership of CI/CD pipelines, ensuring security standards
- Excellent communication skills, both verbal and written
Compensation
$220,000 - $250,000 + Bonus. Restricted Stock Units included.
Skills
Go, Kubernetes, Linux, Terraform, GCP, Argo, CI/CD, Kubernetes Operators, GKE, EKS
Similar jobs
DevOps / SRE jobsLeads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.