Senior Staff Software Engineer, Managed Orchestration
Leads architecture and development of scalable managed Kubernetes and AI orchestration systems, providing technical direction for cloud infrastructure reliability and performance. Requires 10+ years in software engineering with deep expertise in Go, Kubernetes, and large-scale systems.
About the job
What You'll Be Working On
- Drive the development of scalable, resilient, and high-performance software solutions, ensuring alignment with and influence over the strategic objectives outlined in the Crusoe Cloud roadmap
- Provide technical leadership across multiple teams, fostering a culture of innovation, engineering excellence, and accountability while enabling teams to deliver cutting-edge cloud solutions
- Define and evolve architectural standards and best practices, ensuring consistency, scalability, and long-term maintainability across systems
- Continuously stay ahead of emerging trends and technologies in cloud software, proactively shaping Crusoe's technical direction and incorporating innovations that maintain competitive advantage
- Act as a mentor and multiplier for engineering talent, elevating team capabilities through coaching, design reviews, and thought leadership in technical discussions
- Lead cross-functional initiatives and drive alignment between engineering, product, and infrastructure teams to deliver cohesive and impactful solutions
What You'll Bring to the Team
- 10+ years of experience working in software engineering, with deep expertise in Systems Engineering and large-scale distributed systems
- 3+ years of programming experience in GoLang, with a track record of delivering production-grade systems
- Extensive experience with Kubernetes and Linux Engineering, including advanced debugging and performance optimization
- Highly skilled in infrastructure as code and have a strong understanding of complex systems-level challenges at scale
- Experience with Terraform and GCP (preferred), with the ability to influence platform-level decisions
- Strong understanding of Argo, CI/CD, and Automated Testing pipelines, including designing and scaling them for large organizations
- Can architect, build, and evolve Kubernetes operators and controllers, owning critical components that ensure the reliability, scalability, and efficiency of the Kubernetes environment
- Experience designing and operating large-scale systems comparable to leading services like Google Kubernetes Engine (GKE) and Amazon Elastic Kubernetes Service (EKS)
- Can lead and deliver critical, high-impact projects, driving initiatives across networking, quality control, automation, and system reliability at an organizational level
- Can define and own system architecture end-to-end, including CI/CD pipelines, ensuring scalability, security, and long-term sustainability
- Exceptional communication skills, with the ability to influence technical and non-technical stakeholders and drive alignment across the organization
Compensation
Compensation will be paid in the range of up to $237,600 - $288,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.
Skills
Go, Kubernetes, Linux, Terraform, GCP, Argo, CI/CD, Kubernetes Operators, Infrastructure As Code, Distributed Systems
Similar jobs
DevOps / SRE jobsOwns and scales production cloud infrastructure across Kubernetes/EKS, AWS, Terraform, CI/CD, networking, and observability. The role requires 8+ years of infrastructure experience, strong Kubernetes operations expertise, and depth in reliability or scaling challenges.
Leads the technical direction, design, and operation of large-scale multi-cloud network infrastructure, with a focus on connectivity, reliability, performance, and cost efficiency. Requires deep BGP and software-defined networking expertise plus strong software development and production operations experience.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Owns deployment, CI/CD, and integration-testing automation for large-scale multi-node GPU and CPU clusters. The role requires 12+ years of experience, strong Python or Bash skills, and expertise across Linux, Kubernetes, configuration management, GPU ecosystems, and high-performance networking.
Leads software development for diagnostics, observability, automation, and repair tooling across large-scale GPU clusters and data center infrastructure. The role requires distributed systems and cloud-platform expertise, proficiency in Go, Python, Java, or Rust, and hands-on operational problem solving.