Senior Software Engineer, Managed Orchestration (Managed Kubernetes)
Build and scale managed Kubernetes and AI training clusters, developing operators, controllers, and infrastructure using Go, Terraform, and GCP. Design reliable, high-performance systems competing with GKE/EKS, with 5+ years experience required.
About the job
What You'll Be Working On
- Contribute to the development of scalable and robust software solutions, closely aligning with the strategic objectives outlined in the Crusoe Cloud roadmap
- Work collaboratively with tech leads and engineers to create a dynamic environment where creativity and technical excellence are encouraged, leading to the development of cutting-edge cloud solutions
- Continuously stay abreast of the latest trends and techniques in cloud software, incorporating these insights to keep Crusoe's offerings innovative
- Support the development of your peers by sharing knowledge and providing guidance in technical discussions
What You'll Bring to the Team
- 5-7 years of experience working in software engineering, with strong experience in Systems Engineering
- 2+ years of programming experience in GoLang
- Experience with Kubernetes and Linux Engineering and debugging
- Skilled in infrastructure as code and familiar with systems-level challenges
- Experience with Terraform and GCP (preferred)
- Understand Argo, CI/CD, and Automated Testing pipelines
- Build and manage Kubernetes operators and controllers
- Develop scalable systems to compete with GKE and EKS
- Oversee critical projects with broad impact, leading initiatives focused on networking, quality control, and automation
- Design system architecture, taking ownership of system architecture, including CI/CD pipelines, while ensuring adherence to security standards
- Excellent communication skills, both verbal and written
Compensation & Benefits
- Compensation range: $180,000 - $210,000 + Bonus
- Restricted Stock Units included
- Paid time off & paid holidays
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
Skills
Go, Kubernetes, Linux, Terraform, GCP, Argo, CI/CD, Kubernetes Operators, GKE, EKS
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.