Staff Software Engineer
Build diagnostics, automation, observability, and repair tooling for Crusoe’s large-scale GPU fleet and data centers. The role requires software engineering expertise in distributed systems, reliability, cloud platforms, and at least one of Go, Python, Java, or Rust.
About the job
Responsibilities
- Develop deep-level diagnostics and troubleshooting for hardware faults within GPU racks and high-density compute systems.
- Build troubleshooting and automation tooling for NVIDIA A100, H200, GB200, B200, and AMD 350X/355X GPU platforms.
- Develop automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware.
- Partner with data center operations to build tooling and AI agents for managing critical environments.
- Develop post-repair validation and testing tools, including burn-in, PyTorch, and NVIDIA NCCL, to ensure system stability and performance.
- Own deployment, monitoring, and operational support for developed tooling to maximize GPU fleet availability and performance.
- Develop automation and operational tooling for facilities management, power, and direct liquid-cooling hardware systems.
- Identify problems, rapidly develop scalable solutions, and ship them.
- Set technical direction for specific projects and execute independently or collaboratively.
Requirements
- Software engineering experience.
- Expertise in distributed systems, reliability, and cloud platforms.
- Experience with Kubernetes, infrastructure as code, and Google Cloud.
- Proficiency in at least one programming language: Go, Python, Java, or Rust.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently and assist with critical or complex technical initiatives.
Nice-to-Haves
- Experience with Temporal and Kubernetes.
- Experience working directly with hardware vendors.
- Experience operating large-scale GPU fleets or hyperscale data center environments.
Compensation and Benefits
- Compensation range of $215,000–$260,000 plus bonus.
- Restricted Stock Units included in all offers.
- Health insurance options including HDHP and PPO, vision, and dental coverage.
- Employer HSA contributions.
- Paid parental leave.
- Paid life insurance and short- and long-term disability coverage.
- Teladoc.
- 401(k) with a 100% match up to 4% of salary.
- Paid time off and holidays.
- Cell phone and tuition reimbursement.
- Calm app subscription.
- MetLife Legal.
- Company-paid commuter benefit of $300 per month.
Skills
Gpu Infrastructure, Distributed Systems, Reliability Engineering, Kubernetes, Infrastructure As Code, GCP, Go, Python, Java, Rust, Temporal, PyTorch, Nvidia Nccl, AI Agents, Direct Liquid Cooling
Similar jobs
DevOps / SRE jobsLeads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Build and operate Reddit’s internet-scale observability platform across monitoring, logging, and distributed tracing. The role requires 7+ years of infrastructure or software engineering experience, distributed systems expertise, and strong Kubernetes and troubleshooting skills.
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.