Senior Software Engineer on the Developer Experience team building internal tools, libraries, CI/CD pipelines, and paved paths to accelerate engineering productivity and eliminate toil across the full SDLC at Crusoe. Requires strong Go, Kubernetes, DevOps/SRE background and experience creating developer infrastructure.
172k – 209k/yr
On-site5+ YOEDevOps / SRE
About the role
What You’ll Be Working On
Engineering Acceleration: Partner with the broader organization to build paved paths that engineers love to use. Our paved paths are fast, reliable and broadly adopted across all of engineering.
Toil Elimination: Ruthlessly identify and automate away the friction and repetitive tasks that slow down the development process, keeping Crusoe engineers in a state of high productivity.
Library & Environment Creation: Develop the libraries, tools, and pre-production environments necessary for vetting service APIs and complex microservice interactions.
Ecosystem Integration: Unify internal tooling and vendor services to automate workflows, build operational efficiency, and optimize security across the development stack.
Lifecycle Innovation: Innovate across every stage of the development lifecycle, including source code management, build systems, code review, CI/CD pipelines, platform runtimes, and telemetry.
Culture of Quality: Lead efforts to establish a culture of continuous quality delivery that scales seamlessly as our engineering headcount and infrastructure grow.
System Optimization: Work diligently to build efficient systems and processes that serve as a force multiplier for the impact of every engineer around you.
What You’ll Bring to the Team
Previous experience building developer tools and/or infrastructure for engineering teams.
Expert ability to evaluate technical tradeoffs and understand how infrastructure decisions impact the daily productivity of the end-user (the developer).
Fluent knowledge of industry-standard AI tooling, build tooling, containerization, and open-source development frameworks.
A demonstrated passion for building empathetic developer and operator workflows that prioritize human productivity.
Professional experience managing or developing within Kubernetes clusters and a deep understanding of container orchestration.
Proven experience in DevOps, Site Reliability Engineering (SRE), Release Engineering, or a similar productivity-focused discipline.
Deep understanding of automated testing infrastructure and how to integrate it into a seamless CI/CD pipeline.
Expertise in modern programming languages (specifically Go) and advanced proficiency in Git-based workflows (GitLab/GitHub).
A Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, or a related analytical field (or equivalent professional experience).
Bonus Points
Active involvement in the open-source community or a track record of staying current with recent industry advancements in developer productivity.
A background in solving complex, multi-layered technical problems and then successfully automating the resulting solutions.
Benefits
Competitive compensation
Restricted Stock Units
Paid time off & paid holidays
Comprehensive health, dental & vision insurance
Employer contributions to HSA account
Paid parental leave
Paid life insurance, short-term and long-term disability
Professional development & tuition reimbursement
Mental health & wellness support
Commuter benefits (parking & transit)
Cell phone stipend
401(k) Retirement plan with company match up to 4% of salary
Volunteer time off
Compensation Range: $172425 - $209000 + Bonus. Restricted Stock Units are included in all offers.
Validates large-scale multi-node GPU clusters using QEMU and Cloud Hypervisor, focusing on interconnects like NVLink/InfiniBand, collective communications (NCCL/RCCL), and performance in virtualized AI/HPC environments. Requires 5+ years experience, virtualization expertise, and Linux kernel knowledge.
173k – 210k/yrOn-site5+ YOEDevOps / SRE
Senior Production Engineer, Operational Excellence
CrusoeSan Francisco, CA +1
Senior Production Engineer ensures reliability, scalability, and performance of GPU cloud infrastructure powering AI workloads. Drives observability, incident response, automation, and operational improvements in large-scale distributed systems.
172k – 209k/yrOn-site5+ YOEDevOps / SRE
Senior Production Engineer, Compute
CrusoeSunnyvale, CA
Senior Production Engineer responsible for optimizing Crusoe's virtualization, hypervisor, and Linux kernel stack to deliver high-performance AI and HPC compute infrastructure. Requires 5+ years experience with kernel internals, KVM/QEMU, low-level debugging, and performance tuning for GPUs and DPUs.
170k – 205k/yrOn-site5+ YOEDevOps / SRE
Senior Production Engineer, SDN
CrusoeSunnyvale, CA
Senior Production Engineer focused on building automation, self-healing tools, and reliability for Crusoe's SDN infrastructure that powers AI and HPC workloads. Requires 5+ years experience automating network provisioning with deep expertise in SDN platforms, Linux networking, Kubernetes CNIs, and protocols like BGP.
170k – 205k/yrOn-site5+ YOEDevOps / SRE
Senior Production Engineer, Storage
CrusoeSunnyvale, CA
Build and optimize distributed, fault-tolerant cloud storage systems (block, file, object) for Crusoe's AI/HPC infrastructure. Ensure high availability, performance, and reliability through automation, incident response, and collaboration with hardware/kernel teams. Requires 5+ years in storage engineering/SRE with deep Linux and IaC expertise.