Senior Staff Software Engineer, Managed Platform Services
Senior technical leader anchoring distributed systems depth across Crusoe Cloud's Managed Platform Services. Owns performance engineering, operational excellence, and long-term architecture for 10x scale across all platform domains.
About the job
What You’ll Be Working On
Performance Engineering — Platform-Wide
- Own performance at the platform level
- Establish consistent benchmarks across all domains
- Identify systemic bottlenecks before they become incidents
- Drive solutions that scale
Resiliency & Long-Term Architecture
- Ensure every platform domain is architected for long-term scale
- Build in HA, disaster recovery, graceful degradation, and fault isolation from the start
- Identify one-way door decisions early
Operational Excellence
- Set the standard for how the team operates at scale
- Drive down on-call noise
- Automate anything done more than once
- Build shared patterns and frameworks
Cross-Domain Technical Leadership
- Float across platform domains as the team's senior distributed systems authority
- Provide oversight on existing, inflight, and planned systems
- Contribute to technical decisions across Control Plane, Storage & State, Edge & Agents, Data Pipeline, and Async & Metering
Cross-Team Influence
- Serve as a senior technical voice in conversations with adjacent infrastructure teams
- Build credibility across org boundaries to influence outcomes
Roadmap & Product Collaboration
- Engage directly with product and engineering leadership in the earliest stages of scoping
- Help define quarterly roadmaps
- Ensure technical decisions are tied to business outcomes
Mentorship & Culture
- Actively coach senior engineers
- Introduce systems and frameworks that uplevel the entire team
- Identify early signs of burnout
- Create an environment where it is safe to fail
What You’ll Bring to the Team
Distributed Systems Depth
- Deep, hands-on expertise designing and operating distributed systems at scale
- Experience with sharding, replication, consensus, load balancing, and concurrency
Performance Engineering
- Proven track record of impact at the staff level and above
- Proficiency in Go or a modern compiled language (Go strongly preferred)
- Benchmark, profile, and fix issues
Operational Leadership
- On-call experience on a customer-facing team required
- Build runbooks, tune alerts, and reduce noise systematically
Cross-Functional Impact
- Proven ability to influence technical decisions made by other teams
- Resolved root causes of endemic problems that span org boundaries
Roadmap Collaboration
- Experience working directly with product and engineering leadership to scope ambiguous problems
- Plan incremental delivery and tie technical decisions to business outcomes
Mentorship at Scale
- Build systems that make other engineers better without direct involvement
Ownership
- Strong sense of ownership and accountability
- Proactive, solutions-oriented
Communication
- Exemplary communication skills
- Provide the right level of context, stay concise, check for understanding
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
Skills
Go, Distributed Systems, Performance Engineering, Sharding, Replication, Consensus, Load Balancing, Concurrency, High Availability, Disaster Recovery, On-Call, Runbooks, Alert Tuning
Similar jobs
Engineering Management jobsLeads architecture and technical direction for a high-volume recognition platform spanning edge ingestion, real-time decisioning, identity, privacy, and operator systems. Requires 12+ years building distributed production systems and deep expertise in Scala or Java, streaming architectures, and cross-functional technical leadership.
Staff Software Engineer leading technical strategy, complex AI-enabled systems, engineering mentorship, and privacy-focused practices. Requires 6+ years of software engineering experience, team leadership, and familiarity with AWS, Kubernetes, AI coding agents, data security, and HIPAA compliance.
Leads the engineering organization and technical strategy for Pinterest’s AI foundations, proactive experiences, and assistant capabilities. The role requires senior leadership experience managing managers, strong AI/platform depth, cross-functional execution, and a track record of building scalable, high-performing teams.
Leads a small team responsible for the architecture, reliability, security, automation, and delivery of Anthropic’s global corporate campus and edge networks. The role requires 10+ years of enterprise networking experience, people management, and deep expertise in network operations and infrastructure.
Leads a mechanical engineering team developing and integrating flight- and mission-critical aircraft mechanisms and actuation systems for a military VTOL UAV. Requires a mechanical engineering degree, 10+ years of aerospace mechanisms experience, and demonstrated team leadership.