Senior Engineering Manager, Managed Platform Services
Lead the Command Center Insights & Actions team building observability, alerting, and automated remediation systems for Crusoe's AI cloud infrastructure. Own roadmap, mentor engineers, and drive technical excellence in a high-scale environment.
About the job
What You'll Be Working On
- Drive the Insights & Actions Roadmap: Own and execute across alerting infrastructure, control plane APIs, automated action systems, and telemetry-derived insights such as straggler node detection and GPU profiling.
- Influence Strategic Roadmaps: Contribute significantly to the team's roadmap, impacting long-term team goals and operational performance metrics.
- Refine Early Product Requirements: Collaborate with product and engineering leadership to bring clarity to ambiguous problems early in the scoping process.
- Collaborate Cross-Functionally: Partner with product, design, and engineering teams inside and outside the organization to align on goals and deliver integrated solutions.
- Manage Complex Projects: Lead critical initiatives involving multiple engineers, including those outside your direct report structure, ensuring customer outcomes are auditable and decisions are data-driven.
- Drive Technical Excellence: Champion process improvements, operational excellence, and best practices across the team.
- Cultivate Team Growth: Coach and mentor engineers from new grad to Staff level, setting clear performance expectations and defining career paths to build a high-performing, sustainable team.
What You'll Bring to the Team
- Technical Expertise in Observability & Intelligence Systems: Hands-on background in ML, heuristics, or rule-based systems — with the ability to engage deeply on problems like anomaly detection, threshold design, and automated remediation logic.
- Proven Leadership: Demonstrated track record of people management, leading with empathy, and maintaining a sustainable workload for your teams.
- Technical Acumen: Ability to lead effectively in spaces where problems, opportunities, and strategies are not yet fully defined — driving clarity, direction, and execution.
- Cross-Functional Collaboration: Excellent technical communication skills, both verbal and written, to work effectively across diverse roles and functions.
- Project Ownership: Proven experience owning and delivering complex projects end-to-end, with measurable quality and data-driven decision-making.
- Global Scale Experience: Background building and operating global services at scale.
- Organizational Prowess: Highly organized and capable of managing multiple complex initiatives and team priorities in parallel.
Bonus Points
- Background in data platforms and data science
- Background in observability platforms or products
- Familiarity with GPU profiling tools (Nsight, NCCL Inspector) or infrastructure diagnostics at the hardware layer
- Highly motivated and proactive in identifying process improvements and boosting team efficiency
- Passion for coaching and mentoring engineers into high-performing individuals
- Enthusiasm for building team culture with a high quality of life for engineers
- A true "people-person" who thrives in collaborative environments and is energized by teamwork
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Skills
Observability, Machine Learning, Heuristics, Anomaly Detection, Automated Remediation, Alerting Infrastructure, Telemetry, Gpu Profiling, Control Plane Apis, People Management
Similar jobs
Engineering Management jobsLeads a hands-on product engineering team building and operating polished, scalable software products. Requires 10+ years of software engineering experience, including 5+ years managing teams, with depth in full-stack development, distributed systems, reliability, and product quality.
Leads a team of Forward Deployed Engineers delivering and integrating physical AI systems for defense and government customers. Requires 8+ years of industry experience, recent technical leadership, production deployment experience, broad software and hardware expertise, and U.S. security-clearance eligibility.
Leads and grows the ML Platform team responsible for training, evaluating, and serving production machine-learning models. The role requires 8+ years of engineering experience, 4+ years managing engineering teams, strong ML systems expertise, and customer-facing technical leadership.
Leads and scales engineering teams responsible for high-throughput social-data systems, real-time AI agents, analytics, and automation. The role combines people management, product execution, architecture oversight, delivery excellence, hiring, and cross-functional collaboration.
Leads a mixed software and ML engineering team building simulation, scenario authoring, generative modeling, and agentic tools for autonomous vehicle validation. The role owns technical direction, cross-functional delivery, team development, and generative scenario creation at company scale.