Leads and scales an engineering team building a fault-tolerant managed AI platform for LLM workloads, including task queues, model management, scheduling, and agentic execution infrastructure. Requires 5+ years leading engineering teams plus depth in distributed systems, cloud-native platforms, and AI infrastructure.
215k – 260k/yr
On-site8+ YOEEngineering Management
About the role
Responsibilities
Team Leadership & Strategy
Lead, mentor, and grow a team of high-caliber software engineers.
Partner with leadership to define and execute the AI roadmap, set clear goals, and drive accountability.
Cultivate a high-performance, collaborative engineering culture grounded in technical excellence.
Technical Execution
Oversee the architecture and development of core AI services, including fault-tolerant task queues, model management systems, and cost-aware scheduling.
Ensure delivery of scalable systems capable of handling millions of API requests per second.
Deliver an AI platform capable of supporting workloads ranging from model training to agentic execution infrastructure.
Collaboration and Influence
Work cross-functionally with Product, Infrastructure, and GTM stakeholders.
Represent Engineering in strategic discussions to influence AI platform growth and customer adoption.
Promote knowledge sharing, technical mentorship, and the evolution of engineering processes.
Requirements
Leadership Experience
5+ years managing or leading high-performing engineering teams.
Ability to lead teams through ambiguity and align stakeholders on complex technical goals.
Proven success hiring, developing, and retaining talent.
Technical Depth
Hands-on experience with distributed and concurrent systems or AI infrastructure.
Deep knowledge of cloud-native environments, container orchestration, and service-oriented architectures.
Familiarity with CPU and GPU performance, inference frameworks, or LLM systems.
Product and Delivery
Comfortable owning deliverables from design through production.
Strong collaboration skills, prioritizing clarity, context, and customer impact.
Experience in fast-paced startup or growth-stage environments.
Preferred Qualifications
Background in Computer Science, Engineering, or a related technical field.
Proficiency in Python, Go, or Rust.
Experience with Kubernetes, gRPC, and observability stacks.
Familiarity with open-source AI ecosystems such as vLLM, Hugging Face, and Triton.
Benefits and Compensation
Competitive compensation and equity packages.
Restricted Stock Units.
Paid time off, paid holidays, and leave of absence programs.
Comprehensive health, dental, and vision insurance.
Employer contributions to an HSA.
Paid parental leave.
Paid life insurance and short- and long-term disability coverage.
Professional development and tuition reimbursement.
Mental health and wellness support.
Commuter benefits for parking and transit.
Cell phone stipend.
401(k) retirement plan with company match up to 4% of salary.
Volunteer time off.
Global travel insurance and emergency assistance.
Daily meals allowance.
Additional location-specific perks and programs.
Compensation range: $215,000 - $260,000 plus bonus. Restricted Stock Units are included in all offers.
Leads a team building and operating a scalable SDN control plane for multi-tenant VPC networking across large GPU fleets. The role combines hands-on distributed-systems architecture, cloud networking expertise, production reliability, and engineering people leadership.
215k – 260k/yrOn-site8+ YOEEngineering Management
Engineering Manager, Data Platform
CrusoeSan Francisco, CA
Lead and grow a team of data engineers building Crusoe's scalable data platform for AI and cloud intelligence. Define roadmap with cross-functional partners, ensure operational excellence, and drive technical architecture for data lakes, warehouses, and ETL.
215k – 260k/yrOn-site7+ YOEEngineering Management
Senior Engineering Manager
CreditgenieNew York, NY +3
Lead a backend engineering team at a fintech startup as a hands-on Senior Engineering Manager. Set technical direction for scalable financial systems, write production code, mentor engineers, and drive operational excellence in a fast-paced environment.
215k – 275k/yrOn-site8+ YOEEngineering Management
Engineering Manager, Telemetry Agent and Edge
CrusoeSan Francisco, CA
Lead a team of 4-6 engineers building and operating Crusoe's telemetry agent for metrics/logs from hosts and GPUs. Own delivery of next-gen agent releases against fixed deadlines while hiring, coaching staff-level engineers, and maintaining high operational standards for fleet-wide software.
215k – 260k/yrOn-site7+ YOEEngineering Management
Software Engineering Manager, Database
LangChainSan Francisco, CA
Hands-on Engineering Manager to lead a small systems team building SmithDB, LangChain's purpose-built storage and query layer for AI observability at massive scale. Write production Rust, drive architecture and performance, manage the team, and own the technical roadmap.