Engineering Manager - Forward Deployed Engineering (LLM)
Leads and mentors Forward Deployed Engineers building and optimizing LLM inference for customers, while hands-on contributing to product features and customer engagements. Requires 4+ years software engineering with Python/ML expertise and 1+ year leadership.
About the job
Responsibilities
Leadership & Team Management
- Lead, mentor, and grow a team of Forward Deployed Engineers, providing guidance on technical direction, project execution, and professional development.
- Set clear goals and ensure timely, high-quality delivery across multiple customer-facing projects involving LLM deployment and inference optimization.
- Collaborate with leadership to align team priorities with company and customer goals, balancing short-term delivery, widely varying customer priorities, and long-term technical initiatives.
- Player-coach – be a key driver on strategic product initiatives and customer engagements.
Technical Ownership
- Develop and maintain software systems and product features using general-purpose programming languages (Python preferred).
- Drive customer impact by designing, implementing, and deploying Baseten solutions end-to-end (problem framing → evaluation → production deployment → monitoring).
- Deliver with velocity: turn vague objectives into clear specs and well-defined PoCs.
- Optimize and enhance AI/ML projects, contributing to continuous improvement of technical stack.
- Own products and customer projects end-to-end, functioning as engineer, project manager, and product manager.
Requirements
- Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or related field.
- 4+ years of professional software engineering experience, including 1+ year in a leadership or mentorship capacity.
- Strong programming skills in Python, with production experience in building or optimizing ML inference systems.
- Proven experience with LLMs, inference optimization, or serving frameworks (e.g., vLLM, TensorRT, Triton, Hugging Face, Ray Serve).
- Familiarity with observability, profiling, and cost/performance tradeoffs in production ML systems.
- Excellent communication and collaboration skills.
Bonus Points
- Experience leading customer-facing engineering teams or working directly with enterprise partners.
- Deep understanding of GPU infrastructure, distributed inference, or model compression techniques.
Benefits
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents.
- Flexible PTO policy including company wide Winter Break.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups.
Skills
Python, LLMs, vLLM, TensorRT, Triton, Hugging Face, Ray Serve, Gpu Infrastructure, Ml Inference, Distributed Inference
Similar jobs
Engineering Management jobsLeads the Web Infrastructure team responsible for Notion’s web client architecture, performance, reliability, and shared design systems. The role manages senior engineers and managers, drives execution and planning, and partners across the organization on technical and organizational practices.
Leads and scales a customer-facing Applied AI Engineering team serving high-growth startups, guiding customers from experimentation to production while building repeatable deployment mechanisms. Requires technical depth in AI/ML platforms and experience leading teams in startup-focused, ambiguous environments.
Leads multiple engineering teams, combining people management with hands-on technical leadership, architecture, and coding. The role requires at least four years of software engineering experience and a demonstrated ability to build high-performing teams in complex startup environments.
Leads and grows a research engineering team developing and productionizing conversational AI models and agent systems. The role requires substantial machine-learning systems experience, foundation-model expertise, and people-management experience.
Leads architecture and development of machine-learning and generative-AI systems for network analysis, including agentic workflows and conversational interfaces. Manages and mentors ML engineers while coordinating cross-functional delivery, production support, and roadmap execution.