Engineering Manager, AI Platform
Leads the AI Platform organization, owning machine-learning infrastructure and agent runtime capabilities used across product and mission teams. Requires engineering management experience, deep hands-on machine learning expertise, production agentic AI experience, and proficiency with Python, relational databases, cloud services, and distributed systems.
About the job
Responsibilities
- Lead the AI Platform organization, setting technical direction, operating model, and roadmap for enrichments and the agents platform.
- Build standard pipelines, shared services, and operational patterns for reliable, efficient data and model enrichments.
- Scale durable agent execution capabilities, including orchestration, state and workflow management, reliability, observability, evaluation, and quality controls.
- Define interfaces, reusable primitives, and engineering practices that enable product teams to build AI-enabled capabilities.
- Establish measurement, testing, evaluation, and feedback-loop practices that improve agent correctness, safety, and usefulness in production.
- Partner with Product, Design, Mission, Security, and Infrastructure teams to prioritize platform investments and deliver outcomes.
- Balance velocity, reliability, security, and operational risk for AI systems supporting real-world missions.
- Recruit, mentor, and develop engineers and engineering leaders.
Requirements
- 3+ years of experience as an engineering manager, including leading, coaching, and building software engineering teams.
- 5+ years of hands-on machine learning experience.
- 2+ years building agentic AI systems in production.
- Experience building or operating platforms used by multiple internal product or engineering teams.
- Deep familiarity with production AI/ML or agentic systems, including orchestration, evaluation, observability, quality, reliability, and cost.
- Experience leading teams that build workflow, data, ML, or service platforms with strong operational standards.
- Ability to balance speed, reliability, security, and mission impact while leading through ambiguity.
- Strong prioritization, execution, communication, cross-functional leadership, and project ownership skills.
- Data-driven decision-making skills.
- Proficiency with Python, PostgreSQL or other relational databases, AWS or other cloud services, and distributed-system and workflow-engineering patterns.
- Willingness to travel to offsites, team syncs, and customer sites.
- U.S. citizenship, required to access U.S.-only data systems.
Nice to Have
- Experience with LangGraph, Temporal, Agno, or similar agent frameworks and workflow orchestration systems.
- Experience with ML platforms, data pipelines, model evaluation, retrieval, or enrichment systems.
- Experience building secure, observable, highly reliable systems in regulated or mission-critical environments.
- Active security clearance or willingness to obtain one.
- Experience building software for defense, national security, or intelligence use cases.
Compensation and Benefits
- Health, dental, and vision insurance
- Remote-friendly work with WeWork access
- Unlimited PTO, federal-holiday shared downtime, and company-wide year-end time off
- 401(k) match
- Lifestyle and wellbeing stipends
- Salary top-up during military reserve duty
- Fully paid parental leave
- Child and pet care reimbursement during travel
Skills
Machine Learning, Agentic AI, Python, Postgres, AWS, Distributed Systems, Workflow Orchestration, LangGraph, Temporal, Agno, Ml Platforms, Data Pipelines, Model Evaluation, Observability, Retrieval Systems
Similar jobs
Engineering Management jobsLeads a remote-first team of SDETs responsible for product quality, test strategy, automation, and release reliability across SaaS and customer-managed environments. Requires 10+ years of industry experience, technical leadership, and strong expertise in testing, CI/CD, observability, and quality metrics.
Leads multiple service-infrastructure engineering teams and managers, shaping architecture, developer platforms, reliability, and cross-functional delivery. Requires substantial management experience, including managing managers, critical distributed systems, incident response, and geographically distributed teams.
Leads and develops the App Traffic engineering team building reliable, scalable service-mesh and networking infrastructure across multiple clouds. Requires 9+ years of software engineering experience, including engineering leadership and distributed-systems or infrastructure expertise.
Leads the mortgage engineering organization, owning platform architecture, delivery, business-line outcomes, and team development. Requires senior engineering management experience, extensive software engineering experience, large-team leadership, business ownership, and expertise in scalable systems and AI.
Leads a hands-on Shared Services Engineering team building and operating reusable services, SDKs, APIs, and customer-facing systems. Requires 7+ years of software engineering experience, engineering management experience, and strong technical judgment across distributed and full-stack systems.