Head of Engineering, AI Platform
Leads and builds Starburst’s AI agent platform in India, owning its roadmap, architecture, production operations, and evolution from internal infrastructure to enterprise-facing products. Requires 10+ years of engineering experience, team leadership, platform development, and production LLM systems expertise.
About the job
Responsibilities
- Own the agent platform roadmap, from internal foundations through customer-facing capabilities, and deliver it.
- Hire, grow, and set the technical bar for a senior engineering team in India while contributing directly to the code.
- Make architecture decisions for the agent SDK, evaluation frameworks, guardrails, model serving, and observability.
- Evolve internal platform capabilities toward secure, multi-tenant, customer-facing products with stable APIs.
- Run the platform in production, including on-call rotation, incident response, and root-cause analysis.
- Operate autonomously across time zones, documenting decisions and aligning with distributed counterparts.
- Collaborate with product, agent, research, and engineering leaders across the company.
- Engage directly with enterprise customers when platform capabilities are involved.
Requirements
- Typically 10+ years of engineering experience, including 3+ years leading engineering teams; demonstrated evidence is weighted over tenure.
- Experience hiring and growing engineering teams through a full delivery cycle.
- Experience building platform or infrastructure products serving multiple engineering teams.
- Experience running LLM systems in production, including LLM operations, timeout and retry design, token-cost management, latency budgets, capacity planning, and model failure handling.
- Product-oriented approach, including direct engagement with platform users and adapting based on feedback.
- Ability to work across Python and JVM ecosystems.
Nice to Have
- Experience turning an internal platform into a customer-facing commercial product.
- Experience with agentic systems, evaluation methodology, or large-scale LLM serving.
- Experience with data systems, SQL engines, or analytics products.
- Experience operating multiple models behind a single platform, including proprietary APIs, open-weight models, routing, gateways, per-model behaviors, and regulated-environment deployments.
- Active presence in the ML/AI community in India through meetups, conferences, open source, or writing.
Compensation and Benefits
- Competitive pay and total rewards.
- Stock grants.
- Flexible paid time off.
- Additional employee benefits.
Skills
Python, Jvm, Llm Operations, Agent Sdks, Evaluation Frameworks, Guardrails, Model Serving, Observability, Multi-Tenancy, SQL, Data Systems, Model Routing
Similar jobs
Engineering Management jobsLeads India-based Site Reliability Engineering teams responsible for Okta’s cloud platform, databases, networking, Kubernetes, CI/CD, observability, FinOps, and automation. Requires 16+ years in infrastructure or SRE, substantial people-management experience, and expertise in AWS, Kubernetes, Terraform, and reliable SaaS operations.
Leads the engineering organization responsible for scaling GitLab.com through cell-based architecture, customer migrations, routing, and multi-cloud infrastructure. The role requires director-level people leadership, distributed-systems expertise, large-scale database and migration experience, and strong asynchronous communication.
Leads a team of Partner Solution Architects across the APAC partner ecosystem, driving strategic relationships, technical alignment, joint solutions, and business growth. Requires extensive partner-management or solution-architecture experience, substantial leadership experience, and a bachelor’s degree or equivalent.
Leads the organization and strategy for Databricks’ large-scale data infrastructure, including billing correctness, disaster recovery, reliability, deployment automation, and data integration. The role requires extensive distributed-systems experience, infrastructure leadership, and experience managing managers.
Leads engineering organization and developer-platform initiatives, including scalable cloud services, AI-assisted development, technical strategy, hiring, and leadership development. Requires 15+ years building distributed systems and experience managing high-performance engineering teams.