Lead Engineer - AI Studio
Leads hands-on architecture and implementation for an AI-powered React application builder, including LLM pipelines, distributed backend services, generation systems, and deployment infrastructure. Requires 6+ years of backend engineering experience, strong Go and Node/NestJS skills, and expertise in generative AI, reliability, and scalable systems.
About the job
Responsibilities
Architecture & Platform Ownership
- Own architecture and scaling decisions for core AI Studio platform components, including the App Generation Engine, React Build & Render Pipeline, Domain & Publishing Pipeline, and AI Content Systems.
- Lead cross-cutting initiatives to improve system responsiveness, generation accuracy, and platform robustness.
- Build scalable, fault-tolerant LLM pipelines for content generation, layout creation, experimentation, and AI-driven user guidance.
AI & Distributed Systems Engineering
- Build backend services with Go, NestJS, and Node across PostgreSQL, Firestore, vector databases, Cloudflare Workers, and Cloudflare KV.
- Design distributed systems for high-throughput generative and analytical workloads with strong correctness and low latency.
- Build and optimize embeddings pipelines, retrieval-augmented generation (RAG), and multi-agent orchestration frameworks.
Quality, Observability & Reliability
- Improve observability using Prometheus/Grafana, OpenTelemetry, and structured logging.
- Establish and maintain SLOs for generation latency, correctness, model safety, and system uptime.
- Strengthen resiliency and failover strategies for traffic spikes and model-load variations.
AI Guardrails, Safety & Tooling
- Implement hallucination-mitigation strategies, including code constraints, structured outputs, model verification layers, retrieval guards, and test scaffolding.
- Enforce data governance, access control, privacy boundaries, and compliance for user-generated data.
- Use LLMs and AI tools to write, test, and debug code while maintaining reliability and consistency standards.
Technical Leadership & Collaboration
- Mentor engineers through code reviews, design discussions, and pairing.
- Partner with product managers, designers, infrastructure teams, and engineering leaders on roadmap and feature delivery.
- Participate in technical reviews, deep dives, and on-call rotations.
Requirements
Core Experience
- 6+ years of backend engineering experience, including distributed system design and high-scale platform development.
- Strong proficiency in Go (Golang), with Node/NestJS experience.
- Experience operating edge/serverless services with Cloudflare Workers and Cloudflare KV, or comparable technologies.
- Experience with event-driven architectures, asynchronous workflows, and high-throughput data pipelines.
- Strong command of relational and NoSQL data models, query optimization, and complex transactional data.
- Experience with LLM integrations, vector search, embeddings, RAG patterns, and generative AI frameworks.
- Familiarity with frontend architecture using Vue and UI/UX principles.
Operational Excellence
- Experience with production monitoring, alerting, and incident response.
- Strong understanding of scaling, latency optimization, and reliability under real-world load.
AI & LLM Safety
- Experience implementing structured generation, verification layers, retrieval-based grounding, and guardrails to reduce hallucinations.
- Understanding of prompt engineering, context injection, and model evaluation.
Soft Skills
- Exceptional communication and cross-functional leadership capabilities.
- Pragmatic decision-making balancing long-term vision with iterative execution.
Skills
Go, Node.js, Nestjs, Postgres, Firestore, Vector Databases, Cloudflare Workers, Cloudflare Kv, Microservices, Prometheus, Grafana, OpenTelemetry, Embeddings, Retrieval-Augmented Generation, Vue
Similar jobs
ML Engineering jobsBuild and ship autonomous, agentic software development lifecycle capabilities, including AI agents, orchestration, and safety guardrails. The role requires senior software engineering experience, proficiency in Ruby, Go, or Python, distributed systems knowledge, and experience with AI/ML applications.
Build and operate scalable AI/ML systems and pipelines that strengthen Airbnb’s fraud prevention and trust defenses. The role requires 7+ years of backend or platform engineering experience, strong programming and data engineering skills, and production machine learning expertise.
Build production AI agents and the distributed platform that operates large-scale GPU infrastructure. The role requires 5+ years of backend, distributed systems, or infrastructure experience, with expertise in agent systems, knowledge graphs, retrieval, or semantic search.
Build and deploy explainable machine learning, NLP, LLM, and agentic systems that power enterprise go-to-market intelligence products. The role requires 6+ years of production ML experience, strong Python and cloud skills, and end-to-end ownership from modeling through monitoring.
Senior software engineer developing ML-based search relevance and discovery systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 5+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.