Principal Engineer - GenAI Applications & MLOps
Lead the design and scaling of ML infrastructure and GenAI platforms at Weave, enabling product teams to deliver AI-powered healthcare communication features. Requires deep expertise in MLOps, LLMs, distributed systems, and cross-team technical leadership.
About the job
Infrastructure & Delivery
- Design and develop ML infrastructure, tooling, and models to help teams deliver world-class experiences
- Build internal and external products and platforms to enable teams to incorporate AI into their features and customer-facing products
- Translate product goals into actionable engineering plans and build scalable, resilient services for data integration and event processing
Cross-Team Problem Solving
- Help product and development teams understand the data lifecycle and consult with teams on common ML patterns/tradeoffs
- Coach and collaborate inside and outside the team to elevate technical standards
- Write high-quality, performant, sustainable, and testable code while working in a cloud environment
Strategic Technical Leadership
- Monitor the industry landscape, anticipate where technological advances are heading, and ensure Weave stays ahead of the curve
- Cut through noise and hype to identify genuine strategic value; advocate for and lead key initiatives that prepare Weave for emerging challenges
- Shape company-wide standards for engineering excellence, observability, and reliability in distributed systems
Mentorship & Organizational Capability
- Actively mentor Staff and Senior Engineers across the fellowship, developing the next generation of technical leaders
- Elevate architectural thinking across teams through design reviews, documentation standards, and hands-on guidance
- Build organizational capability that persists beyond your individual contributions
What You Will Need
- 12+ years of software engineering experience with progressive technical leadership scope
- Demonstrable experience building and deploying ML-driven B2B multi-tenant applications in production environments at scale
- Deep expertise in distributed systems architecture, including building and operating services that handle hundreds of millions of transactions and terabytes of data
- 8+ years of experience in Machine Learning or AI, preferably with a focus on natural language
- Expertise with modern ML tools and techniques such as LLMs, RAG, Prompt Engineering, Fine Tuning, LLM evaluations, multi-modal models
- Strong background in scalable data stores—both relational (PostgreSQL at scale, Vitess, Spanner) and NoSQL (Bigtable, Redis)
- Operational experience with cloud-native infrastructure on GCP or AWS, including Kubernetes, infrastructure-as-code, and highly available system design
- Track record of leading cross-team technical initiatives that delivered measurable business outcomes
- Demonstrated ability to influence without direct authority, build consensus across organizational boundaries
What Will Make Us Love You
- Expertise with customer-facing GenAI at scale, in production
- Expertise building low-latency high-accuracy AI Agents
- Background in compliance-heavy environments (healthcare, fintech)
- History of external technical leadership: open source contributions, conference speaking, or published technical writing
Skills
Machine Learning, MLOps, LLMs, RAG, Prompt Engineering, Fine Tuning, Kubernetes, GCP, AWS, Postgres, Distributed Systems
Similar jobs
ML Engineering jobsLeads the technical strategy and engineering execution required to achieve driverless freeway operation for autonomous vehicles. Requires 10+ years of software experience, demonstrated freeway autonomy leadership, deep expertise in an autonomy domain, and strong executive communication skills.
Own the architecture and delivery of production AI systems for patient-provider matching, search relevance, personalization, clinical workflows, and engagement. The role requires extensive software engineering, distributed systems, search or recommendation, ML applications, and foundation-model experience in a regulated healthcare setting.
Principal technical leader defining architecture and multi-year strategy for Pinterest’s Homefeed, Search, and AI Assistant experiences. The role requires 15+ years of large-scale systems or machine-learning experience, deep expertise in discovery and generative AI, and hands-on leadership across engineering and product organizations.
Design and scale ML infrastructure and real-time learning systems powering personalization, search, ranking, and ad tech for millions of consumers. The role requires deep distributed-systems and data-pipeline expertise, strong architecture leadership, and experience delivering zero-to-one ML systems.
Leads the technical vision, architecture, and engineering standards for a company-wide ML platform supporting model development, deployment, serving, and monitoring. The role requires principal-level expertise in Python and Java, scalable MLOps, cloud infrastructure, security, and technical leadership across teams.