Build and lead the Snowflake ML Platform for scalable, native machine learning and LLM inference workloads. Requires 7+ years experience with ML serving systems, inference engines (vLLM, TensorRT-LLM), and frameworks like PyTorch.
236k – 339k/yr
On-site7+ YOEML Engineering
About the role
Responsibilities
Help define and own the roadmap, working collaboratively and proactively with senior architects, PMs, and team leadership. Initiatives include platforms and tools that enable customers to do state-of-the-art machine learning on Snowflake natively.
Collaboratively build and execute a vision for incorporating new advances in machine learning in ways that best achieve the team’s business objectives.
Ensure operational excellence of the services and meet the commitments to our customers regarding reliability, availability, and performance.
Collaborate across other ML partner teams to continuously improve ML development velocity and capabilities at Snowflake.
Support team members in delivering a high level of technical quality.
Requirements
7+ years of industry experience designing, building, and supporting Internet serving infrastructure, machine learning platforms, machine learning services, and frameworks.
Strong track record of working with machine learning systems and/or platforms.
Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between them.
Lead design and development of Harvey's Model Infrastructure platform powering all AI requests, including unified model controller, intelligent routing, multi-provider integrations, observability, and capacity management for high reliability, low latency, and efficiency. Requires 7+ years building large-scale distributed systems with strong programming and leadership skills; AI/LLM infrastructure experience preferred.
236k – 290k/yr
On-site7+ YOEML Engineering
Staff AI Engineer - Cortex Code Quality
SnowflakeMenlo Park, CA
Staff AI Engineer builds and owns quality systems for Cortex Code AI coding agents, including agent strategy, experimentation pipelines, failure analysis, and cross-team alignment. Requires 8+ years shipping AI/ML software, proficiency in Python/TypeScript/Go, and expertise in LLM evaluation harnesses.
236k – 339k/yr
On-site8+ YOEML Engineering
Staff Software Engineer
ConfluentMountain View, CA +1
Build and operate backend services for AI and model inference on Confluent's real-time streaming data platform. Own end-to-end feature delivery across model lifecycle, inference routing, and agent execution with strong distributed systems expertise.
236k – 277k/yr
Remote10+ YOEML Engineering
Senior/Staff Software Engineer, Labeling Platform
NuroMountain View, CA
Build and productionize scalable, highly available distributed systems and infrastructure for Nuro's data labeling platform that powers autonomous driving ML models. Requires 5+ years experience with reliable large-scale data systems, technical leadership, and strong programming skills in Python, C++, or Go.
235k – 352k/yr
On-site5+ YOEML Engineering
Senior Staff Software Engineer, AI Model LifeCycle
CrusoeSan Francisco, CA +1
Builds and manages platforms for AI model lifecycles, focusing on fine-tuning, training pipelines, and reinforcement learning for LLMs. Requires 8+ years in AI, advanced degree, and hands-on experience with generative AI techniques.