Senior Software Engineer — LLM Post-Training Platform
Build and scale Snowflake's Cortex Training LLM post-training platform, handling distributed GPU scheduling, orchestration, and productionizing research for enterprise-scale model adaptation.
About the job
Responsibilities
- Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
- Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
- Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
- Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.
Requirements
- 5+ years building and shipping production ML systems
- Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production
- Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
- BS in Computer Science or a related field (MS/PhD a plus)
Nice-to-Haves
- Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition
Skills
PyTorch, Deepspeed, Fsdp, Ray, CUDA, Nccl, vLLM, Kubernetes, Distributed Systems, Llm Post-Training
Similar jobs
ML Engineering jobsBuild and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Own machine learning end to end, from modeling messy clinical data through production deployment, monitoring, and infrastructure. The role requires 7+ years of experience building scalable ML systems, strong software and data engineering skills, and proficiency with Python, SQL, and cloud platforms.
Leads and manages an applied machine learning team developing production fraud detection and identity verification models. The role combines people leadership with hands-on technical work and requires substantial ML experience, production deployment expertise, and experience in risk-focused domains.
Builds and productionizes machine learning systems for trust and safety, including abuse detection, autonomous AI agents, and evaluation frameworks. The role requires 5+ years of applied ML experience, strong Python skills, experience with LLMs and scalable pipelines, and a relevant advanced degree or equivalent background.
Owns the quality, efficiency, and reliability of Cortex Code, Snowflake’s coding agent for data workflows. The role combines production software engineering, AI/LLM evaluation, experimentation, observability, and systems optimization, requiring 6+ years of experience and proficiency in Python, TypeScript, or Go.