Skip to content
SnowflakeSnowflake

Senior Software Engineer — LLM Post-Training Platform

Build and scale Snowflake's Cortex Training LLM post-training platform, handling distributed GPU scheduling, orchestration, and productionizing research for enterprise-scale model adaptation.

About the job

Responsibilities

  • Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
  • Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
  • Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
  • Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.

Requirements

  • 5+ years building and shipping production ML systems
  • Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production
  • Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers
  • Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
  • BS in Computer Science or a related field (MS/PhD a plus)

Nice-to-Haves

  • Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition

Skills

PyTorch, Deepspeed, Fsdp, Ray, CUDA, Nccl, vLLM, Kubernetes, Distributed Systems, Llm Post-Training

Traba

Traba

New York, NY
Senior Software Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.

Metriport

Metriport

San Francisco, CA

Senior AI/ML Engineer
$200k+/yrHybrid7+ YOEML Engineering

Own machine learning end to end, from modeling messy clinical data through production deployment, monitoring, and infrastructure. The role requires 7+ years of experience building scalable ML systems, strong software and data engineering skills, and proficiency with Python, SQL, and cloud platforms.

SentiLink

SentiLink

United States

Applied Machine Learning Manager - Application Fraud
$200k+/yrRemote6+ YOEML Engineering

Leads and manages an applied machine learning team developing production fraud detection and identity verification models. The role combines people leadership with hands-on technical work and requires substantial ML experience, production deployment expertise, and experience in risk-focused domains.

Airbnb

Airbnb

United States

Senior Machine Learning Engineer, Trust
$200k+/yrRemote5+ YOEML Engineering

Builds and productionizes machine learning systems for trust and safety, including abuse detection, autonomous AI agents, and evaluation frameworks. The role requires 5+ years of applied ML experience, strong Python skills, experience with LLMs and scalable pipelines, and a relevant advanced degree or equivalent background.

Snowflake

Snowflake

Senior Software Engineer, Cortex Quality
$200k+/yrOn-site6+ YOEML Engineering

Owns the quality, efficiency, and reliability of Cortex Code, Snowflake’s coding agent for data workflows. The role combines production software engineering, AI/LLM evaluation, experimentation, observability, and systems optimization, requiring 6+ years of experience and proficiency in Python, TypeScript, or Go.