Skip to content
RedditReddit

Senior Machine Learning Systems Engineer

Leads development of large-scale ML platforms, focusing on MLOps, graph ML infrastructure, performance tuning, and distributed training optimization. Requires 5+ years in ML infrastructure with expertise in PyTorch, Kubernetes, Ray, and cloud tools.

About the job

What You’ll Do

  • Design end-to-end model lifecycle patterns (MLOps) to boost velocity of development for ML engineers, including data preparation, model management, experiment tracking, and more
  • Zero-to-one development and support of a graph ML codebase and platform that abstracts away common patterns and enables greater model scalability and iteration
  • Collaborate with ML engineers on performance tuning, including improving model training time, efficiency, and GPU training costs in a large, distributed ML training environment
  • Optimize batch data processing within a data warehouse and with tools such as Apache Beam, Apache Spark, Ray Data, and more
  • Architect pipelines to build and maintain massive graph data structures on the order of billions of nodes and tens of billions of edges

Who You Might Be

  • 5+ years of experience in ML infrastructure, including model training and model deployments
  • Hands-on experience with ML optimization, including memory and GPU profiling
  • Deep experience with cloud-based technologies for supporting an ML platform, including tools like GCP BigQuery, Google Cloud Storage, infrastructure-as-code (Terraform), and more
  • Hands-on experience administering and integrating MLOps tools for experiment tracking, model serving, and model registries (e.g. MLflow or Wandb)
  • Proficiency with the common programming languages and frameworks of ML, such as Python, PyTorch, Tensorflow, etc.
  • Deep experience working with distributed training frameworks, including Ray and Kubernetes
  • Strong focus on scalability, reliability, performance, and ease of use. You are an undying advocate for platform users and have a deep intuition for the machine learning development lifecycle.
  • Strong organizational & communication skills
  • Experience working with graph databases (Neo4j, JanusGraph, TigerGraph) is a big plus
  • Experience working with graph neural networks (GNNs) and associated graph ML frameworks (PyTorch Geometric, Deep Graph Library) is a big plus

Skills

MLOps, PyTorch, TensorFlow, Python, Kubernetes, Ray, Spark, Apache Beam, MLflow, Wandb, Terraform, Gcp Bigquery, Google Cloud Storage, Pytorch Geometric, Deep Graph Library

Reddit

Reddit

Ontario, Canada

Senior Machine Learning Engineer, Ads
$217k+/yrRemote5+ YOEML Engineering

Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.

Fetch

Fetch

United States

Senior Machine Learning Engineer II
$211k+/yrRemote6+ YOEML Engineering

Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.

Dialpad

Dialpad

United States

Senior AI Engineer
$225k+/yrRemote5+ YOEML Engineering

Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.

Ambience Healthcare

Ambience Healthcare

San Francisco, CA

Senior Machine Learning Engineer
$225k+/yrHybrid5+ YOEML Engineering

Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.

Checkr

Checkr

San Francisco, CA

Senior Machine Learning Engineer
$207k+/yrOn-site6+ YOEML Engineering

Build and operate production ML and AI services using Python, LLM APIs, and robust software engineering practices. The role requires 6+ years of professional software experience, including production ML systems, and partners closely with product and engineering teams.