Skip to content
RedditReddit

Staff Machine Learning Systems Engineer

Leads development of large-scale ML platforms, focusing on MLOps, graph ML infrastructure, performance optimization, and distributed training pipelines. Requires 8+ years in ML infrastructure with expertise in Python, PyTorch, Kubernetes, Ray, and cloud tools.

About the job

What You’ll Do:

  • Design end-to-end model lifecycle patterns (MLOps) to boost velocity of development for ML engineers, including data preparation, model management, experiment tracking, and more
  • Zero-to-one development and support of a graph ML codebase and platform that abstracts away common patterns and enables greater model scalability and iteration
  • Collaborate with ML engineers on performance tuning, including improving model training time, efficiency, and GPU training costs in a large, distributed ML training environment
  • Optimize batch data processing within a data warehouse and with tools such as Apache Beam, Apache Spark, Ray Data, and more
  • Architect pipelines to build and maintain massive graph data structures on the order of billions of nodes and tens of billions of edges

Who You Might Be:

  • 8+ years of experience in ML infrastructure, including model training and model deployments
  • Hands-on experience with ML optimization, including memory and GPU profiling
  • Deep experience with cloud-based technologies for supporting an ML platform, including tools like GCP BigQuery, Google Cloud Storage, infrastructure-as-code (Terraform), and more
  • Hands-on experience administering and integrating MLOps tools for experiment tracking, model serving, and model registries (e.g. MLflow or Wandb)
  • Proficiency with the common programming languages and frameworks of ML, such as Python, PyTorch, Tensorflow, etc.
  • Deep experience working with distributed training frameworks, including Ray and Kubernetes
  • Strong focus on scalability, reliability, performance, and ease of use. You are an undying advocate for platform users and have a deep intuition for the machine learning development lifecycle.
  • Strong organizational & communication skills
  • Experience working with graph databases (Neo4j, JanusGraph, TigerGraph) is a big plus
  • Experience working with graph neural networks (GNNs) and associated graph ML frameworks (PyTorch Geometric, Deep Graph Library) is a big plus

Skills

Python, PyTorch, TensorFlow, Kubernetes, Ray, MLflow, Wandb, Terraform, Spark, Apache Beam, Gcp Bigquery, Google Cloud Storage, Pytorch Geometric, Deep Graph Library, Neo4J

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Model Infrastructure
$231k+/yrHybrid7+ YOEML Engineering

Leads the design and operation of reliable, scalable model infrastructure powering AI inference across multiple providers. Requires 7+ years of distributed-systems engineering experience, strong programming skills, and expertise in production reliability and cloud infrastructure.

Shield AI

Shield AI

San Mateo, CA

Staff Software Engineer, Autonomy Capabilities
$234k+/yrOn-site7+ YOEML Engineering

Leads the design, implementation, integration, and field validation of tactical autonomy and multi-agent coordination capabilities for unmanned platforms. Requires 7+ years of relevant experience, production C++, technical leadership, and eligibility for a U.S. Secret clearance.

Snowflake

Snowflake

Bellevue, WA

Staff Software Engineer - Snowflake Feature Store
$236k+/yrOn-site10+ YOEML Engineering

Leads the roadmap and technical vision for Snowflake Feature Store, building reliable, high-performance machine learning platform capabilities and supporting technical execution across partner teams. Requires 10+ years of experience with data-serving infrastructure or ML platforms, plus Java and Python expertise.

Shield AI

Shield AI

Washington, DC
Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site7+ YOEML Engineering

Leads development and integration of advanced maritime autonomy for USVs, UUVs, and cooperating UAVs, including motion planning, localization, safety, and multi-agent coordination. Requires staff-level technical leadership, substantial robotics experience, C++ and Python proficiency, and eligibility for a SECRET clearance.

Ironclad

Ironclad

San Francisco, CA

Senior Staff Software Engineer, Agentic Search
$220k+/yrHybrid10+ YOEML Engineering

Leads architecture and technical direction for agentic search systems combining LLMs, retrieval, and content-understanding pipelines for contract intelligence. The role requires 10+ years building production systems, deep search or LLM expertise, and strong cross-team technical leadership.