Senior Staff Machine Learning Systems Engineer, Ads ML Platform
Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.
About the job
Responsibilities
- Own the technical strategy for the end-to-end Ads ML engineer lifecycle, beginning with feature development, training data, offline experimentation, and model iteration workflows.
- Align Ads ML platform priorities with Reddit’s broader ML Platform vision and translate Ads needs into reusable platform capabilities.
- Define architecture and technical standards for ML feature and training-data systems across batch and streaming computation, backfills, lineage, quality, observability, and online/offline consistency.
- Identify high-leverage friction points for ML engineers and improve development velocity.
- Build platform abstractions and workflow automation that make ML development faster, safer, more reliable, and self-service.
- Extend platform strategy into serving and online experimentation workflows.
- Partner across Ads, ML Platform, Data Platform, modeling, product, and engineering teams to clarify ownership and drive execution.
- Mentor Staff and senior engineers and raise architecture and operational standards.
Requirements
- 8+ years of experience in infrastructure, distributed systems, ML platforms, data platforms, or large-scale backend systems.
- 4+ years building or operating production ML infrastructure, feature platforms, training-data systems, experimentation systems, or large-scale data pipelines.
- Experience leading broad, ambiguous, multi-team platform initiatives from strategy through adoption.
- Experience building platforms used by ML engineers, data scientists, or product teams developing production ML systems.
- Deep experience with ML platforms, feature platforms, training data, experimentation, developer infrastructure, or distributed data infrastructure.
- Ability to balance urgent customer needs with durable architecture and reusable platform patterns.
- Ability to influence senior engineers and leaders through technical reasoning, RFCs, design reviews, decision frameworks, and operating mechanisms.
Nice-to-haves
- Experience with distributed data and compute technologies such as Spark, Flink, Kafka, Ray, Airflow, Iceberg, Kubernetes, BigQuery, Snowflake, or Databricks.
- Interest in shaping how production ML systems are built, scaled, and operated.
Compensation & Benefits
- Base salary range: $292,500–$409,500 USD.
- Equity in the form of restricted stock units; some positions may also be eligible for commission.
- Comprehensive healthcare benefits and income replacement programs.
- 401(k) with employer match.
- Global benefits supporting workspace, professional development, caregiving, and family planning.
- Gender-affirming care.
- Mental health and coaching benefits.
- Flexible vacation and paid volunteer time off.
- Generous paid parental leave.
Skills
Machine Learning, Ml Platforms, Distributed Systems, Backend Systems, Feature Platforms, Training Data Systems, Offline Experimentation, Spark, Flink, Kafka, Ray, Airflow, Iceberg, Kubernetes, BigQuery
Similar jobs
ML Engineering jobsLeads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.