Skip to content
RedditReddit

Engineering Manager, Ads ML Efficiency

Lead a team of ML and systems engineers focused on model optimization, training/inference efficiency, and tooling for Ads ML at Reddit. Drive measurable wins in training time, latency, cost, and launch readiness while partnering with ranking and platform teams.

About the job

What you’ll do

  • Lead & Grow: Hire, mentor, and retain a high-performing team of ML engineers / systems-oriented engineers working on model optimization and ML efficiency.
  • Set Technical Direction: Define the roadmap for training optimization, inference optimization, launch-readiness tooling, and reusable efficiency primitives across Ads ML.
  • Deliver Measurable Wins: Drive reductions in model training time, online latency, serving cost, and infra-driven launch risk.
  • Build Systems and Tooling: Guide the development of profiling, benchmarking, load testing, observability, cost analysis, debugging, and efficiency certification systems.
  • Operate in the Critical Path: Partner with model owners and platform teams to accelerate high-priority launches and remove bottlenecks from the path to production.
  • Shape the Team’s Evolution: Balance near-term white-glove optimization work with medium-term platformization and automation.
  • Build XFN Alignment: Work closely with MLP, AMP, Ranking, and serving teams to clarify boundaries, upstream generic wins, and keep Ads needs on track.
  • Raise the Bar: Establish engineering rigor around measurement, performance debugging, launch safety, and technical decision-making for efficiency work.

What we’re looking for

  • Deep ML Engineering Experience: The candidate should have been close to the models themselves and understand training, serving, debugging, and optimization in depth.
  • Hands-on Optimization Background: Direct experience improving training loops, serving systems, profiling workflows, model/inference efficiency, or GPU utilization.
  • Strong Managerial Ability: Experience building and leading teams, coaching engineers, managing delivery, and making prioritization tradeoffs under ambiguity.
  • Distributed Systems Fluency: Proven ability to reason about production-scale ML systems and the tradeoffs that govern reliability, speed, cost, and scale.
  • Customer and Platform Instincts: Able to work as a service provider to modeling teams while still building reusable systems rather than only heroic one-offs.
  • Strong Communication: Can explain technical tradeoffs clearly to engineers, PMs, and senior stakeholders.
  • Ads experience: Experience in ads ranking, recommender systems, marketplace ML, or adjacent production ML domains is strongly preferred.

Nice-to-have

  • Experience with GPU training and serving migrations.
  • Experience with PyTorch, distributed training frameworks, or kernel/performance optimization.
  • Experience building efficiency benchmarking or launch certification frameworks.
  • Experience working in organizations where ML platform and applied modeling responsibilities are split across multiple teams.

Skills

Ml Engineering, Model Optimization, Training Optimization, Inference Optimization, Gpu Utilization, Distributed Systems, PyTorch, Profiling, Benchmarking, Load Testing

Skydio

Skydio

San Mateo, CA

Senior Manager, Middleware
$230k+/yrHybrid8+ YOEEngineering Management

Leads and develops a software engineering team building middleware and on-device services for autonomous drone platforms. The role combines people management with hands-on Python or C++ development, systems architecture, reliability improvements, and cross-functional technical leadership.

Vercel

Vercel

San Francisco, CA
Member of the Technical Staff - Next.js
$230k+/yrHybrid8+ YOEEngineering Management

Hands-on engineering team lead building and evolving Next.js while coaching a small team and setting technical direction. Requires 8+ years of software engineering experience, strong React or framework architecture expertise, and continued production coding experience.

Crusoe

Crusoe

Bellevue, WA
Senior Manager, Network & Cloud Architecture
$230k+/yrOn-site10+ YOEEngineering Management

Leads the deployment organization responsible for bringing global HPC and GPU infrastructure online, owning people, delivery, standards, vendors, and executive-level program risk. Requires 10+ years in network engineering or data center deployment and substantial engineering management experience.

Genius AI

Genius AI

San Francisco, CA

Senior Engineering Manager, Infrastructure
$230k+/yrOn-site8+ YOEEngineering Management

Leads Compute, Data, and Developer Experience teams, owning infrastructure strategy, reliability, developer enablement, vendor spend, and team growth. The role requires 8+ years of infrastructure or platform experience, 4+ years of engineering leadership, and deep expertise in AWS, Kubernetes, Terraform, ArgoCD, and observability.

Shield AI

Shield AI

Dallas, TX

Senior Manager, GNC
$230k+/yrOn-site8+ YOEEngineering Management

Leads the Guidance and Control team responsible for developing, integrating, qualifying, and flight-testing autonomy and control algorithms for unmanned aerial systems. Requires substantial engineering leadership experience, controls expertise, and a track record of shipping products and developing high-performing teams.