Senior ML engineer focused on improving training and inference efficiency for Reddit's Ads ML models. Own optimization initiatives, diagnose production bottlenecks, build reusable tooling and primitives, and mentor others on measurable efficiency and launch-safety wins.
217k – 303k/yr
Remote5+ YOEML Engineering
About the role
What you’ll do
Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability rather than intuition-first debugging.
Build performance tooling, optimization playbooks, observability hooks, guardrails, or efficiency primitives that help more than one team or workload over time.
Improve launch-safety and efficiency readiness by contributing to load testing, fallback readiness, latency and cost visibility, and operational confidence for heavy models.
Work with model owners and platform teams to land pragmatic fixes while helping the team gradually standardize repeated solutions.
Contribute to the team’s technical direction by surfacing patterns, tradeoffs, and opportunities for reuse or automation.
Mentor less-experienced engineers through code, debugging, measurement rigor, and strong execution habits.
What we’re looking for
Deep ML systems experience close to real production models and workloads, not just generic infra exposure.
Direct hands-on experience improving training or serving efficiency with measurable outcomes.
Strong technical judgment across model-level, runtime-level, and infrastructure-level optimization choices.
Ability to own complex projects end to end and collaborate effectively across team boundaries.
Good customer and platform instincts: can solve concrete bottlenecks while keeping maintainability, adoption, and future reuse in mind.
Strong communication: able to explain tradeoffs clearly to engineers and partner teams.
Nice-to-have
Experience with GPU training or serving migrations.
Experience with PyTorch, distributed training frameworks, or kernel/runtime optimization.
Experience building launch certification, efficiency benchmarking, or cost observability systems.
Experience in organizations where platform and applied modeling responsibilities are split across multiple teams.
Experience with model compression or deployment optimizations such as quantization, pruning, distillation, or checkpoint optimization.
Skills
ml systemsPyTorchDistributed Trainingmodel compressionquantizationpruningdistillationgpu trainingprofilingbenchmarkingObservabilitykernel optimizationruntime optimization
Build, train, and optimize large language models and ML systems to scalably enforce Reddit's community rules and keep users safe. Requires 5+ years MLE experience with NLP, deep learning frameworks, and distributed training.
Senior Machine Learning Engineer building and shipping production ML systems for Ads Content Understanding at Reddit. Focus on LLMs, content taxonomies, opinion mining, product understanding, and operationalizing signals for monetization impact. Requires 5+ years delivering scalable ML in production.
217k – 303k/yr
Remote5+ YOEML Engineering
Senior Machine Learning Systems Engineer
RedditUnited States
Build large-scale ML experimentation and training orchestration platforms, including agentic AI execution systems, to accelerate Ads ML development at Reddit. Requires 5+ years infrastructure experience and 2+ years building production ML platforms.
217k – 303k/yr
Remote5+ YOEML Engineering
Senior Machine Learning Engineer, Ads
RedditUnited States
Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.
217k – 303k/yr
Remote5+ YOEML Engineering
Senior Machine Learning Systems Engineer
RedditUnited States
Design and build scalable ML ranking systems powering Reddit's personalized feeds, search, and recommendations at massive scale. Requires 5+ years building large-scale distributed ML systems.