Skip to content
RedditRedditUnited States

Staff Machine Learning Infrastructure Engineer, Embedding Platform

Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.

253k – 355k/yr
Remote8+ YOEML Engineering

About the role

Responsibilities

  • Architect and lead the development of next-generation, large-scale machine learning techniques.
  • Define and execute ML strategy to improve personalization and recommendation quality.
  • Lead research initiatives on scalable machine learning systems and real-time model adaptation, bringing advancements into production.
  • Partner with ML infrastructure teams to build high-performance, distributed training systems that scale across multiple GPUs and cloud environments.
  • Establish and optimize real-time serving architectures for large-scale embeddings, ensuring low-latency inference and high throughput.
  • Collaborate with Feed Ranking, Ads, Content Understanding, and Core ML teams to integrate models into AI-driven systems.
  • Mentor and guide ML engineers while fostering technical excellence, innovation, and knowledge sharing.
  • Evaluate and introduce new modeling paradigms and contribute to long-term ML planning.
  • Drive technical discussions and present findings to leadership and stakeholders.

Requirements

  • 8+ years of experience in machine learning engineering, focused on large-scale ML systems and recommendation or personalization systems.
  • Expertise in modern deep learning architectures, including sequence models and foundational models.
  • Deep understanding of complex multi-entity relationships and their modeling in large-scale systems.
  • Experience designing, implementing, and optimizing scalable ML architectures, distributed training, and real-time inference.
  • Strong software engineering skills in Python, C++, or similar languages.
  • Experience with ML infrastructure, high-performance computing, and cloud-based ML pipelines.
  • Demonstrated leadership in driving ML strategy, mentoring engineers, and influencing cross-functional teams.
  • Experience with A/B testing, model evaluation frameworks, and real-time feedback loops in large-scale production systems.
  • Excellent communication skills for presenting complex ML concepts to technical and non-technical stakeholders.

Compensation and Benefits

  • Base salary: $253,300–$354,600 USD
  • Equity in the form of restricted stock units; select positions may also be eligible for commission.
  • Comprehensive healthcare benefits and income replacement programs.
  • 401(k) with employer match.
  • Global benefits supporting workspace, professional development, and caregiving.
  • Family planning support.
  • Gender-affirming care.
  • Mental health and coaching benefits.
  • Flexible vacation and paid volunteer time off.
  • Paid parental leave.

Skills

PythonC++Deep Learningsequence modelsFoundation ModelsDistributed Traininggpu computingreal-time inferencemachine learning infrastructurecloud ml pipelinesRecommendation SystemsA/B TestingModel EvaluationEmbeddingshigh-performance computing

Similar roles

ML Engineering jobs
Coinbase

Senior Staff Software Engineer, Finance Automation

CoinbaseUnited States

Founding engineer building an AI-native platform to automate Coinbase's finance workflows (period-close, reconciliation, regulatory filings). Architect governed LLM agents with SOX-compliant controls, integrate with ERP systems, and set technical direction for FP&A/Treasury as an embedded engineer.

254k – 299k/yrRemote12+ YOEML Engineering
Coinbase

Senior Staff Software Engineer, Legal Automation

CoinbaseUnited States

Senior Staff Software Engineer building an AI agent platform and automated workflows to transform Coinbase's Legal organization. Architect production-grade LLM and multi-agent systems that replace manual legal processes such as agreement redlining and governance.

254k – 299k/yrRemote12+ YOEML Engineering
Scale AI

Staff Frontier Agents Engineer

Scale AISan Francisco, CA +2

Build and deploy production AI agents and frontier systems for enterprise customers, combining LLMs with retrieval, memory, multi-agent architectures, and traditional ML. Requires 8+ years experience building production AI systems, strong Python skills, and customer-facing abilities.

252k – 315k/yrHybrid8+ YOEML Engineering
Postman

Member of Technical Staff, AI Agent Development Lead

PostmanSan Francisco, CA +3

Leads design, development, and deployment of AI agents using language models and ML frameworks. Drives scalable, safe AI systems while mentoring teams and collaborating cross-functionally. Requires strong Python skills and AI leadership experience.

256k – 276k/yrHybridML Engineering
Labelbox

Staff ML Engineer, Agent Training & Environments

LabelboxSan Francisco, CA

Build RL environments, verifiers, fine-tuning pipelines, and eval systems for frontier AI agents at Labelbox. Requires deep RL post-training experience (SFT + RL methods), strong Python/systems engineering, and the ability to ship production infrastructure at high velocity.

250k – 280k/yrHybrid7+ YOEML Engineering