Staff Engineer, ML/AI Platform
Staff-level IC building and scaling the ML/AI platform infrastructure that enables training, deployment, and serving of models and agentic systems at massive scale. Focus on high-leverage architecture decisions and technical leadership across Attentive’s AI organization.
About the job
What You’ll Accomplish
- Setting Technical Direction — Architect ML platform strategy spanning data pipelines, training infrastructure, and serving layers using cutting-edge tooling like Ray, MLFlow, Metaflow, Argo, and Spark.
- Uplevel and Innovate Core AI & ML Stack — Build and operate production-grade, low-latency ML serving layers with robust model lifecycle systems including champion/challenger testing, automated rollouts, versioning, and rollback capabilities.
- Uplevel and Innovate Core AI & ML Stack — Define and drive Attentive’s agentic stack.
- Technical Leadership — Provide ML infrastructure perspective in high-level discussions about Attentive’s AI strategy spanning multiple quarters and teams.
- Technical Mentorship — Mentor platform and ML engineers, actively championing team members.
- Being the “Glue” — Build universal interfaces, architectures, and patterns—like data access layers and prediction serving APIs—that bridge platform capabilities with product needs to streamline high-priority ML work across the organization.
Your Expertise
- Experience to know what works, what doesn’t, and why in AI and ML systems.
- 5+ years focused specifically on ML Platform/MLOps, with deep understanding of gold-standard practices and best-in-class tooling.
- Proven track record of owning and building core components of ML platforms using tools like Spark, Ray, MLFlow, Kubeflow, or Metaflow.
- Built and operated a high-throughput agentic stack (MCP / data infrastructure, context store, orchestration, and prompt layer).
- Strong expertise in Python for both batch processing and online service frameworks.
- Experience designing and operating online and offline inference systems, understanding the critical differences and tradeoffs between them.
Sample Projects
- Design and implement inference pipelines with champion/challenger shadow testing and automated model promotion.
- Lead and scale Attentive’s agentic stack from the ground up.
- Scale real-time feature streaming to handle low-latency, high-volume reinforcement learning workloads.
- Build a universal data access layer and prediction serving interface that powers ML capabilities across Attentive’s product suite.
Skills
Python, Ray, MLflow, Metaflow, Spark, Kubeflow, Argo, Ml Platform, MLOps, Agentic Infrastructure, Model Lifecycle Management, Inference Systems, Data Pipelines, Champion/Challenger Testing
Similar jobs
ML Engineering jobsLeads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Staff engineer responsible for designing and scaling the infrastructure, execution environments, verifiers, and tooling used to train and evaluate AI agents. Requires 8+ years of software engineering experience, strong Python and distributed-systems expertise, and familiarity with sandboxing, high-throughput systems, and LLM workflows.
Staff Machine Learning Engineer building and operating production ML systems for causal marketing measurement, optimization, and planning. The role requires deep statistical and machine learning expertise, production programming experience, cross-functional collaboration, and technical mentorship.
Senior Staff ML Engineer fine-tunes and optimizes state-of-the-art LLMs for Airbnb's customer support AI products, including AI assistants and autonomous agents. Partners cross-functionally to productionize models at scale. Requires PhD and 10+ years experience with PyTorch.