Principal Machine Learning Engineer
Leads the architecture, roadmap, and production execution of GenAI and agentic machine-learning systems at enterprise scale. The role requires 10–12+ years of applied ML experience, expertise in retrieval and large-scale systems, and the ability to mentor senior technical teams.
About the job
Responsibilities
- Lead the technical direction of GenAI and agentic ML systems powering enterprise-grade AI agents, including reasoning, retrieval, tool use, and SaaS integrations.
- Architect, design, and implement scalable production pipelines for model training, fine-tuning, retrieval-augmented generation (RAG), agent orchestration, and evaluation.
- Define and own the multi-year ML roadmap for GenAI infrastructure, including agent frameworks, RAG systems, evaluation loops, and MCP, browser, and vision integrations.
- Identify, research, prototype, and integrate advances in deep learning, large models, recommender systems, LLMs, reasoning, memory architectures, multimodal perception, long-context models, and autonomous agents.
- Optimize accuracy, latency, cost, interpretability, and reliability across the agent lifecycle, from prompt design through orchestration and execution.
- Drive observability, reproducibility, versioning, testing, and bias-aware development across ML and agentic systems.
- Mentor senior engineers and researchers while fostering scientific rigor, experimentation, and system-level thinking.
- Collaborate with product, infrastructure, and research teams to align ML innovation with enterprise needs, secure integrations, privacy-aware deployments, and scalable use cases.
- Guide data strategy, including retrieval indices, embeddings, structured and unstructured corpora, and feedback loops.
- Ensure ML agents and RAG pipelines scale across billions of knowledge objects, diverse APIs, and real-time enterprise contexts.
Requirements
- Bachelor's, master's, or PhD in Computer Science, Machine Learning, Statistics, or a related field.
- 10–12+ years of applied machine learning experience, especially in large-scale settings.
- Experience building production ML systems under latency, throughput, and cost constraints.
- Experience with knowledge retrieval and search.
- Exposure to agentic systems and frameworks.
- Proficiency in Python, C++, or Java and ML frameworks such as TensorFlow and PyTorch.
- Strong understanding of the full ML lifecycle, including data pipelines, feature engineering, model training, serving, monitoring, and maintenance.
- Experience designing monitoring, diagnostics, logging, and model-versioning systems.
- Deep knowledge of distributed training and inference optimization techniques such as quantization, pruning, and batching.
- Excellent communication skills and ability to explain complex systems and trade-offs to technical and non-technical stakeholders.
- Experience mentoring senior engineers and leading technical discussions across organizations.
Compensation & Benefits
- Compensation is determined by location, level, job-related knowledge, skills, and experience.
- Certain roles may be eligible for variable compensation, equity, and benefits.
Skills
Python, C++, Java, TensorFlow, PyTorch, Machine Learning, LLMs, Retrieval-Augmented Generation, Agentic Systems, Knowledge Retrieval, Search, Distributed Training, Quantization, Pruning, Model Monitoring
Similar jobs
ML Engineering jobsLeads the development of ML- and NLP-powered search relevance systems, including query understanding, ranking, retrieval, and evaluation. Requires 10+ years of search relevance experience and a bachelor’s degree, with advanced study preferred.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation at scale. Requires 10+ years of search relevance experience and expertise in ML, NLP, or related discovery technologies.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 10+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.
Build and ship autonomous, agentic software development lifecycle capabilities, including AI agents, orchestration, and safety guardrails. The role requires senior software engineering experience, proficiency in Ruby, Go, or Python, distributed systems knowledge, and experience with AI/ML applications.