Latest ML Engineering jobs
Job results
Leads the technical vision, architecture, and engineering standards for a company-wide ML platform supporting model development, deployment, serving, and monitoring. The role requires principal-level expertise in Python and Java, scalable MLOps, cloud infrastructure, security, and technical leadership across teams.
Build production machine learning systems for model customization, post-training, evaluation, and AWS-native API integration. The role requires 7+ years of relevant engineering experience and expertise in deep learning, transformers, LLM fine-tuning, and production ML infrastructure.
Build and deploy production machine-learning models for fraud detection, identity verification, and financial risk products. The role suits new PhD graduates or early-career researchers with strong quantitative foundations, Python experience, and interest in owning the full ML lifecycle.
Build and operate large-scale machine learning infrastructure and models for Reddit’s recommendation and personalization systems. The role requires 5+ years of ML engineering experience, expertise in deep learning and distributed systems, and proficiency with Python and modern ML frameworks.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Build and operate the agentic systems powering an AI tutoring product, including learner modeling, long-horizon planning, tool use, verification, and evaluation. The role requires 3+ years of software engineering experience, production LLM experience, and strong Python or TypeScript/Node skills.
Build and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and maintain tooling, evaluation systems, quality gates, and infrastructure for MongoDB's agent skills and AI platform. Requires 2+ years building production software, developer tools, CLIs, test infrastructure, or CI/CD, with strong fundamentals in API design, testing, and reasoning about nondeterministic AI systems.
Builds and operates the systems, tooling, and deployment workflows that deliver Voyage embedding and reranking models across cloud marketplaces, third-party inference providers, and self-managed environments. The role requires backend or infrastructure experience, cloud and Kubernetes expertise, and familiarity with ML model serving.
Build and operate machine-learning systems that improve search ranking quality across retrieval and later-stage ranking. The role requires deep search or recommender-systems expertise, production ranking ownership, and at least five years of relevant industry experience.
Builds and deploys retrieval, knowledge representation, and ML platform components for enterprise generative AI systems. The role requires 5+ years of production ML/AI experience, strong Python skills, and expertise in RAG, embeddings, vector indexing, and semantic search.
Builds and operates AI platform capabilities including RAG pipelines, semantic retrieval, agentic orchestration, and LLM integrations to power legal tech products. Requires 4+ years in distributed cloud systems, AI/ML experience, and proficiency in modern programming languages.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
Leads and builds a team of Applied Scientists developing production algorithmic systems for healthcare optimization, LLM applications, and member engagement. Requires 6+ years of relevant industry experience, strong technical judgment, and hands-on expertise across machine learning and optimization.
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
Build and maintain machine learning infrastructure for autonomy teams, including model pipelines, observability, inference serving, and compiler platforms. The role requires a relevant degree, at least one year of experience, strong Python skills, and familiarity with C++.
Build and operate ML infrastructure for autonomy teams, including training and deployment pipelines, model observability, inference serving, and compiler platforms across hardware targets. Requires a degree, 3+ years of relevant experience, Python proficiency, and distributed-systems expertise.
Build and operate large-scale infrastructure for autonomous-driving model training, including distributed GPU systems, data pipelines, ML workflows, and reliability tooling. The role requires 3+ years of experience, strong Python and systems-language skills, Kubernetes expertise, and distributed-systems fundamentals.
Build and operate the infrastructure powering large-scale machine-learning training for autonomous-driving systems. The role requires Python proficiency, Kubernetes production experience, distributed-systems expertise, and ownership of reliability, observability, and operational maturity.
Develop and deploy scalable machine learning and AI systems for user-facing products, including forecasting and AutoML capabilities. The role requires strong production ML engineering, modeling, software engineering, statistics, and systems knowledge.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation at scale. Requires 10+ years of search relevance experience and expertise in ML, NLP, or related discovery technologies.
Leads the development of machine-learning search relevance systems, including query understanding, ranking, retrieval, and evaluation pipelines. The role requires 10+ years of search relevance experience and expertise in NLP, LLMs, or related discovery technologies.
The Senior ML and AI Technical Solutions Engineer troubleshoots and optimizes production data, machine learning, and generative AI workloads on Databricks. The role requires 8+ years of production experience with ML/AI systems, distributed computing, cloud platforms, and programming in Python, Scala, and Java.
Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.
Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
Build and scale evaluation, experimentation, and quality systems for Cortex Code’s enterprise coding agents. The role requires 10+ years shipping AI/ML-backed production software, strong programming skills, and Staff-level technical leadership across engineering, modeling, and product.
Build and scale Gusto’s machine learning and AI platform, including MLOps pipelines, model deployment frameworks, infrastructure, and observability. The role requires 5+ years of software engineering experience, proficiency in Python, Ruby, or Java, and experience with ML lifecycle infrastructure and cloud platforms.
Build and ship AI-native quality platform features, integrating and evaluating LLMs in real-world applications. The role requires strong software engineering fundamentals, 3+ years of experience, and hands-on expertise in prompt engineering, LLM observability, fine-tuning, and evaluation systems.
Builds scalable experimentation, training, orchestration, and agentic AI infrastructure that accelerates Reddit’s Ads ML lifecycle. Requires 5+ years in infrastructure or distributed systems and production ML platform experience.
Build and operate production machine learning systems for real-time conversational AI, spanning data pipelines, model training, deployment, monitoring, and MLOps. The role requires strong machine learning, deep learning, NLP, Python, and production systems experience.
Build scalable backend and infrastructure systems for reinforcement-learning environments that train AI agents, partnering with researchers and leading AI labs. The role requires strong engineering judgment, hands-on leadership, and experience with ML/LLM concepts and scalable systems.
Leads the design, governance, evaluation, and production delivery of AI systems for public-sector clients. The role requires 7+ years of engineering experience, production AI/ML ownership, expertise in regulated deployments, and the ability to establish technical standards and advise executive stakeholders.
Leads backend and ML infrastructure architecture across distributed processing, GPU fleets, model serving, and reliability. The role requires 10+ years of large-scale systems experience, strong architectural ownership, and a Computer Science degree or equivalent track record.
Leads the architecture and delivery of scalable machine learning infrastructure and customer-facing voice and audio generative AI solutions. The role requires extensive production ML experience, distributed systems expertise, cloud-native operations, and technical leadership across engineering teams.
AI Engineer Intern working with senior engineers to build and evaluate agentic AI systems, benchmarks, model-training pipelines, and safety evaluations. Requires current study in a quantitative field, hands-on ML experience, strong Python fundamentals, and familiarity with PyTorch.
Build, validate, deploy, and maintain machine and deep learning algorithms for biosignal and brain data used in medical devices and precision medicine. The role requires 4+ years of industry experience, production ML expertise, DSP and statistics knowledge, and proficiency with PyTorch or comparable frameworks.
Build and deploy machine learning models for fraud detection, identity verification, and financial risk products across the full lifecycle. The role targets new PhD graduates or early-career researchers with strong foundations in machine learning, statistics, Python, and production software development.
Build and deploy machine learning models for fraud detection, identity verification, and financial risk products across the full lifecycle. The role targets new PhD graduates or early-career researchers with quantitative training, Python experience, and strong analytical and communication skills.
Leads organization-wide AI evaluation and automation programs, establishing quality standards, metrics, production gates, and scalable evaluation infrastructure for LLM and agentic systems. The role requires 8+ years of relevant experience, strong technical systems expertise, and cross-functional leadership.
Technical leader for Nuro’s Behavior & Planning team, developing and deploying advanced machine learning models for autonomous driving. The role requires 7+ years of ML experience, strong Python or C++ skills, and expertise in areas such as sequential decision-making, generative modeling, or foundation models.
Supports autonomous-vehicle road-rule compliance by building data miners and evaluating LLM/VLM triage pipelines across simulation and fleet data. Requires strong Python, PySpark, and SQL skills, multimodal model evaluation experience, and current enrollment in a relevant bachelor's or master's program.
Build and operate production agentic infrastructure, evaluation systems, guardrails, and enterprise integrations that enable safe, observable AI-assisted engineering. The role requires 3+ years of software development experience, recent LLM or agentic systems depth, strong Python and production operations skills, and sound judgment about automation and trustworthiness.
Builds and productionizes post-trained open-source language models and AI agents for safety and customer-care workflows. The role requires applied ML/AI experience, Python and PyTorch expertise, agentic development, generative AI evaluation, and real-time deployment experience.
Technical Product Engineer advising Digital Native Businesses on integrating Claude API into products. Guides customers from discovery to deployment with expertise in LLMs, prompt engineering, agents, and evaluations; requires 4+ years experience and strong Python/TypeScript skills.
Builds core infrastructure for AI software agents in Devin and Windsurf, focusing on long-horizon task execution, tool use, planning, and reliability at scale. Requires strong Python, systems engineering depth, and AI curiosity; onsite in San Francisco.
Build, evaluate, and deploy machine learning systems and production pipelines, while improving product capabilities through experiments and LLM integrations. The role requires at least 3 years of experience shipping ML systems and strong ownership in ambiguous, collaborative environments.
Builds scalable ML infrastructure services for experimentation, model training, serving, and LLM applications. Requires 2+ years of software development experience plus experience with distributed systems, production ML platforms or MLOps, and high-availability operations.