Latest ML Engineering jobs
Job results
Senior Research Engineer tailoring and deploying machine learning models for partner applications across geospatial and environmental domains. The role requires PyTorch expertise, end-to-end ML deployment experience, geospatial tools knowledge, and strong independent execution.
Staff Machine Learning Engineer building and operating production ML systems for causal marketing measurement, optimization, and planning. The role requires deep statistical and machine learning expertise, production programming experience, cross-functional collaboration, and technical mentorship.
Own the operational reliability, governance, and lifecycle maintenance of production AI solutions for public-sector and enterprise customers. The role combines software engineering, MLOps, incident management, automation, model monitoring, and senior client communication.
Leads Discord’s Safety ML team, setting technical direction and overseeing production machine learning systems for content understanding, account integrity, and platform abuse. Requires substantial machine learning and engineering management experience, hands-on technical depth, and experience delivering ML systems at scale.
Build and ship production AI/ML systems for real-time incident management, including LLM agents, retrieval pipelines, and inference services. The role suits an early-career software engineer with 2+ years of production experience and hands-on experience with modern AI.
Build and operate production AI systems, including LLM agents, retrieval pipelines, and event intelligence, across PagerDuty’s high-scale distributed platform. The role requires 5+ years of software engineering experience, production distributed-systems expertise, and hands-on experience shipping reliable LLM applications.
Owns end-to-end production machine learning systems, including NLP, LLM, agentic, ranking, and recommendation capabilities. Requires 8+ years of industry experience, strong Python and cloud ML expertise, and the ability to deliver explainable AI products with cross-functional and customer impact.
Six-month remote internship for final-year Computer Science students working across AI and data engineering. Residents develop and evaluate machine learning models, build backend and data pipelines, and analyze large datasets while receiving mentorship and a monthly stipend.
Build and improve Magic Patterns’ AI design agent by owning its evaluation infrastructure, self-improvement loop, and model-quality initiatives. The role requires experience with AI models or agents, evaluation systems, post-training, or context engineering, plus strong first-principles problem solving.
The Senior Algorithm Engineer leads development and production deployment of machine and deep learning algorithms for biosignal and medical-device applications. The role requires 5+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated health or similar domains.
Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and operate production AI agent systems that help Sales and Marketing teams with account planning, deal support, competitive intelligence, and content creation. The role requires 6+ years of experience shipping reliable LLM workflows with retrieval, tool use, permissions, evaluation, and observability.
Principal technical leader defining architecture and multi-year strategy for Pinterest’s Homefeed, Search, and AI Assistant experiences. The role requires 15+ years of large-scale systems or machine-learning experience, deep expertise in discovery and generative AI, and hands-on leadership across engineering and product organizations.
Leads the design, deployment, and optimization of agentic and generative AI systems that automate risk and compliance investigations at scale. Requires 8+ years of machine learning modeling experience, production ML expertise, and advanced technical education.
Own North’s evaluation strategy and build feedback systems that connect real enterprise workflows, user needs, and production failures to model improvements. The role requires strong applied machine learning judgment and the ability to translate product signals into rigorous, actionable evaluations.
Build and productionize post-training systems for voice and text agents, including environments, verifiers, synthetic data, evaluations, and model training. The role requires strong Python and end-to-end model development experience, with reinforcement learning and distributed training expertise preferred.
Build and productionize Agnes, an AI supply-chain manager, by developing LLM-powered agents, automations, evaluation pipelines, and observability systems. The role requires hands-on experience with modern LLM APIs, Hugging Face, prompting, system design, and production AI workflows.
Measures and improves healthcare AI agents in production by instrumenting performance, diagnosing quality issues, and designing experiments with clinical experts. The role requires Python, analytical judgment, comfort with ambiguity, and interest in healthcare.
Design and scale ML infrastructure and real-time learning systems powering personalization, search, ranking, and ad tech for millions of consumers. The role requires deep distributed-systems and data-pipeline expertise, strong architecture leadership, and experience delivering zero-to-one ML systems.
Build and ship production Applied AI capabilities, including agent infrastructure, RAG services, evaluation systems, and AI-powered engineering workflows. The role requires 6+ years of software engineering experience, strong backend and distributed-systems skills, and direct experience delivering LLM- or ML-powered products.
The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.
Build and productionize scalable machine learning models and systems for underwriting and portfolio management. The role requires a bachelor's degree and at least two years of experience shipping ML systems, plus expertise in model development, deployment, data pipelines, and deep learning.
Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.
Build and lead production AI products, including agent systems that use tools, retrieve context, and complete complex tasks reliably. The role requires strong full-stack engineering, deep language-model experience, product judgment, and ownership from experimentation through production.
Build reproducible systems for AI model benchmarking, including datasets, evaluation pipelines, containerized environments, scoreboards, and analysis tools. The role requires 4+ years of professional engineering experience, strong Python, dataset rigor, and Docker expertise.
Build and evaluate reinforcement learning environments and AI agents for enterprise workflows. The role focuses on agent training, capability measurement, verifier and reward design, synthetic data pipelines, and improving agent performance through systematic evaluation.
Build and advance multilingual language models through scalable ML systems, research, and large-scale data processing. The role requires deep NLP expertise, strong Python and software engineering skills, and a PhD or equivalent experience, with opportunities to publish research and mentor teammates.
Build production infrastructure for replayable enterprise environments, agent evaluation, and continuous model improvement. The role combines hands-on customer deployment, research experimentation, large-scale data processing, and production software engineering.
Build and operate edge MLOps infrastructure for smart-camera machine-learning systems, including model deployment, TensorRT compilation, fleet updates, telemetry, and reliability. The role requires production MLOps experience, embedded inference optimization, and strong collaboration with data-science and embedded-engineering teams.
Leads the development and production deployment of large-scale ASR and TTS systems for conversational intelligence products. The role requires 5+ years of industry experience, deep speech-model expertise, and strong software engineering and ML operations capabilities.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.
Build and operate AI-driven workflows, integrations, model-routing systems, and agent infrastructure across business and engineering functions. The role also supports model evaluation, AI cost optimization, governance safeguards, and organization-wide training.
Build and ship production AI agents and the platform infrastructure that makes them reliable, steerable, and measurable. The role requires strong backend fundamentals, production LLM or agent experience, and expertise in evaluations, retrieval, orchestration, or tool-use design.
Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.
Builds and productionizes machine learning systems for trust and safety, including abuse detection, autonomous AI agents, and evaluation frameworks. The role requires 5+ years of applied ML experience, strong Python skills, experience with LLMs and scalable pipelines, and a relevant advanced degree or equivalent background.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
Builds secure, vendor-agnostic agent platform primitives, evaluation tooling, safety controls, retrieval systems, and production workflows. Requires 2+ years of software engineering experience plus strong foundations in LLMs, agentic AI, GPU computing, model serving, and distributed systems.
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
Build and ship production algorithmic systems that improve healthcare quality, access, and cost outcomes. The role combines machine learning, optimization, experimentation, and LLM productionization, requiring at least two years of relevant industry or advanced-degree experience.
Leads machine learning strategy and hands-on development for music promotion and royalty products, building scalable production systems and guiding complex technical initiatives. Requires deep machine learning expertise, production-scale implementation experience, and strong technical leadership.
Develops and validates the autonomy stack for adaptive orchestration of heterogeneous UAV, USV, and UUV fleets. The role requires 6+ years in autonomy or robotics, strong C++ and Python skills, experience with planning, navigation, middleware, simulation, and eligibility for a security clearance.
Owns the quality, efficiency, and reliability of Cortex Code, Snowflake’s coding agent for data workflows. The role combines production software engineering, AI/LLM evaluation, experimentation, observability, and systems optimization, requiring 6+ years of experience and proficiency in Python, TypeScript, or Go.
Build and own production AI agent systems spanning orchestration, backend services, integrations, knowledge infrastructure, trust and safety, and evaluation. The role requires 5+ years of backend engineering experience, hands-on LLM or agent experience, and strong ownership in an ambiguous product environment.
Build and operate production machine learning systems for recommendations, search, advertising, content understanding, and LLM-powered experiences at internet scale. The role requires 3–5+ years of production ML experience, strong programming and software engineering fundamentals, and expertise with modern ML frameworks and scalable pipelines.
Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.
Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.
Owns end-to-end post-training for frontier multimodal generative models, spanning reward modeling, preference optimization, distillation, safety tuning, evaluation, and deployment. The role requires prior experience shipping post-training improvements and strong PyTorch expertise.
Build scalable software and tools for frontier model training, research experimentation, and production machine learning systems. The role requires strong Python and distributed-training expertise, experience with ML frameworks and infrastructure, and the ability to optimize and debug large language model systems.
Leads the technical vision, architecture, and roadmap for a company-wide machine learning platform supporting model training, deployment, serving, monitoring, and generative AI. The role requires expert Python and Java skills, large-scale MLOps experience, cloud and Kubernetes expertise, and organization-wide technical leadership.