Latest ML Engineering jobs
Job results
Lead Software Engineer architecting and implementing Model Predictive Control (MPC) systems and vehicle dynamics models for autonomous vehicles. 80% hands-on C++ development and optimization with 20% technical leadership and mentoring of a small controls team. Requires 6+ years experience, deep MPC/optimization expertise, and strong individual contribution.
Lead the architecture and development of a scalable, unified AI/ML platform supporting traditional ML models, LLMs, generative AI, and multi-agent systems. Requires deep expertise in ML engineering, LLMOps, agentic frameworks, Python, cloud infrastructure, and technical leadership.
Steward Roboflow's open-source Inference engine by building an agentic-driven contribution, review, CI/CD, and testing pipeline to enable daily releases from high-volume AI-generated PRs while maintaining quality. Own model integration, community enablement, documentation, and customer education for computer vision inference across cloud and edge.
Leads hands-on architecture and implementation for an AI-powered React application builder, including LLM pipelines, distributed backend services, generation systems, and deployment infrastructure. Requires 6+ years of backend engineering experience, strong Go and Node/NestJS skills, and expertise in generative AI, reliability, and scalable systems.
Lead and mentor an MLOps team to build scalable ML infrastructure, automated pipelines, and low-latency model serving for large tabular models. Requires 7+ years MLOps experience including 3+ years leading teams, deep expertise in Kubernetes, model serving frameworks, and observability tools.
Lead and grow a team developing large-scale Reinforcement Learning and ML models for onboard autonomous vehicle behavior and driving plans. Requires MS/PhD, 5+ years production ML experience (3+ in leadership), and expertise in RL/ML for planning, LLMs, or related areas.
Staff ML Engineer building CustomerLake, Databricks' Customer Data Platform for enterprise ML/AI personalization, recommendations, churn, and LTV modeling. Requires 10+ years shipping production ML/LLM systems with strong product mindset in 0-to-1 environments.
Quantitative Researcher developing cutting-edge ML and statistical models to identify predictive signals in global financial markets. Requires PhD in a STEM field, strong Python skills, and interest in ML/AI; no finance experience needed. Matched to Alpha Research, Data Science, or Strategy Research teams.
Lead a team developing large-scale RL and ML models for autonomous vehicle behavior planning and driving decisions. Requires RL/ML expertise, production ML pipeline experience, and 3+ years in leadership.
Senior technical IC owning the Modeling to ML Serving to API architecture for Airbnb's Host Pricing platform. Lead unified serving stack, backfill/evaluation infrastructure, and domain contracts between ML modeling and serving teams.
Optimize and deploy large multi-modal foundation models (LLMs, VLMs) for real-time inference on power-constrained vehicle SoCs. Focus on quantization, custom CUDA kernels, TensorRT pipelines, resource allocation, and low-latency concurrent C++ code for autonomous systems.
Sr. Staff ML Engineer building backend systems, statistical models, and experiments to optimize Pinterest's ads marketplace and balance long/short-term objectives. Requires MS/PhD (or equiv), 7+ years experience, strong software engineering and math skills.
Research Engineer advancing Claude's computer use capabilities through experiments, RL environments, evaluations, and infrastructure for perception and agentic tasks. Requires Python, ML training/evaluation experience, and a focus on safe AI.
Technical leader on Datadog's APM team building and deploying GenAI/ML models for agentic investigations, automated troubleshooting, and incident triaging. Requires 10+ years experience leading large-scale GenAI initiatives end-to-end in product environments.
Build, train, validate, and deploy computer vision and ML models on 2D/3D CT scan data for customer defect detection, dimensional analysis, and process control. Own full model lifecycle with manufacturing customers (70% model engineering, 30% on-site scanner operation/analysis), translate needs into product features, and contribute to internal AI/ML strategy. Requires 3+ years CV/ML experience and customer-facing skills.
Lead development of computational infrastructure for large-scale neuroscience research, building real-time stimulus/acquisition systems, high-throughput data pipelines, and cloud tools. Requires exceptional full-stack engineering with real-time systems experience; neuroscience background not required.
Senior Engineer building production-grade applied AI systems, LLM-powered tools, retrieval/memory patterns, and human-in-the-loop workflows to create organizational intelligence for home care operations. Requires Python experience, production system design skills, and cross-functional collaboration.
Research Engineer focused on optimizing and stabilizing large-scale GPU training for multimodal generative models. The role combines low-level kernel and precision work, distributed training debugging, profiling, benchmarking, and close collaboration with researchers.
Research engineer focused on post-training LLMs and agents for legal work. Requires hands-on experience training open-weight models and strong Python/research engineering skills.
Lead architect and builder of large-scale ASR/NLP/LLM systems for Otter's conversational intelligence products. Owns end-to-end ML lifecycles from research to production deployment, mentoring engineers and setting technical direction.
Lead projects building and deploying large-scale ASR, NLP, and LLM systems for meeting intelligence. Requires 5+ years building production ML systems with PyTorch/JAX and experience with speech/language models.
Build full-stack AI prototypes and agentic systems to pressure-test venture ideas. Requires 3+ years building production AI applications with strong frontend/backend fluency and frontier coding agent expertise.
Build and evolve auction, bidding, and budgeting ML systems that power Reddit Ads. Design optimization algorithms balancing advertiser performance, user experience, and marketplace efficiency.
Train frontier models to generate polished artifacts (docs, spreadsheets, slides) by owning post-training improvements across RL, data, evals, and alignment. Requires strong ML fundamentals and hands-on LLM/RL experience.
Train frontier models to operate computers, browsers, and desktops. Design experiments, build evals, own post-training pipelines (RL, data, graders), and ship improvements into OpenAI agents.
Train frontier agents to interface with professional software via code, APIs, and structured integrations. Design experiments, own post-training improvements (RL, evals, data), and ship capabilities into major model runs.
Context Researcher on the Agent Post-Training team scaling compute on context for frontier agent models. Designs experiments, owns post-training improvements, builds evals, and ships capabilities into Codex and ChatGPT.
Help shape OpenAI agent personality by turning qualitative collaboration insights into evals, training data, reward signals, and model improvements that reach production.
Improve agentic model capabilities, reliability, and product fit for power users and API developers through evals, training data, and post-training interventions.
Research engineer/scientist building and evaluating vision capabilities for Claude models. Requires 7+ years ML/computer vision experience and work across pretraining, RL, and agentic infrastructure.
Senior engineer building and scaling Mercury's LLM-powered financial assistant Command. Owns full-stack AI product development from system prompts and agentic workflows to eval infrastructure and production reliability.
Build and deploy ML models for entity resolution and knowledge graph expansion on large-scale China-related data. Requires 4+ years clustering ML experience and end-to-end production ML with Python/SQL.
Own the architecture, model design, and production reliability of Taskrabbit's core ranking and matching system for its two-sided marketplace. Requires 8+ years building production ML systems with deep expertise in ranking/recommenders, debiasing, and experimentation at scale.
Build and integrate real-time perception algorithms for autonomous aircraft, including object detection, multi-target tracking, sensor fusion, and state estimation across vision, radar, and inertial sensors. The role requires a relevant engineering degree and 7–10+ years of related experience depending on education.
Research Engineer/Scientist improving model capabilities for personalized AI experiences. Focus on tool-use, instruction following, evaluations, and training improvements. Requires strong ML engineering and research experience.
Improve capabilities, reliability, and product fit of OpenAI's agentic models through research, infrastructure, evals, and training. Work across RL, data, model behavior, and product integration for frontier agents.
Hybrid ML/SRE role owning reliability, security, and safety of a large fleet of generative media model APIs (image, video, audio). Build observability for ML-specific failures, harden deployments, operationalize safety systems, lead incident response, and improve GPU fleet efficiency.
Build and scale real-time TTS serving infrastructure for voice AI models, from GPU inference engines to production APIs. Requires hands-on experience with multinode ML serving frameworks, distributed inference, and cloud/SRE practices.
Technical internships and new graduate opportunities involve tailored work for curious, resourceful builders across Medal and General Intuition, focused on action models and world models for virtual and physical environments.
Build and optimize backend infrastructure that powers high-performance generative AI workloads. The role requires experience scaling enterprise machine-learning systems and working with ML infrastructure such as PyTorch, Vertex AI, or SageMaker.
Build and deploy cutting-edge ML and Generative AI systems to transform Airbnb's customer support experience, focusing on LLM fine-tuning, RAG, and intelligent service automation.
Lead ML engineering on OpenAI's Integrity team to build, deploy, and optimize LLMs and classifiers for content understanding, abuse prevention, and platform safety. Requires advanced degree, deep learning expertise, and LLM fine-tuning experience.
Leads and scales an India-based machine learning organization, owning production ML systems and measurable business impact across risk, growth, personalization, and operations. Requires 7+ years of industry experience, people-management experience, and strong Python, SQL, and applied ML expertise.
Own the training pipeline for search and agent models, building from product usage data through fine-tuning and evaluation to production deployment. Requires deep expertise in transformer fine-tuning, data curation, and training models for ranking, retrieval, and agent behavior.
Own the multi-stage ranking pipeline for web-scale search, balancing precision, recall, latency, and compute cost across retrieval, first-pass ranking, and neural reranking.
Build and productionize multimodal AI systems that turn unstructured healthcare documents, faxes, and transcripts into reliable structured data and automated workflows using LLMs, OCR, and voice AI.
Build and operate production agentic AI systems for Elliptic’s compliance copilot, including LLM workflows, evaluations, retrieval, APIs, and backend services. The role requires senior software engineering experience, production LLM expertise, strong TypeScript/Node.js skills, and experience mentoring engineers.
Staff Engineer building and shipping LLM-powered retirement features end-to-end. Owns architecture decisions, integrates models, and collaborates closely with product and design.
Senior individual contributor architecting and scaling agentic LLM systems that turn messy manufacturing data into reliable root-cause insights. Owns orchestration, retrieval, evaluation, and guardrails for non-deterministic production systems.
Research and implement high-performance vector indexing and retrieval algorithms for Milvus and Zilliz Cloud. Requires 3+ years in vector search or HPC, strong C++ or Rust skills, and a research-driven engineering mindset.