Latest ML Engineering jobs
Job results
Staff Deep Learning Research Engineer designing and training novel neural network architectures from scratch for Generalized Scene Reconstruction tasks including anomaly detection, neural rendering, and 3D reconstruction. Requires PhD/Master's, 6+ years deep learning research experience, expertise in 3D vision or SLAM, and hybrid work in Columbia, MD.
Senior Offline Mapping Engineer owning accuracy and robustness of offline SfM, multi-view stereo, and 3D reconstruction pipelines. Integrates classical geometry with deep learning (learned matching, depth priors) for challenging real-world environments like textureless and reflective scenes. Requires advanced degree, production 3D vision experience, and hybrid work in Columbia, MD.
Build and optimize high-performance ML inference services and APIs that turn frontier research models (FLUX, Stable Diffusion) into production systems serving millions of requests. Requires experience scaling ML serving infrastructure, GPU optimization, and production backend systems.
Build and own ML/LLM systems for internal operations including forecasting, risk flagging, and document extraction. Ship production agentic systems end-to-end with guardrails and partner with data engineering to integrate predictions into tools.
Build and productionize agentic LLM-powered triage systems and ML/DL pipelines to automate failure analysis for autonomous robots. Requires Master's/PhD in STEM + 4+ years production ML/NLP experience with PyTorch, RAG, Databricks, and AWS.
Senior Manager overseeing ML data operations and labeling quality at Coalition. Define guidelines and metrics, drive Label Studio platform requirements, manage vendors, and partner with ML teams to ensure high-quality labeled datasets for models.
Principal Software Engineer building a new Identity Graph platform for real-time identity resolution, fraud, and risk on the Stytch team at Twilio. Own architecture, high-scale distributed systems, complex data pipelines, and synchronous read models while mentoring the team.
Founding member of a new team building foundational evaluation infrastructure and flywheels for Databricks' AI/Genie Agents. Design scalable tooling for benchmarking, regression detection, and quality measurement that drives continuous agent improvement across research, training, and production.
Applied Scientist building optimization, forecasting, and simulation models to solve complex logistics and clinician-patient matching problems for in-home healthcare delivery. Requires strong operations research foundations, Python/SQL/ML expertise, and experience shipping production decision systems.
Build and operate an automated inference optimization platform spanning control planes, secure partner-side runners, compilers, and evaluation systems for new accelerator hardware. Requires 8+ years building large-scale distributed systems with strong performance and systems expertise.
Develop next-generation 3D occupancy and segmentation networks for autonomous vehicles by fusing Lidar, Camera, and Radar data into temporally consistent voxel representations. Requires MS/PhD + 6+ years experience in 3D CV, multi-modal fusion, and PyTorch.
Intern on the engineering team owning and shipping a real scoped project end-to-end (design to production code) in voice AI, ASR, TTS or related systems. Requires self-motivated builder with first-principles reasoning, AI-first mindset, and ability to quickly learn new languages/codebases.
Lead and mentor a team of applied and clinical researchers as a player-coach. Guide ML research in NLP, LLMs, and clinical applications for radiology, translating ideas into production systems while partnering with clinicians and engineers. Requires MS/PhD and 6+ years applied ML research experience.
Lead the new ML Data Operations function at Suno, sourcing and developing training data for AI music models through external vendors, in-product collection, and internal systems. Build and manage the team while partnering closely with Research, Product, and Engineering; requires 6-8+ years in data labeling/operations and director-level leadership building new functions.
Lead ML infrastructure at Reducto by owning the training and inference stack. Hands-on role (80% building/optimizing) focused on GPU utilization, distributed systems, Kubernetes, kernels, and high-performance serving for AI document workflows. Requires 5+ years production ML infra experience and strong systems engineering skills.
Build and productionize agentic AI systems for a cloud-based analytics platform, including routers, tool-calling agents, RAG pipelines, evaluation, and observability. The role requires strong Python and microservices experience plus familiarity with cloud, Kubernetes, and distributed architectures.
Build core AI platform infrastructure at Harvey including model routing, context management, agent architecture, and shared evaluation frameworks that power all agentic legal AI products. Requires 5+ years backend experience with 1+ year in AI/ML, production multi-model systems, and platform-building track record.
Build and shape the Foundation Model API serving layer for large-scale LLM inference (partner and self-hosted models) at Databricks. Requires 8+ years backend/infra engineering experience with distributed systems, ML infrastructure, and a strong product ownership mindset.
Member of Technical Staff on Atlas, Basis's internal team building zero-human AI agent loops for end-to-end knowledge work at an AI-native company. Requires first-principles systems thinking, deep domain understanding, outcome ownership, and in-person work in NYC.
Build and ship production-grade AI-powered features and agentic workflows for analytics and large-dataset use cases. The role requires strong software engineering experience, practical GenAI and LLM expertise, cloud-native exposure, and a rapid experimentation mindset.
Build and productionize scalable, highly available distributed systems and infrastructure for Nuro's data labeling platform that powers autonomous driving ML models. Requires 5+ years experience with reliable large-scale data systems, technical leadership, and strong programming skills in Python, C++, or Go.
Staff Software Engineer owning architecture and delivery of LLM-powered agentic workflows for marketing content creation, publishing, and reliability at scale. Requires 12+ years experience building production LLM systems, human-in-the-loop designs, stakeholder collaboration, and mentoring.
Principal Software Engineer driving technical vision for Upstart's Core Pricing platform. Lead design of large-scale ML-powered systems to optimize loan pricing, borrower-lender matching, and marketplace efficiency for personal and unsecured loans.
Build and productionize internal agentic workflows and tooling on OpenRouter to automate support and go-to-market operations. Requires build-over-buy conviction, reliability focus, domain knowledge in support/GTM, backend systems expertise, security mindset, and quantitative evals.
Senior ML Engineer leading architecture and strategy for LLM- and agent-powered quality evaluation systems at Block. Builds scalable AI frameworks to measure product behavior, detect issues, and generate insights across millions of interactions to improve reliability and decision-making.
Build and ship production AI agents on Cloudflare's edge platform using Workers, Durable Objects, and AI tools. Requires strong TypeScript/Rust experience, observability expertise, and hands-on LLM tooling for evals, safety, and multi-agent systems.
Build and productionize cutting-edge ML models for Airbnb's query intelligence, including autocomplete, query tagging, expansion, intent modeling, and LLM-powered natural language search to understand guest intent.
Build and maintain large-scale distributed inference systems serving Claude to millions of users. Design intelligent routing, autoscaling, and deployment pipelines across diverse AI accelerators while maximizing compute efficiency for production and research workloads. Requires significant distributed systems experience.
Characterize, analyze, and optimize performance of state-of-the-art AI models on Cerebras' wafer-scale hardware. Build performance models, optimize kernels and compilers, debug runtime behavior, and develop visualization tools to influence next-gen AI architecture.
Senior ML Engineer building observability, evaluation frameworks, and improvement loops for production agentic AI systems. Requires 5+ years production ML/LLM experience, strong grounding in agent design or evaluation, and hands-on work taking systems from prototype to scale.
Build and optimize the RL training framework and infrastructure for large-scale workloads at SpaceXAI, from ablations to production runs. Requires experience with distributed systems and proficiency in Python, JAX, Rust, or C++.
Research Engineer on OpenAI's Privacy team designing and prototyping privacy-preserving ML algorithms like differential privacy and federated learning at scale. Requires hands-on PETs experience, fluency in PyTorch/JAX, and a track record implementing or publishing novel privacy work.
Research Engineer building self-improving AI agent systems at Console. Develop eval/optimization loops, fine-tune specialist models, and improve agent reasoning over enterprise context using production data to drive measurable gains in quality, latency, and reliability.
Build and optimize scalable AI infrastructure for real-time inference, evaluation, and continuous improvement of LLMs, LVMs, computer vision, and multimodal models on large-scale video data. Requires 4+ years production ML systems experience, strong Python skills, and expertise in inference optimization and model serving.
Build and own validation pipelines, CI/CD infrastructure, and platform integrations to launch frontier models and inference features reliably across AWS, GCP, and Azure. Requires strong large-scale distributed systems experience and track record improving release velocity.
Build and scale the shared AI platform foundations at Notion, enabling fast and safe shipping of AI products. Requires experience with LLM/ML platforms, strong ownership, and comfort across backend, infrastructure, and product code.
Senior Staff Software Engineer building an AI agent platform and automated workflows to transform Coinbase's Legal organization. Architect production-grade LLM and multi-agent systems that replace manual legal processes such as agreement redlining and governance.
Research Engineer advancing RL for silicon chip design at Anthropic. Design RL environments for RTL generation, verification, and physical optimization; requires deep ASIC/FPGA expertise from spec to tapeout.
Machine Learning Engineer building statistical models, optimization systems, and experiments for mobile ad tech economics on the Revenue Engine team. Requires PhD in CS/ML/Economics and industry experience applying ML or economics at scale.
Build synthetic data pipelines and tasks to train frontier AI agents. Requires Python, Docker, Linux, and experience creating realistic, scalable synthetic training data for models and evals.
Build QC automation systems for RL training data and agent evals at HUDHUD. Design quality standards, validation pipelines, experiments and metrics without heavy LLM reliance; partner with vendors to debug and improve data generation. Requires Python, Docker, Linux and experience building scalable QA/QC systems end-to-end.
Applied Research Engineer owning technical deployment requests, troubleshooting, building tools/pipelines, and resolving ambiguous problems for frontier AI labs and data vendors at HUDHUD. Requires strong Python/Docker/Linux skills, eval/benchmark experience, independent problem-solving in fast-paced ambiguous settings.
Staff Machine Learning Engineer developing state-of-the-art visual encoders and multimodal models at Pinterest Labs. Prototype visual reasoning tools, train billion-scale models on rich visual-text data, ship to production for recommender systems and VLMs, publish research, and mentor juniors. Requires strong CV/ML background, publications, and PhD or equivalent.
Build and own ML models, fine-tuning, evaluation harnesses, and routing for Kepler's AI agent harness in finance. Requires 5+ years production software experience and shipped ML systems focused on correctness, evals, and real-world reliability.
Build and deploy machine-learning perception systems for L4 autonomous trucks, spanning multimodal 3D detection, sensor fusion, tracking, localization, and safety validation. Requires 5+ years of perception, computer vision, or robotics experience with strong Python and C++ skills.
Machine Learning Engineer owning the full ML lifecycle for multimodal video datasets at Sieve. Fine-tune VLMs, build evaluation/QA pipelines with frontier models, design filtering systems over internet-scale data, and ship production improvements for top AI labs. Requires strong Python, PyTorch, and production ML experience.
Early-career Quantitative Researcher developing ML models and trading signals to predict asset movements and build market-neutral portfolios. Requires STEM degree, Python fluency, and passion for machine learning; embedded in Alpha, Data Science, or Strategy teams.
Build and own evaluation systems, metrics, pipelines, and datasets to rigorously measure the quality of Firecrawl's LLM-ready web data outputs at massive scale, driving model and product improvements.
Build production ML systems for measuring, predicting, and scaling data quality for frontier AI models. Requires 3-6 years experience in applied ML or related production systems (ranking, recommendations, data quality, fraud) plus strong software engineering skills.
Develop synthetic sensor simulation models and algorithms using ML techniques like NeRF and Gaussian splatting to generate photorealistic images and realistic lidar/radar data for autonomous vehicles. Requires advanced degree plus 3-5+ years experience, strong ML fundamentals, and Python/deep learning expertise.