Latest ML Engineering jobs
Job results
Builds benchmarks and runs experiments in document AI, then publishes high-velocity technical content like blogs and analyses to drive developer awareness and adoption. Requires strong Python/ML engineering and rapid technical writing skills.
Technical Lead owning architecture, execution, and evolution of an AI-driven telehealth platform built on GCP. Hands-on player-coach role integrating production LLMs, leading a small senior team, and driving scalable distributed systems in a healthcare setting.
Staff ML Engineer on Spotify's Personalization team, owning ML models and systems for the Home feed and Shortcuts experience. Build and productionize personalized recommendation systems and agentic AI experiences using LLMs, PyTorch, and large-scale infrastructure; drive experimentation, optimization, and technical direction while mentoring engineers.
Build and operate the platform, runtime environments, and evaluation and observability systems powering Commure’s fleet of autonomous AI agents. The role requires strong Python, Linux, cloud, Kubernetes, and production-reliability experience, plus hands-on experience with LLM-powered agents.
Build state-of-the-art end-to-end speech and audio generation systems with a focus on joint audio-video modeling. Own audio representations (VAEs, neural codecs), generative backbones (diffusion/flow-matching transformers), conditioning, alignment for voice cloning and sync with video, plus data flywheel, evaluation, and inference optimization for large-scale multimodal models.
Build and optimize prompts, tools, skills, and memory systems to shape how Perplexity's AI models respond, use tools, and leverage context across products. Requires strong software engineering skills and experience with LLM behavior design.
Build and ship AI-powered analytics agents, LLM-driven workflows, and intelligence features for Databricks' internal GTM platform (Customer Zero). Requires 7+ years software/AI engineering experience, strong Python/SQL, hands-on LLM/RAG/agents experience, and fluency with AI coding tools like Claude.
Senior Data/ML Engineer owning end-to-end data pipelines, feature platforms, and production ML models that power Sardine's real-time fraud, KYC, and compliance decisions. Requires 8+ years building production data and ML systems with deep Python, distributed frameworks, GCP cloud stack, and fraud/risk domain knowledge.
Build and own core layers of a scalable Agentic AI platform for autonomous engineering agents at Okta, including agent identity/auth, knowledge/memory, observability, governance/safety, and orchestration. Requires 5+ years backend engineering experience plus AI/agent exposure; strong distributed systems and security knowledge.
Research Engineer turning efficient foundation model training research into robust high-performance production systems at Together AI. Optimize large-scale training infrastructure, profile bottlenecks, integrate new models, and productionize novel methods in close partnership with scientists.
Builds production-grade machine learning services and orchestration for conversational AI, coordinating vendor and internal LLM systems. Requires at least two years of ML/software engineering experience, strong Python skills, and familiarity with modern AI architectures and tooling.
Lead the ML infrastructure layer with focus on optimizing GPU inference performance for real-time onboard and high-throughput offboard autonomous driving applications. Requires deep experience with PyTorch, C++, GPUs, distributed systems, and performance optimization.
Build RL environments, verifiers, fine-tuning pipelines, and eval systems for frontier AI agents at Labelbox. Requires deep RL post-training experience (SFT + RL methods), strong Python/systems engineering, and the ability to ship production infrastructure at high velocity.
Build and deploy production ML models and intelligent services (including LLMs and RAG) that automate healthcare workflows at Plenful. Requires 5+ years ML/software engineering experience, strong Python skills, and familiarity with modern MLOps and cloud infrastructure.
Lead the architecture and development of Dialpad's autonomous Agentic AI platform, building multi-agent orchestration, memory systems, real-time reasoning, and tool execution for enterprise workflows. Requires 10+ years experience, technical leadership at Staff/Principal level, and deep expertise in LLM platforms, agent frameworks, and production AI infrastructure.
Lead the architecture and development of Dialpad's autonomous Agentic AI platform, building multi-agent orchestration, memory systems, real-time reasoning, and tool execution for enterprise workflows. Requires 10+ years experience, prior technical leadership at Staff/Principal level, and deep expertise in LLM platforms, agent frameworks, and production AI infrastructure.
Lead the data quality team at HUDHUD to build QC systems, validation methods, and experiments that measure and improve training data for frontier AI agents and RL environments. Requires deep data quality intuition, Python/Docker/Linux proficiency, and experience turning research insights into production evaluation pipelines.
Design and build core infrastructure for AI inference on Cloudflare's global network of GPUs and accelerators. Optimize scheduling, routing, reliability, and observability for low-latency, serverless LLM and model serving at the edge. Requires expert Rust and distributed systems experience.
Lead development of state-of-the-art vision, VLM, and VLA models for autonomous perception systems. Own challenging ML problems, build data pipelines and deployment infrastructure for production autonomy, and mentor engineers while bridging research to real-world systems. Requires 7+ years experience, strong CV/ML expertise, and C++/Python proficiency.
Develop and deploy advanced vision, VLM, and VLA machine learning models for autonomous systems perception. Own models from training through optimization and deployment on embedded hardware, collaborating with research and engineering teams to deliver production capabilities for complex real-world environments.
Develop and deploy state-of-the-art vision, VLM, and VLA models for autonomous perception systems. Requires 7+ years experience (with PhD), mastery of ML and computer vision, PyTorch/TensorFlow, TensorRT/ONNX, C++/Python, and ability to obtain SECRET clearance.
Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.
Build foundational AI agent infrastructure at Rippling, owning agent creation, invocation, skill packaging, plugin connectivity, and automations. Lead platform and distributed systems work with 8+ years experience, technical leadership, and cross-functional impact.
Build scalable AI platform infrastructure including agent orchestration frameworks, evaluation systems, and high-performance serving layers to accelerate AI development and deployment at Mixpanel and for customers. Requires 2+ years software engineering experience with hands-on LLM/agent integration and full-stack fundamentals.
Senior MLOps Engineer building scalable ML infrastructure and pipelines for autonomous battery-electric rail vehicles. Lead design of distributed training, experiment tracking, deployment, and monitoring systems for safety-critical perception and autonomy models. Requires 5+ years building large-scale systems with 2+ years in ML infrastructure.
Senior Software Engineer building automated evaluation pipelines, test infrastructure, and monitoring systems to validate quality of Deepgram's speech, audio, LLM, and multimodal AI models before production release. Requires 5+ years building test/evaluation frameworks, strong analytical skills, and backend experience in Python/Rust/Go.
Develop metrics and large-scale evaluation pipelines to assess fidelity of Zoox's GenAI-powered autonomous vehicle simulator. Requires 5+ years in quantitative evaluation or robotics, strong Python/data analysis skills, and statistics knowledge.
Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.
Build and deploy production-grade AI agents for high-stakes fraud, risk, compliance, and AML decisions at enterprise customers. Own the full agent lifecycle from scoping and tool engineering through evals, tuning, and platform impact. Requires strong LLM experience, software engineering fundamentals, and customer-facing skills.
Build and ship AI agent orchestration workflows, integrations, and shared infrastructure that power automated growth across paid media, lifecycle, organic, CRO, incentives and more at Kraken. Requires 3+ years building agentic systems or automation with LLM APIs, strong API integration skills, and growth/marketing context.
Build and maintain production ML systems including training/inference pipelines, model serving via APIs/batch, monitoring for drift, and automated retraining. Productionize models from prototypes with strong Python, MLOps, and reliability focus.
Lead Garner Health Intelligence as Director of Applied Science, owning doctor-ranking algorithms, data assets, and the enterprise Insights platform. Build and manage teams of scientists, researchers, and PMs while driving both technical innovation and commercial business results in healthcare.
Develop and productionalize ML and optimization models for Lyft's pricing and ETA systems. Requires advanced degree, production ML/algorithms experience, and Python proficiency to solve large-scale marketplace problems.
Build and improve production AI agent systems for OpenAI's GTM workflows. Own the end-to-end improvement loop using feedback, evaluation, experimentation and backend services to drive measurable gains in customer engagement, pipeline and team productivity. Requires 4+ years building reliable LLM-powered production systems plus strong product judgment.
Staff Software Engineer driving reinforcement learning infrastructure for Claude's coding capabilities at Anthropic. Design APIs/frameworks, embed with research teams to build and hand off maintainable systems, improve research code reliability, and ensure production RL run health. Requires deep Python expertise, API design track record, and failure-mode intuition.
Build and own Python frameworks, APIs, and infrastructure for Anthropic's RL environments and agent runtimes. Embed with research teams to productionize their work, design for correctness in stateful distributed systems, and create self-service tooling for production debugging.
Build AI-powered tools, quantitative models, autonomous agents, and analytics to automate treasury workflows, enhance risk management, and deliver insights at Stripe. Requires 6-8 years experience, Python proficiency, AI/ML frameworks, and treasury/finance expertise.
Senior Machine Learning Engineer building and deploying production GenAI, agentic systems, LLMs, and computer vision models for mission-critical public sector and defense applications at Scale AI. Requires active security clearance, extensive production ML experience, and strong Python/TF/PyTorch skills.
Quant Developer designing, implementing, deploying, and owning production vault strategies for onchain finance at Gauntlet. Full ownership of quantitative systems from research through live operation, risk management, and on-call monitoring. Requires production quant systems experience, strong Python, and applied statistics/optimization skills.
Build and optimize LLM inference infrastructure at enterprise scale for partner and self-hosted frontier models. Requires 8+ years backend/infrastructure engineering experience with distributed systems, real-time serving, and ML/GPU orchestration.
Build and own AI-powered decision systems and internal tools that augment the compute production team at Fluidstack, leveraging frontier LLMs, agents, and custom integrations to drive measurable productivity gains across operations.
Build and deploy production AI automation that improves global marketing workflows, data quality, content generation, and operational efficiency. The role requires 5–7 years of engineering experience, Python or JavaScript proficiency, API integration skills, and practical LLM application expertise.
Build and ship high-fidelity RL environments, verifiers, and evaluation pipelines from enterprise workflow data to train and assess frontier AI models. Requires prior experience shipping environments or agentic evals plus strong full-stack engineering skills with a bias for action and detail.
Senior Staff Engineer developing state-of-the-art discrete optimization algorithms for task allocation, scheduling, and mission planning of heterogeneous autonomous vehicle teams. Requires PhD/Master's in Optimization or Operations Research, deep expertise in integer programming and solvers, and strong C++ implementation skills.
Build and scale eval systems, fine-tuning pipelines, agent-first infrastructure, and data platforms that power frontier AI labs at Labelbox. Requires full-stack prototyping expertise, strong architecture judgment, daily use of coding agents, and deep TypeScript/Python proficiency (7+ years implied by Staff level).
Senior Backend Engineer building scalable data processing, ML-integrated workflows, and high-volume automation for healthcare referrals and insurance paperwork at Tennr. Requires 5+ years backend experience with JavaScript/TypeScript, production ML integration, and systems architecture skills.
Build AI agents and systems that automate end-to-end data science workflows including hypothesis formation, querying, analysis, and recommendations at Perplexity. Requires 6+ years in data roles, strong SQL/analytics judgment, production Python, hands-on LLM experience, and product sense to create scalable AI-native data infrastructure.
Build and productionize ML measurement, causal inference, and platform tooling at Pinterest. Translate research into scalable pipelines, develop self-serve causal tools, and create centralized systems for feature importance, model evaluation, and infrastructure efficiency.
Lead AI adoption and transformation across a 500+ engineer organization to create "Bionic Engineers" and implement Virtual Engineer capabilities, targeting significant productivity gains. Requires 10+ years software engineering experience, deep expertise with AI coding tools like Copilot and Claude, change management, and executive-level influence.
Develop deep learning models using imitation and reinforcement learning for autonomous driving agents and planning. Requires advanced degree plus experience with RL, transformers, and production ML pipelines.