Latest ML Engineering jobs
Job results
Senior Data Scientist building and improving production LLM agents, workflows, and internal AI applications for Rippling's GTM teams (Sales, RevOps, Customer Success). Combines applied AI development, full-stack engineering, data pipelines, ML experimentation, and evaluation infrastructure.
Senior Application Engineer building and deploying AI Agents with Salesforce Agentforce, intelligent automations, and custom Salesforce applications. Requires 5+ years experience, specific Salesforce certifications (Platform Developer I/II, Advanced Admin, Agentforce Specialist), strong coding and DevOps skills, and mentoring ability.
Lead internal AI automation initiatives by embedding with business teams to map workflows, independently build and deploy Claude-powered agents and tools, measure impact, and enable non-technical staff to create their own automations. Requires 2+ years hands-on Claude/LLM agent building experience in a scrappy, self-directed manner.
Build and deploy production GenAI/LLM applications and intelligent agents to automate workflows across Procurement, Supply Chain, Legal, Finance, HR, and Marketing. Requires 5+ years engineering experience including 2+ years production LLMs, Python, LangChain/LlamaIndex, and cloud AI/vector DB tools.
Senior Applied AI Engineer building the core intelligence layer for Roger, an AI platform for home health clinicians. Responsibilities include training/fine-tuning LLMs on proprietary clinical data, building rigorous eval and monitoring systems, and shipping reliable agentic LLM workflows that improve patient care.
Develop and integrate real-time AI/ML perception algorithms and sensor fusion software for autonomous vehicles across land, air, sea, and space domains. Requires MS/PhD or 5+ years experience with multi-modal sensors, ML deployment, and Linux/Docker; US citizenship and security clearance eligibility mandatory.
Senior AI Builder responsible for designing and shipping production AI agents, agentic workflows, evaluation harnesses, and reusable AI tooling that transform EarnIn's product development lifecycle. Requires 4+ years software engineering experience with strong LLM/agent expertise.
Principal engineer on OpenAI's Codex Cyber team building AI-powered security products for trusted defenders. Define roadmaps, shape model training and safeguards, build evaluations for cyber capabilities and risks, and translate frontier research into production tools while collaborating across product, research, safety, and external partners.
Senior ML Engineer optimizing and productionizing LLMs and other models on Cloudflare's global serverless inference platform. Focus on inference performance, benchmarking, evaluation, and deployment at scale across heterogeneous GPUs and accelerators.
Build and run evaluations to measure Claude's capabilities, safety, and performance. Design metrics, implement scalable distributed eval infrastructure and dashboards, debug training runs, and partner with researchers to characterize and improve AI systems.
Build and deploy production machine learning systems across pricing, marketplace optimization, fraud detection, and agentic AI for Lyft Business. The role requires end-to-end ML ownership, experience with generative AI and LLM ecosystems, and the ability to independently scope high-impact projects.
Build and own the core AI platform powering an AI-native trading copilot. Develop high-performance Rust backend for streaming, tool execution, and safe trading actions; design robust APIs with observability and security. Requires 8+ years systems programming experience.
Own reliability and quality for an AI copilot in a trading platform. Design evaluation systems, benchmarks, quality gates, model improvement loops, and AI monitoring for correctness, safety, and performance in market analysis and trading workflows. Requires 8+ years production software experience and strong ML eval expertise.
Build high-performance Rust backend and core AI platform powering an AI-native trading copilot with streaming, low-latency tools, safe trading actions, and robust APIs. Requires 8+ years experience in systems programming, production services, and architecture.
Senior engineer responsible for building production-grade Python connectors and AI/ML integrations that enable ClickHouse in RAG, feature-store, and LLM application workflows. Requires 7+ years of software development experience and hands-on Data Scientist or ML Engineer experience.
Senior engineer owning Python connectors and AI/ML integrations that connect ClickHouse to RAG, vector search, feature-store, and LLM application workflows. Requires 7+ years of software development experience and hands-on Data Scientist or ML Engineer experience.
Own the research-to-production pipeline at Deepgram, turning experimental speech ML models into reliable, scalable production services. Partner with researchers on robust workflows, automated release gates, inference optimization, and feedback loops across hybrid GPU infrastructure.
Backend Engineer building and optimizing Deepgram's core inference services for speech processing, including networking, audio transcoding, latency/memory optimization, and distributed compute orchestration. Requires 3+ years experience with Rust (or C/C++) and Python.
Lead technical vision for Pinterest's Ads Conversion Core Modeling team. Build state-of-the-art large-scale DNN models for user action prediction, mine multi-modal signals for intent understanding, and mentor engineers while leveraging AI tools to accelerate development.
Build and operate a hosted AI training platform spanning Kubernetes GPU orchestration, Python control-plane services, developer-facing APIs, and monitoring interfaces. The role requires depth across AI infrastructure, distributed training, cloud operations, and full-stack platform development.
Build and optimize large-scale LLM inference and serving infrastructure across cloud GPU fleets, integrating inference systems with RL training. Requires 3+ years operating ML/LLM services, strong distributed systems and GPU expertise, and hands-on experience with modern inference frameworks.
Build and optimize infrastructure for frontier-scale reinforcement learning and distributed model training, including kernels, runtimes, parallelism, and asynchronous rollouts. The role requires strong AI/ML systems experience, PyTorch expertise, and GPU performance optimization skills.
Conducts frontier research and builds scalable synthetic-data and distributed reinforcement-learning infrastructure for large AI models. Requires strong AI/ML engineering experience, distributed inference expertise with tools such as vLLM or SGLang, and MLOps knowledge.
Research Engineer building and optimizing distributed infrastructure for frontier-scale model training and reinforcement learning. The role requires strong AI systems experience, PyTorch and distributed-training expertise, GPU performance optimization, and familiarity with parallelism and large-scale clusters.
Develops reinforcement learning, post-training, and agent systems that advance model reasoning and support real-world workflows. The role combines applied research with scalable training infrastructure, evaluations, and production deployment.
Customer-facing applied research role focused on building AI agents, evaluation systems, and post-training workflows for frontier models. The role combines reinforcement learning, distributed infrastructure, applied data, and close collaboration with customers and research teams.
Forward Deployed Scientist partnering with biotech/pharma customers to build biomedical environments and evaluation pipelines for AI agents. Requires PhD, industry computational biology experience, strong engineering skills, and deep expertise in drug discovery, preclinical, or clinical domains.
Build and operate AI-assisted platform infrastructure, developer tooling, and production workflows across Kubernetes, AWS, GitOps, observability, and FinOps. Requires 4+ years of platform, infrastructure, or backend engineering experience and hands-on experience with MCP servers and agentic AI systems.
Build and productionize vision-language models for document understanding at LlamaIndex. Focus on training, fine-tuning, synthetic data, benchmarking, and turning research prototypes into accurate, low-latency production systems for real-world PDFs, tables, and enterprise docs. Requires 3+ years ML engineering/applied research experience with strong PyTorch skills.
Research Engineer/Scientist shaping personalities and behaviors of personalized AI models like ChatGPT using RL, reward modeling, synthetic data, and post-training methods. Requires strong ML engineering and research experience with large models.
Senior technical leader building, productionizing and operating large-scale ML models and Agentic AI systems to fight fraud, ensure safety and build trust across Airbnb's platform. Requires 12+ years applied ML experience and deep expertise in LLMs/GenAI.
Generalist Research Engineer working across Exa's search and retrieval stack including crawling, parsing, ML performance, and retrieval algorithms to improve search quality and performance for customers.
Scientist building ensemble-aware protein representations and ML models that integrate dynamic structural data with PLMs, ligands, and functional info to advance dynamic structural biology. Requires PhD and experience with large bioinformatic pipelines.
Research Engineers at Distyl build and productionize post-training techniques (fine-tuning, RLHF, reward models, evals) to improve reliability and behavior of compound AI systems for enterprise customers. Requires strong applied ML experimentation skills and ownership of real-world outcomes.
Research Engineers at Distyl build and productionize reliable agentic AI systems and compound architectures for enterprise workflows. They design agents, develop evaluation frameworks, run experiments on reasoning and failure modes, and integrate into customer environments.
Tech Lead Manager to hands-on lead the inference platform team at Luma AI. Own the full serving stack for multimodal models across thousands of GPUs, spending 50%+ time as IC on architecture, optimization, and debugging while growing the team and setting technical direction.
Build and scale production multi-agent systems for customer support, integrating LLMs with internal tools and APIs. The role requires 3+ years of production ML/AI experience, strong Python and microservices expertise, and familiarity with orchestration, RAG, evaluation, and vector databases.
Build and scale AI-powered workflow automation, including prompt-based actions, custom agents, RAG systems, knowledge bases, and advanced data features. The role requires strong JavaScript and Node.js expertise, AI platform experience, and at least four years of software development experience.
Build and optimize the high-performance inference platform serving Grok at massive scale. Design distributed serving infrastructure, low-level GPU optimizations, quantization, speculative decoding, and CI/CD for production reliability and low latency.
Be the first dedicated owner of Confido's ML platform, owning end-to-end ML pipelines, infrastructure for training/inference/agentic workloads, and providing reproducible environments for the AI/ML team in a fast-growing CPG AI startup.
Software Engineer on the Logistics Optimization team designing and implementing algorithms for clinician routing, scheduling, dispatch, simulations, and predictive models to optimize in-home healthcare delivery at national scale. Requires 2-3 years software engineering experience with optimization, forecasting or simulation systems, preferably in TypeScript/Python.
Build and ship applied machine learning features powered by generative and multimodal models, taking projects from experimentation and evaluation through production. The role requires backend Python experience, familiarity with PyTorch or JAX, and the ability to improve quality, latency, or cost.
Early-career engineer building and productionizing AI-powered features at Notion using LLMs and embeddings. Less than 2 years experience; strong fundamentals in algorithms, data structures, and distributed systems required.
Research Engineer building platforms to customize open-source LLMs via fine-tuning, RL, and evaluation. Focus on integrating post-training with inference engines (vLLM, SGLang, TensorRT-LLM), optimizing for RL workloads, and ensuring production reliability. Requires 2+ years ML production experience and strong Python/Go skills.
Staff ML Engineer owning end-to-end lifecycle for enterprise AI at Rippling: design novel architectures (LLMs, RAG, RLHF), build evaluation and self-improving systems, and ship production ML leveraging proprietary data graph. Requires 8+ years engineering with 5+ in ML.
Engineer optimizing RL inference stack for workloads from ablations to production training. Requires experience with large-scale distributed systems, LLM inference, and proficiency in Python/C++/Rust with PyTorch/JAX/CUDA.
Senior Applied AI/ML Engineer responsible for designing, developing, training, and deploying AI/ML products focused on life sciences and drug discovery acceleration. Requires advanced degree, 10+ years AI/ML experience (or 5+ in regulated life sciences), end-to-end product ownership, and expertise in RL, fine-tuning, or LLM agents.
Member of Technical Staff conducting hands-on LLM inference research at Modal. Own end-to-end bets on techniques like speculative decoding, quantization, KV-cache management, and disaggregation to improve cost per token and tail latency on production workloads. Requires strong LLM serving stack expertise and a track record shipping research or systems.
Research Engineer developing novel evaluation frameworks and training strategies for AI systems in life sciences and biology. Requires experience training/evaluating LLMs, Python/ML proficiency, and data pipeline expertise; biology background preferred but not required.
Research Engineer implementing novel ML models for clinical AI, focusing on self-supervised learning, survival analysis, multi-modal data, causality, and interpretability to predict patient outcomes in precision medicine. Requires strong Python/PyTorch skills, deep learning experience, and statistics foundations; publications are a plus.