Latest ML Engineering jobs
Job results
Develops and deploys cutting-edge AI models and LLM-based solutions for internal tools and customer products. Collaborates with product/engineering teams to prototype, iterate, and shape technical direction in a fast-paced startup environment. Requires 3+ years experience with Python, TypeScript, and AI frameworks.
Designs, builds, and maintains ML training and serving infrastructure, providing support to research teams. Requires 4+ years in ML infrastructure, cloud platforms like Kubernetes and Google Cloud, and GPU experience.
Leads GPU inference engineering for Sora, optimizing model serving efficiency, kernel-level performance, and scalability. Collaborates with research and product teams to build reliable infrastructure for multimodal AI models.
Develops advanced LLM-based knowledge graphs, RAG techniques, agents, and NLP features to enhance Onyx's AI knowledge retrieval platform. Requires 3+ years AI/ML experience with PyTorch/TensorFlow and strong software engineering skills.
Designs, develops, and deploys scalable ML systems and backend infrastructure for healthcare applications, translating LLM research into production while handling large-scale clinical data. Requires 3+ years ML backend experience, 5+ years software development, and backend languages like Python.
Develops novel AI agent applications using language models, manages model alpha program with OpenAI, architects risk AI workflows, and conducts experiments to evaluate model capabilities for B2B SaaS products.
Build and deploy retrieval, ranking, classification, and LLM-based systems that improve large-scale search quality. The role requires deep search or recommender-systems expertise and at least five years of relevant project experience.
Designs, develops, and deploys AI-driven agentic systems and integrates emerging AI research to enhance platform capabilities for software organizations. Requires 2+ years AI/ML experience, proficiency with LLMs, and Bachelor's/Master's in CS or related field.
Designs evaluation frameworks and benchmarks to test AI agents' autonomy, reasoning, and reliability in data pipelines and warehouses. Requires experience in LLM benchmarking, reinforcement learning, Python, PyTorch/JAX, and data engineering tools.
Build and own backend infrastructure for an AI-powered wealth manager, including multi-agent LLM systems, financial-data platforms, and agentic financial-planning capabilities. Requires 5+ years of backend engineering experience and proficiency with Node.js, TypeScript, Python, PostgreSQL, AWS, and distributed-systems technologies.
Develops RL environments and fine-tunes language models using PPO, DPO, and KTO to enhance agentic capabilities for data infrastructure tasks. Requires deep RL expertise, LLM fine-tuning knowledge, and strong problem-solving skills.
Build and maintain anti-abuse and content moderation infrastructure to ensure AI safety. Collaborate with engineers on AI alignment techniques, incident response, and risk mitigation using Python and cloud tools.
Optimizes and extends ML model serving infrastructure for LLMs, speech, and vision models, focusing on high-throughput, low-latency inference using frameworks like VLLM and SGLang. Requires deep PyTorch expertise, systems programming, and performance engineering for reliable production deployment.
Staff Engineer architects scalable agentic AI systems for property management platform, integrating LLMs to automate workflows. Requires hands-on expertise in agentic frameworks, React/Python/Node.js, and leading AI feature development in startup environment.
Builds and optimizes distributed training infrastructure for large-scale multimodal AI models across thousands of GPUs. Requires deep expertise in PyTorch, CUDA, parallelization techniques, and GPU clusters.
Develops and deploys ML models for NLP, retrieval, ranking, reasoning, dialog, and code-generation systems. Requires Master's/PhD, 2+ years experience with production ML, deep NLP expertise, Python, and frameworks like PyTorch/TensorFlow.
As an Autonomy Engineer, you will develop, integrate, and test core path planning capabilities for aerial platforms. This role involves writing software for real autonomous aircraft systems and collaborating with DoD experts to integrate autonomy software onto OEM hardware.
Build and optimize scalable infrastructure for training large frontier AI models, bridging research and production. The role requires strong software engineering, Python and ML framework expertise, and hands-on experience with distributed training at scale.
Optimizes training performance for advanced language models by developing scalable software, GPU kernels, distributed training systems, and profiling tools. The role requires strong software engineering skills, Python and ML framework proficiency, and experience with CUDA or Triton.
Develops LLM-powered tools to boost developer productivity 10x, including code assistance, automated reviews, and task automation for autonomy software. Requires 7+ years experience as full-stack/backend developer fluent in Python or C++ with LLM optimization knowledge.
Optimizes large AI models for high-volume, low-latency production and research environments. Collaborates with researchers and engineers on inference stack performance, requiring 5+ years experience with PyTorch, GPUs, CUDA, and distributed systems.
Pioneers post-training techniques to enhance LLMs for agentic systems, focusing on tool-use, continuous updates, synthetic data infrastructure, and capability evaluations. Requires Python/PyTorch proficiency, post-training expertise, and proven research impact.
Designs and implements core Python framework components for building real-time AI agents that see, hear, and speak. Owns features end-to-end with strong API design skills and experience in production Python systems.
Builds and productionizes AI models and systems for life sciences document generation, focusing on LLMs, NLP robustness, and reliable deployment pipelines. Bridges ML research, software engineering, and product needs in a high-stakes domain.
Designs, develops, and deploys ML models focused on fine-tuning multimodal LLMs for fraud detection and application automation in financial profiles. Requires 3+ years experience with Python, PyTorch, and expertise in information extraction from financial documents.
Build and ship LLM-powered product features, evaluation frameworks, and retrieval systems end-to-end. The role requires production experience with multi-provider LLM solutions, large-context architectures, and TypeScript, React.js, and Node.js, with regular in-person work in London.
Owns end-to-end ML lifecycle from prototyping clinical prediction models to productionizing and deploying them using production-grade Python/SQL. Requires PhD +3yrs or Master's +5yrs experience with MLOps tools like SageMaker/MLFlow.
Engineers optimize ML systems for performance at scale, focusing on GPU utilization, inference engines, and container runtime to boost throughput and reduce latency for language and diffusion models. Requires 5+ years experience with PyTorch, CUDA, and performance debugging.
Builds state-of-the-art document processing infrastructure using LLMs, including QA agents, optimizers, multimodal models, and self-correcting systems. Monitors production models, runs experiments, and owns large product areas for real-world customer impact.
As a Software Engineer on the Behavior Capabilities team, you will develop and implement algorithmic advancements to expand the robot's driving abilities in complex scenarios, focusing on improving trip progress and vehicle uptime.
Hands-on AI Engineer prototyping and refining LLM-based features, upgrading capabilities with prompt engineering and fine-tuning, while creating evaluations and migration processes for Fathom's meeting AI product. Requires Python proficiency, analytics skills, and Master's degree.
Leads development of ML algorithms for robot perception, including scene understanding, tracking, segmentation, and multi-modal foundation models using sensor data. Requires deep expertise in deep learning, computer vision, and production ML pipelines, collaborating across autonomy teams.
Develops deep learning models using imitation and reinforcement learning to generate safe, efficient driving trajectories for autonomous vehicles. Collaborates with Perception, Planning, and Simulation teams; requires ML expertise, Python fluency, and transformer experience.
Build and scale production AI systems while researching and experimenting with novel modeling ideas. The role requires strong software engineering, Python and ML framework proficiency, GPU kernel development, distributed training experience, and familiarity with Transformer-based sequence models.
Develops APIs and optimizes large-scale ML model inference using Python, Rust, C++, PyTorch, and CUDA. Benchmarks performance, improves reliability, and implements LLM optimizations on GPU architectures.
Develops compiler algorithms and optimization passes to lower and optimize deep learning and high-performance computing workloads for Quadric’s edge-focused neural processing architecture. Requires advanced computer science education, eight or more years of industry experience, and expertise in optimization, graphs, and machine-learning algorithms.
Develops and deploys state-of-the-art ML models for AI music generation, owning full research projects from data engineering to evaluation. Requires 5+ years training large generative models like LLMs/diffusion with distributed PyTorch.
Designs and deploys AI-driven agentic systems for retrieval, code generation evaluation, and agent backend architecture. Requires 2+ years AI/ML experience with LLMs, Bachelor's or Master's in CS/AI, and ability to prototype innovative solutions onsite in San Francisco.
Leads development and deployment of ML models for NLP, retrieval, ranking, reasoning, dialog, and code-generation systems. Requires Master's/PhD, production ML experience, deep NLP expertise, Python proficiency, and MLOps knowledge in a fast-paced startup.
Software Engineer optimizes ML model inference performance using techniques like quantization and speculative decoding. Requires backend experience with PyTorch, TensorRT, CUDA, and deep GPU knowledge for LLMs.
Develop and deploy retrieval and search algorithms for OpenAI's API and ChatGPT, collaborating with research teams on production systems for millions of users. Requires experience with ML systems, vector databases, and large-scale search.
Develops and optimizes deep neural networks for Quadric's GPNPU architecture, focusing on algorithmic lowering, graph-based execution, and performance extraction. Requires MS/PhD, 8+ years experience in optimization, ML algorithms, and graphs.
Develops alignment algorithms, data pipelines, and sampling methods to optimize post-training AI models for performance and efficiency. Requires PhD or equivalent, ML expertise including reinforcement learning and transformers, and production code experience.
Build, deploy, and maintain scalable speech recognition pipelines, while experimenting with model architectures and modern ASR techniques. Requires 2+ years of experience, strong machine-learning fundamentals, Python and PyTorch expertise, and a bachelor's degree.
Leads ML performance optimization for training and inference platforms in autonomous driving, collaborating across teams to enhance efficiency using PyTorch, TensorRT, and profiling tools. Requires strong Python/C++ skills, GPU expertise, and 4+ years in large-scale ML platforms; leads engineering team.
As a Software Engineer specializing in Mapping and Localization, you will design and implement state-of-the-art systems for autonomous vehicles and mobile robots. This role involves building scalable solutions, fusing live data with maps, and continuously improving system performance using data-driven methods.
Develop, deploy, and optimize production-scale ML models using deep learning frameworks like PyTorch/TensorFlow. Requires 5+ years experience, expertise in Python, distributed systems, and focus areas like computer vision or NLP.
Develops, deploys, and optimizes state-of-the-art ML models for production at scale, handling terabyte datasets and neural networks. Requires 8+ years experience, expertise in PyTorch/TensorFlow/Python, and focus in CV/NLP.
Develops and deploys perception algorithms for autonomous driving, focusing on object detection, sensor fusion, localization, and tracking using LiDAR, cameras, and radar. Requires Master's degree and 4+ years non-academic experience in computer vision.
Builds product workflows and agentic systems using language models for research tasks like evidence synthesis and experiment planning. Combines ML fluency with strong software engineering to create reliable, trustworthy AI tools for scientific decision-making.