Latest ML Engineering jobs
Job results
Designs threat models and experiments for agentic AI security risks, builds prototypes with fine-tuned models and analysis tools, and turns research into scalable product defenses. Requires MS/PhD in CS/ML, production coding skills, and security mindset.
Build and deploy scalable AI agents for customer support that handle complex interactions across industries like finance and healthcare. Requires 5+ years experience with Python and TypeScript, focusing on system reliability and model integration.
Builds and trains large-scale multimodal agentic models involving reasoning, planning, coding, and tool calling. Requires strong ML foundations, PyTorch expertise, and experience with distributed training on massive datasets.
Leads applied AI investments, owning technical direction for high-impact AI products from design to production deployment. Requires 8+ years AI/ML experience, Python expertise, foundation models, LLM patterns, and evaluation systems.
Leads architecture of AI agents that convert natural language to production full-stack apps, driving multi-provider LLM strategies, tool integration, evaluation standards, and cross-team AI initiatives. Requires deep LLM expertise, prompt engineering, and scalable systems design.
Leads development of agentic AI systems for public sector, including guardrails, data processing, and fleet orchestration for federal datasets. Mentors engineers, defines technical strategy, and communicates with stakeholders to ensure reliable, secure solutions.
Develops advanced Vision-Language-Action models for robotaxi scene understanding, detecting hazards and enabling safe driving. Leads data strategies, post-training of large models, and deployment using PyTorch and production ML pipelines. Requires MS/PhD in CS and deep learning expertise.
Build benchmarks, datasets, and evaluation systems to measure and improve AI model quality for fraud, identity, and risk judgment tasks. Collaborate across research, engineering, and product to drive rigorous experimentation and iteration in high-stakes environments.
Research Engineer designs evaluations, studies model failures, and builds research loops to improve AI agents for high-stakes fraud detection and judgment tasks. Requires ML training experience, experimental rigor, and strong engineering skills in adversarial environments.
Designs and implements AI agents that transform natural language into production-ready full-stack applications using state-of-the-art LLMs. Integrates multiple LLM providers, orchestrates workflows, and continuously improves agent performance through data analysis and experimentation. Requires TypeScript proficiency and hands-on LLM experience.
Senior ML Engineer optimizes inference for voice AI models (STT, TTS, speech-to-speech) using engines like TensorRT-LLM and SGLang on GPUs. Requires 5+ years ML engineering with serving/inference expertise, Python/PyTorch proficiency, and production ML experience.
Conducts foundational research in spatial AI for residential construction, developing novel models using reinforcement learning, computer vision, LLMs, and 3D geometry. Requires 5+ years software engineering with 2+ years LLM experience, Master's degree, and expertise in PyTorch and RAG systems.
Build and validate an AI Chief of Staff prototype that automatically extracts, sorts, assigns, and follows up on meeting action items using LLMs, RAG, and agentic systems. Requires strong GenAI, NLP, Python backend, and productivity tool integration experience.
Software Engineer enabling production AI workloads on new hardware platforms through porting, benchmarking, stress testing, and performance optimization. Requires 5+ years in ML systems, distributed training, PyTorch, and RDMA/NCCL expertise.
Builds scalable backend systems and deploys ML models in production for client engagements, working embedded with top clients 3-4 days/week in New York. Requires 8+ years experience in ML engineering, Python, LLMs, cloud platforms, and client-facing work.
Develops systems for LLM interpretability and deterministic governance by working directly with model weights, activations, and architectures. Implements mechanistic interpretability techniques like activation patching and control vectors for enterprise policy enforcement in production.
Staff Data Scientist owns end-to-end development of ML and Generative AI solutions for the RiskOS fraud prevention platform, from data exploration and modeling to production deployment and monitoring. Requires 6+ years experience in data science with fraud/risk focus, Python/SQL proficiency, and GenAI expertise.
Develops and improves Codex AI agents for real-world software engineering tasks, focusing on performance, reliability, and integration with research and product teams. Requires strong Python, ML/LLM experience, and skills in evaluation, prompting, and debugging production failures.
Builds and owns production ML platform systems, turning AI research into reliable, scalable features. Partners with CTO on prototyping, observability, and end-to-end deployment of AI capabilities. Requires 3+ years experience with Python and modern ML frameworks; onsite in NYC.
The role optimizes large-scale distributed training and inference for foundation models, focusing on profiling, parallelization, memory efficiency, and productionization. It requires strong Python skills, multi-GPU training experience, and expertise in modern ML architectures.
Develops post-training pipelines, RLVR experiments, synthetic data generation, and large-scale LLM evaluation systems to enhance frontier language model performance in tool use, agentic behavior, and reasoning. Requires strong ML experience, coding skills, and research background.
Designs, develops, and optimizes core runtime infrastructure for distributed AI training and inference using PyTorch-based stack. Requires 8+ years in systems engineering, deep learning runtimes, Python/C++, and multi-node GPU workloads.
Build in-house tooling for post-training custom ML models using advanced techniques like RL and finetuning. Requires deep expertise in transformer training, PyTorch distributed systems, parallelism strategies, GPU performance optimization, and HPC platforms.
Senior AI Engineer builds, trains, deploys, and operates character AI agents at scale, managing LLM/SLM pipelines, social platform integrations, and feedback loops. Requires 5+ years in backend/ML engineering with production AI deployment experience.
Research Engineer on the Code RL team advancing AI models' ability to write efficient code for accelerators. Requires deep expertise in accelerators like CUDA/ROCm and ML frameworks like JAX/PyTorch, plus experience across kernels, model code, and distributed systems.
Develops data generation, post-training, and evaluation methods to improve the safety, fairness, robustness, and security of large language models that can take actions in the world. The role requires strong statistics and software engineering skills, distributed LLM training experience, and expertise in data collection and ML evaluation.
Develops and trains advanced AI models to tackle critical challenges in truth-seeking AI systems. Requires deep passion for building useful models and power user experience with AI; prior large-scale training experience preferred.
Develops and deploys production machine learning models for underwriting small business loans, focusing on risk assessment, pricing optimization, and automated systems. Requires strong statistical reasoning, ML expertise, engineering skills, and business acumen.
Develops multimodal AI models focused on image, video, and audio for high-fidelity generation, understanding, and agentic systems. Drives data curation, training, evaluation, and production integration of cutting-edge models.
Builds and manages platforms for AI model lifecycles, focusing on fine-tuning, training pipelines, and reinforcement learning for LLMs. Requires 8+ years in AI, advanced degree, and hands-on experience with generative AI techniques.
Builds and maintains platforms for fine-tuning, training, and managing LLMs including reinforcement learning pipelines and multi-node orchestration. Requires 4+ years in AI, hands-on LLM experience, and advanced degree in CS/Engineering.
Leads development of large-scale ML platforms, focusing on MLOps, graph ML infrastructure, performance optimization, and distributed training pipelines. Requires 8+ years in ML infrastructure with expertise in Python, PyTorch, Kubernetes, Ray, and cloud tools.
Leads development of large-scale ML platforms, focusing on MLOps, graph ML infrastructure, performance tuning, and distributed training optimization. Requires 5+ years in ML infrastructure with expertise in PyTorch, Kubernetes, Ray, and cloud tools.
Deploys state-of-the-art ML models and fine-tunes LLMs in production to enhance platform integrity and safety against adversarial threats. Requires Master's/PhD, deep learning expertise, PyTorch/TensorFlow proficiency.
Staff ML Engineer owns foundational model research and end-to-end quality improvements for clinical AI products like coding models, adaptive scribing, and chart understanding. Requires 5+ years ML experience with deep RL/deep learning expertise and production shipping track record.
Develop voice AI models for natural, low-latency spoken interactions on the Grok team. Handle data pipelines, model training with JAX/PyTorch, evaluations, and product integrations. Requires Python expertise, large-scale data processing, and distributed systems experience.
Builds scalable ML training systems and infrastructure for speech AI models (STT/TTS), prototypes novel ideas with researchers, and creates internal tools for cross-functional teams. Requires strong ML research pipeline experience, especially in speech domains, plus orchestration tools expertise.
Senior ML Engineer builds scalable ETL pipelines, develops ML infrastructure for training/evaluation/inference on cloud/edge, and integrates CV models for analyzing farm images from tractor cameras. Requires 5+ years experience with Python, PyTorch/TensorFlow, data engineering tools, and cloud platforms.
Lead development of Context Hub, an open-source CLI for AI agents to access up-to-date API docs. Own technical direction, build infrastructure for intelligent retrieval and community features, requiring 3+ years experience, TypeScript/Node.js proficiency, and strong LLM/AI agent knowledge.
Forward Deployed Engineer embeds with customer teams to build and own production-grade AI-powered coding workflows using Cursor, from discovery to iteration and integration back into the core product. Requires strong Python/TS skills, end-to-end ownership, and production reliability experience.
Build and maintain fraud detection models and financial risk products through full ML lifecycle, including production code. Requires 6+ years experience (or 8+ with Masters), PhD preferred, strong end-to-end DS/ML skills, and domain interest in fraud/identity.
Build and maintain fraud detection models and financial risk products through full data science lifecycle, including production code. Requires 6+ years experience (or 8+ with Masters), advanced ML/stats skills, and PhD preferred.
Build and maintain fraud detection ML models and financial risk products with end-to-end ownership. Requires 4+ years experience (or 6+ with Masters), advanced degree, strong ML/stats skills, production coding, and domain interest in fraud/identity.
The Model Developer owns end-to-end delivery of production-grade Python and predictive models, partnering with customer actuarial teams to solve complex modelling problems. The role also contributes to engineering standards, testing, code review, and mentoring junior developers.
Build core AI platform infrastructure at Harvey including model routing, context management, evaluation frameworks, and shared abstractions for agentic AI products. Requires 5+ years backend experience with 1+ year in AI/ML, production multi-model systems, and strong platform-building skills.
Develops machine-learning models and systematic trading strategies by analyzing large financial datasets, researching alpha signals, and backtesting hypotheses. Requires a STEM degree, 3+ years in systematic trading, Python proficiency, and strong English communication skills.
Leads technical direction for bio AI research programs, architects AI/computational infrastructure, prototypes modeling systems, and builds high-caliber engineering team. Requires deep ML expertise in generative models, large-scale training, and bio applications with hands-on leadership.
Build and deploy scalable machine learning infrastructure, models, and AI platforms for voice, speech, and natural-language products. The role requires 5+ years in ML or AI, production experience with large-scale data and multi-tenant applications, and expertise in distributed cloud systems.
Lead personalization and classification algorithm productionization, building ML model-serving infrastructure and translating experimental outputs into scalable production systems. Requires 5+ years experience, backend/infrastructure specialization, and proficiency in AWS, Java, Python, and TypeScript.
Forward Deployed Engineer deploys AI agents for financial services customers, handles onsite integrations in regulated environments, builds connectors and frameworks, and translates customer needs into product insights. Requires 2+ years software engineering with Python, APIs, data pipelines, cloud, and AI experience.