Latest remote AI Research jobs
Job results
Conduct applied research on AI agents, designing experiments and evaluation systems to improve reliability, context retention, and multi-step task completion. The role requires strong AI/ML research, engineering, experimental design, and communication skills.
AI Research Intern researching agentic AI applications for customer-facing products and developing working prototypes. The role requires current pursuit of a technical bachelor's degree, prior software engineering or substantial project experience, and interest in LLMs or generative AI.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.
Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.
Own applied multimodal ML research from clinical problem definition through production, developing and rigorously evaluating computer vision, NLP, and deep learning systems for radiology. Requires strong Python and PyTorch expertise, 4+ years of relevant experience, and an MS, PhD, or equivalent practical experience.
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.
Owns the end-to-end design, build, deployment, and optimization of AI workflows inside ORA for Customer Success, spanning call preparation, guidance, grading, and follow-up. Requires 7+ years in applied AI or related solution delivery, strong LLM and data-flow expertise, and the ability to translate business problems into shipped systems.
Research fellows propose, build, validate, and publish benchmarks or evaluation methodologies for measuring frontier AI performance on economically valuable professional and scientific work. The fellowship requires a specific research pitch, relevant technical or adjacent-field background, and a commitment of at least 20 hours per week.
Leads Deepgram’s end-to-end TTS research program, setting technical direction, training and evaluating large-scale speech-generation models, and turning breakthroughs into production systems. The role combines hands-on technical leadership with building and developing a high-performing research organization.
Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.
Senior software engineer responsible for building AI agents, developer tooling, and automation that improve planning, coding, testing, code review, CI, and local development workflows. Requires 6–8 years of software engineering experience, hands-on agentic coding tool experience, and proficiency in Python or JavaScript/TypeScript, AWS, and PostgreSQL.
Conduct speech technology research on text-to-speech, ASR, speech-to-speech, and speech analysis models during a six-month project. The role requires a relevant master's qualification or PhD study, neural-network development experience, and Python skills.
Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.
Conduct research on long-horizon, multi-agent AI behavior by designing agent environments, analyzing large-scale data, and running experiments. The role requires strong research judgment, rapid execution, independence, and familiarity with current AI developments.
Build, optimize, and evaluate long-running and multi-agent AI systems, along with tools for monitoring and analyzing their real-world behavior. The role requires software engineering experience with coding agents, strong independence, and familiarity with current AI developments.
Research Scientist developing and deploying machine learning models for fraud detection, identity verification, and financial risk. The role targets new PhD graduates or early-career researchers with strong quantitative foundations, Python experience, and interest in owning the full production ML lifecycle.
Conduct machine learning and statistical research to improve unsecured underwriting models, evaluating enhancements through rigorous experimentation and validation. The role requires a graduate degree in a quantitative field, Python modeling experience, and 0–2 years of applied research experience.
Research Advisors apply deep finance, legal, medical, or related expertise to evaluate advanced generative AI systems, shape model-governance frameworks, and collaborate on research and client engagements. Candidates need at least five years of relevant experience, strong analytical skills, and hands-on AI experience.
The fellowship engages experienced software engineers or technical researchers in designing evaluations, datasets, and expert analyses for advanced generative AI systems. Fellows contribute to applied AI research and publications with flexible remote project work.
STEM Fellows apply academic and professional expertise to design evaluation datasets, assess generative AI systems, and contribute research insights and publications. The fully remote, six-month independent contractor opportunity is suited to PhDs, postdoctoral researchers, and professors with relevant domain expertise.
Medical fellows apply clinical expertise to design scenarios, evaluate generative AI decision-making, and provide structured feedback for safer, more accurate healthcare systems. The role requires an MD or DO, board certification, strong clinical reasoning and writing skills, and a relevant medical specialty.
Sr AI Architect leading Twilio's conversational AI strategy, including memory, knowledge, and behavioral intelligence systems. Requires 15+ years software engineering experience (6+ in production ML at platform scale), deep LLM/LLMOps expertise, and a Master's or PhD in a quantitative field.
Conduct applied research and engineering to improve language-model behavior in real-time voice conversations. The role focuses on fine-tuning, rigorous evaluation, production failure analysis, data strategies, and safely deploying improvements.
Conduct original research on LLM evaluation, routing optimization, and model behavior using billions of real-world generations. Design novel benchmarks, run large-scale empirical studies, and develop statistical foundations for intelligent routing. Requires MS/PhD, publication track record, deep stats/ML expertise, and Python/SQL skills.
Research Scientist advancing generative audio models (diffusion/flow matching) for music, focusing on vocal synthesis, post-training alignment (DPO/RLHF), or audio editing. Requires PhD, top publications, and PyTorch expertise to turn research into artist-first Spotify products.
Graduate research intern (MS/PhD) building Customer World Models and proactive intelligence systems using representation learning, RL, and agentic decision-making. Own research end-to-end from framing to production deployment.
As a Research Scientist II on the Video team, you will drive core research initiatives, deliver reproducible experimental results, and help translate machine learning models into real-world product solutions, focusing on real-time video processing and deepfake detection.
This is a research role focused on building models that continuously evolve with the world, with a focus on efficiency, gradient-free exploration, real-time learning, and interface design. The role requires strong programming skills and expertise in model optimization techniques.
Join the alignment research team to work on high-impact, under-resourced projects focused on AI alignment. This role requires strong ML research and Python skills to build, train, or evaluate deep learning models.
Lead Protege's DataLab research organization, defining strategy for AI training data quality, evaluation systems, and marketplace optimization while managing researchers and partnering with Product, Engineering, and GTM.
Provides expert advisory on AI model behavior in finance, legal, or medical domains, collaborates on research tasks and publications, and engages in executive sales and GTM activities. Requires 5+ years top-tier experience and advanced degree (PhD/Masters/MD/JD).
Conducts and publishes cutting-edge machine learning research, building and training large language models and contributing to applied AI product initiatives. Applicants should be pursuing a PhD or demonstrate exceptional equivalent experience, with expertise in ML systems, Transformers, programming, and modern ML frameworks.
Build and own the agent runtime, orchestration layer, and long-horizon coding agent workflows for an AI-driven consumer social platform.
Postdoctoral researcher leading high-impact AI projects on Mixture-of-Experts and long-context language models, training/releasing models, building open-source tools, publishing papers, and mentoring juniors. Requires recent PhD in CS/ML with strong publication record and PyTorch expertise.
AI Researcher develops and fine-tunes large multimodal models for real-time conversational avatars, modeling verbal/non-verbal behaviors with low latency. Requires PhD or equivalent, hands-on experience with VLMs, PyTorch, and deep learning.
Develops next-generation multimodal LLMs integrating speech, text, tools, and real-time reasoning for conversational AI agents. Requires strong background in LLMs, multimodal models, fast experimentation, and production deployment experience.
Conducts foundational research and develops scalable ML models for speech-to-text, text-to-speech, and neural audio codecs in real-time voice AI agents. Requires deep expertise in voice modeling, self-supervised learning, and production deployment at enterprise scale.
Leads research in computer vision, multimodal understanding, and visual generation. Develops novel models and methodologies, translates research to production, and mentors teams. Requires PhD preferred, 8+ years experience, and expertise in PyTorch, TensorFlow, transformers.
Fellows conduct research on AI safety, collaborating with Anthropic researchers to develop evaluation methods and alignment techniques. Requires strong interest in AI safety, CS/math background, and ML experience.
Develops state-of-the-art Visual Question Answering systems for medical records using advanced NLP and Computer Vision techniques. Requires expertise in these fields, strong software engineering skills, and ability to work with noisy data to achieve production-scale model performance.
Conducts research on efficient AI systems focusing on real-time adaptation, model efficiency, and cross-stack optimization. Requires PhD or equivalent, 4-5+ years industry experience, and deep ML expertise including PyTorch/JAX and optimization techniques.
Part-time AI Trainer annotates mortgage conversation data, provides feedback on AI responses, and tests system quality for mortgage servicing calls. Requires 5+ years mortgage customer service experience and industry knowledge.
The Senior Applied Researcher develops and evaluates AI/ML model systems for healthcare SaaS products, working with structured and unstructured clinical data. The role requires advanced graduate education, industry healthcare ML experience, and hands-on expertise in LLMs, cloud platforms, data engineering, and production deployment.
Senior Applied Researcher develops AI/ML solutions for healthcare challenges, applying GenAI, LLMs, and techniques like RAG to build production-ready models in SaaS environments. Requires Master's in relevant field, healthcare data experience, Python proficiency, and ML frameworks expertise.
Collaborate on high-impact AI research projects, design and review deep learning models in PyTorch, analyze GPU performance, and co-author publications. Requires PhD in CS/AI/ML and hands-on expertise with Transformers, CNNs, and diffusion models.
Develops and optimizes GPU-accelerated kernels and algorithms for ML/AI applications, co-designing with modeling, hardware, and software teams. Requires strong GPU programming expertise in CUDA/Triton and knowledge of ML models.
Leads the technical vision and implementation of Docker’s containerized AI agent platform, including runtime infrastructure, distributed systems, evaluation, and operational excellence. The role requires 10+ years of software engineering experience, principal-level technical leadership, and practical experience with Go, Docker, and LLM-based agent development.
Conduct cutting-edge research on Large Language Models, focusing on transformer optimization, distributed training, data curation, and RL. Collaborate on experiments, deploy models to production, and drive voice AI innovations.