Latest AI Research jobs
Job results
Conduct foundational research on LLMs and multimodal systems, designing architectures and training methods and helping move prototypes into production. The role targets PhD researchers graduating by December 2026 with strong machine-learning research and programming experience.
Research Engineer building large-scale AI capability evaluations, telemetry, data pipelines, and analysis tools for Anthropic’s Takeoff Intel team. The role requires hands-on large language model experimentation, rapid prototyping, data expertise, and strong research collaboration.
AI Research Intern researching agentic AI applications for customer-facing products and developing working prototypes. The role requires current pursuit of a technical bachelor's degree, prior software engineering or substantial project experience, and interest in LLMs or generative AI.
Research Intern working on next-generation video generation models through experimentation in distillation, inference efficiency, reward modeling, preference optimization, and scalable training infrastructure. Applicants should be pursuing a PhD or final-year master’s degree with relevant research experience and strong Python and machine learning framework skills.
Paid internship for quantitative students contributing to AI-powered software creation, systems optimization, and developer tooling. Candidates should demonstrate strong mathematical ability, curiosity about AI, and autonomous cross-functional collaboration.
Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.
Conduct applied research on foundation models for fraud detection using large-scale behavioral and financial-risk data. The role spans experimentation, evaluation, production deployment, and cross-functional work on model governance, requiring 4+ years of applied ML experience and strong Python and SQL skills.
Evaluates and improves AI-generated clinical outputs, partnering with product and engineering teams to establish safety, accuracy, and clinical-quality standards. Requires an MD, DO, or equivalent clinical doctorate, substantial patient-care experience, strong clinical judgment, and the ability to learn AI evaluation techniques.
Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.
Applied AI Research Engineer who tests model capabilities, builds demos and evaluations, supports strategic customer implementations, and translates field insights into product and research direction. Requires 6+ years of technical experience, programming proficiency, LLM development experience, and strong communication skills.
Research and develop computer vision and deep learning algorithms for autonomous drones, taking ownership of projects that advance Skydio’s autonomy capabilities. Requires PhD-level study or background, strong C++/Python and PyTorch skills, and foundational expertise in software, mathematics, and relevant autonomy fields.
Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.
Researcher focused on training and evaluating frontier AI agents, mining incidents, and building scalable safety measurement systems. The role requires strong research or ML engineering execution, quantitative judgment, and the ability to own ambiguous projects end to end.
Sets company-wide architecture and strategy for data and applied AI, connecting governed data foundations to production intelligence and measurable business outcomes. The role requires 14+ years of experience, strong production engineering judgment, executive partnership, and hands-on delivery.
The applied scientist will lead GenSim’s methodology for creating realistic, production-grade simulated environments and high-quality post-training data for Datadog agents. The role requires deep LLM and agent experience, evaluation expertise, Python, distributed systems, and the ability to set technical direction.
ML research internship focused on improving search and retrieval through deep learning, representation learning, and RAG systems. The role requires strong PyTorch and distributed-training skills, plus familiarity with search evaluation and AI/ML research publications.
Researcher developing and publishing mechanistic interpretability techniques, building infrastructure to study model internals, and guiding alignment-focused research. Requires research experience in machine learning or a related field, strong engineering skills, and proficiency in Python or similar languages.
Conduct research and engineering on large-scale post-training of diffusion and language models, focusing on aesthetics, preference optimization, reinforcement learning, reward modeling, and evaluation. The role requires strong PyTorch, distributed training, low-precision inference, and VLM experience.
Researcher developing and scaling image and video diffusion models on large GPU clusters. The role requires deep PyTorch and distributed-training expertise, proficiency in low-precision computation, and experience profiling and debugging large-scale model training.
Leads the research agenda and hands-on development of replayable enterprise environments, agent evaluations, and post-training systems. The role requires deep AI research experience, a PhD or equivalent track record, and the ability to translate open-ended questions into production systems.
Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.
Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.
Build and expand customer-facing agentic AI products, MCP integrations, and automated reconciliation workflows for private fund management. The role requires senior-level software engineering, strong systems thinking, product judgment, and hands-on experience building and evaluating AI systems.
Conducts deep technical research into cloud- and AI-native environments to identify novel risks and attack vectors, then translates findings into product capabilities with Product and Engineering teams. Requires at least five years of security research experience and strong scripting and telemetry-analysis skills.
Research and engineer security measures for LLM-powered web agents and chatbots, including adversarial testing, secure architectures, model evaluation, and security-focused training. The role requires strong Python and ML expertise, AI security experience, and preferably a computer science Ph.D.
Develops experimental AI techniques and prototypes for agentic marketing applications, with emphasis on image and video generation. The role requires strong backend or probabilistic systems expertise, quantitative thinking, creativity with LLM applications, and product intuition.
Research Scientist developing and evaluating health-focused AI models, large language models, and agentic systems for clinical applications. The role requires advanced research experience, strong coding skills, healthcare or clinical-data experience, and top-tier AI/ML publications.
Research role focused on improving agentic coding capabilities through reinforcement-learning training, synthetic data, coding environments, reward design, and evaluations. Requires strong Python engineering, scalable distributed-training experience, and a bachelor’s degree or equivalent; research experience and a PhD are preferred.
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.
Own applied multimodal ML research from clinical problem definition through production, developing and rigorously evaluating computer vision, NLP, and deep learning systems for radiology. Requires strong Python and PyTorch expertise, 4+ years of relevant experience, and an MS, PhD, or equivalent practical experience.
Researcher developing internal evaluations and research signals for AI post-training, including usability, correctness, auditing, agentic systems, and nuanced model behaviors. Requires evaluation experience, strong research judgment, Python, and familiarity with deep learning frameworks.
Research Scientist responsible for measuring and improving how frontier models learn from tasks, designing post-training experiments, validating data quality, and creating datasets and evaluation systems. Requires production reinforcement-learning experience at a frontier lab and end-to-end LLM post-training experience.
Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.
Research Scientist focused on evaluating frontier language and multimodal models, diagnosing failure modes, and building rigorous benchmarks. The role requires advanced training in AI or a related field, post-training expertise, and published machine learning research.
Research novel post-training methods for large language models, focusing on preference optimization, data curation, evaluation, alignment, and robustness across text and multimodal systems. Requires advanced academic training and experience with deep learning, reinforcement learning, and post-training techniques.
Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.
Research and engineer AI systems, optimizing and evaluating machine-learning models and on-device performance across desktop and mobile products. The role requires software development experience, systems and networking knowledge, and expertise with modern ML frameworks, privacy, and security.
Conduct research and develop foundation models for robotic manipulation and high-precision manufacturing, taking projects from data curation through deployment on industrial robots. The role requires current PhD study, strong Python and deep learning expertise, robotics simulation experience, and research in foundation models.
Build agent-driven chatbots and generative AI workflows for financial-wellness products, owning features from design through impact measurement. The role requires at least three years of software engineering experience, strong system design, maintainable coding practices, and a bachelor’s degree or equivalent experience.
Leads experiments investigating working-memory circuits in behaving mice through multiregional Neuropixels recordings, optical perturbation, and large-scale neural-data analysis. Requires PhD-level neuroscience training, mouse survival surgery, awake-rodent electrophysiology, and Python or MATLAB expertise.
Build and evolve the agent harness powering Perplexity’s flagship answer experience, improving orchestration, context management, performance, reliability, observability, and evaluation. The role requires strong software engineering skills, Python proficiency, and experience shipping large-scale AI systems.
Applied research scientists develop deep-learning and generative media systems for video, audio, and multimodal editing features that ship to millions of users. The role requires strong PyTorch or TensorFlow skills, rapid experimentation, and evidence of impactful research or production machine-learning work.
Conduct rigorous people research and applied data science to evaluate talent programs, organizational health, and employee experiences. The role requires advanced expertise in research design, experimentation, measurement, causal inference, statistical modeling, and responsible handling of sensitive employee data.
Conduct AI safety research across data curation, post-training, evaluations, synthetic data, and red-teaming to improve model reliability on harmful and dual-use requests. The role requires AI safety experience, Python, deep learning frameworks, and scalable technical research skills.
Leads the architecture and production deployment of real-time computer vision and perception systems, combining classical techniques with deep learning. Requires principal-level technical leadership, 6+ years of industry experience, Python, PyTorch, and real-time optimization expertise.
Own and optimize an enterprise customer-support AI agent by improving prompts, conversational quality, guardrails, intent recognition, and performance. The role requires deep conversational AI experience, strong analytical skills, platform expertise, and strategic collaboration across support teams.
Owns the end-to-end design, build, deployment, and optimization of AI workflows inside ORA for Customer Success, spanning call preparation, guidance, grading, and follow-up. Requires 7+ years in applied AI or related solution delivery, strong LLM and data-flow expertise, and the ability to translate business problems into shipped systems.
Research fellows propose, build, validate, and publish benchmarks or evaluation methodologies for measuring frontier AI performance on economically valuable professional and scientific work. The fellowship requires a specific research pitch, relevant technical or adjacent-field background, and a commitment of at least 20 hours per week.
Leads the research agenda for humanoid robotics, developing foundation-model and reinforcement-learning methods for dexterous manipulation and deploying them on real robotic systems. Requires a PhD, strong robotics research publications, and senior-level technical leadership.
The Research Engineer will apply advances in agents and language models to build and evaluate multi-agent systems for automated code validation and review. The role requires a computer science or equivalent background, research experience, strong programming skills, and product intuition.