Skip to content
Rad AIRad AI

Staff ML Research Scientist

Leads end-to-end applied ML research in NLP, LLMs, retrieval, and multimodal models for healthcare AI, driving from experimentation to production deployment with rigorous evaluation and clinician collaboration. Requires 7+ years experience, MS/PhD, and depth in ML areas like PyTorch tooling.

About the job

What You'll Do

  • Own end-to-end applied research: frame the problem, design experiments, ship to production, and monitor impact against real-world metrics.
  • Set technical direction across LLMs, retrieval, and multimodal; run ablations/error analysis that change product decisions.
  • Build evaluation that matters: link offline metrics to online outcomes; define thresholds, monitoring, and rollback.
  • Partner to deliver with engineering and product—and, when relevant, clinicians/domain experts—to align data, success criteria, and timelines.
  • Raise the bar by mentoring peers and codifying standards for reliability, safety, and documentation.
  • Improve the platform (data, training, serving, observability) to speed iteration and ensure reproducibility.
  • Explore new directions, with computer vision/vision-language work as a nice-to-have for future strategic initiatives.

What We're Looking For

  • MS or PhD (or equivalent research experience) in Computer Science, Electrical Engineering, Computational Linguistics, Biomedical Informatics, or related quantitative field.
  • 7+ years of applied ML research experience (or PhD + 5 years, or equivalent evidence of Staff-level impact).
  • Depth in one or more areas: LLMs and NLP, computer vision, speech, recommendation/ranking, retrieval, or multimodal modeling.
  • Strong experimental rigor: clear hypothesis framing, offline→online linkage, calibration and stratified analyses, ablations that influence decisions.
  • Proven ability to take models to production.
  • Hands-on with modern tooling: PyTorch and common experiment/ops tools (for example MLflow, Databricks, Ray, or similar).
  • System thinking: can choose methods based on constraints, design for observability and rollback, and document decisions clearly.
  • Collaborative communicator who writes crisp design docs and explains complex ideas to non-specialists; comfortable mentoring peers.

Preferred Qualifications

  • Health data familiarity, including EHR or imaging.
  • Experience in one or more areas: clinical NLP or LLMs, computer vision, speech, retrieval or multimodal modeling.
  • Shipped, measured models in production with monitoring and clear rollback; external or multi-site validation is a plus.
  • Workflow integration with EHR, RIS, PACS, or reporting systems; PowerScribe or Dragon exposure helpful.
  • Strong evaluation practices: calibration, slice analysis, and ablations.
  • Safety and governance in sensitive domains, including PHI handling and HIPAA or FDA-adjacent environments.
  • Technical mentorship and contributions to team research culture; publications or impactful open-source work.
  • Practical tooling: PyTorch plus modern ML ops tools such as MLflow, Databricks, Ray, or Triton.

Skills

PyTorch, LLMs, NLP, Computer Vision, Retrieval, Multimodal Modeling, MLflow, Databricks, Ray, Triton

Shield AI

Shield AI

Washington, DC
Senior Staff Engineer, Autonomy Capabilities – Maritime
$221k+/yrOn-site10+ YOEAI Research

Leads technical direction and develops maritime autonomy capabilities for unmanned surface and underwater vehicles, including motion planning, localization, safe behaviors, and heterogeneous multi-agent collaboration. Requires deep robotics and unmanned-systems experience, strong C++/Python skills, and senior technical leadership.

Upstart

Upstart

United States

Staff Machine Learning Model Risk Specialist
$140k+/yrRemote7+ YOEAI Research

Evaluates model and Generative AI risks across Upstart Bank’s model inventory, conducting risk assessments, monitoring reviews, quantitative analyses, and governance activities. Requires a quantitative master’s degree, 4+ years of relevant experience, and coding skills in Python, R, or similar languages.

Shield AI

Shield AI

San Mateo, CA

Senior Staff Software Engineer, Autonomy Capabilities
$281k+/yrOn-site10+ YOEAI Research

Leads design, implementation, integration, and field validation of tactical autonomy software for unmanned systems and multi-agent missions. Requires extensive autonomy or robotics experience, strong C++ and Python skills, and the ability to obtain a SECRET clearance.

Anthropic

Anthropic

San Francisco, CA

Staff+ Researcher, Cybersecurity Products
$405k+/yrHybrid7+ YOEAI Research

Research and evaluate frontier AI capabilities for cybersecurity, rapidly prototyping tools, designing rigorous benchmarks, and helping operationalize reliable capabilities into products. Requires deep security expertise, strong technical communication, and at least seven years of relevant experience.

Order.co

Order.co

United States

Staff Applied AI Scientist
No salary listedRemote10+ YOEAI Research

Own the architecture, delivery, evaluation, and production operations of AI capabilities embedded in procurement and finance workflows. The role requires 10+ years in applied AI or machine learning, deep LLM and agent expertise, and experience delivering measurable production outcomes.