Senior Data Scientist, AI Retrieval Systems
Build and operate biomedical retrieval, ranking, and concept-mapping systems that ground LLM applications for rare disease research. The role requires at least five years of production software or data-systems experience, including two years delivering LLM-powered applications, plus strong retrieval and Kubernetes skills.
About the job
Responsibilities
- Model biomedical knowledge for rare disease research by ingesting disease and phenotype ontologies and controlled vocabularies into PostgreSQL, reconciling identifiers, and maintaining release and refresh workflows.
- Build retrieval-augmented services that map everyday language to clinical concepts using embeddings, candidate retrieval, contextual model disambiguation, structured outputs, and validation.
- Tune keyword and vector search over large biomedical corpora, balancing recall and latency.
- Build ranking and relevance layers with domain-aware weighting and graceful degradation outside curated coverage.
- Deliver user interfaces with Next.js, React, and TypeScript, including question and confirmation flows, result presentation, and live pipeline status.
- Deploy and operate services on NIH on-premises and high-performance computing Kubernetes environments, including Helm charts, StatefulSets, secrets, ingress, GPU scheduling, and scheduled jobs.
- Build evaluation and regression systems using golden sets and retrieval metrics such as Recall@K and MRR.
- Instrument systems with request identifiers, latency, errors, and selected concepts for reviewable AI-assisted results.
- Translate researcher, clinician, and patient needs into data models, retrieval behavior, and interface design.
- Contribute to manuscripts, conference abstracts, and posters with NIH investigators.
Requirements
- Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related field; equivalent professional experience may be considered.
- At least 5 years building and operating production software or data systems, including at least 2 years shipping LLM-powered applications.
- End-to-end retrieval systems experience, including indexing, query construction, and retrieval-quality measurement.
- Experience evaluating systems without a single correct answer using golden sets, offline regression suites, Recall@K, MRR, or similar metrics.
- Experience with structured output and tool or function calling using typed schemas and validation.
- Ability to own services from schema design through deployment and operation.
- Ability to obtain and maintain a Public Trust Security clearance.
- Proficiency with Python, FastAPI, Pydantic, pytest, PostgreSQL, vector and full-text search, embedding pipelines, containers, Kubernetes, Next.js, React, TypeScript, Git, and CI/CD.
Nice-to-haves
- Biomedical ontologies and controlled vocabularies, including MONDO, HPO, UMLS, MeSH, and OBO Foundry resources.
- Entity linking, concept normalization, ontology alignment, or embedding-based domain grounding.
- Helm and deployment to on-premises or HPC Kubernetes environments.
- Production serving of open-weight models with Ollama or vLLM, LiteLLM, and domain embedding models such as MedCPT.
- Rare disease, clinical genetics, or translational research experience.
- Contributions to biomedical resources, standards, or consortia such as OBO Foundry or GA4GH.
- Published or presented engineering work, open-source contributions, or prior NIH experience.
Compensation and Benefits
- Annual salary: $130,000–$150,000.
- 100% employee medical, dental, and vision coverage.
- Paid time off and holidays.
- 401(k) match up to 5%.
- Educational benefits, employee referral bonus, and flexible spending accounts.
Skills
Python, FastAPI, Pydantic, Pytest, Postgres, Pgvector, Vector Search, Full-Text Search, Embeddings, Kubernetes, Helm, React, TypeScript, Next.js, Llm Engineering
Similar jobs
Data Science jobsSenior Data Scientist who will own growth forecasting, operating models, scenario analysis, and strategic planning for executive leadership. Requires 5+ years in business data science or analytics, advanced SQL, quantitative modeling, and experience with modern cloud data tools.
Senior data scientist leading marketing measurement, incrementality experiments, attribution, and growth modeling. The role partners with Growth leadership and requires 5+ years of relevant experience, advanced SQL, machine learning deployment, and modern cloud data-stack expertise.
Senior Data Scientist focused on product analytics, experimentation, predictive modeling, and advisor behavior insights for a dynamic travel marketplace. The role partners with Product, Engineering, and Design and requires 5+ years of relevant experience plus advanced SQL skills.
Leads operational analysis, modeling, simulation, and wargaming for autonomous aircraft systems, translating military mission insights into design decisions and capability improvements. Requires 5+ years of operational analysis or military operations research experience and expertise in defense modeling tools and modern simulation workflows.
The Senior Data Scientist will measure and forecast organic growth across SEO, ASO, and content, designing incrementality tests and attribution frameworks to guide investment. Requires 5+ years of organic growth or marketing analytics experience, advanced SQL, Python or R, and strong experimentation skills.