Data Scientist II
Builds and maintains reproducible multi-omics data pipelines and analysis workflows supporting NCI cancer research. Requires strong Python/R skills and experience with at least two omics data types.
About the job
Key Responsibilities
Bioinformatics Workflow and Data Pipeline Development
- Design, build, and maintain reproducible pipelines for genomic, transcriptomic, single-cell, spatial, proteomic, metagenomic, metabolomic, and clinical datasets.
- Develop reusable transformation logic and curated datasets supporting analytics, dashboards, APIs, notebooks, and downstream research workflows.
Multi-Omics Analysis
- Support bulk RNA-seq (QC, DEG, GSEA), single-cell RNA-seq (clustering, UMAP/t-SNE, cell type annotation, DEG), and Digital Spatial Profiling (annotation, QC, normalization, spatial deconvolution, volcano plots, heatmaps).
Data Integration and Lifecycle Support
- Enable reliable data movement from source systems into structured, analysis-ready formats.
- Support ingestion, curation, metadata capture, source-to-target mapping, schema management, provenance tracking, and long-term maintainability of data products.
Statistical Modeling and Machine Learning
- Apply statistical and ML methods including hypothesis testing, regression, clustering, PCA, UMAP, t-SNE, and classification to biomedical datasets.
- Incorporate AI/LLM-based extraction where appropriate.
Researcher-Facing Applications and Visualization
- Build and support interactive dashboards (Shiny, Streamlit), notebooks, reports, and APIs enabling researchers to explore multi-omics and clinical data.
- Support figure generation for QC, differential expression, pathway, and spatial analyses.
Collaboration
- Partner with data scientists, bioinformaticians, researchers, developers, and government stakeholders to translate scientific needs into technical specifications, data models, and reusable workflows.
Required Qualifications
- Bachelor's degree in Data Science, Bioinformatics, Computer Science, Biological Sciences, or a related field (advanced degree preferred), or equivalent experience.
- Demonstrated experience in a data-intensive role supporting biomedical research or scientific computing.
- Strong proficiency in Python and R for analysis, scripting, and visualization.
- Hands-on experience with at least two omics data types (e.g., bulk RNA-seq, scRNA-seq, spatial transcriptomics, proteomics, metagenomics, GWAS).
- Solid understanding of statistical modeling, dimensionality reduction, clustering, differential expression, and pathway analysis.
- Ability to work with structured, semi-structured, and unstructured data across relational and data lake environments.
- Strong problem-solving skills with the ability to communicate effectively across technical and non-technical audiences.
- Genuine interest in biomedical and translational research; awareness of data governance, privacy, and compliance requirements.
Preferred Qualifications
- Experience building analytics solutions in platforms such as Snowflake, Databricks, or cloud data warehouses.
- Experience with workflow and reproducibility tools: Galaxy, Terra, Nextflow/WDL, Snakemake, Singularity, CWL.
- Familiarity with the scverse Python ecosystem (Scanpy, Squidpy, SCIMAP, AnnData) and spatial single-cell analysis methods (PhenoGraph, Louvain/Leiden clustering, UMAP, Ripley's L statistic).
- Experience preparing curated datasets for dashboards, APIs, and web applications; familiarity with Posit Connect, R/Shiny, Streamlit, Jupyter.
- Experience with AWS (EC2, S3, Lambda), object storage, relational databases, scheduled jobs, API integrations, secure data movement; familiarity with HPC environments, SLURM/SGE, or NIH Biowulf.
- Background in biomedical research, clinical research, or healthcare analytics; familiarity with HL7/FHIR, CDISC, or OMOP standards.
- Experience with metadata management, data lineage, open-source code release, containerized analyses, and secure handling of de-identified or access-controlled research datasets.
- Experience creating documentation, training materials, or workshops for researchers and non-coder audiences.
Skills
Python, R, Rna-Seq, Scrna-Seq, Spatial Transcriptomics, Proteomics, Metagenomics, Gwas, Statistical Modeling, Machine Learning, Pca, Umap, T-Sne, Clustering, Differential Expression Analysis
Similar jobs
Data Science jobsData Scientist II developing graph-based algorithms, entity-resolution models, and large-scale data pipelines for identity and deceased-monitoring products. Requires advanced academic training or equivalent experience, strong Python/Scala and SQL skills, and hands-on Spark and AWS experience.
Analyzes proprietary private-market transaction and pricing data to identify valuation, liquidity, and investor trends, then translates findings into research publications and investor-facing thought leadership. Requires a bachelor's degree, strong quantitative skills, Excel and Python proficiency, and 3–7 years of relevant experience.
Analyzes product and business data, develops metrics and dashboards, and communicates actionable recommendations to cross-functional teams and leadership. Candidates should be pursuing a quantitative bachelor's or master's degree, with SQL proficiency and familiarity with Python or R.
The Forward Deploy Data Scientist will partner with health systems and internal teams to investigate healthcare data, build and operationalize ML/LLM pipelines, and deliver actionable insights. The role requires 2–3 years of data science experience, strong Python and ML/NLP skills, and customer-facing communication ability.
Data Science Intern supporting product and engineering teams through statistical analysis, experimentation, analytical modeling, and data-driven recommendations. The role requires quantitative academic study or equivalent project experience, plus familiarity with Python and SQL.