Skip to content
AxleAxle

Data Scientist II

Builds and maintains reproducible multi-omics data pipelines and analysis workflows supporting NCI cancer research. Requires strong Python/R skills and experience with at least two omics data types.

About the job

Key Responsibilities

Bioinformatics Workflow and Data Pipeline Development

  • Design, build, and maintain reproducible pipelines for genomic, transcriptomic, single-cell, spatial, proteomic, metagenomic, metabolomic, and clinical datasets.
  • Develop reusable transformation logic and curated datasets supporting analytics, dashboards, APIs, notebooks, and downstream research workflows.

Multi-Omics Analysis

  • Support bulk RNA-seq (QC, DEG, GSEA), single-cell RNA-seq (clustering, UMAP/t-SNE, cell type annotation, DEG), and Digital Spatial Profiling (annotation, QC, normalization, spatial deconvolution, volcano plots, heatmaps).

Data Integration and Lifecycle Support

  • Enable reliable data movement from source systems into structured, analysis-ready formats.
  • Support ingestion, curation, metadata capture, source-to-target mapping, schema management, provenance tracking, and long-term maintainability of data products.

Statistical Modeling and Machine Learning

  • Apply statistical and ML methods including hypothesis testing, regression, clustering, PCA, UMAP, t-SNE, and classification to biomedical datasets.
  • Incorporate AI/LLM-based extraction where appropriate.

Researcher-Facing Applications and Visualization

  • Build and support interactive dashboards (Shiny, Streamlit), notebooks, reports, and APIs enabling researchers to explore multi-omics and clinical data.
  • Support figure generation for QC, differential expression, pathway, and spatial analyses.

Collaboration

  • Partner with data scientists, bioinformaticians, researchers, developers, and government stakeholders to translate scientific needs into technical specifications, data models, and reusable workflows.

Required Qualifications

  • Bachelor's degree in Data Science, Bioinformatics, Computer Science, Biological Sciences, or a related field (advanced degree preferred), or equivalent experience.
  • Demonstrated experience in a data-intensive role supporting biomedical research or scientific computing.
  • Strong proficiency in Python and R for analysis, scripting, and visualization.
  • Hands-on experience with at least two omics data types (e.g., bulk RNA-seq, scRNA-seq, spatial transcriptomics, proteomics, metagenomics, GWAS).
  • Solid understanding of statistical modeling, dimensionality reduction, clustering, differential expression, and pathway analysis.
  • Ability to work with structured, semi-structured, and unstructured data across relational and data lake environments.
  • Strong problem-solving skills with the ability to communicate effectively across technical and non-technical audiences.
  • Genuine interest in biomedical and translational research; awareness of data governance, privacy, and compliance requirements.

Preferred Qualifications

  • Experience building analytics solutions in platforms such as Snowflake, Databricks, or cloud data warehouses.
  • Experience with workflow and reproducibility tools: Galaxy, Terra, Nextflow/WDL, Snakemake, Singularity, CWL.
  • Familiarity with the scverse Python ecosystem (Scanpy, Squidpy, SCIMAP, AnnData) and spatial single-cell analysis methods (PhenoGraph, Louvain/Leiden clustering, UMAP, Ripley's L statistic).
  • Experience preparing curated datasets for dashboards, APIs, and web applications; familiarity with Posit Connect, R/Shiny, Streamlit, Jupyter.
  • Experience with AWS (EC2, S3, Lambda), object storage, relational databases, scheduled jobs, API integrations, secure data movement; familiarity with HPC environments, SLURM/SGE, or NIH Biowulf.
  • Background in biomedical research, clinical research, or healthcare analytics; familiarity with HL7/FHIR, CDISC, or OMOP standards.
  • Experience with metadata management, data lineage, open-source code release, containerized analyses, and secure handling of de-identified or access-controlled research datasets.
  • Experience creating documentation, training materials, or workshops for researchers and non-coder audiences.

Skills

Python, R, Rna-Seq, Scrna-Seq, Spatial Transcriptomics, Proteomics, Metagenomics, Gwas, Statistical Modeling, Machine Learning, Pca, Umap, T-Sne, Clustering, Differential Expression Analysis

Socure

Socure

San Francisco, CA
Data Scientist II - Big Data R&D, Identity Graph & Deceased Monitoring
$140k+/yrHybrid2+ YOEData Science

Data Scientist II developing graph-based algorithms, entity-resolution models, and large-scale data pipelines for identity and deceased-monitoring products. Requires advanced academic training or equivalent experience, strong Python/Scala and SQL skills, and hands-on Spark and AWS experience.

Forge

Forge

San Francisco, CA

Private Markets Research Associate
$142k+/yrOn-site7+ YOEData Science

Analyzes proprietary private-market transaction and pricing data to identify valuation, liquidity, and investor trends, then translates findings into research publications and investor-facing thought leadership. Requires a bachelor's degree, strong quantitative skills, Excel and Python proficiency, and 3–7 years of relevant experience.

Notion

Notion

San Francisco, CA

Data Science Intern
$114k+/yrHybridData Science

Analyzes product and business data, develops metrics and dashboards, and communicates actionable recommendations to cross-functional teams and leadership. Candidates should be pursuing a quantitative bachelor's or master's degree, with SQL proficiency and familiarity with Python or R.

Layer Health

Layer Health

Boston, MA
Forward Deploy Data Scientist
$150k+/yrHybrid2+ YOEData Science

The Forward Deploy Data Scientist will partner with health systems and internal teams to investigate healthcare data, build and operationalize ML/LLM pipelines, and deliver actionable insights. The role requires 2–3 years of data science experience, strong Python and ML/NLP skills, and customer-facing communication ability.

Coinbase

Coinbase

San Francisco, CA

Data Science Intern
$50+/hrHybridData Science

Data Science Intern supporting product and engineering teams through statistical analysis, experimentation, analytical modeling, and data-driven recommendations. The role requires quantitative academic study or equivalent project experience, plus familiarity with Python and SQL.