Skip to content
SocureSocure

Data Scientist II - Big Data R&D, Identity Graph & Deceased Monitoring

Data Scientist II developing graph-based algorithms, entity-resolution models, and large-scale data pipelines for identity and deceased-monitoring products. Requires advanced academic training or equivalent experience, strong Python/Scala and SQL skills, and hands-on Spark and AWS experience.

About the job

Responsibilities

  • Design and implement machine-learning, data-mining, statistical, and graph-based algorithms for large-scale identity verification and anomaly detection.
  • Analyze large datasets to develop and refine entity-resolution and identity-matching algorithms for deceased monitoring and compliance solutions.
  • Build and maintain ETL, feature-generation, normalization, and other data-processing pipeline components using Spark/PySpark and AWS services such as EMR and S3.
  • Support senior data scientists with feature engineering, data exploration, error analysis, and A/B test setup.
  • Evaluate third-party and internal data sources, including data-quality profiling, offline experiments, and impact analysis on coverage and model performance.
  • Maintain SQL and Python/R code for data extraction, transformation, and validation; contribute to code reviews and basic testing.
  • Provide analytical support to compliance and regulatory product teams through investigations, dashboards, and data deep dives.
  • Communicate findings and trade-offs to Product, Engineering, Client Analysis, and other cross-functional partners.

Requirements

  • Master’s degree with 2+ years of data science or analytics experience, Ph.D. with 1+ years of experience, or equivalent practical experience.
  • Proficiency in Python or Scala.
  • Experience writing and optimizing SQL for large datasets and working in data lake or data warehouse environments.
  • Hands-on experience with Spark or PySpark and common machine-learning libraries.
  • Familiarity with Unix environments and AWS.
  • Working knowledge of supervised and unsupervised machine learning and basic statistics.
  • Ability to break down loosely defined problems, ask clarifying questions, and iterate quickly.

Nice-to-haves

  • TensorFlow or PyTorch.
  • Databricks.
  • Graph techniques or graph databases such as Neo4j, AWS Neptune, or GraphFrames.
  • Elasticsearch or DynamoDB.
  • Airflow for automating data pipelines.

Compensation

  • Annual salary: $140,000–$170,000.

Skills

Python, Scala, SQL, Spark, Pyspark, AWS, Amazon Emr, Amazon S3, Machine Learning, scikit-learn, Xgboost, Graphframes, Neo4J, Airflow

Forge

Forge

San Francisco, CA

Private Markets Research Associate
$142k+/yrOn-site7+ YOEData Science

Analyzes proprietary private-market transaction and pricing data to identify valuation, liquidity, and investor trends, then translates findings into research publications and investor-facing thought leadership. Requires a bachelor's degree, strong quantitative skills, Excel and Python proficiency, and 3–7 years of relevant experience.

Layer Health

Layer Health

Boston, MA
Forward Deploy Data Scientist
$150k+/yrHybrid2+ YOEData Science

The Forward Deploy Data Scientist will partner with health systems and internal teams to investigate healthcare data, build and operationalize ML/LLM pipelines, and deliver actionable insights. The role requires 2–3 years of data science experience, strong Python and ML/NLP skills, and customer-facing communication ability.

Notion

Notion

San Francisco, CA

Data Science Intern
$114k+/yrHybridData Science

Analyzes product and business data, develops metrics and dashboards, and communicates actionable recommendations to cross-functional teams and leadership. Candidates should be pursuing a quantitative bachelor's or master's degree, with SQL proficiency and familiarity with Python or R.

Coinbase

Coinbase

San Francisco, CA

Data Science Intern
$50+/hrHybridData Science

Data Science Intern supporting product and engineering teams through statistical analysis, experimentation, analytical modeling, and data-driven recommendations. The role requires quantitative academic study or equivalent project experience, plus familiarity with Python and SQL.

Solace

Solace

Redwood City, CA

Associate Data Scientist
No salary listedHybridData Science

Associate Data Scientist role for a 2027 college graduate, contributing to forecasting, matching, experimentation, and machine-learning systems. Requires quantitative academic training, Python and SQL proficiency, strong statistical foundations, and clear communication.