# Data Scientist II - Big Data R&D, Identity Graph & Deceased Monitoring

**Company:** [Socure](https://hotfix.jobs/companies/socure)
**Location:** San Francisco, CA, Seattle, WA, New York, NY, Miami, FL
**Role:** Data Science
**Salary:** $140k – $170k/yr
**Experience:** 2+ years
**Skills:** Python, Scala, SQL, Spark, Pyspark, AWS, Amazon Emr, Amazon S3, Machine Learning, scikit-learn, Xgboost, Graphframes, Neo4J, Airflow
**Posted:** 2026-09-02

> Data Scientist II developing graph-based algorithms, entity-resolution models, and large-scale data pipelines for identity and deceased-monitoring products. Requires advanced academic training or equivalent experience, strong Python/Scala and SQL skills, and hands-on Spark and AWS experience.

## Job Description

## Responsibilities
- Design and implement machine-learning, data-mining, statistical, and graph-based algorithms for large-scale identity verification and anomaly detection.
- Analyze large datasets to develop and refine entity-resolution and identity-matching algorithms for deceased monitoring and compliance solutions.
- Build and maintain ETL, feature-generation, normalization, and other data-processing pipeline components using Spark/PySpark and AWS services such as EMR and S3.
- Support senior data scientists with feature engineering, data exploration, error analysis, and A/B test setup.
- Evaluate third-party and internal data sources, including data-quality profiling, offline experiments, and impact analysis on coverage and model performance.
- Maintain SQL and Python/R code for data extraction, transformation, and validation; contribute to code reviews and basic testing.
- Provide analytical support to compliance and regulatory product teams through investigations, dashboards, and data deep dives.
- Communicate findings and trade-offs to Product, Engineering, Client Analysis, and other cross-functional partners.

## Requirements
- Master’s degree with 2+ years of data science or analytics experience, Ph.D. with 1+ years of experience, or equivalent practical experience.
- Proficiency in Python or Scala.
- Experience writing and optimizing SQL for large datasets and working in data lake or data warehouse environments.
- Hands-on experience with Spark or PySpark and common machine-learning libraries.
- Familiarity with Unix environments and AWS.
- Working knowledge of supervised and unsupervised machine learning and basic statistics.
- Ability to break down loosely defined problems, ask clarifying questions, and iterate quickly.

## Nice-to-haves
- TensorFlow or PyTorch.
- Databricks.
- Graph techniques or graph databases such as Neo4j, AWS Neptune, or GraphFrames.
- Elasticsearch or DynamoDB.
- Airflow for automating data pipelines.

## Compensation
- Annual salary: $140,000–$170,000.

## Similar jobs

- [Private Markets Research Associate](https://hotfix.jobs/jobs/a89c224e-f20f-41f0-9619-48318a86d964) - Forge - San Francisco, CA - $142k – $180k/yr
- [Forward Deploy Data Scientist](https://hotfix.jobs/jobs/09af735d-01e0-427f-bfd3-8bf43b3b8c6c) - Layer Health - Boston, MA - $150k – $180k/yr
- [Data Science Intern](https://hotfix.jobs/jobs/71f696af-e67a-4bc5-b321-7ca8332cfe98) - Notion - San Francisco, CA - $114k – $123k/yr
- [Data Science Intern](https://hotfix.jobs/jobs/f5e5a495-739d-4996-aba9-a0de5d630d5a) - Coinbase - San Francisco, CA - $50 – $50/hr
- [Associate Data Scientist](https://hotfix.jobs/jobs/0905303c-677a-47e2-95dd-d83d1df084aa) - Solace - Redwood City, CA

**Apply:** https://hotfix.jobs/jobs/19c05636-52d6-4dda-86ec-9fb2bcca6334
**Canonical:** https://hotfix.jobs/jobs/19c05636-52d6-4dda-86ec-9fb2bcca6334