# Data Scientist

**Company:** [Arena](https://hotfix.jobs/companies/arena)
**Location:** San Francisco, CA
**Role:** Data Science
**Experience:** 6+ years
**Skills:** Python, pandas, NumPy, Spark, Statistical Modeling, Causal Inference, Experimental Design, A/B Testing, Llm Outputs, Embeddings
**Posted:** 2026-08-31

> Analyzes large-scale datasets from AI model evaluations to uncover patterns, biases, and causal relationships. Designs experiments, builds pipelines with Python/Pandas/Spark, and collaborates with ML teams on metrics and insights. Requires 6+ years in data science/ML analytics.

## Job Description

## Responsibilities
- Explore and analyze large, complex datasets to uncover patterns, biases, and causal relationships in model behavior and system performance.
- Formulate hypotheses about data quality, evaluation outcomes, and model performance — then design experiments to validate or refute them.
- Build reproducible analysis pipelines using Python, Pandas, NumPy, and Spark to process and interrogate large-scale data.
- Partner with ML researchers and engineers to design metrics and analyses that evaluate how models perform across domains, prompts, and tasks.
- Develop causal reasoning frameworks and statistical methods that help explain why models behave as they do — not just how well they perform.
- Communicate insights (for example, via blog posts) clearly to technical and non-technical partners, informing both research direction and infrastructure improvements.

## Requirements
- 6+ years of experience in data science, ML analytics, or applied research, preferably in AI, ML, or large-scale data environments.
- Strong proficiency in Python, with deep experience in **Pandas**, **NumPy**, and distributed frameworks like **Spark**.
- Expertise in statistical modeling, **causal inference**, and experimental design.
- Experience reasoning about data distributions, sample quality, and the effects of data distribution shifts.
- Strong communication skills and the ability to collaborate closely with ML researchers and engineers.

## Nice-to-haves
- Background in AI model evaluation.
- Experience working with LLM outputs (for example, LLM-as-a-judge), embeddings, or other large-scale model artifacts.
- Experience with A/B testing.

## Similar jobs

- [Senior Data Scientist, Organic Growth](https://hotfix.jobs/jobs/790d357e-5d4f-4070-ada3-a38c3020abc7) - Chime - San Francisco, CA - $133k – $185k/yr
- [Senior Data Scientist, Risk and Support](https://hotfix.jobs/jobs/3a233707-7146-45b6-b6bc-c31731150278) - Square - Remote - $168k – $297k/yr
- [Senior Engineer, Operations Analyst](https://hotfix.jobs/jobs/b469fc3b-2fbd-4158-b432-5f0de488c864) - Shield AI - Washington, DC - $128k – $192k/yr
- [Senior Data Scientist](https://hotfix.jobs/jobs/e9c87139-0f0f-41cc-a442-28a4958bb090) - Bestow - Remote - $125k – $140k/yr
- [Senior Data Scientist, Business Analytics](https://hotfix.jobs/jobs/db971610-ecfe-4488-a2b3-84577cefbc27) - Fora - New York, NY - $130k – $185k/yr

**Apply:** https://hotfix.jobs/jobs/8d0f48ab-dcc6-45ba-bf2b-d47a5dec49db
**Canonical:** https://hotfix.jobs/jobs/8d0f48ab-dcc6-45ba-bf2b-d47a5dec49db