Data Scientist, Core Experimentation
Leads evolution of OpenAI's core experimentation platform, driving statistical strategy, designing methodologies, and building scalable Python/Spark pipelines to ensure reliable, trustworthy experiments at massive scale. Requires deep stats expertise, causal inference, and production experimentation experience.
About the job
Responsibilities
- Drive the statistical direction and technical strategy for OpenAI’s experimentation platform
- Design and improve experimentation methodologies used across product and research teams
- Build pragmatic solutions to real-world experimentation challenges, balancing rigor with operational simplicity
- Improve the reliability and trustworthiness of experiment results, including detection and prevention of bias, logging issues, and data quality failures
- Develop scalable analytical systems and pipelines in Python and distributed compute environments
- Partner with engineers and product teams to improve experiment design, metric quality, and decision-making practices
- Lead investigations into complex experimentation anomalies and measurement failures
- Establish best practices for experimentation governance, interpretation, and statistical correctness
- Mentor other data scientists and raise the overall technical bar for experimentation and causal inference
Requirements
- Experience building, scaling, or operating experimentation platforms at a large technology company
- Deep expertise in statistics, causal inference, and online experimentation methodology
- Strong understanding of practical experimentation challenges in production systems
- Experience with areas such as variance reduction, CUPED, sequential testing, SRM detection, metric design, or heterogeneous effects
- Strong coding and systems skills in Python and large-scale data processing frameworks (e.g. Spark)
- Experience designing analytical data models and scalable experimentation pipelines
- Ability to communicate complex statistical concepts clearly to technical and non-technical audiences
- Track record of influencing technical strategy through hands-on technical leadership
Nice-to-Haves
- Experience in large-scale product experimentation, ML experimentation, ranking systems, marketplace systems, or similar high-scale experimentation domains
Compensation
$293K - $325K USD
Skills
Python, Spark, Statistics, Causal Inference, Online Experimentation, Cuped, Sequential Testing, Srm Detection, Variance Reduction, Metric Design
Similar jobs
Data Science jobsThis senior data scientist will define measurement frameworks for AI-agent security, evaluate controls and security findings, and improve detection and response outcomes. The role requires 5+ years of quantitative experience, strong SQL and Python skills, and experience with cybersecurity or other adversarial-risk domains.
The Data Scientist will shape analytics strategy and deliver forecasting, experimentation, optimization, dashboards, models, and decision tools for Real Estate & Workplace operations. The role requires strong applied statistics, causal inference, SQL, Python, stakeholder communication, and comfort working with ambiguous operational data.
People Data Scientist focused on AI fairness and bias testing for People systems. Designs algorithmic audits, validation studies, and fairness infrastructure across the employee lifecycle. Requires deep expertise in fairness metrics, statistical modeling, and Python/R/SQL.
Develops scientifically rigorous sustainability methodologies that become product capabilities for corporate climate and ESG data. The role combines climate expertise, GHG accounting, standards interpretation, data reasoning, and hands-on collaboration with engineers and product teams.
Senior Data Scientist supporting Plaid’s Credit product area, translating ambiguous product questions into analytics, metrics, experiments, and strategic insights. The role partners closely with product and engineering teams and requires 5+ years of data science or related analytics experience.