The Data Scientist will turn large-scale LLM usage data into routing improvements, analytics products, and user-facing features. The role requires 4+ years of data science, applied ML, or quantitative product experience, plus strong SQL, Python, statistical, and AI expertise.
180k – 240k/yr
Remote4+ YOEData Science
About the role
Responsibilities
Build and improve the rules, heuristics, and models that drive model and provider routing systems.
Build analytics products and conduct deep-dive analyses that help users understand usage, diagnose issues, and optimize workflows while informing product strategy.
Design and ship user-facing data features, including task-type classification taxonomies and public rankings pages analyzing model performance by use case.
Define and track product metrics such as feature adoption, user engagement, and platform health, tying insights to roadmap and architectural decisions.
Contribute to the data warehouse using dbt and ClickHouse.
Requirements
Experience building and shipping production-level data products, including models, algorithms, classifiers, scoring systems, or automated processes.
Ability to partner with engineering and product teams to turn insights and analytics into customer-facing features.
Strong SQL skills, including complex, performant queries against large-scale analytical databases such as ClickHouse or BigQuery.
Deep fluency with AI and agents, including model evaluation, prompt engineering, and the agentic ecosystem.
Excellent statistical and mathematical foundations, including distributions, causal inference, and experimental design.
Proficiency in Python or a similar scripting language for prototyping models, running simulations, and building data pipelines.
4+ years of relevant experience in data science, applied machine learning, or quantitative product roles at a high-growth company.
Product orientation and the ability to build for developers and translate ambiguous systems into shippable solutions.
Ability to work independently in fast-moving, low-structure environments.
Strong problem-solving and written communication skills.
Nice-to-haves
Experience with LLM evaluation and benchmarking, including online or offline evaluations, human-in-the-loop assessment, and safety or quality guardrails.
Experience with developer platforms, APIs, SDKs, or usage-based products.
NLP or machine learning engineering experience with classifiers, embeddings, or ranking models.
TypeScript experience and ability to contribute to routing, API, or front-end codebases.
Builds and owns fraud detection and financial-risk models end to end, from data acquisition and feature engineering through production deployment and monitoring. The role requires strong practical machine learning and statistics expertise, production coding experience, and either 4+ years with a relevant master’s degree or 2+ years with a relevant PhD.
180k – 220k/yrRemote4+ YOEData Science
Data Scientist
Layer HealthBoston, MA +1
Develops and validates LLM/ML models for healthcare chart abstraction from unstructured/structured data, evaluates cutting-edge NLP/AI techniques for clinical problems, and collaborates with customers to optimize workflows. Requires 5-7+ years data science experience including 3-4 years in healthcare.
180k – 200k/yrHybrid5+ YOEData Science
Data Scientist, Product
ReplitFoster City, CA
Analyzes user behavior and product experiments to drive insights for growth, retention, and enterprise strategy at Replit. Requires 5+ years in product analytics, strong SQL/Python skills, and expertise in A/B testing and causal inference.
180k – 250k/yrHybrid5+ YOEData Science
Data Scientist
RetoolSan Francisco, CA
Data Scientist analyzes customer behavior, product usage, and business performance to drive strategic decisions across Product, GTM, and leadership. Requires 5+ years in data science/product analytics, strong SQL, and experience in high-growth B2B SaaS environments.
183k – 250k/yrOn-site5+ YOEData Science
Data Scientist - Mapping
ZooxFoster City, CA
Design and build large-scale HD mapping datasets, data pipelines, and benchmarking tools to improve ML model accuracy for autonomous driving. Requires 3+ years experience, strong Python and statistical skills.