Member of Technical Staff
Build specialized evals and automated pipelines to measure and improve answer quality for Perplexity's LLM-powered search engine, focusing on retrieval, tool calls, and visual rendering. Requires 4+ years in data science/ML, strong Python/SQL, and cloud experience (MS/PhD preferred).
About the job
Responsibilities
- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness.
- Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality.
- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices.
- Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements.
- Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality.
Requirements
- PhD or MS in a technical field or equivalent experience.
- 4+ years of experience in data science or machine learning.
- Strong proficiency in Python and SQL (expected to write production-grade code).
- Experience building within a modern cloud data stack, specifically AWS and Databricks.
- Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster.
Preferred Qualifications
- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups.
- Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale.
- A strong research background, with experience applying research methods to real-world ML problems.
- Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets.
Skills
Python, SQL, AWS, Databricks, LLMs, Vlm, Machine Learning, Data Science, Llm-As-A-Judge
Similar jobs
Data Science jobsPeople Data Scientist focused on AI fairness and bias testing for People systems. Designs algorithmic audits, validation studies, and fairness infrastructure across the employee lifecycle. Requires deep expertise in fairness metrics, statistical modeling, and Python/R/SQL.
Develops scientifically rigorous sustainability methodologies that become product capabilities for corporate climate and ESG data. The role combines climate expertise, GHG accounting, standards interpretation, data reasoning, and hands-on collaboration with engineers and product teams.
Senior Data Scientist supporting Plaid’s Credit product area, translating ambiguous product questions into analytics, metrics, experiments, and strategic insights. The role partners closely with product and engineering teams and requires 5+ years of data science or related analytics experience.
The Data Scientist will turn large-scale LLM usage data into routing improvements, analytics products, and user-facing features. The role requires 4+ years of data science, applied ML, or quantitative product experience, plus strong SQL, Python, statistical, and AI expertise.
Builds and owns fraud detection and financial-risk models end to end, from data acquisition and feature engineering through production deployment and monitoring. The role requires strong practical machine learning and statistics expertise, production coding experience, and either 4+ years with a relevant master’s degree or 2+ years with a relevant PhD.