P
Perplexity
Data Scientist

Member of Technical Staff (Data Scientist, Evals)

On-siteMidData Scientistposted 3mo ago
✦Role summaryAI-generated

Builds specialized, automated evaluation frameworks to measure LLM-driven search and answer quality at Perplexity, focusing on tool-call impact and visual rendering metrics for high-scale, real-world use cases.

Skills required

About this role

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases. In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.

Responsibilities

  • Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness

  • Design evaluation sets and methods specifically to measure the impact of tool calls (particularly web search retrieval) on the final answer's quality

  • Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices

  • Continuously review public benchmarks and academic evaluations for their applicability to the Perplexity product, adapting and incorporating them into our regular performance measurements

  • Operate within a small, high-impact team where your evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve Answer Quality

Qualifications

  • PhD or MS in a technical field or equivalent experience

  • 4+ years of experience in data science or machine learning

  • Strong proficiency in Python and SQL (expected to write production-grade code)

  • Experience building within a modern cloud data stack, specifically AWS and Databricks

  • Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster

Preferred Qualifications

  • 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups

  • Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale

  • A strong research background, with experience applying research methods to real-world ML problems

  • Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets

✕ position closed

This role is no longer accepting applications. It’s kept here for reference — check out the similar open roles below.

Similar open roles

UP
NEW

Customer First Analytics Data Science Manager - FLEX Location - Atlanta Preferred

United Parcel Service·340 MACARTHUR BLVD, MAHWAH, NJ 07430, United States of America
HybridSeniorData Scientist
$109k USD
yesterday
RB
NEW

Staff Data Scientist, AI Evaluations Platform

Royal Bank of Canada·MONTRÉAL, Quebec, Canada
On-siteStaffData Scientist
yesterday
P
NEW

Quantitative Analytics & Model Analyst Senior - Credit Loss Forecasting (CLF)

PNC·Pittsburgh, Pennsylvania, United States of America; Washington, District of Columbia, United States of America; Charlotte, North Carolina, United States of America; Cleveland, Ohio, United States of America; Columbus, Ohio, United States of America; Tysons Corner, Virginia, United States of America
On-siteSeniorData Scientist
$86k – $173k USD
yesterday
R
NEW

Senior Data Scientist (AI-assisted Clinical Development)

Roche·Boston, Massachusetts, United States of America
On-siteSeniorData Scientist
$143k – $265k USD
yesterday
JC
NEW

Quant Analytics

JPMorgan Chase·Columbiana, OH, United States; San Antonio, TX, United States
On-siteSeniorData Scientist
yesterday
J
NEW

Senior Data Scientist

JDI·Saint John, NB, Canada
On-siteSeniorData Scientist
yesterday