Member of Technical Staff (Data Scientist, Evals)

USD 200,000-300,000 per year
MIDDLE
✅ Hybrid

Tech Stack

AI @ 3 AWS @ 3 Data Science @ 5 Databricks @ 3 LLM @ 3 Machine Learning @ 5 Python @ 6 SQL @ 6 Technical Leadership

Details

Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. This role focuses on building specialized evaluations to improve answer quality across Perplexity, including search-based LLM answers and other scenarios popular with users.

Responsibilities

  • Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness.
  • Design evaluation sets and methods to measure the impact of tool calls, particularly web search retrieval, on final answer quality.
  • Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices.
  • Continuously review public benchmarks and academic evaluations for applicability to the Perplexity product, adapting and incorporating them into regular performance measurements.
  • Work within a small, high-impact team where evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve answer quality.

Requirements

  • PhD or MS in a technical field, or equivalent experience.
  • 4+ years of experience in data science or machine learning.
  • Strong proficiency in Python and SQL, including the ability to write production-grade code.
  • Experience building within a modern cloud data stack, specifically AWS and Databricks.
  • Comfort with agentic coding workflows and AI-assisted development tools.

Preferred Qualifications

  • 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups.
  • Experience working on customer-facing web products or consumer apps with real user traffic at scale.
  • Strong research background and experience applying research methods to real-world machine learning problems.
  • Experience defining evaluation metrics such as factual consistency, hallucination rate, and retrieval precision, and building ground-truth datasets.

Benefits

  • U.S. full-time employees receive benefits including equity, health, dental, vision, retirement, fitness, commuter, and dependent care accounts.
  • International employees receive a comprehensive benefits program tailored to their region of residence.
  • USD salary ranges apply only to U.S.-based positions. Final offer amounts depend on factors including experience and expertise.

More jobs at Perplexity AI

Similar jobs