Member of Technical Staff (Data Scientist, Evals)
USD 200,000-300,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
AWS @ 3
Data Science @ 5
Databricks @ 3
LLM @ 3
Machine Learning @ 5
Python @ 6
SQL @ 6
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. This role focuses on building specialized evaluations to improve answer quality across Perplexity, including search-based LLM answers and other scenarios popular with users.
Responsibilities
- Architect and maintain automated evaluation pipelines to assess answer quality across Perplexity's products, ensuring high standards for accuracy and helpfulness.
- Design evaluation sets and methods to measure the impact of tool calls, particularly web search retrieval, on final answer quality.
- Develop VLM-based solutions to programmatically evaluate how final answers render visually across different platforms and devices.
- Continuously review public benchmarks and academic evaluations for applicability to the Perplexity product, adapting and incorporating them into regular performance measurements.
- Work within a small, high-impact team where evaluation metrics directly shape product changes, collaborating closely with technical leadership to measure and improve answer quality.
Requirements
- PhD or MS in a technical field, or equivalent experience.
- 4+ years of experience in data science or machine learning.
- Strong proficiency in Python and SQL, including the ability to write production-grade code.
- Experience building within a modern cloud data stack, specifically AWS and Databricks.
- Comfort with agentic coding workflows and AI-assisted development tools.
Preferred Qualifications
- 1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups.
- Experience working on customer-facing web products or consumer apps with real user traffic at scale.
- Strong research background and experience applying research methods to real-world machine learning problems.
- Experience defining evaluation metrics such as factual consistency, hallucination rate, and retrieval precision, and building ground-truth datasets.
Benefits
- U.S. full-time employees receive benefits including equity, health, dental, vision, retirement, fitness, commuter, and dependent care accounts.
- International employees receive a comprehensive benefits program tailored to their region of residence.
- USD salary ranges apply only to U.S.-based positions. Final offer amounts depend on factors including experience and expertise.
More jobs at Perplexity AI
Member of Technical Staff (Design Systems)
Perplexity AI · New York City, United States, San Francisco, United States, Seattle, United States
USD 220,000-405,000 per year
Member Of Technical Staff (Software Engineer, Backend API)
Perplexity AI · Serbia, Berlin, Germany, London, United Kingdom, New York City, United States
USD 200,000-350,000 per year
Legal and Operations Engineering Lead
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 150,000-250,000 per year
Member of Technical Staff (Android Engineer, Computer Growth)
Perplexity AI · New York City, United States, San Francisco, United States
USD 220,000-405,000 per year
Member of Technical Staff (iOS Engineer, Computer Growth)
Perplexity AI · New York City, United States, San Francisco, United States
USD 220,000-405,000 per year
Similar jobs
Principal Data Scientist - Cloud Gaming and AI
Nvidia · Santa Clara, United States
USD 248,000-379,500 per year
Member Of Technical Staff (Software Engineer, Agent Capabilities)
Perplexity AI · New York City, United States, San Francisco, United States
USD 220,000-405,000 per year
Member of Technical Staff (Software Engineer, Data Flywheel)
Perplexity AI · New York City, United States, San Francisco, United States
USD 200,000-350,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Lead Data Scientist, Platform Product
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 285,000-380,000 per year
Member of Technical Staff (Software Engineer, Computer Growth)
Perplexity AI · New York City, United States, San Francisco, United States
USD 220,000-405,000 per year
Data Scientist, GTM
Anthropic · New York City, United States, San Francisco, United States
USD 285,000-380,000 per year
Customer Success Insights Engineer
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year