PhD Data Scientist, Intern
📍 New York City, United States
📍 South San Francisco, United States
📍 Seattle, United States
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 5
API
Data Analysis
Data Pipelines @ 3
Data Science @ 3
Debugging @ 3
Machine Learning
Mathematics @ 3
Python @ 3
R @ 3
SQL @ 3
Spark
Statistics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Who We Are
About Stripe
Stripe is a technology company focused on improving the conditions for economic growth and prosperity. We build programmable financial infrastructure, rethinking from first principles how financial services should work, to make it easier and cheaper for any business to start and scale. More than 10 million businesses build on Stripe, spanning the economic frontier—from solo founders to established enterprises—united by a practical focus on growth.
Stripe maintains reliable APIs and builds financial infrastructure, risk and fraud systems, and products that support businesses around the world.
About the Team
The Data Science team partners deeply with teams across Stripe to ensure that users, products, and the business have the models, data products, and insights needed to make decisions and grow responsibly. The Fraud, Losses, and Financial Crime Data Science team builds models and data products that protect Stripe and its users from fraud, account takeover, and financial crime.
The team owns the full fraud and loss modeling stack, including account takeover detection, card fraud classification, merchant-level loss estimation, unsupervised anomaly detection, and financial crime risk modeling. It partners with Fraud Engineering, Financial Crimes Engineering, and Risk Operations to bring these systems into production and ensure measurable impact on Stripe’s financial integrity and user trust.
Responsibilities
As an intern, you will work on projects across Stripe’s technology stack that directly impact millions of businesses. You will own problems end to end with support from your manager and teammates, partner with teams across Stripe, extract insights from complex data, and deliver actionable business recommendations through analysis and data storytelling.
- Apply probability distributions, statistical inference, and hypothesis testing to quantify uncertainty and evaluate business outcomes.
- Use Python or R for data analysis, data processing, visualizations, statistical modeling, machine learning, predictive analytics, automation, causal inference, and experimental analyses.
- Build, train, and evaluate predictive models across regression and classification tasks, including bias-variance trade-offs and model selection.
- Model temporal dependencies, seasonality, and trend decomposition to generate and evaluate time-series predictions.
- Identify structural patterns, clusters, and outliers in unlabeled data.
- Deploy models in production and adjust model thresholds to improve performance.
- Design, run, and analyze complex experiments using causal inference designs.
- Use SQL and Spark to create, transform, and analyze large datasets.
- Learn quickly, ask effective questions, work productively with mentors and teammates, and communicate work status clearly.
- Present work to the Data Science team, partner teams, and fellow interns.
The internship program is competitive and expectations are high.
Requirements
Minimum Requirements
- Enrollment in a quantitative PhD program, such as Data Science, Statistics, Economics, or Mathematics, with an expected graduation date of December 2027 or spring/summer 2028.
- Experience with SQL and a scientific computing language such as Python or R.
- Proficiency with AI tools to accelerate model development, analysis, and coding.
- Experience communicating and collaborating with multidisciplinary stakeholders in a team environment.
Preferred Qualifications
- Experience writing and debugging data pipelines.
- Demonstrated ability to evaluate and receive feedback from mentors, peers, and stakeholders through previous internships or other multi-person projects.
- Ability to learn new systems and develop an understanding of those systems through independent research and collaboration with mentors and subject-matter experts.
Additional Attributes
- Ambitious builder who is energized by creating solutions without clear precedent and solving problems with far-reaching consequences.
- Rigorous thinker who enjoys working on complex problems that have not been tackled before.
- Adaptable problem solver who treats obstacles as opportunities and can act boldly in the absence of consensus.