Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API
Audit @ 4
Azure @ 7
Compliance @ 4
Data Pipelines
Flink @ 6
Machine Learning
Security @ 4
Spark @ 6
Trino @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Stripe is a financial infrastructure platform used by millions of businesses. The Data Lake team builds and maintains foundational data access and governance infrastructure for safe, fast, and compliant access to Stripe's critical big data assets. The team supports developers, data engineers, analysts, ML and AI teams, security teams, and business users across the company.
The team is undertaking a major architectural transition to make Stripe's data lake a first-class citizen of the modern data ecosystem. This includes a multi-year migration to open-source solutions such as the Apache Iceberg REST Catalog, ownership of the object storage abstraction layer, and the design of access control, IAM policy, lifecycle management, and compliance architecture for petabyte-scale data infrastructure.
Responsibilities
- Architect the unified Iceberg platform and lead the technical design of a metastore service serving as the single source of truth for Iceberg table management across Spark, Trino, Flink, and PyIceberg.
- Define API contracts, authorization models, per-table credential vending, and integration patterns for the company's data pipelines.
- Own the metastore migration strategy, including sequencing, backward compatibility, rollback planning, and coordination with dozens of consuming teams while maintaining production operations.
- Define the object storage abstraction layer, including bucket provisioning, access control policy design, and developer-facing client libraries.
- Lead compliance architecture by partnering with security and compliance teams to implement audit logging, access review infrastructure, data segregation, and lifecycle enforcement as preventative technical controls.
- Drive cost and efficiency improvements at petabyte scale across storage layout, snapshot retention, and data lifecycle management.
- Design automated, self-service tooling that scales without ongoing manual intervention.
- Lead critical design reviews, establish reliability, security, and developer experience standards, and mentor senior engineers through complex architectural decisions.
Requirements
Minimum Requirements
- 10+ years of professional software engineering experience.
- Demonstrated experience designing, building, and operating large-scale distributed storage or data infrastructure systems.
- Deep experience with object storage such as S3, Azure Blob, or equivalent, including IAM, access control policy design, lifecycle management, and petabyte-scale operations.
- Experience leading complex, multi-quarter infrastructure projects end-to-end, including cross-team dependency management and migrations across many consuming teams.
- Strong background in authorization and access control design for distributed data systems.
Preferred Requirements
- Deep expertise in Apache Iceberg, including table format internals, the REST Catalog specification, snapshot lifecycle management, compaction, and integration with Spark, Trino, Flink, and PyIceberg.
- Experience with compliance-sensitive infrastructure, including SOX, ICFR, or equivalent regulatory frameworks, and translating audit and access review requirements into preventative technical controls.
- Experience executing large-scale data migrations with careful sequencing, blast-radius reduction, rollback planning, and data integrity validation.
- Strong developer experience sensibility, including building ergonomic, well-documented abstractions that reduce toil for engineering teams.