Staff Software Engineer, Stream Compute
📍 New York City, United States
📍 South San Francisco, United States
📍 Seattle, United States
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS
Communication @ 7
Debugging @ 4
Distributed Systems @ 4
Flink @ 4
Go @ 6
Hadoop @ 6
Java @ 6
Kafka @ 4
Payments
SQL @ 4
Scala @ 6
Spark @ 4
Spark Streaming @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Who We Are
About Stripe
Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world's largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Stripe's mission is to increase the GDP of the internet.
About the Team
The Stream Compute team at Stripe builds and operates the infrastructure, tooling, and systems behind Stripe's Apache Flink-powered stream processing systems. The team supports core asynchronous workflows involving critical financial operations and real-time analytics, handling sensitive financial data at significant scale.
The team operates globally distributed systems with high reliability and performance requirements. It also invests in automation and self-service tooling for upgrades, maintenance, and daily operations. The team is distributed between Seattle, Toronto, and remote locations.
The team focuses on ensuring that no event is dropped, state integrity is preserved, and exactly-once processing is supported as a first-class feature.
Responsibilities
- Design, build, and operate stream compute infrastructure centered on Apache Flink, alongside technologies such as Kafka, Temporal, and AWS services.
- Partner with product and platform teams across Stripe to understand requirements, unblock Flink adoption, and improve end-to-end stream-processing infrastructure usage.
- Define and implement operational best practices, including shuffle sharding, cellular architecture, load shedding, and automated state recovery, to improve resilience and reliability at scale.
- Drive fleet-level automation and standardization through self-service workflows, safer rollouts, and self-healing systems that reduce manual operations.
- Lead initiatives to improve Flink availability and state durability, including multi-region strategies, disaster-recovery readiness, operational-readiness reviews, and incident learning.
- Evaluate and productionize Flink ecosystem capabilities, including SQL, connectors, and state backends, to improve developer experience and scalability without compromising reliability.
- Work with the open-source community to identify opportunities to adopt new open-source features and contribute back to open source.
Requirements
- This is a Staff-level role, typically requiring 10+ years of experience building, operating, and evolving large-scale production systems.
- Experience as a technical lead for teams working on distributed systems, including scaling them in fast-moving environments.
- Hands-on experience with big data technologies such as Flink, Spark, Kafka, Pulsar, or Pinot.
- Experience developing, maintaining, and debugging distributed systems built with open-source tools.
- Experience building and scaling infrastructure as a product.
- Strong software engineering skills and a passion for big-data distributed systems.
- Ability to write high-quality code in programming languages such as Go, Java, or Scala.
- Comfort working with high autonomy and ownership.
- A growth mindset and willingness to learn quickly, explore ambiguous problem spaces, and dive deep when needed.
- Strong written and verbal communication skills, including the ability to produce clear technical documentation.
Preferred Qualifications
- Experience operating streaming infrastructure as a platform, such as Flink clusters, Kafka, or Pulsar, for internal customers at scale.
- Deep hands-on experience authoring, optimizing, and operating real-time processing frameworks such as Flink, Spark Streaming, Storm, or Kafka Streams in production.
- Experience building or operating control planes for managing large-scale infrastructure.
- Open-source contributions to data processing or big-data systems such as Hadoop, Spark, Celeborn, or Flink.