Staff Software Engineer
📍 Toronto, Canada
📍 New York City, United States
📍 South San Francisco, United States
📍 Seattle, United States
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Distributed Systems @ 7
Payments
Scoping @ 6
Software Development @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Who We Are
About Stripe
Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Stripe’s mission is to increase the GDP of the internet.
About the Team
You will join the High Availability and Disaster Recovery team. Availability is a core feature of Stripe’s products, and the team designs and builds solutions that enable latency-critical, stateful applications to survive any type of disaster. The team builds distributed systems on top of unreliable architecture to provide highly available and resilient customer solutions.
This is a distributed team with many remote engineers. Candidates who meet the minimum requirements are encouraged to apply if they can work from anywhere in the United States or Canada.
Responsibilities
- Work with leaders and engineers across the company to identify opportunities to improve Stripe’s reliability posture.
- Scope, design, implement, and deploy robust distributed services, balancing reliability, throughput, latency, resiliency, engineering velocity, and cost.
- Design and implement new products and prototypes to improve service resiliency, engineering velocity, and management at scale.
- Mentor and develop the next generation of technical leaders at Stripe.
- Contribute to engineering strategy, tooling, processes, and culture.
- Uphold high engineering standards and improve the codebase and engineering processes.
- Develop global architecture by combining less-available components and data centers into a highly available and resilient whole.
- Work on latency-critical solutions where every millisecond matters and data redundancy is a hard requirement.
- Investigate Mongo write concerns, minimize cross-region TLS handshakes, and develop systems to automate disaster detection and failovers.
Requirements
- At least 10 years of distributed and cloud-native software development experience at a highly scaled service or hyperscaler.
- Deep understanding of scalable, resilient, distributed systems, cloud-native architectures, and mission-critical systems.
- Experience influencing, planning, scoping, and leading large projects across multiple teams.
- Ability to thrive in a collaborative environment involving diverse stakeholders and subject-matter experts.
- High autonomy and responsibility, with an entrepreneurial, proactive, and self-driven approach.
- Strong history of serving as a technical lead and mentor across several engineering teams.