Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 7
Azure @ 7
Compliance
Debugging @ 6
Kubernetes @ 6
Linux @ 4
Networking @ 6
Observability
Payments
Security @ 4
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Who We Are
About Stripe
Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises to the most ambitious startups—use Stripe to accept payments, grow their revenue, and accelerate new business opportunities. Stripe’s mission is to increase the GDP of the internet.
About the Team
The Compute Foundation group builds and operates the core systems that enable every Stripe engineer to ship code, configuration, and infrastructure changes safely and at high velocity. Embedded on the Imaging team, you will design, build, and distribute Linux machine images; manage host bootstrapping and OS-configuration systems; oversee OS packages and security-patching infrastructure; and provide platforms that empower service teams to safely run workloads across cloud providers and regions.
What You’ll Do
- Own the foundation of Stripe’s compute fleet, including host infrastructure that boots reliably, receives accurate configurations, and operates safely.
- Lead architectural transformations, including modernizing the OS configuration layer, containerizing workloads, scaling machine image distribution across multi-cloud regions, and streamlining OS upgrades and critical security patching.
- Drive end-to-end host lifecycle ownership across image construction, package distribution, kernel and OS configuration, early boot/systemd, storage, networking initialization, and workload readiness.
- Design validation, canarying, rollout, health-evaluation, and rollback mechanisms that contain blast radius and make foundational infrastructure changes safer.
- Set technical direction for the team’s systems, author designs spanning multiple teams, mentor engineers, and provide technical leadership to engineering managers and engineers.
Responsibilities
- Lead large, ambiguous infrastructure projects from initial design through production launch, broad adoption, and long-term operation.
- Define architecture and success criteria, sequence work, resolve technical ambiguity, coordinate cross-team dependencies, and shepherd projects to measurable impact.
- Set technical direction for Stripe’s host infrastructure, including how machine images are built, validated, distributed, and operated.
- Guide decisions about when to containerize workloads and when host-based approaches remain appropriate.
- Establish validation, staged rollout, observability, and rollback capabilities for broad infrastructure changes.
- Build guardrails that prevent unhealthy changes from progressing and provide operators with clear, actionable information when failures occur.
- Lead incident response for the systems owned by the team and reduce operational toil.
- Make reliability, security, debuggability, and maintainability first-class properties of the platform.
- Partner with technical leaders across infrastructure, security, compliance, and service teams.
- Lead critical design reviews and establish standards for the safety, reliability, operability, and usability of Stripe’s host infrastructure.
- Mentor senior engineers through high-stakes architectural decisions.
- Improve designs and implementations through hands-on coding, code review, debugging, and operational leadership.
Requirements
Minimum Requirements
- 10+ years of professional software engineering experience, with a demonstrated track record of designing and shipping production infrastructure systems of significant scale and complexity.
- Proven ability to lead large, ambiguous infrastructure projects end-to-end, from technical design through delivery, including managing cross-team dependencies and coordinating migrations across many consuming teams.
- Deep expertise in Linux systems, including boot and initialization, systems, processes, networking, filesystems, storage, packages, and production debugging.
- Strong foundations in cloud compute infrastructure, including machine images, virtual machines, autoscaling fleets, launch configuration, block storage, IAM, health evaluation, and multi-region operation in AWS, Azure, or an equivalent environment.
- Experience designing or operating fleet-scale image, host-management, configuration-management, or infrastructure-provisioning systems.
- Experience with safe infrastructure rollout practices, including automated validation, canarying, progressive rollout, health evaluation, rollback, and blast-radius containment.
- Strong background in service reliability and operational excellence, including incident response, toil reduction, and building reliable, debuggable, and maintainable systems.
- Track record of broad technical impact across multiple large systems, including fluency across a complex codebase, code review and mentorship, and the ability to set technical direction for a team.
Preferred Qualifications
- Experience building modern image-based, immutable, or declarative infrastructure platforms.
- Expertise with containers and Kubernetes, including migrating VM-based workloads to managed platforms.
- Experience operating large-scale Linux fleets across AWS, Azure, or multiple regions.
- Familiarity with progressive delivery, automated validation, fleet health signals, and rollback systems.
- Experience modernizing legacy configuration-management or host-provisioning systems.
- Knowledge of OS security patching, software supply-chain security, and artifact provenance.
- Strong developer-platform instincts, with experience building safe abstractions and self-service workflows.