Staff Engineer, Release Engineering

at Stripe
USD 224,000-336,000 per year
SENIOR
✅ Hybrid

Tech Stack

API AWS @ 4 Audit @ 4 Azure @ 4 Change Management @ 4 Distributed Systems @ 6 IaC Kafka @ 3 Kubernetes @ 4 Observability @ 3 Security Terraform @ 4

Details

Stripe is a financial infrastructure platform for businesses. The Core Change Management group owns the systems that allow Stripe engineers to ship code, configuration, and infrastructure changes safely and at high velocity.

The role is embedded primarily on the Service Deployments team, which owns Stripe's end-to-end code deployment platform, with regular collaboration with the Resource Automation and Feature Deployments teams. Service Deployments operates systems in the critical path of every engineer's daily workflow and maintains a meaningful on-call rotation.

The role covers deployment orchestration, container scheduling, developer-facing internal platforms, reliability, developer experience, performance, and deployment safety. Current platform initiatives include containerizing host-based services at scale, building intelligent multi-service deployment pipelines, extending real-time anomaly detection to earlier traffic-shift stages, and rebuilding deployment event infrastructure on a durable message bus.

Responsibilities

  • Own end-to-end technical delivery of large, ambiguous infrastructure projects, from initial design through production launch and long-term reliability.
  • Architect the next generation of Stripe's deployment platform, including multi-service dependency-aware autodeploy pipelines, Kubernetes-native deployment primitives, and fleetwide container migration.
  • Define API contracts, rollout strategies, and operational models used by hundreds of teams.
  • Extend deploy anomaly detection by expanding blue-green traffic analysis, designing API- and method-based regression detection, and building self-service onboarding for supported service types.
  • Lead the migration of host-based services to containerized, fleetwide deployments while maintaining production operations.
  • Own reliability and operational excellence for the deployment platform, including incident response, toil reduction, security, and maintainability.
  • Build deployment event infrastructure, including event schemas, durability models, and integration contracts for downstream observability and automation systems.
  • Collaborate with Resource Automation on cloud resource management, IAM, account provisioning, and infrastructure automation.
  • Collaborate with Feature Deployments on feature flags, configuration management, and change audit logs.
  • Partner with the service mesh team on routing capabilities for canary rollouts and merchant-priority traffic shaping.
  • Lead critical design reviews, establish deployment safety and developer experience standards, mentor senior engineers, and advocate for effective abstractions.
  • Break down complex platform challenges into scoped, parallelizable work and provide technical guidance on difficult engineering decisions.

Requirements

Minimum Requirements

  • 10+ years of professional software engineering experience.
  • Demonstrated experience designing and shipping production infrastructure systems of significant scale and complexity.
  • Proven ability to lead large, ambiguous infrastructure projects from technical design through delivery, including cross-team dependencies and migrations across many consuming teams.
  • Deep expertise in distributed systems and deployment orchestration, including service construction, scheduling, operation at scale, rollout strategies, staged delivery, and failure modes.
  • Hands-on experience with Kubernetes and container-based deployments, including service lifecycle management, workload scheduling, and migration of large fleets from VM-based to containerized infrastructure.
  • Strong background in service reliability and operational excellence, including incident response, toil reduction, and building reliable, debuggable, and maintainable systems.
  • Broad technical impact across multiple large systems, including fluency across complex codebases, effective code review and mentorship, and the ability to set technical direction.

Preferred Requirements

  • Experience with deployment safety systems such as anomaly detection, automated rollback, or progressive delivery.
  • Familiarity with event-driven architectures such as Kafka for deployment lifecycle observability and notification.
  • Experience with Infrastructure as Code at scale, including Terraform or equivalent, cloud resource governance, and IAM management in AWS or Azure.
  • Experience building developer platforms or internal tooling with a strong focus on developer experience and toil reduction.
  • Experience with change management and feature rollout systems, including feature flags, configuration distribution, or audit-log infrastructure.
  • Familiarity with service mesh concepts such as canary deployments, weighted routing, and traffic splitting.

Office Expectations

Office-assigned employees in most locations are expected to spend at least 50% of each month in their local office or with users. This expectation may vary by role, team, and location. Some teams may have greater in-office attendance requirements.

More jobs at Stripe

Similar jobs