Senior Site Reliability Engineer (For Independent Contractors)

EUR 60-135 per hour
SENIOR
✅ Hybrid ✅ On-site
✅ Contract / Freelance

Tech Stack

AWS @ 4 Airflow @ 3 Data Pipelines @ 4 Debugging @ 7 Experimentation Flink @ 4 Java @ 7 Kafka @ 4 Kubernetes @ 4 Observability @ 7 Parquet @ 4 Protobuf @ 4 Python @ 7 Spark @ 3

Details

Booking.com runs large-scale experiments as a vital part of its software development cycle. The Experimentation Platform enables product teams to make data-driven decisions by safely assigning experiments, collecting tracking data, and providing statistically reliable metrics across user interactions.

This hands-on contractor role focuses on treating operations as a software problem and strengthening the reliability of the Experimentation Data Platform. The engineer will work across services, streaming infrastructure, and data pipelines, with a particular focus on a high-risk platform migration. Production experience with the relevant programming languages, streaming systems, and cloud infrastructure is required.

Responsibilities

  • Build software and automation in Java and/or Python to improve availability, scalability, latency, efficiency, and operational safety.
  • Write readable, reusable code and guide less-experienced engineers in these practices.
  • Design reliable solutions for distributed services and data pipelines, evaluating cost, business requirements, non-functional requirements, technology options, and future extensibility.
  • Own systems end to end by monitoring application health and performance, setting service and business metrics, supporting deployment and operations, and maintaining runbooks and operational documentation.
  • Lead incident response for team-owned issues, mitigate customer impact within agreed service levels, contribute to postmortems, and deliver long-term fixes through root-cause analysis.
  • Reduce toil and operational cost by removing bottlenecks, addressing technical debt, preparing for scale, and automating repetitive operational work.
  • Improve monitoring and alerting by reviewing observability metrics, business KPIs, and capacity signals, and partnering with development teams to define useful reliability indicators.
  • Advise product and engineering teams on architecture, communicate clearly with stakeholders, challenge assumptions constructively, and coach colleagues on reliability practices.

Requirements

  • Approximately 5–8 years of relevant engineering experience or equivalent practical experience. A master's degree or equivalent professional experience is welcome.
  • Strong professional programming experience in Java and/or Python, including debugging, testing, refactoring, and safely changing production code. Experience with both languages is preferred.
  • Hands-on experience operating distributed production systems or data platforms, including Apache Kafka and Apache Flink.
  • Practical experience with Kubernetes and AWS, including EKS, S3, and IAM, or equivalent cloud infrastructure.
  • Experience operating streaming pipelines and reliable file-based handoffs, including Parquet on object storage.
  • Strong experience with observability, alerting, incident response, root-cause analysis, postmortems, and production documentation.
  • Experience designing or supporting high-risk production migrations, including rollout, rollback, reconciliation, or backfill strategies.
  • Demonstrated ability to design complex technical solutions, identify underlying issues, improve engineering processes, and communicate decisions clearly.
  • Experience coaching, guiding, or enabling less-experienced engineers and partner teams.

Particularly Valuable Experience

  • Experience with the Flink Kubernetes Operator or large-scale streaming data pipelines.
  • Experience with data freshness, completeness, and correctness monitoring in addition to service availability.
  • Familiarity with downstream Spark/Airflow or governed data platforms. The team owns the upstream pipeline and its S3/Parquet handoff rather than BDX itself.
  • Experience with capacity planning, resilience testing, performance engineering, or infrastructure-cost optimization.
  • Experience with on-premises service sidecars, Protobuf/stateful event processing, or Java/Python SDKs and libraries.

Engagement Outcomes

  • During the first month, establish an ordered view of infrastructure needs, reliability risks, and migration dependencies, together with a committed mitigation plan.
  • Improve the team's ability to detect, diagnose, and recover from service and pipeline failures.
  • Deliver practical automation, observability, and documentation improvements that reduce operational toil and strengthen the migration.
  • Become a trusted technical point of contact for the group's infrastructure and operational needs while leaving behind maintainable practices and knowledge.

More jobs at Booking.com

Similar jobs