Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Airflow @ 3
Data Pipelines @ 4
Debugging @ 7
Experimentation
Flink @ 4
Java @ 7
Kafka @ 4
Kubernetes @ 4
Observability @ 7
Parquet @ 4
Protobuf @ 4
Python @ 7
Spark @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Booking.com runs large-scale experiments as a vital part of its software development cycle. The Experimentation Platform enables product teams to make data-driven decisions by safely assigning experiments, collecting tracking data, and providing statistically reliable metrics across user interactions.
This hands-on contractor role focuses on treating operations as a software problem and strengthening the reliability of the Experimentation Data Platform. The engineer will work across services, streaming infrastructure, and data pipelines, with a particular focus on a high-risk platform migration. Production experience with the relevant programming languages, streaming systems, and cloud infrastructure is required.
Responsibilities
- Build software and automation in Java and/or Python to improve availability, scalability, latency, efficiency, and operational safety.
- Write readable, reusable code and guide less-experienced engineers in these practices.
- Design reliable solutions for distributed services and data pipelines, evaluating cost, business requirements, non-functional requirements, technology options, and future extensibility.
- Own systems end to end by monitoring application health and performance, setting service and business metrics, supporting deployment and operations, and maintaining runbooks and operational documentation.
- Lead incident response for team-owned issues, mitigate customer impact within agreed service levels, contribute to postmortems, and deliver long-term fixes through root-cause analysis.
- Reduce toil and operational cost by removing bottlenecks, addressing technical debt, preparing for scale, and automating repetitive operational work.
- Improve monitoring and alerting by reviewing observability metrics, business KPIs, and capacity signals, and partnering with development teams to define useful reliability indicators.
- Advise product and engineering teams on architecture, communicate clearly with stakeholders, challenge assumptions constructively, and coach colleagues on reliability practices.
Requirements
- Approximately 5–8 years of relevant engineering experience or equivalent practical experience. A master's degree or equivalent professional experience is welcome.
- Strong professional programming experience in Java and/or Python, including debugging, testing, refactoring, and safely changing production code. Experience with both languages is preferred.
- Hands-on experience operating distributed production systems or data platforms, including Apache Kafka and Apache Flink.
- Practical experience with Kubernetes and AWS, including EKS, S3, and IAM, or equivalent cloud infrastructure.
- Experience operating streaming pipelines and reliable file-based handoffs, including Parquet on object storage.
- Strong experience with observability, alerting, incident response, root-cause analysis, postmortems, and production documentation.
- Experience designing or supporting high-risk production migrations, including rollout, rollback, reconciliation, or backfill strategies.
- Demonstrated ability to design complex technical solutions, identify underlying issues, improve engineering processes, and communicate decisions clearly.
- Experience coaching, guiding, or enabling less-experienced engineers and partner teams.
Particularly Valuable Experience
- Experience with the Flink Kubernetes Operator or large-scale streaming data pipelines.
- Experience with data freshness, completeness, and correctness monitoring in addition to service availability.
- Familiarity with downstream Spark/Airflow or governed data platforms. The team owns the upstream pipeline and its S3/Parquet handoff rather than BDX itself.
- Experience with capacity planning, resilience testing, performance engineering, or infrastructure-cost optimization.
- Experience with on-premises service sidecars, Protobuf/stateful event processing, or Java/Python SDKs and libraries.
Engagement Outcomes
- During the first month, establish an ordered view of infrastructure needs, reliability risks, and migration dependencies, together with a committed mitigation plan.
- Improve the team's ability to detect, diagnose, and recover from service and pipeline failures.
- Deliver practical automation, observability, and documentation improvements that reduce operational toil and strengthen the migration.
- Become a trusted technical point of contact for the group's infrastructure and operational needs while leaving behind maintainable practices and knowledge.
More jobs at Booking.com
Senior SAP FICO Specialist
Booking.com · Bengaluru, India
EUR 10-38 per hour
Software Engineer II
Booking.com · Amsterdam, Netherlands
EUR 50-100 per hour
Data Analytics Engineer
Booking.com · Amsterdam, Netherlands
EUR 50-103 per hour
Senior Machine Learning Scientist
Booking.com · Amsterdam, Netherlands
EUR 60-120 per hour
Senior Software Engineer (For Independent Contractors)
Booking.com · Amsterdam, Netherlands
EUR 60-120 per hour
Similar jobs
Software Engineer, Product Security Data Platforms
Stripe · United States, New York City, United States, Seattle, United States
USD 158,800-238,200 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior AI and HPC Observability Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Staff Software Engineer, Risk Data Engineering
Stripe · Canada, United States, Toronto, Canada, South San Francisco, United States, Seattle, United States
USD 224,000-336,000 per year
Senior Software Engineer - ClickPipes (CDC/Streaming)
ClickHouse · United States
USD 147,000-238,000 per year
Senior Software Engineer - ClickPipes (CDC/Streaming)
ClickHouse · Canada
USD 147,000-238,000 per year
Senior Staff Machine Learning Systems Engineer, Ads ML Platform
Reddit · United States
USD 292,500-409,500 per year
Senior Software Engineer (AI/ML), Trust
Airbnb · Bengaluru, India
INR 4,500,000-6,500,000 per year