Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS
Communication @ 3
Design Patterns
Experimentation
Kubernetes
Observability @ 3
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Site Reliability Engineer II specialists treat operations as a software problem, focusing on the availability, performance, scalability, latency, observability, and efficiency of systems and services. The role aims to reduce operational toil and complexity through automation and improve system reliability.
SRE II engineers implement technical solutions based on business requirements, estimate effort and impact, deliver high-quality work, collaborate with partner teams, and participate in incident response. Depending on the business area, they may act as part of a business service owner team, infrastructure owner, or consultant to product development teams.
Responsibilities
Building Software Applications
- Build software applications using relevant development languages and business-area systems, services, and tools.
- Refactor and simplify code using design patterns where appropriate.
- Ensure application quality by following testing techniques and the test strategy.
- Write readable and reusable code using standard patterns and libraries.
- Maintain data security, integrity, and quality according to company standards and best practices.
Software Systems Design
- Evaluate architecture solutions considering cost, business requirements, technology requirements, and emerging technologies.
- Understand and describe the implications of changing or adding systems within the infrastructure and architecture.
- Apply engineering techniques such as prototyping, spiking, and vendor evaluation.
- Design adaptable solutions that meet current and future business requirements.
End-to-End System Ownership
- Own services end to end by monitoring application health and performance and acting on relevant metrics.
- Reduce business continuity risks and bus factor through appropriate practices, tools, documentation, runbooks, and OpDocs.
- Use continuous delivery and experimentation frameworks to reduce risk and obtain customer feedback.
- Independently manage applications or services through deployment and production operations.
- Maintain data security, integrity, and quality.
Incident Management, Automation, and Observability
- Resolve live production issues and mitigate customer impact within SLA.
- Improve system reliability through root cause analysis, long-term solutions, postmortems, and incident tracking.
- Reduce technical debt, identify bottlenecks, and prepare infrastructure for scaling.
- Reduce operational costs and human labor through new technologies, automation, and software features addressing availability, scalability, latency, and efficiency.
- Monitor production systems and network infrastructure using observability metrics, business KPIs, and capacity planning.
- Partner with development teams to establish appropriate observability metrics.
Additional Responsibilities
- Apply critical thinking to identify underlying issues and develop logical solutions.
- Identify and implement continuous quality, process, system, and structural improvements.
- Communicate clearly with different audiences and use active listening to reach mutually agreeable solutions.
- Advise product teams on technical solutions meeting functional, nonfunctional, and architectural requirements.
- Help define technical capability direction by evaluating target architecture improvements and aligning architectural decisions.
Requirements
Knowledge and skills in:
- Building software applications
- Software system design
- End-to-end system ownership
- Technical incident management
- Operations, automation, and toil reduction
- Observability, monitoring, and alerting
- Critical thinking
- Continuous quality and process improvement
- Effective communication
- Architectural guidance
Tech Stack
- Kubernetes
- AWS
- Any coding language
Contract Details
- Independent contractor engagement
- Project duration: 3 months
- Start date: September 21, 2026
- End date: December 20, 2026
More jobs at Booking.com
Similar jobs
Machine Learning Engineer - Energy Management
Eneco · Rotterdam, Netherlands
EUR 84,000-117,000 per year
DevOps Engineer · Mid-Senior · Infra
Nord Security · Vilnius, Lithuania, Kaunas, Lithuania
EUR 43,200-88,800 per year
Senior Platform Engineer, GitLab Orbit
GitLab · Canada, United States
USD 139,200-235,200 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff Site Reliability Engineer - AI Platform Runtime
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Site Reliability Engineer - US
Teleport · United States
USD 222,000-326,000 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Germany, Spain, United Kingdom, Sweden
EUR 109,700-131,700 per year