Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 3
Communication @ 6
Customer Support
Debugging @ 3
Grafana @ 3
Kubernetes @ 3
Networking
Observability @ 3
Prometheus @ 3
Puppet
Security
Terraform
VictoriaMetrics @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
Cluster operations & maintenance
- Plan and execute Kubernetes version upgrades across EKS and on-premises clusters, coordinating with internal teams to minimise disruption
- Perform routine maintenance: add-on upgrades, storage and networking configuration, component upgrades for monitoring, security and other tooling
- Act on cluster health issues across the fleet and proactively work on degradation signals before they become incidents
Internal customer support
- Be the first point of contact for engineering teams running workloads on BKS: triaging issues, diagnosing failures, and guiding teams to resolution
- Help teams understand platform capabilities, quota management, and best practices for running reliable workloads. Evaluate quota requests and current usage requirements and weigh them against cluster capabilities.
- Contribute to runbooks and FAQs so common questions are answered before they become support requests
Toil reduction & automation
- Identify repetitive manual tasks and reduce them through scripting and automation
- Flag and address technical debt that slows down operations or increases risk
- Partner with the wider BKS team on tooling improvements that reduce the operational burden across the fleet
Requirements
- Kubernetes operational experience: node management, upgrades, debugging workloads, cluster health
- Experience with managed Kubernetes (EKS or equivalent) or on-premises cluster operations
- Comfortable with observability tooling: Prometheus/VictoriaMetrics, Grafana, alerting pipelines
- Strong written and verbal communication: clear in tickets, runbooks, and support threads
- Able to work independently, prioritise effectively, and keep commitments
Nice to have
- Terraform for infrastructure provisioning
- Puppet or similar configuration management for on-premises environments
- AWS experience
- Experience supporting internal developer platforms or infrastructure teams
What we expect
- Think Customer First: the engineering teams relying on BKS are your customers; their problems are your problems
- Own It: take issues through to resolution, raise blockers early, and follow through without needing to be chased
- Succeed Together: share knowledge, document what you learn, and make the team stronger for it
- Learn Forever: whether it’s a new technology or a Booking-specific tool, knowing how to tackle something unknown by asking the right questions
- Do The Right Thing: when something looks risky, an unsafe upgrade path, a coverage gap, a missed dependency, say so and help address it
What this role is not
- A pure software development role, the focus is operational excellence and customer support, not feature building
- A solo heroics role, we escalate, pair, and work in the open
- A reactive-only role, proactive maintenance and toil reduction are just as important as incident response
Similar jobs
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Site Reliability Engineer - Us
Teleport · United States, Oakland, United States
USD 222,000-342,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer - Cloud Infrastructure
ClickHouse · United States
USD 141,000-208,000 per year
Senior Field Engineer
Grafana Labs · Netherlands
EUR 96,000-116,000 per year
Senior Devops Engineer, GeForce Now
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year