Senior Site Reliability Engineer - Remote

USD 141,000-230,000 per year
SENIOR
✅ Remote

Tech Stack

AWS @ 7 Ansible @ 7 Azure @ 7 ClickHouse @ 1 Cloud Computing @ 7 Communication @ 6 Debugging @ 7 Distributed Systems Docker @ 4 Go @ 4 Kubernetes @ 4 Puppet @ 7 Python @ 4 SQL @ 1 SRE Security Terraform @ 7

Details

ClickHouse is expanding its central Site Reliability Engineering team to provide reliable and secure services for ClickHouse Cloud. The role is responsible for building and leading processes that ensure the reliability, availability, scalability, and performance of cloud infrastructure. You will collaborate with the Control Plane, Data Plane, Core, Security, Support, and Operations teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems.

You will own incident management and response, blameless post-mortem analysis, and continuous improvement of ClickHouse Cloud services. You will also leverage software engineering expertise to develop platforms and tools that improve the operational and engineering efficiency of ClickHouse Cloud.

Responsibilities

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems.
  • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  • Ensure infrastructure components, including the Data Plane, Control Plane, and ClickHouse Core, have monitoring and alerting for timely incident detection and resolution.
  • Improve incident response and post-mortem processes for ClickHouse Cloud outages, including communication with impacted customers in collaboration with the Support team.
  • Continuously improve the reliability and performance of ClickHouse services.
  • Plan, enable, and drive chaos initiatives across Engineering teams based on internal priorities.
  • Manage on-call processes and establish best practices for escalation and issue resolution to minimize downtime.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL; ClickHouse experience is a major plus.
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Strong problem-solving and production debugging skills.
  • Passion for efficiency, availability, scalability, and data governance.
  • Ability to work effectively in a fast-paced environment and partner with the business.
  • High level of responsibility, ownership, and accountability.
  • Excellent communication and interpersonal skills.

Benefits

  • Flexible work environment at a globally distributed, remote-friendly company operating in more than 20 countries.
  • Employer contributions toward healthcare.
  • Stock options for every new team member.
  • Flexible time off in the United States and generous entitlement in other countries.
  • $500 home office setup allowance for remote employees.
  • Opportunities to participate in company-wide offsites and global gatherings.

ClickHouse provides equal employment opportunities and prohibits discrimination and harassment based on protected characteristics under applicable laws.

More jobs at ClickHouse

Similar jobs