Senior Site Reliability Engineer

USD 150,000-230,000 per year
SENIOR
✅ Remote

Tech Stack

AWS @ 7 Ansible @ 7 Azure @ 7 ClickHouse @ 4 Cloud Computing @ 7 Communication @ 6 Debugging @ 7 Distributed Systems Docker @ 4 Go @ 4 Kubernetes @ 4 Puppet @ 7 Python @ 4 SQL @ 1 SRE Security Terraform @ 7

Details

We are committed to providing customers with reliable and secure services and are building a newly formed Site Reliability Engineering team. As one of the first members of the Reliability Engineering Team at ClickHouse, you will build and lead processes that ensure the reliability, availability, scalability, and performance of the cloud infrastructure running ClickHouse databases.

You will collaborate with the Control Plane, Dataplane, Core, Security, Support, and Operations teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems. You will own incident management and response, post-mortem analysis, blameless postmortems, and continuous improvement of ClickHouse services. You will also use your software engineering expertise to develop platforms and tools that improve the operational and engineering efficiency of ClickHouse Cloud.

Responsibilities

  • Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
  • Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
  • Ensure infrastructure components across ClickHouse Cloud, including the Dataplane, Control Plane, and ClickHouse Core, have monitoring and alerting for timely incident detection and resolution.
  • Improve incident response and post-mortem processes for outages, including communicating with impacted customers in collaboration with the Support team.
  • Continuously improve the reliability and performance of ClickHouse services.
  • Plan, enable, and drive chaos initiatives across Engineering teams based on internal priorities.
  • Manage on-call processes for performance and reliability issues.
  • Establish best practices for coordinating escalations, resolving issues, and minimizing downtime.

Requirements

  • Bachelor's or Master's degree in Computer Science or a related field.
  • At least 8 years of experience in Site Reliability Engineering or a related field.
  • Previous production experience using ClickHouse.
  • Hands-on experience with Go and/or Python.
  • Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
  • Excellent understanding of distributed databases and SQL; ClickHouse experience is a major plus.
  • Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
  • Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
  • Strong problem-solving and production-debugging skills.
  • Passion for efficiency, availability, scalability, and data governance.
  • Ability to thrive in a fast-paced environment and partner with the business to move it forward.
  • High level of responsibility, ownership, and accountability.
  • Excellent communication and interpersonal skills.

Benefits

  • Flexible work environment at a globally distributed, remote-friendly company operating in more than 25 countries.
  • Employer contributions toward healthcare.
  • Stock options for every new team member.
  • Flexible time off in the United States and generous entitlement in other countries.
  • USD 500 home office setup allowance for remote employees.
  • Opportunities to engage with colleagues at company-wide global gatherings.

ClickHouse provides equal employment opportunities and prohibits discrimination and harassment based on protected characteristics under applicable law.

More jobs at ClickHouse

Similar jobs