Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 7
Ansible @ 7
Azure @ 7
ClickHouse @ 1
Cloud Computing @ 7
Communication @ 6
Debugging @ 7
Distributed Systems
Docker @ 4
Go @ 4
Kubernetes @ 4
Puppet @ 7
Python @ 4
SQL @ 1
SRE
Security
Terraform @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
ClickHouse is expanding its central Site Reliability Engineering team to provide reliable and secure services for ClickHouse Cloud. The role is responsible for building and leading processes that ensure the reliability, availability, scalability, and performance of cloud infrastructure. You will collaborate with the Control Plane, Data Plane, Core, Security, Support, and Operations teams to design and implement scalable, secure, highly available, and fault-tolerant distributed systems.
You will own incident management and response, blameless post-mortem analysis, and continuous improvement of ClickHouse Cloud services. You will also leverage software engineering expertise to develop platforms and tools that improve the operational and engineering efficiency of ClickHouse Cloud.
Responsibilities
- Collaborate with engineering teams to design and implement scalable, secure, and highly available systems.
- Establish and manage service level objectives (SLOs) and service level agreements (SLAs) for ClickHouse Cloud.
- Ensure infrastructure components, including the Data Plane, Control Plane, and ClickHouse Core, have monitoring and alerting for timely incident detection and resolution.
- Improve incident response and post-mortem processes for ClickHouse Cloud outages, including communication with impacted customers in collaboration with the Support team.
- Continuously improve the reliability and performance of ClickHouse services.
- Plan, enable, and drive chaos initiatives across Engineering teams based on internal priorities.
- Manage on-call processes and establish best practices for escalation and issue resolution to minimize downtime.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- At least 8 years of experience in Site Reliability Engineering or a related field.
- Hands-on experience with Go and/or Python.
- Strong knowledge of cloud computing platforms such as AWS, Azure, or Google Cloud Platform.
- Excellent understanding of distributed databases and SQL; ClickHouse experience is a major plus.
- Hands-on experience with container orchestration tools such as Kubernetes or Docker Swarm.
- Strong experience with automation and configuration management tools such as Ansible, Terraform, or Puppet.
- Strong problem-solving and production debugging skills.
- Passion for efficiency, availability, scalability, and data governance.
- Ability to work effectively in a fast-paced environment and partner with the business.
- High level of responsibility, ownership, and accountability.
- Excellent communication and interpersonal skills.
Benefits
- Flexible work environment at a globally distributed, remote-friendly company operating in more than 20 countries.
- Employer contributions toward healthcare.
- Stock options for every new team member.
- Flexible time off in the United States and generous entitlement in other countries.
- $500 home office setup allowance for remote employees.
- Opportunities to participate in company-wide offsites and global gatherings.
ClickHouse provides equal employment opportunities and prohibits discrimination and harassment based on protected characteristics under applicable laws.