Cloud Software Engineer - Observability Platform

USD 141,000-230,000 per year
SENIOR
✅ Remote

Tech Stack

AWS @ 4 Azure @ 4 ClickHouse @ 4 Communication @ 6 Debugging @ 6 Distributed Systems @ 6 GCP @ 4 Go @ 7 Grafana @ 4 Helm @ 4 Kubernetes @ 4 Observability @ 4 OpenTelemetry @ 4 Prometheus @ 4 Terraform @ 4 TypeScript @ 4

Details

ClickHouse is looking for an experienced software engineer to join its Observability organization, across the Observability Platform and Internal Observability teams. These teams build and operate the systems behind ClickHouse’s internal observability and ClickStack Cloud, processing trillions of events per day at throughput of hundreds of millions of events per second.

This role sits at the intersection of distributed systems, cloud infrastructure, and production operations. You will build systems, operate what you build, respond to incidents, and turn recurring operational problems into durable software and automation.

Responsibilities

  • Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
  • Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
  • Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.
  • Build software and automation that eliminate repetitive operational work and make the platform easier to operate.
  • Identify architectural bottlenecks and help shape the roadmap for the next stage of scale.
  • Work closely with product, infrastructure, and service teams across ClickHouse.
  • Contribute to architecture and design reviews and help raise engineering quality across the team.

Requirements

  • 5+ years of experience building and operating production systems at scale.
  • Strong proficiency in Go.
  • Experience building and operating services on Kubernetes.
  • Experience with infrastructure-as-code and GitOps tooling such as Terraform, Helm, and Argo CD.
  • Production experience with at least one major cloud provider: AWS, GCP, or Azure.
  • Hands-on experience with telemetry systems such as OpenTelemetry, Prometheus, Grafana, or comparable technologies.
  • Ability to solve ambiguous production problems, make sound engineering tradeoffs, and take ownership from design through operation.
  • Comfortable debugging unfamiliar distributed systems in production.
  • Ability to make pragmatic tradeoffs between reliability, performance, delivery speed, and cost.
  • Clear communication in a remote, async-friendly environment.

Bonus Points

  • Experience with ClickHouse.
  • Experience with high-throughput ingestion, streaming, queueing, or storage systems.
  • Experience building multi-tenant cloud services.
  • Experience optimizing infrastructure for both performance and cost.
  • Experience with TypeScript.

Compensation

The typical starting salary for this role in the United States is $141,000–$220,000 USD. In US premium markets, such as the San Francisco Bay Area and the New York City Metro Area, the typical starting salary range is $185,000–$230,000 USD. Actual compensation depends on factors including education, qualifications, certifications, experience, skills, location, performance, and business needs.

Perks

  • Flexible work environment at a globally distributed, remote-friendly company.
  • Employer contributions toward healthcare.
  • Stock options for every new team member.
  • Flexible time off in the United States and generous entitlement in other countries.
  • $500 home office setup for remote employees.
  • Opportunities to engage with colleagues at company-wide offsites.

More jobs at ClickHouse

Similar jobs