Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
ClickHouse @ 6
Debugging
Docker @ 4
Go @ 3
Kubernetes @ 4
LLM @ 6
Node.js @ 7
Observability @ 4
OpenTelemetry @ 4
Python @ 6
Rust @ 3
SQL @ 6
SRE @ 4
Scoping @ 4
TypeScript @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
ClickHouse is hiring an AI Product Engineer to build agentic capabilities on ClickStack, its open-source observability platform for logs, metrics, traces, and session replays. The role focuses on developer experience and building production-ready AI agents that investigate incidents, identify root causes, and summarize findings using a petabyte-scale observability platform.
Responsibilities
- Build agents that investigate incidents, surface anomalies, answer why production is broken, and use ClickStack as their substrate.
- Build reusable skills that encode debugging, root-cause analysis, ClickHouse query writing, and incident-response workflows.
- Own the agent stack end to end, including context engineering, tool design, evaluations, tracing, and cost management.
- Build MCP servers, SDKs, and integrations that allow customer agents to read telemetry, take action, and remain observable.
- Collaborate with open-source contributors and customers, debug problems, and feed learnings back into the product.
- Address production challenges including latency, cost, context-window limits, evaluation coverage, and hallucinations on real telemetry.
Requirements
- 5+ years of software engineering experience, including 1–2 years working on LLM-powered systems or agents in production.
- Strong backend skills in TypeScript/Node.js and/or Python; comfort working in both languages.
- Hands-on experience building and shipping agents with multi-step tool use, planning, memory, and error recovery.
- Experience designing skills using Markdown-based workflow encodings, Anthropic-style approaches, or similar methods.
- Experience building MCP servers and designing tools with authentication, scoping, and observability considerations.
- Strong evaluation practices, including golden sets, LLM-as-judge, and regression detection.
- SQL proficiency, including the ability to write ClickHouse queries directly.
- Experience with Docker and Kubernetes.
- Active participation in open source and the developer community.
- Ability to work with ambiguity, take ownership, move quickly, and learn from failures.
- Interest in developer tools and a strong understanding of good developer experience.
Bonus Qualifications
- Experience building or operating production agents for observability, incident response, or SRE.
- Experience with agent observability, including tracing, cost attribution, evaluation pipelines, or OpenTelemetry for agents.
- Experience with prompt caching, context compaction, or techniques for running agents on production telemetry volumes.
- Experience with columnar databases and event-ingestion pipelines.
- Contributions to or maintenance of an open-source AI or agent project.
- Familiarity with Go, Rust, or other systems languages for integrations and high-throughput infrastructure.
Compensation
The typical starting salary in the United States is $130,000–$208,000 USD. In US premium markets, including the San Francisco Bay Area and New York City Metro Area, the typical starting salary range is $141,000–$230,000 USD. Actual compensation depends on factors including education, qualifications, certifications, experience, skills, location, performance, and business needs.
Benefits
- Flexible, remote-friendly work environment.
- Employer healthcare contributions.
- Stock options for new team members.
- Flexible time off in the United States and generous entitlement in other countries.
- $500 home-office setup allowance for remote employees.
- Opportunities to participate in company-wide global gatherings.