Senior Cloud Infrastructure Engineer
๐ Germany
๐ Spain
๐ France
๐ United Kingdom
๐ Netherlands
๐ Zurich, Switzerland
๐ Munich, Germany
๐ Berlin, Germany
๐ Paris, France
๐ London, United Kingdom
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 โ basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 โ daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 โ you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 โ exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
CI/CD
ClickHouse @ 4
CloudFormation @ 4
Compliance
Datadog @ 1
Docker @ 4
Helm @ 4
IaC
Kubernetes @ 4
LLM
Networking @ 4
Next.js
Observability @ 7
PostgreSQL
Redis
SRE @ 7
Security
Terraform @ 4
TypeScript
- 1-2 โ basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 โ daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 โ you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 โ exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Langfuse is an open-source LLM engineering platform for tracing, evaluation, and prompt management, now part of ClickHouse. The platform processes more than a billion trace events per month and supports thousands of self-hosted deployments.
This role will own uptime, performance, cost efficiency, and reliability across Langfuse Cloud and self-hosted infrastructure. The team operates in European time zones and expects approximately one week per month in the Berlin office.
Responsibilities
- Operate Langfuse Cloud production environments on AWS ECS Fargate and ClickHouse Cloud.
- Manage deployments, autoscaling, capacity planning, and cost optimization.
- Own Datadog dashboards, alerts, and service-level objectives.
- Improve monitoring and observability so that issues are identified before they affect customers.
- Own and evolve the Helm chart, Docker Compose configuration, and deployment documentation for self-hosted deployments.
- Automate CI/CD pipelines, infrastructure provisioning, scaling, and zero-downtime deployments.
- Plan for future scale, including long-running agent observability and real-time evaluation workloads.
- Improve security and compliance across cloud and self-hosted deployments.
- Help users debug self-hosted deployments.
Requirements
- Strong infrastructure or SRE engineering experience operating systems at scale.
- Experience operating production workloads on AWS, including ECS/Fargate, networking, IAM, and S3, or comparable hyperscale cloud platforms.
- Experience with container orchestration, including Kubernetes and/or ECS, Helm charts, and Docker.
- Experience with infrastructure as code using Terraform, Pulumi, CloudFormation, or similar tools.
- Strong monitoring and observability skills, including building effective dashboards and alerts. Datadog experience is a plus.
- Strong focus on reliability, automation, and safely shipping infrastructure changes.
- Interest in open-source software and helping users debug self-hosted deployments.
- Ability to thrive in a small, accountable, engineering-focused team.
- A computer science or quantitative degree is preferred.
Bonus Qualifications
- Experience with ClickHouse Cloud or other managed analytical databases.
- Experience operating high-throughput event-processing or observability infrastructure.
- Contributions to open-source infrastructure tooling, such as Helm charts or Terraform modules.
- Former founder.
Tech Stack
The company uses a TypeScript monorepo with Next.js on the frontend, Express workers for background jobs, PostgreSQL for transactional data, ClickHouse for tracing at scale, S3 for file storage, and Redis for queues and caching. Cloud infrastructure includes AWS ECS Fargate, ClickHouse Cloud, Datadog, Helm, Docker Compose, and infrastructure-as-code tooling.
Process
The full hiring process can be completed through the offer letter in less than seven days.
How We Ship
The team emphasizes ownership, RFC-based solution design, collaboration, maker schedules, mentorship through code reviews, and the use of AI tooling in workflows.