Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
AWS @ 3
CI/CD
Communication @ 6
GCP @ 3
Go @ 3
Hiring @ 3
IaC
Kubernetes @ 3
Networking
Observability @ 2
Ruby @ 3
SRE
System Architecture
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve production infrastructure.
This is a single application for Site Reliability Engineering opportunities across GitLab's Infrastructure Platforms department. Candidates are evaluated holistically and matched with opportunities from Intermediate through Senior Staff across multiple teams.
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient.
- Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows.
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling.
- Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps.
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately.
- Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early.
- Participate in incident response and post-incident reviews, turning learnings into improvements in automation and processes.
- Document runbooks, architecture decisions, and reviews so findings become repeatable practices.
Requirements
- Experience keeping production systems reliable, combining an operations mindset with software engineering practice.
- Experience building new infrastructure tooling and automation, such as Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratch.
- Ability to read, debug, and reason about code, including its behavior, performance, and failure modes. Most teams work in Go; some work in Ruby.
- Experience with infrastructure as code, Kubernetes, and its ecosystem at a depth appropriate to the candidate's level.
- Hands-on experience with at least one major cloud provider, such as GCP or AWS.
- Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisions.
- Comfort participating in on-call and incident response, with a structured approach to troubleshooting under pressure.
- Strong written communication and the ability to operate as a manager-of-one in an asynchronous, distributed environment.
- A track record of using automation and increasingly AI to reduce toil and improve team workflows.
- Alignment with GitLab's values.
Hiring Process
- Recruiter screen.
- Core technical assessment covering source code, system architecture, and incident review.
- Peer technical interview focused on the prospective team's work.
- Hiring manager interview covering ownership, judgment, execution, collaboration, and growth.
- Skip-level interview covering values alignment and cross-team collaboration.
Final level and team placement are determined based on interview performance, experience, and current hiring priorities.
About the Team
Infrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab's user-facing services, including GitLab.com. The department includes Production Engineering and Dedicated, with teams responsible for the production fleet, networking platform, observability, incident response, and GitLab's single-tenant Dedicated offering. The team is globally distributed, fully remote, asynchronous, and focused on automation, monitoring, and metrics.
Benefits
- Benefits supporting health, finances, and well-being.
- Flexible paid time off.
- Team member resource groups.
- Equity compensation and employee stock purchase plan.
- Growth and development fund.
- Parental leave.