Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 7
Algorithms
Azure @ 7
Compliance
Distributed Systems @ 6
Docker @ 4
IaC
Kubernetes @ 4
Networking @ 4
Performance Optimization
SRE @ 4
Security @ 4
Software Development @ 6
Technical Leadership
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Glean is seeking a Site Reliability Engineering Lead to foster engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. The role is responsible for ensuring services meet stringent Service Level Objectives (SLOs), building resilient and automated production environments in the cloud, and providing technical leadership for globally used products.
The SRE team builds infrastructure to scale operations in a hybrid cloud environment and eliminates manual work through automation. The role involves managing complex scale and growth challenges while applying expertise in coding, algorithms, problem-solving, and site reliability engineering practices.
Responsibilities
- Provide technical leadership and mentorship across engineering teams.
- Establish best practices for incident management, performance optimization, and automation.
- Drive cross-team collaboration and contribute to key engineering objectives.
- Shape architectural decisions and ensure the delivery of reliable, high-quality systems.
- Implement and maintain resilient cloud architectures.
- Monitor system performance and proactively identify and resolve bottlenecks and points of failure.
- Participate in the primary on-call rotation.
- Foster a blameless postmortem culture and continuously improve the on-call process.
- Develop and maintain automation scripts, tools, and processes for system deployment, monitoring, and management.
- Optimize cloud infrastructure and applications for performance, scalability, and cost-effectiveness.
- Collaborate with security engineers to implement best practices and ensure compliance with security standards and policies.
- Design and configure monitoring and alerting systems, dashboards, and production on-call playbooks.
- Participate in system design and launch reviews, providing SRE guidance throughout the software development lifecycle.
Requirements
- Bachelor's degree in Computer Science, a related field, or equivalent practical experience.
- 8+ years of experience in a senior-level Site Reliability Engineering or similar role, particularly managing cloud-based services and infrastructure.
- 5+ years of software development experience in one or more programming languages.
- 3+ years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems running in the cloud.
- Strong knowledge of cloud platforms such as Google Cloud Platform, AWS, or Azure.
- Practical experience with Docker and Kubernetes.
- Essential familiarity with infrastructure as code tools such as Terraform.
- Solid understanding of networking, security principles, and SRE and security best practices.
- Proficiency with monitoring and alerting tools.
Location
- Hybrid role requiring four days per week in the Mountain View office.
Compensation And Benefits
- Standard base salary: $200,000–$260,000 annually.
- Certain roles may be eligible for variable compensation, equity, and benefits.
- Medical, vision, and dental coverage.
- Generous time-off policy.
- 401(k) plan.
- Home office improvement stipend.
- Annual education and wellness stipends.
- Regular company events and healthy lunches daily.
Glean is committed to building and sustaining a diverse, inclusive workplace. The interview process includes a brief AI-focused exercise or discussion addressing how candidates think about, design, and use AI. Prior Glean experience is not required.