Lead Site Reliability Engineer

at Glean
USD 200,000-260,000 per year
SENIOR
✅ Hybrid

Tech Stack

AI AWS @ 7 Algorithms Azure @ 7 Compliance Distributed Systems @ 6 Docker @ 4 IaC Kubernetes @ 4 Networking @ 4 Performance Optimization SRE @ 4 Security @ 4 Software Development @ 6 Technical Leadership Terraform @ 3

Details

Glean is seeking a Site Reliability Engineering Lead to foster engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. The role is responsible for ensuring services meet stringent Service Level Objectives (SLOs), building resilient and automated production environments in the cloud, and providing technical leadership for globally used products.

The SRE team builds infrastructure to scale operations in a hybrid cloud environment and eliminates manual work through automation. The role involves managing complex scale and growth challenges while applying expertise in coding, algorithms, problem-solving, and site reliability engineering practices.

Responsibilities

  • Provide technical leadership and mentorship across engineering teams.
  • Establish best practices for incident management, performance optimization, and automation.
  • Drive cross-team collaboration and contribute to key engineering objectives.
  • Shape architectural decisions and ensure the delivery of reliable, high-quality systems.
  • Implement and maintain resilient cloud architectures.
  • Monitor system performance and proactively identify and resolve bottlenecks and points of failure.
  • Participate in the primary on-call rotation.
  • Foster a blameless postmortem culture and continuously improve the on-call process.
  • Develop and maintain automation scripts, tools, and processes for system deployment, monitoring, and management.
  • Optimize cloud infrastructure and applications for performance, scalability, and cost-effectiveness.
  • Collaborate with security engineers to implement best practices and ensure compliance with security standards and policies.
  • Design and configure monitoring and alerting systems, dashboards, and production on-call playbooks.
  • Participate in system design and launch reviews, providing SRE guidance throughout the software development lifecycle.

Requirements

  • Bachelor's degree in Computer Science, a related field, or equivalent practical experience.
  • 8+ years of experience in a senior-level Site Reliability Engineering or similar role, particularly managing cloud-based services and infrastructure.
  • 5+ years of software development experience in one or more programming languages.
  • 3+ years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems running in the cloud.
  • Strong knowledge of cloud platforms such as Google Cloud Platform, AWS, or Azure.
  • Practical experience with Docker and Kubernetes.
  • Essential familiarity with infrastructure as code tools such as Terraform.
  • Solid understanding of networking, security principles, and SRE and security best practices.
  • Proficiency with monitoring and alerting tools.

Location

  • Hybrid role requiring four days per week in the Mountain View office.

Compensation And Benefits

  • Standard base salary: $200,000–$260,000 annually.
  • Certain roles may be eligible for variable compensation, equity, and benefits.
  • Medical, vision, and dental coverage.
  • Generous time-off policy.
  • 401(k) plan.
  • Home office improvement stipend.
  • Annual education and wellness stipends.
  • Regular company events and healthy lunches daily.

Glean is committed to building and sustaining a diverse, inclusive workplace. The interview process includes a brief AI-focused exercise or discussion addressing how candidates think about, design, and use AI. Prior Glean experience is not required.

More jobs at Glean

Similar jobs