Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 7
Azure @ 7
Communication @ 7
Data Pipelines @ 4
Distributed Systems @ 4
GCP @ 7
LLM @ 4
Machine Learning
Networking @ 7
Observability
Performance Optimization
SRE @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Glean is seeking a Senior Infrastructure Technical Program Manager (TPM) to lead large-scale, cross-functional initiatives that define, scale, and optimize its infrastructure platform. The role sits at the intersection of infrastructure engineering, reliability, cost efficiency, and AI systems, partnering with Infrastructure, AI/ML, and Platform Engineering teams to design orchestration systems, streamline deployments, and build the foundations that power Glean's search and AI capabilities.
Responsibilities
- Drive the company's infrastructure roadmap across Setup & Deployment, Runtime, Storage, and AI Infrastructure.
- Lead cross-functional programs that improve scalability, reliability, cost efficiency, and developer velocity.
- Define and orchestrate how Glean instances are deployed, upgraded, and monitored at scale.
- Partner with AI and Data teams to evolve ML pipelines, model training infrastructure, and LLM serving systems.
- Lead initiatives to improve observability, configuration management, and resource utilization.
- Coordinate capacity planning, infrastructure migrations, and performance optimization programs.
- Build visibility into infrastructure cost drivers and partner with finance and engineering leaders on optimization initiatives.
- Lead end-to-end infrastructure programs spanning compute, networking, storage, orchestration, and AI workloads.
- Partner with Engineering to define standards for environment provisioning, deployment automation, and configuration governance.
- Develop and operationalize frameworks for runtime health, scaling, and disaster recovery.
- Drive consistency and automation across deployment orchestration systems.
- Establish metrics for reliability, performance, and cost efficiency.
- Coordinate cross-team delivery of programs such as data pipeline scalability, LLM infrastructure expansion, and infrastructure observability improvements.
- Communicate program status and technical risks to leadership and stakeholders.
- Identify process and system bottlenecks and drive automation to improve the speed and reliability of infrastructure operations.
Requirements
- Bachelor's or master's degree in Computer Science, Engineering, or a related technical field.
- 8–10+ years of experience in technical program management, infrastructure, or SRE, including at least 3–5 years managing infrastructure- or platform-scale programs.
- Proven success delivering cross-functional infrastructure programs in B2B or enterprise environments where scalability, uptime, and performance are critical.
- Experience working with Infrastructure, SRE, and ML/AI teams on distributed systems or data infrastructure.
- Strong understanding of cloud infrastructure, including compute, networking, storage, and orchestration systems, on AWS, GCP, or Azure.
- Understanding of data pipelines, ML training workflows, and LLM runtime infrastructure is a plus.
- Ability to structure complex, multi-quarter infrastructure programs with clear milestones and measurable impact.
- Strong written and verbal communication skills, with the ability to manage through ambiguity, anticipate scaling challenges, and align teams across priorities.
- Builder mindset focused on automation, reliability, and efficiency.
Location
- Hybrid role requiring four days per week in the San Francisco office.
Compensation And Benefits
- Standard base salary range: $198,000–$235,500 annually.
- Certain roles may be eligible for variable compensation, equity, and benefits.
- Medical, vision, and dental coverage.
- Generous time-off policy.
- 401(k) plan.
- Home office improvement stipend.
- Annual education and wellness stipends.
- Regular company events and healthy lunches daily.
- Diverse and inclusive workplace.
Interview Process
Candidates complete a brief AI-focused exercise or discussion covering how they think about, design, and use AI to drive impact. Prior Glean experience is not required.
More jobs at Glean
Senior/Staff Applied Scientist
Glean · San Francisco, United States
USD 180,000-330,000 per year
Senior/Staff Applied Scientist
Glean · Mountain View, United States, San Francisco, United States
USD 180,000-330,000 per year
Software Engineer, Intern (Summer 2027)
Glean · Mountain View, United States, San Francisco, United States
USD 57-69 per hour
Machine Learning Engineer, Search Quality
Glean · San Francisco, United States
USD 140,000-265,000 per year
Designated Technical Support Engineer - Central/East
Glean · Nashville, United States, New York City, United States
USD 120,000-170,000 per year
Similar jobs
Senior Technical Program Manager, Infrastructure
Glean · Mountain View, United States
USD 198,000-235,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Engineering Manager – AI Platform & SRE
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Staff Site Reliability Engineer - AI Platform Runtime
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year