Member of Technical Staff (TPM, Inference)
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API
Agile @ 3
Distributed Systems @ 6
GPU
LLM @ 3
Machine Learning
Observability
Product Management @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Perplexity is looking for a technical program manager to serve as the connective tissue between model providers, engineering, and product teams, driving the core inference platform forward. The role supports a high-throughput inference stack serving Ask, Computer, and API traffic across a shifting portfolio of first-party and third-party models. It sits at the intersection of product, engineering, and finance, coordinating model onboarding and capacity while executing the inference platform roadmap.
Our Mission
Perplexity's mission is to power curiosity through a continuous cycle of learning, building, and integrating.
Responsibilities
- Execute the roadmap for the inference platform, including request handling, rate limits and quotas, usage controls, reliability, and observability.
- Coordinate onboarding, launch readiness, and rollout for new models and capacity across model providers and internal engineering and product teams.
- Drive latency, throughput, uptime, and cost-efficiency as core execution metrics, surfacing tradeoffs between them.
- Run the operating model for model-release and optimization programs, including day-zero launches, across performance engineering, infrastructure, and product teams.
- Lead cross-functional delivery for inference-stack changes from planning through launch and post-launch validation.
- Build mechanisms that make releases predictable, including rituals, dashboards, and launch checklists.
- Partner with GPU capacity and compute teams to reconcile execution decisions with cost, capacity, and vendor constraints.
Requirements
- Strong experience in technical program management or product management for infrastructure, distributed systems, or ML/model-serving products.
- Direct experience with production LLM or ML inference, including understanding how to make serving fast, reliable, and cost-effective.
- Ability to coordinate external partners and internal engineering teams with competing priorities and timelines.
- Experience working with data and metrics, with judgment to surface tradeoffs between latency, throughput, uptime, and cost.
- Ability to thrive in a small, agile team with initiative and ownership in an environment with little precedent.
- 6+ years of combined technical program management or product management experience.
Benefits
Full-time U.S. employees receive benefits including equity, health, dental, vision, retirement, fitness, commuter and dependent care accounts, and more. International employees receive a benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions, and final offers vary based on factors including experience and expertise.