Member of Technical Staff (Applied AI Engineer, Agent Capabilities)
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS
Agentic AI
Data Analysis @ 4
Data Science @ 4
Distributed Systems @ 7
Go
LLM @ 4
Machine Learning
Observability @ 4
PostgreSQL
Product Management @ 4
Python
Reinforcement Learning @ 4
Rust
Technical Leadership
TypeScript
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Perplexity Computer is one of the defining products of the new era of agentic AI. Millions of people use Perplexity to transform knowledge into action, and the Agent Capabilities team sits at the intersection of frontier AI research and product innovation, building the foundations that shape how users and agents solve increasingly complex tasks.
The Agent Capabilities team turns frontier AI breakthroughs into reusable product capabilities. The team evaluates emerging model capabilities, determines where they create real user value, and transforms them into reliable, scalable, high-quality experiences for users and agents. This role has broad ownership across frontier AI research, agent systems, platform engineering, and product innovation.
Tech Stack
Python, Go, Rust, PostgreSQL, DynamoDB, AWS, and TypeScript.
Responsibilities
- Evaluate frontier models against real user tasks, identify useful behaviors and failure modes, and turn promising advances into production agent systems. Own the lifecycle from rapid prototyping and evaluation through launch, monitoring, and iteration.
- Improve agents' ability to plan, use tools, manage context, recover from errors, and complete long-running tasks reliably.
- Apply state-of-the-art machine learning and large language model techniques to design scalable agent capabilities such as skills, plugins, artifact generation, tool integration and use, auto-research, and multi-agent collaboration.
- Shape the architecture, abstractions, and product experiences that enable users and agents to compose increasingly sophisticated solutions for real-world tasks.
- Own agent behavior and capabilities end-to-end, from user-facing products and interfaces to backend services.
- Define offline and online evaluations for task completion, correctness, safety, latency, cost, and user satisfaction. Iteratively improve models, prompts, harnesses, and products for different problem spaces.
- Build secure, observable, and reliable agent systems, including permissions and safeguards for sensitive actions.
- Develop tracing, replay, and monitoring infrastructure that makes agent failures reproducible and actionable.
- Collaborate with Product Management, Data Science, and Research to identify high-impact opportunities, validate emerging model capabilities, and turn complex agent behaviors into simple, reliable product experiences.
- Apply advances in models, inference, evaluation, and agent architecture when they produce measurable improvements in production performance.
- Set technical direction on ambiguous problems and raise the bar through design reviews, mentorship, and technical leadership.
Requirements
- Typically 6+ years of professional software engineering experience, with a track record of building and owning robust AI-powered, large-scale, user-facing, or data-intensive products. Exceptional candidates with less experience and an outstanding record of impact are encouraged to apply.
- Strong software engineering fundamentals, with experience building and operating AI/ML products, backend services, or distributed systems at scale.
- Experience owning the AI product lifecycle, including data analysis, rigorous evaluation, production monitoring, and iterative improvement.
- Ability to define metrics and use production data and user feedback to guide decisions.
- Practical experience in one or more relevant areas, such as agent harnesses, tool use, context engineering, model evaluation, browser automation, or long-running task execution.
- Strong product judgment and execution, including the ability to translate ambiguous user needs into applied AI or ML problems and ship durable solutions with measurable user impact.
- Genuine interest in frontier AI capabilities and agent systems, with enthusiasm for rapidly exploring, evaluating, and productizing new model behaviors.
Preferred Qualifications
- Experience with LLM context engineering or harness engineering, subagents, coding assistants, or long-running or autonomous task execution.
- Deep familiarity with the strengths and limitations of current model families across reasoning, tool use, context management, and long-horizon tasks.
- Experience building agent permissions, safeguards, evaluation infrastructure, or production observability systems.
- Experience with mid-training, post-training, or reinforcement learning for frontier or open-source models.
- AI/ML research experience demonstrated through publications, open-source contributions, or other meaningful research impact.
- Experience at a fast-growing startup or on a high-ownership engineering team.
Benefits
Full-time U.S. employees receive a comprehensive benefits program including equity, health, dental, vision, retirement, fitness, commuter, and dependent care accounts. Final offer amounts may vary based on experience and expertise.