Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 6
Agentic Systems
Communication @ 7
Compliance
Data Pipelines
GPU
LLM @ 6
Leadership @ 7
RAG
Security @ 7
Vector Databases
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Senior Engineering Manager for Agentic Systems and Platform Architecture, you will lead the strategy and execution for NVIDIA’s agentic developer platform. You will understand how teams across the company build, evaluate, and improve autonomous agents, and turn evolving patterns into scalable platform capabilities. You will identify gaps and friction, drive rapid proofs of concept on emerging agent constructs and ecosystem tools, and operationalize effective approaches into reusable building blocks, integrations, and governance mechanisms that accelerate developer productivity and agent quality.
Responsibilities
- Track and deeply understand evolving agent development patterns across NVIDIA and the broader ecosystem.
- Identify gaps and friction in current agent architectures and translate insights into a platform strategy that improves developer velocity and agent quality through evaluations, benchmarking, and feedback loops.
- Assess and integrate open-source and third-party tools where they add leverage, and drive clear build-versus-use decisions.
- Architect and integrate high-performance data pipelines, retrieval-augmented generation (RAG) systems, vector databases, and GPU-optimized training and inference workflows.
- Lead integration of the AI Data Platform into NVIDIA’s on-premises AI Factory, optimizing GPU-to-storage throughput, data locality, and distributed inference performance.
- Establish and enforce Agent Governance policies covering model and tool usage, data lineage, compliance, and Responsible AI frameworks.
- Design, implement, and maintain a centralized Agent Safety Toolkit with pre-vetted components for input and output guardrails and prompt-injection defenses.
- Lead and grow a high-performing team and a multifunctional community to standardize procedures and scale adoption.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
- More than 10 years of overall software engineering experience, including more than 4 years managing high-performing teams.
- Strong hands-on experience with evolving agent architectures and open-source libraries.
- Deep expertise in large language model (LLM) and agent architectures, including leading proofs of concept and integrating them into real business use cases with measurable adoption and impact.
- Ability to turn fast-moving, ambiguous problem spaces into a clear platform strategy, roadmap, and outcomes.
- Proven track record building multi-team developer platforms, including APIs, SDKs, reusable components, and reference implementations.
- Experience building evaluation and benchmarking systems for agent workflows, including metrics, regression testing, and feedback loops.
- Strong judgment integrating open-source and third-party tools, with clear build-versus-use decision-making and integration strategy.
- Product-oriented approach to safety and governance, including controls, auditability, monitoring, and risk management.
- Strong leadership and executive communication skills across engineering, product, security, and research.
Ways to Stand Out
- Experience implementing enterprise-grade governance for agent systems in production autonomous workflows, including controls, auditability, monitoring, and policy enforcement.
- Demonstrated success taking new or open-source agent constructs from proof of concept to production adoption with measurable business impact, such as improved cycle time, quality, cost, or reliability.
- Experience building and scaling an agent platform or agent developer experience used by multiple teams, including SDKs, templates, reference applications, and reusable building blocks.
- Clear perspective and practical examples regarding build-versus-use decisions, including when to adopt open-source or third-party tools versus building internal primitives.
- Deep experience with agent evaluation at scale, including long-horizon tasks, tool correctness, reliability testing, automated regressions, and offline and online feedback loops.
Compensation and Benefits
The base salary range is USD 272,000–431,250 per year. The role is also eligible for equity and benefits.
Applications will be accepted at least until February 27, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.