Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Agentic AI @ 4
Audit @ 4
CI/CD @ 4
Claude Code
Distributed Systems @ 8
GPU @ 4
Go @ 8
Kubernetes @ 4
LangChain @ 7
Networking @ 4
Observability @ 7
Python @ 8
RAG @ 4
Security
Vector Databases @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Join NVIDIA IT’s Enterprise AI & Automation team to develop and expand enterprise-grade agentic AI systems. NVIDIA’s Enterprise AI Platform drives production AI agents that securely link with enterprise systems to boost employee efficiency and accelerate business results across engineering, IT, supply chain, finance, HR, and sales.
This role is for a Principal- or Distinguished Engineer-level architect who defines systems through direct construction. The position requires a deeply involved technical leader who writes code daily in Python and/or Go and rapidly develops prototypes using modern code-generation tools such as Cursor, Claude Code, and Claude Cowork. The candidate must understand infrastructure ranging from Kubernetes to GPU inference stacks and translate emerging agent development patterns into scalable platform capabilities.
You will build NVIDIA’s enterprise agent architecture by delivering functional systems, developing reference implementations, and elevating technical standards across the organization. This is not a strategy-only or governance-only role; architecture authority is earned through production systems, measurable impact, and technical depth.
The role covers the full agent development process: creating, sandboxing, launching, observing, controlling, and continuously enhancing agents through data-driven cycles. You will develop systems incorporating persistent memory, controlled runtime environments, rigorous assessment, and GPU-powered performance to ensure agents are intelligent, trackable, protected, and production-ready.
Responsibilities
- Develop and deliver production-quality agentic AI systems end to end using Python and/or Go, including Kubernetes deployment, agent runtimes, memory systems, orchestration, tool integration, and evaluation pipelines.
- Define and advance NVIDIA’s Enterprise Agentic AI architecture through practical implementations, reference systems, and production deployments.
- Build and implement multi-agent orchestration patterns, including planner, executor, reviewer, and tool agents, using frameworks such as LangChain, LangGraph, or similar systems, with strong regression coverage and observability.
- Develop high-quality proofs of concept for emerging agent architectures and harden successful patterns into reusable platform services, APIs, SDKs, and developer templates.
- Architect and implement data flywheels that continuously improve agent quality through telemetry, benchmarking, automated evaluation, and structured feedback loops.
- Embed security, guardrails, sandbox isolation, auditability, and policy enforcement directly into agent runtimes in partnership with security and governance teams.
- Evaluate, integrate, and extend open-source and third-party agent platforms, making build-versus-use decisions based on performance, scalability, control, and long-term platform ownership.
- Collaborate with engineering, infrastructure, product, and business stakeholders to align architectural direction with enterprise priorities and accelerate adoption.
Requirements
- Bachelor’s degree in Computer Science or a related field, or equivalent experience; a Master’s degree or PhD is preferred.
- 15+ years of experience building and shipping large-scale distributed systems, with significant hands-on coding in Python, Go, or similar systems languages.
- Proven ability to transition quickly from an idea to a functional prototype and then to a robust, scalable platform solution.
- Proven experience building agentic AI systems, including RAG pipelines, long-term memory models, multi-agent management such as LangChain or LangGraph, tool frameworks, and evaluation infrastructure.
- Expert-level knowledge of Kubernetes, containerized workloads, networking, APIs, and secure enterprise integration patterns.
- Experience developing benchmarking, regression testing, telemetry, and observability systems that measure agent quality, latency, cost, reliability, and safety.
- Comprehensive knowledge of performance tuning in hybrid environments, including GPU-based inference systems.
- Excellent collaboration skills, including the ability to influence cross-functional stakeholders, build positive relationships, and communicate complex architectural concepts clearly to technical and business audiences.
Preferred Qualifications
- Experience delivering reusable developer-acceleration components such as SDKs, APIs, templates, reference implementations, and CI/CD automation.
- Experience integrating enterprise vector databases and retrieval systems, and working with agentic search and orchestration platforms such as Glean, Microsoft Copilot Studio, Google Agentspace, or similar enterprise AI ecosystems.
- Experience embedding fine-grained policy enforcement, access controls, sandbox isolation, and audit trails directly into AI runtimes.
- Experience optimizing model inference, batching strategies, memory utilization, and efficiency on NVIDIA hardware.
- Meaningful open-source contributions, including core commits, maintainership, widely adopted libraries, or public technical artifacts demonstrating system-level depth.
Compensation and Benefits
The base salary range is 272,000 USD to 431,250 USD, determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until March 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to fostering a diverse, equal-opportunity work environment.