Principal Engineer, Local AI - Agents and Systems

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 4 GPU @ 4 LLM @ 4 LangChain @ 4 Python @ 6 Security @ 7 TensorRT @ 4 vLLM @ 6

Details

Artificial intelligence is shifting from passive help to autonomous, always-on workflows. The mission is to make this change seamless, efficient, and secure for millions globally. This role will lead technical efforts in deploying advanced AI agent frameworks and local runtimes on Windows and NVIDIA GeForce RTX GPUs. The engineer will guide development so open-source AI agents, such as Nemoclaw and OpenClaw, operate locally, safely, and efficiently on consumer PCs. By combining local inference using Nemotron models with privacy routers and sandboxed execution, the role will help develop the foundation of a desktop AI operating system.

Responsibilities

  • Act as the lead engineer for developing agent frameworks natively on Windows.
  • Build the technical roadmap for bringing always-on, self-evolving AI assistants to GeForce RTX PCs and laptops.
  • Lead efforts to optimize agent runtimes for Windows.
  • Ensure autonomous agents operate within policy-based privacy and security frameworks, including filesystem access, secure inference routing, and network egress.
  • Partner with internal AI research teams, driver teams, and the open-source OpenClaw community.
  • Help ensure consumer hardware provides an excellent ecosystem for autonomous agents.
  • Mentor other engineers, establish guidelines for AI agent deployment, and write reliable, production-ready code.

Requirements

  • 15+ years of relevant professional software engineering experience, including at least 3+ years in a Staff or Lead Architect role.
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related technical field, or equivalent experience.
  • Deep understanding of Windows operating system internals, process isolation, sandboxing technologies, and system-level security architecture.
  • Proven understanding of LLM inference pipelines, including Ollama, Llama.cpp, and vLLM.
  • Experience with GPU-accelerated computing using CUDA and TensorRT.
  • Experience running local models on consumer-grade hardware.
  • Practical experience with modern AI orchestration and agentic frameworks, such as OpenClaw, Hermes, and LangChain.
  • Understanding of how multi-agent systems plan, act, and use tools.
  • Proficiency in multiple programming languages, particularly C++ for performance-critical systems and operating system integration, and Python for AI and blueprint logic.
  • Experience building virtualization, containerization, or robust sandboxing tools natively for the Windows ecosystem.

Benefits

  • Competitive salary.
  • Comprehensive benefits package.
  • Eligibility for equity and benefits.

The base salary range is 272,000 USD - 431,250 USD, determined by location, experience, and the pay of employees in similar positions. Applications will be accepted at least until July 31, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive, equal-opportunity work environment.

More jobs at Nvidia

Similar jobs