Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
C @ 7
C++ @ 7
CUDA @ 4
Communication @ 6
Deep Learning @ 3
GPU @ 6
LLM
Performance Analysis @ 6
Profiling @ 6
Python
Software Development @ 6
TensorRT
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA’s TensorRT team is building an AI-native development initiative to make TensorRT the default entry point for out-of-framework inference globally. The initiative uses swarms of AI agents to produce high-performance, high-quality, modern C++ software at scale. The role focuses on scaling an agentic development framework, applying deep learning advances, improving model onboarding, and optimizing inference performance for LLM, Diffusion, Audio, Vision, and multimodal models.
Responsibilities
- Architect and build an AI-native framework and codebase that supports large numbers of AI agents working in parallel to generate, test, and validate production-grade software.
- Improve compute-to-software output through AI-native tools, multi-agent orchestrators, and codebase harnesses.
- Identify state-of-the-art industry and academic breakthroughs, such as new attention mechanisms and KV cache strategies, and dispatch AI agent swarms to prototype and integrate them.
- Deliver a seamless, high-performance path to production for the latest LLM, Diffusion, Audio, Vision, and multimodal model families.
- Optimize performance at the intersection of Python orchestration and C++ engine-level development to achieve latency and throughput improvements.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science, Computer Engineering, AI, or equivalent experience.
- 4+ years of relevant software development experience.
- Strong modern C++ skills, including C++11/14/17 or newer and the STL, with an emphasis on clean, maintainable, performant code.
- Familiarity with deep learning, modern inference frameworks, and the architectural nuances of LLMs, Diffusion, and multimodal models.
- Interest in evolving software architecture to support automated, agent-driven development and indefinitely scaling codebases.
- Ability to translate high-level customer needs into technical requirements and user-centric solutions.
- Ability to deliver production-quality software from customer requests on tight timelines.
- Excellent communication skills and comfort working across internal organizations and with customers.
Preferred Qualifications
- Hands-on experience with AI agent orchestrators, multi-agent coding frameworks, or custom agentic coding harnesses for production software.
- CUDA programming experience or exposure to kernel generation and autotuning.
- Experience rapidly turning state-of-the-art research papers into working prototypes.
- Expertise in CPU and/or GPU performance analysis, profiling, and optimization, including the use of tooling to drive measurable improvements.
Benefits
- Equity and benefits are provided in addition to the base salary.
- NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer.
Applications will be accepted at least until April 25, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
Senior Software Engineer, Deep Learning Inference - Automotive Safety
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Engineering Manager, Deep Learning Inference
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Software Engineer - Autonomous Driving
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year