Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
Algorithms @ 7
CUDA @ 6
Communication @ 6
Data Structures @ 7
Debugging @ 7
GPU @ 6
GenAI
Generative AI @ 4
Linux
Machine Learning @ 4
Mathematics @ 4
PyTorch @ 6
TensorRT @ 4
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Local AI Systems Software Engineering at NVIDIA
There is a growing emphasis on running AI models locally, closer to the source of data to reduce latency, improve real-time processing, and address privacy concerns by minimizing the need for sending data to centralized servers.
Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI challenges on Windows and Linux PCs with limited resources.
Responsibilities
- Partnering with NVIDIA software, research, architecture, and product teams to align strategies and technical needs for fostering the ecosystem of AI on RTX and DGX PCs.
- Collaborate closely with industry partners to advance AI across critical domains—including graphics, web browsers, and edge devices—by driving innovation in both open and closed source technologies with emphasis on system level support.
- Improving performance on current and next-generation GPU architectures by conducting in-depth analysis and end-to-end optimization of AI models, data processing pipelines, and inference runtime features.
- Identifying, evaluating, and implementing compute and memory optimization techniques—such as quantization, distillation, and pruning—for large AI models; fine-tuning and compressing models to fit edge devices.
Requirements
- Bachelor’s, Master’s, or PhD in Computer Science, Software Engineering, Mathematics, or a related field (or equivalent experience).
- Excellent C++ programming and debugging skills with a strong understanding of data structures and algorithms.
- 5+ years of experience with proficiency in AI inferencing pipelines and applications using ML/DL frameworks, including ONNX RT, PyTorch, Tensor RT, llama.cpp and vLLM.
- Strong analytical and problem-solving abilities, with the ability to multitask effectively in a dynamic environment.
- Outstanding written and oral communication skills enabling effective collaboration with management and engineering teams.
Ways To Stand Out from The Crowd
- Understanding modern techniques in Machine Learning, Deep Neural Networks, and Generative AI with relevant contributions to major open-source projects will be a plus.
- Consistent track record of delivering end-to-end products with geographically distributed teams in multinational product companies.
- Proficiency in lower-level system/GPU programming, CUDA, developing high-performance systems.
- Hands-on experience with building applications using APIs like ONNX RT, DirectX, PyTorch, TensorRT, Vulkan, llama.cpp.
Compensation and Applications
- Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.
- The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
- You will also be eligible for equity and benefits.
- Applications for this job will be accepted at least until August 1, 2026.
More jobs at Nvidia
Senior Engineer, Local AI - Agents And Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Engineer, Local Ai - Agents And Systems
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Package Design Methodology Engineer
Nvidia · Santa Clara, United States
USD 136,000-264,500 per year
Principal Engineer, Local AI - Agents And Systems
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Similar jobs
Senior Software Engineer, CUDA Deep Learning Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Developer Technology Engineer - Edge Agentic Ai
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Metropolis Vision Ai
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
ML Solutions Architect (Early Talent)
Nebius · United States
USD 102,000-126,000 per year
Senior Software Engineer - Autonomous Driving Simulation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Metropolis Vision AI
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year