Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Agile @ 4
Communication @ 7
Distributed Systems @ 7
GPU
LLM @ 4
Machine Learning @ 4
Networking @ 7
People Management @ 6
Python @ 6
Rust @ 6
SGLang @ 4
TensorRT @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Software Engineering Manager for its applied research team within the Networking Systems & Software Architecture group. Based in Santa Clara, California, the team addresses infrastructure challenges in AI by building systems-level software that moves data between GPUs, nodes, and storage. Its work spans low-level transport optimization, hardware-software co-design, communication frameworks for production AI stacks, and emerging domains such as quantum computing interconnects.
Responsibilities
- Lead and develop a team of systems and networking engineers building distributed AI communication systems, including libraries, frameworks, and system integrations.
- Build the technical roadmap in partnership with principal engineers and architects, balancing near-term delivery with long-term research initiatives.
- Establish a culture of technical excellence and open collaboration.
- Manage project planning, resource allocation, and delivery timelines across concurrent workstreams.
- Own execution and set technical direction for the team.
Requirements
- 8+ years of software engineering experience, with advanced knowledge of systems software, networking, or distributed systems.
- 3+ years of direct people management experience.
- Bachelor's, master's, or PhD degree, or equivalent experience, in Computer Science, Computer Engineering, or a related field.
- Ability to scope problems, create plans, and deliver results in a fast-paced research and development environment.
- Strong communication skills, including public speaking, technical writing, and providing candid feedback.
- Good understanding of computer architecture, memory hierarchies, DMA engines, and networking.
- Proficiency in C, C++, Rust, and Python.
- Understanding of machine learning systems concepts, including transformer architectures, KV cache mechanics, model parallelism, and distributed training and inference patterns.
Preferred Qualifications
- Knowledge of machine learning inference frameworks such as vLLM, SGLang, and TensorRT-LLM, including their communication requirements.
- Familiarity with hardware and software ecosystems.
- Experience with agile methodologies adapted for engineering teams focused on research.
Compensation And Benefits
- Level 2 base salary: $184,000–$287,500 USD per year.
- Level 3 base salary: $224,000–$356,500 USD per year.
- Eligible for equity and benefits.
- Applications will be accepted at least until July 25, 2026.
- NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior System Software Engineer - Dynamo-Triton Inference Server
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal Architect, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Deep Learning Algorithm Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Software Engineering Intern, Dynamo – Fall 2026
Nvidia · Santa Clara, United States
USD 20-71 per hour
Senior System Software Engineer, Agentic Inference – Dynamo
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year