Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
GPU
InfiniBand @ 7
Leadership @ 7
Networking @ 7
System Architecture @ 7
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a highly technical, motivated manager to lead and manage the team responsible for rack-scale system software architecture, spanning firmware, kernel drivers, operating systems, networking, fabrics, and associated user mode drivers, plus manageability software.
Responsibilities
- Drive the software end-to-end architecture for NVIDIA's rack-scale products
- Maintain deep understanding of the product portfolio and roadmap; translate forward-looking plans into clear, formal software requirements that anchor execution across the organization
- Ensure high quality and reliable software; serve as a trusted architectural partner to teams requiring guidance or oversight
- Work directly with major customers to understand their requirements and align their roadmap with NVIDIA’s roadmap
- Using strong communication skills, present the team vision to senior NVIDIA and external leaders
- Provide technical leadership and career mentorship to the team
- Make key technical decisions even when faced with ambiguity
Requirements
- BS or MS degree in Computer Engineering, Computer Science, or related degree, or equivalent experience
- 15+ overall years of experience in system architecture and design with 8+ years of proven experience in management
- Deep experience designing architecture for scalable and performant server systems, particularly at the SW/HW interface
- Proven leadership skills and strong ownership on past projects involving a large scale sophisticated code base
- Previous experience working with complex system software for accelerators such as GPUs, DPUs, or FPGAs
- Strong managerial, problem solving, and critical thinking skills
- Comfortable operating in highly matrixed organizations while holding a leadership position
- Known for strong interactive, verbal and written communications skills
Ways to Stand Out from the Crowd
- Knowledge of large-scale cloud and cluster level deployment and management systems; experience with designing robust, resilient and performant scale-up fabrics
- Demonstrated track record of leading data center products across the entire lifecycle, spanning inception, pre-silicon development, post-silicon bring-up, manufacturing, and deployment
- Strong understanding of networking technology and protocols (e.g. Ethernet, Infiniband)
- Familiarity with CXL, UCIE and other C2C technology architectures
- Knowledge in storage and networking technologies
Compensation
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.
You will also be eligible for equity and benefits.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal System Software Engineer - Av Platform
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Director, Software Tpm - Server Firmware And System Software
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Principal Security Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 347,000-490,000 per year
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Software Architect - Deep Learning And Hpc Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Nvidia · Santa Clara, United States
USD 292,000-442,800 per year