Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Cumulus Linux
Data Modeling @ 4
Debugging @ 4
HPC
Leadership @ 8
Linux @ 8
Networking @ 8
People Management @ 4
Python
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a software engineering leader to join the Cumulus Linux team and lead development of the network operating system software that powers accelerated, disaggregated, and software-defined data centers for AI and high-performance computing. Cumulus Linux is a Debian-based distribution. The role is responsible for leading a software development team focused on core infrastructure services, Reliability, Availability and Serviceability features, and telemetry, while collaborating with multifunctional engineering and product teams.
Responsibilities
- Lead a team developing and delivering Cumulus Linux operating system telemetry and infrastructure features.
- Partner with engineering teams to scope and develop solutions that improve system security, performance, and reliability.
- Develop and debug C and Python code for system monitoring, reliability, and serviceability features as needed.
- Collaborate with product, architecture, and engineering teams on end-to-end integration of infrastructure features into Linux and the Cumulus Linux distribution.
- Work with project management on feature estimation and planning.
- Support team growth by sourcing and interviewing candidates, participating in conferences and events, and onboarding new employees.
- Help engineers develop their careers by assigning projects aligned with their current skills and long-term development goals.
- Work with upstream communities and monitor technology trends and emerging standards.
- Guide problem-solving efforts, reduce recurring issues, and proactively prevent problems.
Requirements
- Master of Science in Electrical Engineering, Computer Science, Computer Engineering, or a related field; or a bachelor's degree or equivalent experience.
- 10+ years of proven leadership experience in Linux systems and data center networking technologies.
- At least 2 years of people management experience in an enterprise environment.
- Familiarity with cloud-native concepts.
- Strong background in Linux operating system feature development.
- Knowledge of data modeling concepts, OpenConfig, and streaming telemetry protocols such as gNMI.
- Experience driving projects from concept through production.
- Excellent written, verbal, and interpersonal communication skills.
- Ability to articulate value propositions to customers and influence internal teams.
- Experience with embedded software on network switches.
- Experience bringing up and troubleshooting Ethernet interfaces and modules.
- Familiarity with data center protocols.
- Ability to work independently with minimal direction.
Additional Qualifications
- Strong background in Linux systems and Linux kernel networking.
- Hands-on laboratory experience with system bring-up and debugging.
- Collaborative working style.
Compensation and Benefits
- Base salary for Level 3: USD 224,000–356,500 per year.
- Base salary for Level 4: USD 272,000–431,250 per year.
- Eligible employees may also receive equity and benefits.
- Applications will be accepted at least until June 4, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior System Software Engineer, Platform - OpenBMC
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Security Engineer, RTOS and Virtualization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVIDIA Warp
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Software And System Architect
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Similar jobs
Senior Software Engineer, Platforms
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Systems Software Engineer, Data Center Platform Enablement
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Platform Security Engineer, OpenBMC
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-405,000 per year
Service Reliability Engineer
Nvidia · United States
USD 168,000-333,500 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year