Senior Software Engineer - NVLink Rack Scale Stability and Reliability
at Nvidia
USD 152,000-287,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 1
CI/CD @ 4
CUDA @ 4
Communication @ 7
Debugging @ 7
Distributed Systems @ 6
GPU @ 4
HPC @ 4
InfiniBand @ 6
NVLink @ 4
Networking @ 6
Performance Analysis @ 6
Python @ 1
SRE
Stress Testing @ 4
System Architecture @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for highly motivated Senior Software Engineers to join its Fabric Networking team, focusing on NVLink Rack-Scale Systems Stability and Reliability. The role involves transforming next-generation NVLink and NVSwitch platforms into stable, reliable, volume production-ready systems and contributing to the software foundation for large-scale data center deployments.
Responsibilities
- Drive platform bring-up, feature enablement, end-to-end software validation, and debugging for next-generation NVLink-based GPU and rack-scale systems.
- Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.
- Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.
- Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.
- Collaborate with architecture, hardware, firmware, software, and customer engagement teams to improve system quality and reliability.
- Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.
- Create automation, dashboards, runbooks, and debugging workflows to improve root-cause analysis and operational efficiency.
Requirements
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.
- Strong programming skills in C/C++ and Python; Bash or shell scripting experience is a plus.
- Strong system-level debugging skills across software, firmware, hardware, and networking layers.
- Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.
- Experience with large-scale AI systems, including platform bring-up, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.
- Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.
- Strong communication and collaboration skills across engineering, customer, and operations teams.
- Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.
Preferred Qualifications
- Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI or HPC clusters such as NVIDIA GB200 NVL72.
- Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training and inference systems.
- Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.
- Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.
Compensation and Benefits
- Base salary range for Level 3: USD 152,000–241,500 per year.
- Base salary range for Level 4: USD 184,000–287,500 per year.
- Eligibility for equity and benefits.
- Applications will be accepted at least until July 31, 2026.
- NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Senior DevTech Compute Engineer, Compression and Data Processing
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year