Senior Software Engineer - Nvlink Rack Scale Stability And Reliability
at Nvidia
USD 152,000-287,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 1
Communication @ 7
Debugging @ 7
Distributed Systems @ 6
GPU
InfiniBand @ 6
NVLink
Networking @ 6
Performance Analysis @ 6
Python @ 1
SRE
Stress Testing @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for highly motivated Senior Software Engineers to join its Fabric Networking team with a targeted focus on NVLink Rack-Scale Systems Stability & Reliability.
In this role, you will partner closely with architects and developers building next-generation NVLink and NVSwitch systems, helping transform first-of-their-kind platforms into stable, reliable, and volume production-ready systems. You will work on complex system-level challenges spanning resiliency, diagnostics, recovery, and large-scale AI infrastructure.
Responsibilities
- Drive platform bringup, feature enablement, end-to-end software validation, and debug for next-generation NVLink-based GPU and rack-scale systems.
- Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.
- Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.
- Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.
- Collaborate with architecture, hardware, firmware, software, and Customer engagement teams to improve system quality and reliability.
- Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.
- Create automation, dashboards, runbooks, and debug workflows that improve root-cause analysis and operational efficiency.
Requirements
- BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
- 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.
- Strong programming skills in C/C++ and Python; Bash/Shell scripting experience is a plus.
- Strong system-level debugging across software, firmware, hardware, and networking layers.
- Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.
- Experience with large-scale AI systems, including platform bringup, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.
- Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.
- Strong communication and collaboration skills across engineering, customer, and operations teams.
Benefits
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Principal Developer, AI Networking
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
System Software Engineer, Dynamo-Triton Inference Server - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior GPU and HPC Infrastructure Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Nvidia · Santa Clara, United States
USD 292,000-442,800 per year
Senior System Software Engineer – Data Center GPU Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year