Senior Software Engineer - NVLink Rack Scale Stability and Reliability

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Bash @ 1 CI/CD @ 4 CUDA @ 4 Communication @ 7 Debugging @ 7 Distributed Systems @ 6 GPU @ 4 HPC @ 4 InfiniBand @ 6 NVLink @ 4 Networking @ 6 Performance Analysis @ 6 Python @ 1 SRE Stress Testing @ 4 System Architecture @ 7

Details

NVIDIA is looking for highly motivated Senior Software Engineers to join its Fabric Networking team, focusing on NVLink Rack-Scale Systems Stability and Reliability. The role involves transforming next-generation NVLink and NVSwitch platforms into stable, reliable, volume production-ready systems and contributing to the software foundation for large-scale data center deployments.

Responsibilities

  • Drive platform bring-up, feature enablement, end-to-end software validation, and debugging for next-generation NVLink-based GPU and rack-scale systems.
  • Develop tools, diagnostics, automation, and infrastructure for system validation, regression testing, and fleet support.
  • Lead reliability and MTBI validation through stress testing, telemetry analysis, failure injection, and issue resolution.
  • Triage complex software, firmware, networking, and platform issues across validation, deployment, and production environments.
  • Collaborate with architecture, hardware, firmware, software, and customer engagement teams to improve system quality and reliability.
  • Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and operational readiness.
  • Create automation, dashboards, runbooks, and debugging workflows to improve root-cause analysis and operational efficiency.

Requirements

  • BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
  • 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.
  • Strong programming skills in C/C++ and Python; Bash or shell scripting experience is a plus.
  • Strong system-level debugging skills across software, firmware, hardware, and networking layers.
  • Solid networking fundamentals, including TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.
  • Experience with large-scale AI systems, including platform bring-up, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging.
  • Ability to triage complex multi-domain issues using logs, telemetry, experiments, and structured debugging methods.
  • Strong communication and collaboration skills across engineering, customer, and operations teams.
  • Passion for building reliable next-generation AI infrastructure and solving complex system-level challenges at scale.

Preferred Qualifications

  • Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI or HPC clusters such as NVIDIA GB200 NVL72.
  • Strong understanding of large-scale AI system architecture, including PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training and inference systems.
  • Experience with server management technologies, data center operations, cluster provisioning, scaling, and fleet monitoring.
  • Proven experience building diagnostics, automation, CI/CD pipelines, dashboards, and reliability tooling.

Compensation and Benefits

  • Base salary range for Level 3: USD 152,000–241,500 per year.
  • Base salary range for Level 4: USD 184,000–287,500 per year.
  • Eligibility for equity and benefits.
  • Applications will be accepted at least until July 31, 2026.
  • NVIDIA is an equal opportunity employer.

More jobs at Nvidia

Similar jobs