Senior Network Reliability Engineer - DGX Cloud

at Nvidia
USD 136,000-264,500 per year
SENIOR
✅ On-site

Tech Stack

AI AWS @ 4 Azure @ 4 BGP @ 7 Change Management Communication @ 6 GCP @ 4 Grafana @ 3 InfiniBand @ 4 Linux @ 4 Prometheus @ 3 Python @ 4 System Administration @ 4

Details

NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain cloud and datacenter network infrastructures serving the company's software stack, including Graphics Drivers, Autonomous Vehicles, and Artificial Intelligence.

The role involves remediating critical alerts within defined SLAs, triaging production-impacting network incidents, addressing internal customer network issues, coordinating with external vendors on hardware and software problems, and participating in network device upgrades and capacity augmentations.

Responsibilities

  • Participate in 24/7 global shift rotations to provide remote support for network repairs and changes.
  • Collaborate across teams and update customers on status and ticket information.
  • Drive operational improvements in change management and daily operations by following established procedures.
  • Manage and operate large-scale IP network technologies and infrastructures.
  • Work with peering and datacenter interconnect technologies, including PNI, transit, exchange, passive DWDM, and wave circuits.
  • Monitor and support network health across on-premises and cloud infrastructures.
  • Collaborate on workflow enhancements and document best practices.
  • Respond to alerts within defined SLAs and participate in incident management.
  • Contribute to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.

Requirements

  • Deep knowledge and experience with TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.
  • 5+ years of experience in network operations.
  • Strong network troubleshooting and creative problem-solving skills.
  • Experience with one or more cloud service provider environments: AWS, Azure, GCP, or OCI.
  • Familiarity with Arista, Fortinet, and Juniper.
  • Bachelor's degree in Computer Science, a related technical field, or equivalent experience.
  • Excellent verbal and written communication skills.

Preferred Qualifications

  • Solid understanding of Mellanox/Cumulus OS and InfiniBand technology.
  • Unix/Linux system administration experience.
  • Ability to write and understand Python and Shell scripts to improve efficiency in hyperscale environments.
  • Familiarity with NetBox/Nautobot, Prometheus, Grafana, and Panoptes for monitoring and managing global networks.

Compensation And Benefits

The base salary is determined based on location, experience, and the pay of employees in similar positions. The base salary ranges are USD 136,000–224,250 for Level 3 and USD 168,000–264,500 for Level 4. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs