Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Azure @ 4
BGP @ 7
Change Management
Communication @ 6
GCP @ 4
Grafana @ 3
InfiniBand @ 4
Linux @ 4
Prometheus @ 3
Python @ 4
System Administration @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Network Reliability Engineer to support and maintain our cloud and datacenter network infrastructures. This network serves the needs across the whole software stack for NVIDIA.
In this role, the Senior Network Operations Engineer will remediate critical alerts within defined SLAs, triage production impacting network incidents, and interact with internal customers on network related issues. They will also be responsible for engaging with external vendors to remediate hardware and software issues, and participate in project related work such as network device upgrades and capacity augmentations. An ideal candidate will possess a wide range of skills, including alert monitoring & resolution in large-scale networks and CSP environments, outstanding troubleshooting skills, understanding of L3 underlay networks, and network protocol knowledge in large multi-vendor infrastructures.
Responsibilities
- Engage in 24/7 global shift rotations to provide remote support for network repairs and changes while collaborating across teams and updating customers on status and ticket information.
- Drive operational improvements in change management and daily operations by following procedures.
- Manage and operate large scale IP network technologies and infrastructures.
- Utilize Peering and Datacenter interconnect technologies: PNI, Transit, Exchange, Passive DWDM, Wave circuits.
- Monitor and support the network health of on-premises and cloud infrastructures.
- Collaborate and develop workflow enhancements while documenting best practices.
Requirements
- Deep knowledge and experience of TCP/IP, BGP, OSPF, MPLS, IS-IS, VxLAN, EVPN, QoS, GRE, IPsec, DNS, and MACsec.
- 5+ years of experience in network operations.
- Skilled in network troubleshooting techniques and demonstrating creative problem-solving abilities.
- Strong track record of alert response within defined SLAs and Incident management.
- Experience with one or more of the following CSP environments: AWS, Azure, GCP, OCI.
- Familiarity with Arista, Fortinet and Juniper.
- Hands-on experience with contributing to tooling and automation for provisioning, monitoring, and managing complex network infrastructures.
- Bachelor’s degree in Computer Science, related technical field, or equivalent experience.
- Excellent verbal and written communication skills.
Ways To Stand Out From The Crowd:
- Solid understanding of Mellanox/Cumulus OS and Infiniband technology.
- Skilled in Unix/Linux system administration, with the ability to write and understand Python/Shell scripts to improve efficiency in hyperscale environments.
- Familiarity with leveraging tools such as Netbox/Nautobot, Prometheus, Grafana, Panoptes to monitor and manage a global network.
Compensation
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 136,000 USD - 224,250 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4.
Applications for this job will be accepted at least until July 19, 2026. This posting is for an existing vacancy.
Note: NVIDIA uses AI tools in its recruiting processes.