Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Ansible @ 4
BGP @ 4
Cumulus Linux @ 4
Debugging @ 4
Distributed Systems @ 8
GDPR @ 6
GPU @ 4
Grafana @ 4
IaC
InfiniBand @ 6
Linux @ 4
Machine Learning
Networking @ 8
Observability
Prometheus @ 4
SRE @ 6
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
GeForce NOW is a cloud gaming platform that uses NVIDIA data centers to stream games at high resolution and frame rates across devices, including smartphones and VR headsets.
The Senior Manager will lead the design, scaling, and operations of high-performance networking for GPU-based cloud infrastructure. The role supports cloud gaming workloads, AI/ML training, and inference platforms by delivering ultra-low-latency, high-throughput, and highly reliable interconnects across data centers and cloud environments.
Responsibilities
- Build and mentor a specialized team of network architects focused on high-performance GPU infrastructure.
- Oversee intra-cluster and inter-cluster connectivity using RoCE, Ethernet-based AI fabrics, and high-bandwidth data center interconnects.
- Drive technical tuning to reduce latency and jitter, increase throughput, and implement congestion-control and packet-loss mitigation strategies.
- Define networking roadmaps supporting gaming, AI/ML training, and real-time inference at scale.
- Engage with ISPs to optimize low-latency edge networks and ensure seamless connections between data centers and end clients.
- Implement Infrastructure as Code and observability frameworks for automated provisioning, scaling, and real-time cluster health monitoring.
- Work with AI platform teams, hardware vendors, and SRE groups to influence technology direction and vendor selection.
- Establish fault-tolerance protocols and lead incident response and root-cause analysis for complex network issues.
Requirements
- 12+ years of experience in networking, cloud infrastructure, or distributed systems, including 5+ years directly managing technical teams.
- Mastery of data center networking, including Clos/spine-leaf architectures and high-performance fabrics such as RDMA, RoCE, or InfiniBand.
- Hands-on experience with BGP, EVPN/VXLAN, and kernel-level development for routing and switching.
- Experience using Ansible or Terraform for infrastructure automation and monitoring tools such as Prometheus and Grafana.
- Practical experience designing large-scale configurations using SR-IOV, Xen virtualization, or Open vSwitch.
- Bachelor's or Master's degree in Computer Science or a related engineering field, or equivalent experience.
- Ability to ensure infrastructure meets rigorous internal policies and regulatory standards such as GDPR.
Preferred Qualifications
- Experience managing networking for large-scale GPU clusters or hyperscale cloud environments.
- Familiarity with optical networking and high-speed 400G or 800G interconnects.
- Experience debugging and improving code for Mellanox/Cumulus Linux or managing Palo Alto and Netscaler appliances.
- Strong understanding of streaming telemetry and operational signals, including SNMP and Syslog.
- Relevant certifications such as CCIE or specialized cloud networking designations.
Benefits
NVIDIA offers competitive salaries, equity, and a generous benefits package. The base salary range is USD 256,000–414,000, determined by location, experience, and the pay of employees in similar positions.
More jobs at Nvidia
Senior Software Engineer, Unified Access Management Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AV System Integration
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
Senior System Software Engineer, Automotive Performance
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Embedded System Software Engineer – Platform Execution Lead
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year