Senior Software Engineer - Cluster Networking

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Cloudflare Communication @ 6 GPU Go @ 6 HPC @ 4 InfiniBand @ 4 Kubernetes @ 7 Linux @ 7 Machine Learning Networking @ 7 Python @ 6 Slurm @ 3

Details

NVIDIA's Managed AI Research Superclusters (MARS) team builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop next-generation AI/ML systems. The platform runs frontier model training across tens of thousands of GPUs on multiple clouds and is scaling toward clusters of ten thousand nodes and beyond.

The team is seeking a senior networking engineer to lead the network architecture of GPU superclusters and own how clusters communicate internally and externally, including the CNI data plane, overlay mesh, and nodes connecting clusters across regions and providers.

Responsibilities

  • Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale.
  • Design, operate, and scale overlay networks, including CNI, mesh and VPN topologies using Tailscale and WireGuard, as well as gateways connecting control and data planes.
  • Design, operate, and scale L7 gateways, load balancers, and tunnels using Envoy and Cloudflare.
  • Identify and eliminate scale ceilings, including packet loss under load, control-plane saturation, IP address management exhaustion, and failure modes that emerge above a few thousand nodes.
  • Build scale-test environments and validation suites to detect networking regressions before they reach production.
  • Diagnose complex problems across the stack, including issues where symptoms in Slurm or training jobs trace back to mark collisions, stale routes, or saturated tunnels.
  • Partner with cloud and neocloud providers on network topology, requirements, and capabilities when bringing up new clusters.
  • Provide senior technical judgment to a distributed team, with deep expertise in cluster networking.

Requirements

  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 6+ years of professional experience in systems, network, or infrastructure software engineering.
  • Deep knowledge of Kubernetes networking architecture and CNI standards; production experience operating Calico is strongly preferred.
  • Proficiency designing and maintaining modern mesh and VPN networking topologies, including Tailscale, WireGuard, or equivalent technologies.
  • Strong Linux networking fundamentals, including routing, netfilter, iptables/nftables, packet marking, network namespaces, and their interaction with container runtimes.
  • Demonstrated ability to debug distributed network problems at scale using packet capture, tracing, and correlation of behavior across many hosts to identify root causes.
  • Proficiency in Go, Python, C, or a comparable systems programming language.
  • Clear written and verbal communication skills and the ability to work effectively with engineers across multiple time zones.

Preferred Qualifications

  • Direct experience architecting and operating massive-scale Kubernetes topologies across thousands of concurrent nodes.
  • Experience with high-performance fabrics in AI or HPC environments, including InfiniBand, RoCE, or RDMA over converged networks.
  • Upstream contributions to Calico, Cilium, Tailscale, or Kubernetes networking SIGs.
  • Experience operating networking across multiple public clouds and on-premises environments simultaneously.
  • Familiarity with Slurm or other HPC schedulers running on Kubernetes.

Compensation and Benefits

The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until September 19, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs