Principal Software Engineer, AI Networking

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 6 Debugging @ 6 Distributed Systems @ 7 GPU InfiniBand @ 7 Leadership @ 8 NCCL @ 6 Networking @ 7 Technical Leadership @ 8

Details

NVIDIA is advancing AI, computer graphics, and accelerated computing through innovative GPU and networking technologies.

Join NVIDIA as a Principal Software Engineer to lead the transformation of AI networking systems. You will apply deep technical expertise to manage complex customer engagements and help shape product and architecture direction.

Responsibilities

  • Lead the technical strategy for AI Factory networking deployments at strategic customers, including architecture reviews, risk assessments, and multi-phase execution plans.
  • Serve as the principal-level technical authority for embedded networking products such as BlueField and ConnectX, as well as the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and InfiniBand.
  • Lead deep technical engagements with hyperscalers and AI Factory customers, covering design-in, coding, bring-up, performance tuning, failure analysis, and production hardening.
  • Partner with internal engineering, product, and architecture teams to translate customer needs into product features, reference architectures, tooling, and guidelines.
  • Drive performance, reliability, and debuggability improvements across customer stacks.
  • Translate technical findings into actionable product, firmware, and software roadmap items.

Requirements

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 15+ years of relevant industry experience, including technical leadership across complex systems.
  • Deep knowledge of networking protocols and distributed systems, including RoCE, InfiniBand, L1–L4 fundamentals, and performance and latency tradeoffs.
  • Low-level software expertise with proficiency in C/C++ and experience debugging across firmware, driver, and user-space environments.
  • Experience in high-performance networking and system-level debugging, including packet drops, retransmissions, congestion, QoS, ordering, and buffer management.
  • Excellent interpersonal skills, with the ability to explain complex topics to engineers, product managers, and customer collaborators and align cross-organizational teams toward decisions.

Preferred Qualifications

  • Customer-facing technical leadership experience with hyperscalers, cloud service providers, AI factories, or similarly complex production environments.
  • Hands-on expertise with DPDK, DOCA, RDMA verbs, NCCL, CUDA-aware networking, congestion control, and performance tuning at scale.
  • Experience building internal tools, telemetry, and automation to improve triage speed and operational excellence.
  • Demonstrated innovation through patents, publications, hackathons, rapid prototyping, or shipping new architectures or features end to end.
  • Experience leading multi-team initiatives across geographic regions and time zones, including influencing without authority.
  • Experience using AI-powered tools to accelerate debugging, documentation, and engineering efficiency while maintaining sound engineering judgment.

Benefits

  • Competitive salary with a base salary range of USD 272,000–431,250 per year.
  • Equity and benefits.
  • Inclusive and equal-opportunity work environment.

Applications for this job will be accepted at least until July 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs