Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Debugging @ 6
Distributed Systems @ 7
GPU
InfiniBand @ 7
Leadership @ 8
NCCL @ 6
Networking @ 7
Technical Leadership @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is advancing AI, computer graphics, and accelerated computing through innovative GPU and networking technologies.
Join NVIDIA as a Principal Software Engineer to lead the transformation of AI networking systems. You will apply deep technical expertise to manage complex customer engagements and help shape product and architecture direction.
Responsibilities
- Lead the technical strategy for AI Factory networking deployments at strategic customers, including architecture reviews, risk assessments, and multi-phase execution plans.
- Serve as the principal-level technical authority for embedded networking products such as BlueField and ConnectX, as well as the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and InfiniBand.
- Lead deep technical engagements with hyperscalers and AI Factory customers, covering design-in, coding, bring-up, performance tuning, failure analysis, and production hardening.
- Partner with internal engineering, product, and architecture teams to translate customer needs into product features, reference architectures, tooling, and guidelines.
- Drive performance, reliability, and debuggability improvements across customer stacks.
- Translate technical findings into actionable product, firmware, and software roadmap items.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- 15+ years of relevant industry experience, including technical leadership across complex systems.
- Deep knowledge of networking protocols and distributed systems, including RoCE, InfiniBand, L1–L4 fundamentals, and performance and latency tradeoffs.
- Low-level software expertise with proficiency in C/C++ and experience debugging across firmware, driver, and user-space environments.
- Experience in high-performance networking and system-level debugging, including packet drops, retransmissions, congestion, QoS, ordering, and buffer management.
- Excellent interpersonal skills, with the ability to explain complex topics to engineers, product managers, and customer collaborators and align cross-organizational teams toward decisions.
Preferred Qualifications
- Customer-facing technical leadership experience with hyperscalers, cloud service providers, AI factories, or similarly complex production environments.
- Hands-on expertise with DPDK, DOCA, RDMA verbs, NCCL, CUDA-aware networking, congestion control, and performance tuning at scale.
- Experience building internal tools, telemetry, and automation to improve triage speed and operational excellence.
- Demonstrated innovation through patents, publications, hackathons, rapid prototyping, or shipping new architectures or features end to end.
- Experience leading multi-team initiatives across geographic regions and time zones, including influencing without authority.
- Experience using AI-powered tools to accelerate debugging, documentation, and engineering efficiency while maintaining sound engineering judgment.
Benefits
- Competitive salary with a base salary range of USD 272,000–431,250 per year.
- Equity and benefits.
- Inclusive and equal-opportunity work environment.
Applications for this job will be accepted at least until July 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Engineering Manager, Data Labeling Platform
Nvidia · Santa Clara, United States
USD 200,000-391,000 per year
Engineering Manager, Local AI Agents
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Deep Learning Software Engineer, DLSim
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, Fleet Intelligence Backend
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Business Systems Analyst
Nvidia · Santa Clara, United States
USD 144,000-270,200 per year
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Seattle, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Austin, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Austin, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Distinguished Software Architect - Deep Learning and HPC Communications
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year