Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 6
Debugging @ 6
Distributed Systems @ 7
GPU
InfiniBand @ 7
Leadership @ 8
NCCL @ 6
Networking @ 7
Technical Leadership @ 8
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is advancing AI, computer graphics, and accelerated computing through innovative GPU and networking technologies.
Join NVIDIA as a Principal Software Engineer to lead the transformation of AI networking systems. You will apply deep technical expertise to manage complex customer engagements and help shape product and architecture direction.
Responsibilities
- Lead the technical strategy for AI Factory networking deployments at strategic customers, including architecture reviews, risk assessments, and multi-phase execution plans.
- Serve as the principal-level technical authority for embedded networking products such as BlueField and ConnectX, as well as the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and InfiniBand.
- Lead deep technical engagements with hyperscalers and AI Factory customers, covering design-in, coding, bring-up, performance tuning, failure analysis, and production hardening.
- Partner with internal engineering, product, and architecture teams to translate customer needs into product features, reference architectures, tooling, and guidelines.
- Drive performance, reliability, and debuggability improvements across customer stacks.
- Translate technical findings into actionable product, firmware, and software roadmap items.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- 15+ years of relevant industry experience, including technical leadership across complex systems.
- Deep knowledge of networking protocols and distributed systems, including RoCE, InfiniBand, L1–L4 fundamentals, and performance and latency tradeoffs.
- Low-level software expertise with proficiency in C/C++ and experience debugging across firmware, driver, and user-space environments.
- Experience in high-performance networking and system-level debugging, including packet drops, retransmissions, congestion, QoS, ordering, and buffer management.
- Excellent interpersonal skills, with the ability to explain complex topics to engineers, product managers, and customer collaborators and align cross-organizational teams toward decisions.
Preferred Qualifications
- Customer-facing technical leadership experience with hyperscalers, cloud service providers, AI factories, or similarly complex production environments.
- Hands-on expertise with DPDK, DOCA, RDMA verbs, NCCL, CUDA-aware networking, congestion control, and performance tuning at scale.
- Experience building internal tools, telemetry, and automation to improve triage speed and operational excellence.
- Demonstrated innovation through patents, publications, hackathons, rapid prototyping, or shipping new architectures or features end to end.
- Experience leading multi-team initiatives across geographic regions and time zones, including influencing without authority.
- Experience using AI-powered tools to accelerate debugging, documentation, and engineering efficiency while maintaining sound engineering judgment.
Benefits
- Competitive salary with a base salary range of USD 272,000–431,250 per year.
- Equity and benefits.
- Inclusive and equal-opportunity work environment.
Applications for this job will be accepted at least until July 10, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Math Libraries Engineer - LLM Integration and Developer Experience
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Architect, Networking AI
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Manager, CUDA Driver
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Custom SoC IP Verification Engineer
Nvidia · Santa Clara, United States
USD 168,000-310,500 per year
Research Intern, Fundamental Generative AI - 2027
Nvidia · Santa Clara, United States
USD 38-94 per hour
Similar jobs
Senior Software Engineer, DGX Cloud AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Seattle, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Austin, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Networking
Nvidia · Austin, United States
USD 184,000-356,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Deep Learning Engineer – Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year