Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
CUDA @ 4
Communication @ 6
Customer Support
Debugging
HPC @ 4
InfiniBand @ 4
Linux @ 7
MPI @ 4
NCCL @ 4
NVLink
Networking @ 4
Python @ 4
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's NVIDIA Enterprise Experience (NVEX) Solutions Engineering team is seeking a senior computer or software engineer to establish expertise in networking technologies used in AI clusters. The role connects customer support teams and R&D, focusing on solving complex production issues involving high-speed interconnect technologies such as InfiniBand, NVLink, and Spectrum-X.
The successful candidate will have a software development background in the networking industry, experience with production network operations, and the ability to root-cause customer issues down to the source-code level, primarily using C and Python. Relevant areas include network operating systems, Linux network drivers and internals, network hardware, NIC software, SmartNICs, DPUs, embedded firmware, Software-Defined Networking, and infrastructure management technologies. The role involves investigating IPC, race conditions, finite state machines, event-processing loops, queue management, network traffic and flow analysis, and software design gaps.
Responsibilities
- Assist network and AI cluster support teams in reproducing, resolving, and root-causing sophisticated customer issues.
- Work with R&D teams to develop bug fixes, workarounds, and solutions for critical customers using NVIDIA networking technologies.
- Become an authority in NVIDIA networking technologies used in AI clusters, including InfiniBand, NVLink, and Spectrum-X.
- Analyze network performance metrics and make tuning recommendations for high-performance, lossless networks.
- Develop support and analysis tools to investigate and root-cause field issues.
- Use AI tools for software development, log and trace analysis, and source-code debugging.
- Occasionally work on weekends or holidays to support customers.
- Interact with internal and external customers and provide detailed explanations of findings.
Requirements
- Bachelor's degree in Computer, Electrical, or Software Engineering, or equivalent experience.
- At least 8 years of experience programming in C on Linux and embedded systems.
- Proficiency in Python.
- At least 8 years of experience developing software for one or more of the following:
- Linux NIC drivers
- Switch ASICs and SDKs
- Embedded network-device firmware
- Linux-based network equipment, including routers, switches, and gateways
- Network operating systems
- Virtual routers
- Software-Defined Networking stacks
- Virtual switching
- DPDK
- SR-IOV stacks
- At least 3 years of experience directly supporting end customers, partners, or integrators for network equipment and infrastructure.
- Strong system software expertise, including firmware, BIOS, kernels, drivers, and operating systems.
- Professional communication skills, including the ability to adjust communication to the audience's technical level and remain calm and focused in negative situations.
- Passion for learning innovative technologies and working on advanced products.
Preferred Qualifications
- Background in AI infrastructure and high-performance computing networking.
- Experience programming switch and NIC ASICs and SDKs.
- Experience with InfiniBand or other non-Ethernet networking technologies.
- Experience developing or supporting DPUs or SmartNICs.
- Knowledge of HPC performance-testing tools and NVIDIA AI stacks, including NCCL, MPI, DOCA, and CUDA.
Benefits
- Base salary determined by location, experience, and compensation for similar positions.
- Base salary range of USD 168,000–270,250 for Level 4 or USD 200,000–322,000 for Level 5.
- Equity and benefits.
- Comprehensive benefits package.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until August 11, 2026.