Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
GitHub
HPC
InfiniBand @ 4
Kubernetes @ 3
LLM
Linux @ 6
MPI @ 4
Networking @ 4
Performance Optimization @ 4
PyTorch @ 4
Python @ 4
Slurm @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is expanding its AI Networking software development and codesign team. The team develops and benchmarks full-stack systems for AI training and inference at data-center scale, builds automation and production tools, adopts community frameworks, and contributes tools to public GitHub repositories. The goal is to identify bottlenecks and improve the real-world performance of large-scale systems across hardware and software stacks.
Responsibilities
- Develop AI networking communication frameworks and production applications for large supercomputers and data centers.
- Develop production tools and benchmarks used by teams inside and outside NVIDIA.
- Enable new AI models within the benchmarking infrastructure and provide insights through end-to-end analysis of large-scale workloads across hardware and software stacks.
- Design and implement automation systems, including large-scale parameter searches to identify optimal configurations across complex systems.
- Collaborate with networking and hardware teams to co-design new features and software interfaces.
Requirements
- Bachelor's or master's degree in Computer Science, Software Engineering, or equivalent experience.
- More than 5 years of experience.
- Professional Python development experience, with a focus on building maintainable, long-lived tools.
- Solid Linux expertise and a willingness to work extensively in command-line environments.
- Ability and motivation to work across a broad, evolving stack, from hardware and networking to large-scale AI systems running across clusters.
Preferred Qualifications
- Knowledge of the modern AI ecosystem, including PyTorch, large language models, inference, and training.
- Familiarity with cluster orchestration systems such as Slurm or Kubernetes.
- Knowledge of MPI, high-performance computing, InfiniBand, Ethernet, and networking.
- Experience with performance optimization.
Benefits
- Equity and employee benefits.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications will be accepted at least until July 19, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Data Center Performance Engineer - Benchmarking and Optimization
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software SDET Test Development Engineer
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior HPC Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior HPC Performance Engineer
Nvidia · Germany
PLN 221,200-507,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year