Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Bash @ 3
CI/CD
DevOps @ 5
Docker @ 3
GPU @ 3
InfiniBand @ 3
JSON @ 3
Linux @ 6
Networking @ 6
Python @ 3
SRE @ 5
Software Development @ 6
gRPC @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an engineer to solve software integration challenges for next-generation data center platforms. The role supports GPU architectures and AI infrastructure projects involving high-speed communication, virtualization, Ethernet, and InfiniBand. You will provide first-tier support to R&D teams and help bridge hardware development with stable software deployments.
Responsibilities
- Fix and prioritize complex system issues during high-stakes bring-ups and proof-of-concept activities for next-generation computing architectures.
- Manage the integration of large-scale products involving GPUs, network stacks, firmware, and drivers.
- Create, recreate, and redeploy software artifacts.
- Fix code, update builds, and develop workarounds to unblock development.
- Serve as the primary technical point of contact for R&D teams addressing infrastructure and integration blockers.
- Work with R&D, verification, and DevOps teams to streamline CI/CD pipelines for specialized high-speed interconnect and system management projects.
- Lead technically during critical system failures in a fast-paced environment.
Requirements
- Bachelor's degree in Computer Science or a similar discipline, or equivalent experience.
- Software engineering experience and a strong understanding of software development methodologies, modern Linux-based operating systems, and computer networking.
- At least 5 years of experience in DevOps, SRE, or systems integration roles.
- In-depth knowledge of Linux distributions, including Ubuntu and RHEL.
- Experience with Docker containerization.
- Coding skills in C/C++, Python, and Bash for automation and system-level fixes.
- Experience with GitLab and GitLab CI for managing complex build pipelines.
- Ability to multitask and self-manage.
- Excellent problem-solving and critical-thinking skills.
Preferred Qualifications
- In-depth knowledge of high-performance networking, including InfiniBand and Ethernet.
- Practical experience with gRPC, gNMI, REST, and JSON for system management and telemetry.
- Experience working on large-scale hardware and software converged systems, such as rack-scale computing or GPU clusters.
Benefits
The role includes competitive compensation, equity, and benefits. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Engineer System Software, SDN Operations
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Software QA Test Development Engineer - Diagnostics
Nvidia · Santa Clara, United States
USD 140,000-270,200 per year
Senior Site Reliability Engineer, BCM - DGX Cloud
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Networking Verification Engineer
Nvidia · Austin, United States
USD 124,000-241,500 per year
Principal Engineer, Cloud Site Reliability Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer - NVLink Rack Scale Stability and Reliability
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year