Used Tools & Technologies
Not specified
Required Skills & Competences ?
System Administration @ 4 Linux @ 4 Python @ 4 Machine Learning @ 4 Bash @ 4 Communication @ 7 Mathematics @ 4 GPU @ 4Details
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to solve, that only we can tackle, and that matter to the world. This is our life’s work, to amplify human imagination and intelligence.
We are looking for an outstanding hands-on architect/engineer for a Senior HPC architect role to support deployment and bringup of large-scale GPU compute clusters. You will enable deployment and bring-up of large-scale GPU compute clusters, provide at-scale system administration and tuning mechanisms for large-scale compute runs, and work with accelerated computing and deep learning software and hardware platforms. You will collaborate with scientific researchers, developers, customers, and specialists in HPC, OS, GPU compute, and systems to architect, develop and bring up large scale performance platforms.
Responsibilities
- Provide engineering solutions to operationalize the latest GPU computing products and software stacks.
- Ensure technical relationships with internal and external engineering teams and assist systems, machine learning/deep learning engineers in building creative solutions based on NVIDIA technology.
- Act as an internal reference for system administration, at-scale system analysis, and other datacenter and large-scale GPU-accelerated system solutions within the NVIDIA technical community.
- Support deployment, bring-up, tuning, and operationalization of large-scale GPU compute clusters and related workflows.
Requirements
- 5+ years of experience using accelerated computing for datacenter/HPC-based enterprise computing solutions.
- Solid understanding of accelerated computing scheduling and I/O stacks.
- Programming/scripting experience in C/C++/Python/Bash.
- Experience working with engineering or academic research communities supporting high performance computing or deep learning.
- Experience with parallel filesystems.
- Strong teamwork and communication skills (verbal and written).
- Ability to multitask effectively in a dynamic environment.
- Action-driven with strong analytical skills.
- Desire to be involved in multiple diverse and innovative projects.
- BS (or equivalent experience) in Engineering, Mathematics, Physics, or Computer Science. MS or PhD desirable.
Ways to stand out
- Deep learning framework skills.
- Exposure to using and deploying telemetry and visualization pipelines.
- Exposure to container technology and Linux performance tools.
Compensation & Additional Information
- Base salary ranges by level are provided: Level 3: 148,000 USD - 235,750 USD; Level 4: 184,000 USD - 287,500 USD. Base salary will be determined based on location, experience, and comparable pay.
- Eligible for equity and benefits.
- Applications accepted at least until July 29, 2025.
Equal Opportunity
NVIDIA is committed to fostering a diverse work environment and is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.