Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Algorithms @ 6
CUDA @ 7
Communication @ 7
GPU @ 7
HPC
LLM @ 4
Machine Learning @ 4
Mathematics
Parallel Programming @ 7
Prioritization @ 7
TensorRT @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a software engineer with a strong background in parallel processing and GPU architecture to push the limits of performance at the intersection of AI, high-performance computing, and financial markets. In this role, you will work with parallel algorithms, GPUs, and sophisticated systems to identify and eliminate bottlenecks and unlock the capabilities of advanced processing hardware.
You will collaborate with experts across industry and academia, influence next-generation platforms, and share insights with the global developer community.
Responsibilities
- Design and develop techniques to accelerate high-performance workloads at the intersection of AI, mathematics, and financial systems.
- Analyze, optimize, and scale complex AI and HPC workloads for modern CPU and GPU architectures.
- Profile and eliminate performance bottlenecks across the stack, from algorithms and kernels to system-level behavior.
- Publish and present work in conferences, talks, and blogs to educate and inspire the developer community.
- Influence the design of future hardware architectures, system software, libraries, and programming models by collaborating with NVIDIA research, hardware, compiler, and tools teams.
Requirements
- Strong hands-on experience with CUDA and parallel programming.
- Deep understanding of CPU and GPU architecture fundamentals and their impact on performance.
- Master's or PhD in Computer Science, Computer Engineering, Electrical and Computer Engineering, or a related field.
- Fluency in C/C++ and a solid foundation in algorithms and software design.
- At least 5 years of relevant work or research experience.
- Proven experience improving the performance of large-scale computational applications on GPUs.
- Excellent understanding of linear algebra.
- Strong communication and organizational skills, with a logical approach to problem-solving and solid prioritization abilities.
Preferred Qualifications
- Experience with inference optimization techniques and deploying optimized AI models in production.
- Experience with TensorRT, TensorRT-LLM, and cuTile.
- Experience parallelizing and optimizing machine learning methods such as decision trees, time-series models, and Monte Carlo simulations.
Compensation and Benefits
- Base salary range of USD 152,000–241,500 for Level 3.
- Base salary range of USD 184,000–287,500 for Level 4.
- Salary is determined based on location, experience, and compensation for employees in similar positions.
- Eligible for equity and benefits.
- Full-time position.
More jobs at Nvidia
Senior DevTech Compute Engineer, Compression and Data Processing
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
AI and FSI Developer Technology Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Performance Compiler Engineer - Triton
Nvidia · Redmond, United States
USD 184,000-287,500 per year
Senior HPC Performance Engineer - AI for Science at Scale
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior AI Developer Technology Engineer, Financial Sector
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year