Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Agile @ 1
Algorithms
CUDA @ 4
Debugging @ 1
Deep Learning @ 1
GPU @ 1
GenAI
Generative AI
LLM
NCCL @ 4
Profiling @ 1
PyTorch @ 6
Python @ 1
SGLang @ 6
Software Development @ 6
vLLM @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Software Engineer specializing in deep learning inference. The role involves designing, building, and optimizing GPU-accelerated software for advanced AI applications, including high-performance deep learning frameworks such as SGLang and vLLM. You will help improve large-scale model serving and inference platforms and support the deployment of language models.
You will collaborate with the deep learning community to implement the latest algorithms for public release in SGLang, vLLM, and other deep learning frameworks. The work focuses on performance improvements for state-of-the-art large language models, multimodal models, and generative AI models across NVIDIA accelerators, including data center GPUs and edge SoCs. Technologies and tools include CUTLASS, OAI Triton, NCCL, CUDA kernels, FlashInfer, and NVIDIA inference libraries.
Responsibilities
- Optimize, analyze, and tune deep learning models across large language model, multimodal, and generative AI workloads.
- Scale deep learning model performance across different architectures and types of NVIDIA accelerators.
- Contribute features and code to NVIDIA inference libraries, vLLM, SGLang, FlashInfer, and other large language model software solutions.
- Collaborate with teams working on frameworks, NVIDIA libraries, and innovative inference optimization solutions.
Requirements
- Master's degree, PhD, or equivalent experience in a relevant field such as Computer Engineering, Computer Science, EECS, or AI.
- 5+ years of relevant software development experience.
- Excellent C/C++ programming and software design skills.
- Agile software development skills are helpful, and Python experience is a plus.
- Experience training, deploying, or optimizing deep learning model inference in production is a plus.
- Experience with performance modeling, profiling, debugging, code optimization, or CPU and GPU architecture is a plus.
Preferred Qualifications
- Contributions to deep learning software projects such as PyTorch, vLLM, or SGLang.
- Experience with multi-GPU communications, including NCCL or NVSHMEM.
- Experience building and shipping products to enterprise customers.
- GPU programming experience with CUDA, OAI Triton, or CUTLASS.
Benefits
NVIDIA offers competitive salaries, an extensive benefits package, and a work environment that promotes diversity, inclusion, and flexibility. NVIDIA is an equal opportunity employer committed to fostering a supportive and empowering workplace.
For Poland, the base salary range is PLN 221,250–383,500 for Level 3 and PLN 292,500–507,000 for Level 4.