Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Communication @ 7
Data Analysis
Debugging
LLM @ 4
Machine Learning
PyTorch @ 3
Python @ 3
SGLang @ 4
vLLM @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
- Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
- Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
- Develop data analysis tooling and visualizations for numerics debugging
- Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction
- Participate in code reviews and incorporate feedback
Requirements
- Proficient in Python; familiarity with C++
- Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling
- Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time
- Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework
- Experience reading, modifying, or contributing to a large open-source codebase
- MS/PhD in Computer Science or related field, or equivalent experience
- 4+ years in a relevant software engineering role
- Demonstrated ability to move fast with ambiguous requirements, with strong written and verbal communication
Ways to stand out from the crowd
- Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development
- Track record of debugging numerical issues across mixed-precision boundaries
- Deep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsity
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Applied Scientist, Efficient LLM Inference & Model Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
System Software Engineer, Dynamo-Triton Inference Server - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
ML Solutions Architect (Early Talent)
Nebius · United States
USD 102,000-126,000 per year
Senior Performance Architect, Nemotron
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Software Engineer, RL Post-Training Frameworks
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - Dynamo-Triton Inference Server
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year