Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
API @ 4
Agentic AI @ 6
Claude Code @ 6
Codex @ 6
Communication @ 6
Deep Learning @ 4
GPU @ 7
GenAI
Generative AI
LLM @ 4
Marketing
OSS
Profiling @ 4
PyTorch @ 4
SGLang
Software Development @ 4
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Inference Benchmarking team focuses on optimizing inference server performance for large language models (LLMs). This role involves pushing the boundaries of GPU hardware and software performance across technologies such as disaggregated serving, data parallel attention, mixture-of-experts (MoE), Qwen3.5, DeepSeek, and GPT-OSS.
Responsibilities
- Characterize workloads for the latest LLMs and inference servers, including vLLM, SGLang, and TRT-LLM.
- Collaborate with the performance marketing team to create technical content, including blog posts and InferenceX updates.
- Work with engineers from AI startup companies to establish standard benchmarking methodologies.
- Develop and maintain an evolving inference performance data results website.
- Create end-to-end profiling and analysis tools for generative AI workloads.
- Contribute to deep learning software projects such as PyTorch, TRT-LLM, vLLM, and SGLang.
- Verify that new GPU product launches deliver industry-leading performance.
- Collaborate with software, research, and product teams to guide the direction of inference serving.
- Use current coding agents and inference technologies to improve team efficiency.
Requirements
- Master's or PhD degree in Computer Science, Computer Engineering, a related field, or equivalent experience.
- 6 or more years of relevant software development experience.
- Detailed knowledge of deep learning inference serving, PyTorch programming, profiling, and compiler optimizations.
- Experience developing client-server LLM applications using the OpenAI API or MCP and identifying performance bottlenecks.
- Strong understanding of CPU and GPU microarchitecture and performance characteristics.
- Experience with complex software projects such as frameworks, compilers, or operating systems.
- Proficiency with AI coding agents such as Claude Code, Codex, and Cursor.
- Excellent written and verbal communication skills, with the ability to work independently and collaboratively in a fast-paced environment.
Preferred Qualifications
- A demonstrated drive to continuously improve software and hardware performance.
- Examples of novel workplace use cases for agentic AI tools.
- Experience with databases and visualization tools.
Compensation
The base salary depends on location, experience, and the pay of employees in similar positions. The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Employees are also eligible for equity and benefits.
Applications will be accepted at least until August 15, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.