Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Algorithms
Communication @ 6
Docker @ 4
GPU
GenAI
Generative AI @ 4
Kubernetes @ 4
LLM @ 4
MLOps @ 4
Machine Learning
NLP @ 4
Profiling
Prompt Engineering
Python
Rust
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a System Software Engineer specializing in Agentic Retrieval to develop pipelines for indexing and querying multimodal content. The role focuses on Generative AI, LLMs, VLLMs, and Agentic Retrieval using NVIDIA hardware and software platforms to build powerful, flexible multimodal retrievers and agents driven by large language models.
The role is pivotal in accelerating containerized pipelines for high-quality multimodal datasets and improving retrieval efficacy. Responsibilities include deduplicating, filtering, and classifying training corpora; optimizing system cost, speed, and accuracy through micro-optimization, prompt engineering, fine-tuning, and applied research; and evaluating AI models and frameworks for acceleration and capability enhancement.
Responsibilities
- Develop and optimize Rust-based data-processing frameworks for efficiently handling large datasets in GPU-accelerated environments for LLM training.
- Lead the development and iterative optimization of Agentic Retrieval pipeline components, including GPU acceleration and model performance improvements for better total cost of ownership.
- Collaborate with LLM and machine learning researchers to develop full-stack, GPU-accelerated data-preparation pipelines for multimodal models.
- Implement benchmarking, profiling, and optimization of algorithms in Python across various system architectures, specifically for LLM applications.
- Work with cross-functional teams to understand requirements, build and evaluate proofs of concept, and develop roadmaps for production tools and library features within the LLM ecosystem.
- Build products that improve employee productivity through generative AI and copilot experiences.
- Develop integrated systems that provide unified experiences across applications and generate insights for end-to-end user experiences.
- Help build and maintain the Continuous Delivery pipeline to move changes to production faster and more safely while maintaining operational standards.
- Perform peer reviews and provide feedback on performance, scalability, and correctness.
- Contribute to the adoption of frameworks, standards, and new technologies.
Requirements
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- At least 6 years of experience in a similar or related role.
- Experience delivering software in a cloud context and familiarity with managing cloud infrastructure.
- Knowledge of MLOps technologies, including Docker Compose, containers, Kubernetes, and data center deployments.
- In-depth hands-on understanding of NLP, LLMs, VLMs, Generative AI, and Agentic Retrieval workflows.
- Self-motivation, enthusiasm for continuous learning, and a willingness to share findings across the team.
- Excellent communication skills, including the ability to explain sophisticated topics clearly and effectively.
- Ability to work successfully with multifunctional teams, principals, and architects across organizational boundaries and geographies.
Benefits
- Equity and benefits are included.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Applications for this job will be accepted at least until October 6, 2026. This posting is for an existing vacancy.