Senior System Software Engineer, Agentic Retrieval

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Algorithms Communication @ 6 Docker @ 4 GPU GenAI Generative AI @ 4 Kubernetes @ 4 LLM @ 4 MLOps @ 4 Machine Learning NLP @ 4 Profiling Prompt Engineering Python Rust

Details

NVIDIA is looking for a System Software Engineer specializing in Agentic Retrieval to develop pipelines for indexing and querying multimodal content. The role focuses on Generative AI, LLMs, VLLMs, and Agentic Retrieval using NVIDIA hardware and software platforms to build powerful, flexible multimodal retrievers and agents driven by large language models.

The role is pivotal in accelerating containerized pipelines for high-quality multimodal datasets and improving retrieval efficacy. Responsibilities include deduplicating, filtering, and classifying training corpora; optimizing system cost, speed, and accuracy through micro-optimization, prompt engineering, fine-tuning, and applied research; and evaluating AI models and frameworks for acceleration and capability enhancement.

Responsibilities

  • Develop and optimize Rust-based data-processing frameworks for efficiently handling large datasets in GPU-accelerated environments for LLM training.
  • Lead the development and iterative optimization of Agentic Retrieval pipeline components, including GPU acceleration and model performance improvements for better total cost of ownership.
  • Collaborate with LLM and machine learning researchers to develop full-stack, GPU-accelerated data-preparation pipelines for multimodal models.
  • Implement benchmarking, profiling, and optimization of algorithms in Python across various system architectures, specifically for LLM applications.
  • Work with cross-functional teams to understand requirements, build and evaluate proofs of concept, and develop roadmaps for production tools and library features within the LLM ecosystem.
  • Build products that improve employee productivity through generative AI and copilot experiences.
  • Develop integrated systems that provide unified experiences across applications and generate insights for end-to-end user experiences.
  • Help build and maintain the Continuous Delivery pipeline to move changes to production faster and more safely while maintaining operational standards.
  • Perform peer reviews and provide feedback on performance, scalability, and correctness.
  • Contribute to the adoption of frameworks, standards, and new technologies.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • At least 6 years of experience in a similar or related role.
  • Experience delivering software in a cloud context and familiarity with managing cloud infrastructure.
  • Knowledge of MLOps technologies, including Docker Compose, containers, Kubernetes, and data center deployments.
  • In-depth hands-on understanding of NLP, LLMs, VLMs, Generative AI, and Agentic Retrieval workflows.
  • Self-motivation, enthusiasm for continuous learning, and a willingness to share findings across the team.
  • Excellent communication skills, including the ability to explain sophisticated topics clearly and effectively.
  • Ability to work successfully with multifunctional teams, principals, and architects across organizational boundaries and geographies.

Benefits

  • Equity and benefits are included.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

Applications for this job will be accepted at least until October 6, 2026. This posting is for an existing vacancy.

More jobs at Nvidia

Similar jobs