Senior AI Engineer, NeMo Retriever - Model Optimization and MLOps

at Nvidia

📍 Santa Clara, United States

USD 184,000-356,500 per year

SENIOR

✅ On-site

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ^?

Docker @ 4 Kubernetes @ 4 Python @ 4 Machine Learning @ 4 MLOps @ 4 Helm @ 4 Performance Optimization @ 4 Microservices @ 4 API @ 4 NLP @ 8 LLM @ 4 PyTorch @ 4 OpenAPI @ 4 GPU @ 4

Details

NVIDIA's technology is at the heart of the AI revolution, touching people across the planet by powering everything from self-driving cars and robotics to co-pilots and more. Join us at the forefront of technological advancement in intelligent assistants and information retrieval.

NVIDIA NIM provides containers to self-host GPU-accelerated inferencing microservices for pre-trained and customized AI models across clouds, data centers, RTX™ AI PCs, and workstations. NIM microservices expose industry-standard APIs for simple integration into AI applications, development frameworks, and workflows. Built on pre-optimized inference engines from NVIDIA and the community, including NVIDIA TensorRT and TensorRT-LLM, NIM microservices optimize response latency and throughput for each combination of foundation model and GPU.

NVIDIA NeMo Retriever is a collection of NIMs for building multimodal extraction, re-ranking, and embedding pipelines with high accuracy and maximum data privacy. It delivers quick, context-aware responses for AI applications like advanced retrieval-augmented generation (RAG) and Agentic AI workflows. The NeMo Retriever team is looking for an AI Engineer to join our team, focusing on the intersection of machine learning development, performance optimization, and MLOps. This role requires a unique blend of technical expertise in ML model development, system optimization, and operational excellence. We are looking for someone with a passion for working with the world’s most complicated problems in Generative AI, LLM, MLLM, and RAG spaces using our innovative hardware and software platforms. You will leverage and augment existing tools that enable building NIMs, which power flexible, multi-modal retrievers and agents. If you’re creative & passionate about solving real-world conversational AI problems, come join us.

Responsibilities

Develop and maintain NIMs that containerize optimized models using OpenAPI standards using Python or an equivalent performant language.
Work closely with partner teams to understand requirements, build & evaluate POCs, and develop roadmaps for production-level tools.
Enable development of integrated systems - AI Blueprints that provide a unified, turnkey experience.
Help build and maintain our Continuous Delivery pipeline with the goal of moving changes to production faster and safer while ensuring key operational standards.
Provide peer reviews to other specialists, including feedback on performance, scalability, and correctness.

Requirements

Bachelor’s or Master’s Degree program in Computer Science, Computer Engineering, or a related field (or equivalent experience).
8+ years of demonstrated experience in a similar or related role.
Python programming expertise with Deep Learning (DL) frameworks such as PyTorch.
Experience delivering software in a cloud context and is familiar with the patterns and processes of handling cloud infrastructure.
Knowledge of MLOps technologies such as Docker-Compose, Containers, Kubernetes, Helm, data center deployments, etc.
Familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
Excellent in-depth hands-on understanding of NLP, LLM, MLLM, Generative AI , and RAG workflows.
Self-starter with a passion for growth, enthusiasm for continuous learning, and sharing findings across the team.
Extremely motivated, highly passionate, and curious about new technologies.

Benefits

With competitive salaries and a generous benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. Due to unprecedented growth, our exclusive engineering teams are rapidly growing. The base salary range is 184,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. You will also be eligible for equity and benefits.