Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Bash @ 6
CI/CD @ 3
CUDA @ 4
Communication @ 6
Data Engineering @ 7
Data Science @ 7
Deep Learning @ 4
Go @ 6
InfiniBand @ 4
Kubernetes @ 4
Leadership @ 6
Linux @ 4
MLOps @ 3
Machine Learning @ 7
NVLink @ 4
Networking @ 4
Observability @ 3
Python @ 6
Rust @ 6
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior MLOps Engineer to join the DSX Enablement team and collaborate closely with strategic customers to implement and enhance AI workloads. The team partners with innovative AI companies and open-source communities to address challenging technical problems.
Responsibilities
- Develop innovative solutions that advance AI infrastructure capabilities.
- Advise infrastructure experts on the demands of machine learning workloads.
- Help practitioners diagnose and solve full-stack AI and machine learning system problems.
- Build and deploy custom AI solutions on NeoCloud platforms and NVIDIA Cloud Partners, including distributed training, inference optimization, and MLOps pipelines.
- Act as a primary technical contact for internal and external customers and partners, guiding joint engagements and ensuring the success of initiatives on DGX Cloud.
- Solve complex production problems.
- Collaborate with teams building infrastructure software and accelerated frameworks for AI applications.
- Profile and tune large-scale training and inference workloads on NVIDIA Cloud Partner platforms to reduce latency, cost, and operational risk.
- Develop open-source tools and reference architectures for building and managing machine learning and AI workloads, pipelines, and systems at scale.
Requirements
- BS, MS, or Ph.D. in Computer Science, Computer or Electrical Engineering, a related technical field, or equivalent experience.
- 8+ years of experience in technical roles such as data science, data engineering, or machine learning engineering, ideally involving large-scale production systems.
- Demonstrated AI and machine learning experience across multiple phases of the machine learning lifecycle, from exploratory analysis through production systems.
- Experience with Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale.
- Solid scripting and programming skills in Bash and Python.
- Systems programming skills in C++, Go, or Rust.
- Experience using machine learning or deep learning frameworks for training and inference.
- Excellent communication and technical presentation skills, including the ability to explain architectures, trade-offs, and recommendations to engineering and leadership audiences.
- A strong record of engineering discipline and execution on individual and collaborative projects.
Preferred Qualifications
- Experience contributing to and working in open-source communities.
- Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, InfiniBand, NVLink, and RoCE.
- Experience building machine learning systems in security-critical environments.
- Experience with distributed training and inference frameworks.
- Familiarity with cloud-native MLOps practices, including containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows.
- Deep systems knowledge for diagnosing and fixing performance or correctness problems spanning hardware, networking, accelerators, hypervisors or operating systems, compilers or runtimes, application code, and libraries.
Benefits
NVIDIA offers competitive salaries, equity, and a generous benefits package. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 24, 2026. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Nvidia · Warsaw, Poland
PLN 221,200-507,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year