Senior MLOps Engineer - DSX Enablement

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Bash @ 6 CI/CD @ 3 CUDA @ 4 Communication @ 6 Data Engineering @ 7 Data Science @ 7 Deep Learning @ 4 Go @ 6 InfiniBand @ 4 Kubernetes @ 4 Leadership @ 6 Linux @ 4 MLOps @ 3 Machine Learning @ 7 NVLink @ 4 Networking @ 4 Observability @ 3 Python @ 6 Rust @ 6 Security @ 4

Details

NVIDIA is seeking a Senior MLOps Engineer to join the DSX Enablement team and collaborate closely with strategic customers to implement and enhance AI workloads. The team partners with innovative AI companies and open-source communities to address challenging technical problems.

Responsibilities

  • Develop innovative solutions that advance AI infrastructure capabilities.
  • Advise infrastructure experts on the demands of machine learning workloads.
  • Help practitioners diagnose and solve full-stack AI and machine learning system problems.
  • Build and deploy custom AI solutions on NeoCloud platforms and NVIDIA Cloud Partners, including distributed training, inference optimization, and MLOps pipelines.
  • Act as a primary technical contact for internal and external customers and partners, guiding joint engagements and ensuring the success of initiatives on DGX Cloud.
  • Solve complex production problems.
  • Collaborate with teams building infrastructure software and accelerated frameworks for AI applications.
  • Profile and tune large-scale training and inference workloads on NVIDIA Cloud Partner platforms to reduce latency, cost, and operational risk.
  • Develop open-source tools and reference architectures for building and managing machine learning and AI workloads, pipelines, and systems at scale.

Requirements

  • BS, MS, or Ph.D. in Computer Science, Computer or Electrical Engineering, a related technical field, or equivalent experience.
  • 8+ years of experience in technical roles such as data science, data engineering, or machine learning engineering, ideally involving large-scale production systems.
  • Demonstrated AI and machine learning experience across multiple phases of the machine learning lifecycle, from exploratory analysis through production systems.
  • Experience with Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale.
  • Solid scripting and programming skills in Bash and Python.
  • Systems programming skills in C++, Go, or Rust.
  • Experience using machine learning or deep learning frameworks for training and inference.
  • Excellent communication and technical presentation skills, including the ability to explain architectures, trade-offs, and recommendations to engineering and leadership audiences.
  • A strong record of engineering discipline and execution on individual and collaborative projects.

Preferred Qualifications

  • Experience contributing to and working in open-source communities.
  • Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, InfiniBand, NVLink, and RoCE.
  • Experience building machine learning systems in security-critical environments.
  • Experience with distributed training and inference frameworks.
  • Familiarity with cloud-native MLOps practices, including containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows.
  • Deep systems knowledge for diagnosing and fixing performance or correctness problems spanning hardware, networking, accelerators, hypervisors or operating systems, compilers or runtimes, application code, and libraries.

Benefits

NVIDIA offers competitive salaries, equity, and a generous benefits package. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Applications will be accepted at least until August 24, 2026. NVIDIA is an equal opportunity employer and is committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs