Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 3 CUDA @ 4 Debugging @ 7 GPU @ 4 Linux @ 7 Machine Learning Microservices NCCL NVLink Performance Optimization @ 4 Profiling @ 4 Python @ 6 TensorRT @ 4

Details

DGX Station is NVIDIA's next-generation personal AI supercomputer, built on the NVIDIA Grace Blackwell GB300 Superchip. It is a deskside workstation designed to provide data-center-class AI capabilities to researchers, developers, and AI engineers.

This hands-on role owns full-stack operating system enablement for DGX Station, with a primary focus on Windows and strong coverage of Linux. The engineer will work across NVIDIA's GPU driver, CUDA, firmware, BMC, and AI software teams, collaborate with Microsoft and ODM/OEM partners, and ensure a polished, production-ready experience across both operating systems.

Responsibilities

  • Own end-to-end Windows enablement for DGX Station, from initial bring-up through WHQL certification and customer-ready shipping quality.
  • Drive Linux bring-up and continuous enablement on DGX OS and Ubuntu, including kernel module integration, device tree and ACPI configuration, systemd services, initramfs, and DKMS packaging.
  • Partner with DGX OS and kernel teams to land platform support upstream and in NVIDIA's distribution.
  • Enable and validate BIOS/UEFI, BMC, and system-level firmware for Windows and Linux on the Grace Arm and Blackwell GB300 architecture.
  • Ensure ACPI tables, SMBIOS, Secure Boot, measured boot, power management, and hardware abstraction layers work correctly on both operating systems.
  • Coordinate GPU, display, and compute driver bring-up and validation on Windows using WDDM and MCDM, and on Linux using open GPU kernel modules and DRM/KMS.
  • Work with NVIDIA's driver team and Microsoft to resolve compatibility issues, achieve WHQL certification, and ensure driver stability across Windows Update and Linux kernel revisions.
  • Ensure the CUDA Toolkit, cuDNN, TensorRT, NCCL, and NVIDIA AI SDK stack function correctly on DGX Station under Windows and Linux.
  • Validate AI and deep-learning workload performance, including training, fine-tuning, and inference, and resolve platform gaps on the Arm and GB300 architecture.
  • Validate NVIDIA AI applications, including NIM microservices, NemoClaw, AI Workbench, and developer tools.
  • Define and implement test plans covering single-user and multi-user scenarios, container runtimes, application installation flows, and developer workflows.
  • Drive functional, stress, power, thermal, sleep/resume, S-state cycle, Windows Update, Linux kernel-upgrade compatibility, and long-duration reliability testing.
  • Own bug triage and resolution across firmware, BMC, driver, and operating system layers.
  • Serve as the primary technical interface with Microsoft and ODM/OEM partners, coordinating schedules and resolving cross-company technical blockers.
  • Profile and optimize boot time, GPU compute throughput, NVLink-C2C and memory bandwidth utilization, power efficiency, and thermal behavior.
  • Create and maintain bring-up guides, known-issue documentation, driver compatibility matrices, recovery and re-imaging procedures, and developer setup instructions.
  • Enable field and support teams for customer deployments.

Requirements

  • Bachelor's or master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 12 or more years of confirmed experience in systems software engineering, with deep expertise in Windows platform enablement, driver development, or operating system integration, along with hands-on experience bringing up Linux on new hardware platforms.
  • Strong hands-on experience with Windows internals, including kernel-mode drivers, ACPI, power management, Secure Boot, UEFI, WDM/WDF driver frameworks, and the WHQL certification process.
  • Solid understanding of Linux platform enablement, including kernel modules, device tree and ACPI on Arm, systemd, initramfs, DKMS, and packaging for Ubuntu or DGX OS.
  • Experience with GPU, display, or compute driver stacks on Windows and/or Linux.
  • Familiarity with DirectX, WDDM, DRM/KMS, and GPU compute APIs is a strong plus.
  • Experience enabling hardware platforms, including bring-up, driver integration, validation, and certification for shipping products on Windows and Linux.
  • Strong debugging and root-cause analysis skills across firmware, driver, and operating system boundaries.
  • Experience with WinDbg, kernel debugging using kd or kgdb/crash, crash dump analysis, ftrace/ETW, and performance profiling tools.
  • Ability to coordinate across GPU driver, CUDA, firmware, BMC, and AI software teams, as well as external partners such as Microsoft and ODM/OEMs.
  • Proficiency in C/C++ and Python.
  • Experience with Arm architecture is a plus.

Preferred Qualifications

  • Experience with Windows on Arm platforms, including driver enablement, performance optimization, or application compatibility.
  • Hands-on experience with CUDA, TensorRT, or AI/ML frameworks on Windows and Linux, especially on Arm and NVIDIA GPU systems.
  • Experience working with OEM/ODM partners or silicon vendors on Windows and Linux platform certification for workstation- or server-class hardware.
  • Track record of shipping workstation or server hardware products from bring-up through general availability with Windows and Linux support.
  • Experience with BMC, Redfish, out-of-band management, or platform manageability software on high-end workstations or servers.
  • Experience with GPU-accelerated applications, including AI training and inference, content creation tools, or scientific computing.

Benefits

  • Equity and employee benefits are provided.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs