Senior Linux Kernel Systems Software Engineer – CSP Engagements

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI CUDA @ 4 Communication @ 4 Debugging @ 4 Deep Learning @ 4 GPU @ 4 HPC @ 6 Kubernetes @ 4 Linux @ 4 Machine Learning Marketing Performance Analysis @ 4 Performance Optimization @ 7 Python @ 7 Software Development @ 4

Details

NVIDIA is seeking a Senior Software Engineer to join the CSP Engagements team, focusing on system software for data center products such as GB200. This role combines deep technical expertise in embedded firmware, Linux kernel development, and middleware development with customer-facing responsibilities to enable cloud service providers with next-generation computing platforms. You will work at the intersection of hardware and software, driving technical solutions from concept through deployment.

Responsibilities

  • Design and develop software solutions for data center servers, including Linux kernel modifications, device drivers, and system optimizations for GB200 and next-generation platforms.
  • Lead hardware bring-up activities, BSP development, and hardware-software co-design for Cloud Service Provider deployments.
  • Partner directly with CSPs to deliver technical solutions, co-develop and co-debug features and optimizations, and provide support during new product introductions.
  • Collaborate with cross-functional teams to design end-to-end solutions spanning firmware, operating systems, middleware, and applications, with a focus on AI/ML and HPC workloads.
  • Perform advanced system debugging, root cause analysis, and performance optimization for large-scale data center environments.
  • Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation.
  • Work with internal leaders in Software, Hardware, Firmware, Marketing, and Operations to support the successful introduction of next-generation GPU/CPU-based products.
  • Use knowledge of driver, firmware, diagnostics, and software-stack development processes to help keep complex projects on track.

Requirements

  • Deep expertise in data center server architectures, HPC systems, and hardware-software co-design.
  • Expert knowledge of Linux kernel internals and device drivers.
  • Knowledge of communication protocols, including PCIe, USB, and Ethernet.
  • Deep understanding of computer architecture and microprocessor concepts.
  • Expert knowledge of ARM (AArch64) and x86 architectures.
  • Ability to debug kernel crash dumps and lock-up issues.
  • Hands-on experience with GDB, kdump, and eBPF tracing for debugging multiprocessor systems.
  • Good understanding of ARM and Intel assembly.
  • Strong understanding of PCIe virtualization and IOMMU.
  • Deep understanding of NUMA architectures, including memory topology, processor-memory locality, and performance optimization for multi-CPU systems in data center environments.
  • Strong programming skills in C/C++ and Python.
  • Experience with virtualization, Kubernetes, and cloud-native architectures.
  • Experience with complex system-level debugging, performance analysis, and test design.
  • Bachelor’s or master’s degree in Computer Engineering, Computer Science, or a related field, or equivalent experience.
  • 10+ years of system software development experience.

Preferred Qualifications

  • Experience with GPU computing, CUDA, and deep learning workloads.
  • Expertise in out-of-band and in-band management architectures.
  • Knowledge of memory fabric and CXL architectures.

Benefits

  • Equity eligibility.
  • Employee benefits.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Applications for this job will be accepted at least until August 3, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs