Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale – DGX Cloud

at Nvidia
USD 152,000-241,500 per year
SENIOR
✅ Remote ✅ Hybrid

Tech Stack

AI @ 7 API AWS @ 6 Azure @ 6 CI/CD Communication @ 6 Distributed Systems @ 7 GCP @ 6 GPU Go Kubernetes @ 3 Networking @ 6 Performance Optimization @ 7 Python @ 6

Details

NVIDIA’s DGX Cloud organization develops accelerated computing infrastructure for AI workloads. The team is seeking a Senior Systems Software Engineer with deep expertise in distributed systems, Kubernetes, containers, systems performance, and scalability. The role focuses on scaling AI infrastructure, minimizing total cost of ownership, reducing cost per token, and enabling future AI innovation and AI factories.

Responsibilities

  • Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack, including control and data planes and NVIDIA components such as GPU Operator, Network Operator, node-feature-discovery, topograph, dra-driver-nvidia-gpu, and nvsentinel.
  • Track performance and scalability issues from orchestration down to the underlying hardware.
  • Design and contribute upstream architectural changes to the Kubernetes control plane and related projects for reliable operation at hyperscale cluster sizes.
  • Improve container startup and cold-start latency for low-latency inference scaling across thousands of GPU nodes.
  • Assess, improve, and contribute to open-source projects supporting Kubernetes AI workloads, including Grove and gateway-api-inference-extension.
  • Advance the scalability and performance of confidential containers (CoCo) on Kubernetes for encrypted inference workloads.
  • Use DSX and related large-scale simulation infrastructure to model AI-factory deployments and validate scalability across thousands of simulated GPUs.
  • Collaborate with AI researchers, developers, customers, and upstream communities to design workload tests, including production agent-trace replay.
  • Build monitoring and analysis tooling and integrate continuous performance and scale testing into CI/CD workflows.
  • Document methods and results and present findings internally and at industry events such as KubeCon and GTC.
  • Engage with Kubernetes SIG Scalability, CNCF, and NVIDIA open-source communities.

Requirements

  • Bachelor’s or Master’s degree in Engineering, or equivalent experience, ideally in Electrical Engineering, Computer Engineering, or Computer Science.
  • 5+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms.
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
  • Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads.
  • Experience with performance modeling and benchmarking for large-scale systems.
  • Proficiency in Golang and/or Python.
  • Strong familiarity with the NVIDIA software stack across training and inference.
  • Expertise with at least one major public cloud provider, such as AWS, Azure, GCP, or OCI.

Preferred Qualifications

  • Strong operational experience with a Kubernetes distribution.
  • Experience scaling Kubernetes clusters to ultra-large node and object counts.
  • Demonstrated history of working in the open-source community.
  • Excellent communication and interpersonal abilities.
  • PhD or equivalent experience in relevant areas.

Benefits

  • Base salary range of USD 152,000–241,500, determined by location, experience, and comparable employee compensation.
  • Eligibility for equity and benefits.
  • NVIDIA offers a preference for hybrid work while remaining open to remote arrangements.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs