Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud

at Nvidia
📍 Spain
PLN 292,500-650,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 6 AWS @ 6 Azure @ 6 CI/CD Communication @ 6 Distributed Systems @ 7 GCP @ 6 GPU @ 4 Go Kubernetes @ 7 Networking @ 7 Python @ 6

Details

The DGX Cloud organization at NVIDIA brings together cutting-edge hardware and software innovation to deliver industry-leading accelerated computing for the world's most adventurous AI workloads. The team is dedicated to solving major challenges, driving advancements, and impacting millions of lives worldwide.

The organization is seeking an outstanding Senior Systems Software Engineer with deep experience in distributed systems, open-source technologies such as Kubernetes and containers, and systems performance and scalability. The ideal candidate has broad, end-to-end experience across the stack, from GPU operators and device plugins to distributed inference serving and cloud platforms, along with the technical depth to investigate and resolve real-world problems at scale. This role focuses on scaling AI infrastructure, optimizing total cost of ownership, and reducing cost per token to support the next generation of AI innovation and AI factories.

Responsibilities

  • Drive end-to-end performance and scale characterization for the NVIDIA DGX Cloud software stack, from Kubernetes control and data planes through NVIDIA components such as GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving.
  • Collaborate with AI researchers, developers, and customers to develop automated tests that simulate real user workloads using custom-built and open-source tools and frameworks.
  • Investigate performance and scale issues in complex distributed systems, including interactions between Kubernetes and the NVIDIA software stack, to identify and resolve root causes.
  • Design and develop monitoring, reporting, and analysis tools for performance and scale testing across software, GPU, and CPU resources.
  • Triage, debug, and identify root causes of issues related to operating Kubernetes clusters at ultra-large scale, ensuring reliability and efficiency.
  • Build and maintain a high-velocity framework for continuous, always-on performance and scale testing through a modern CI/CD pipeline.
  • Document research, methodologies, and results clearly and concisely, and present findings at internal and external venues, including KubeCon and GTC.
  • Engage with upstream communities, including Kubernetes, CNCF, and NVIDIA open-source projects, to validate AI workload performance and scalability and help shape design and development decisions.

Requirements

  • 8+ years of experience in computer architecture, networking, storage systems, and accelerators.
  • Bachelor's or master's degree in engineering, preferably electrical engineering, computer engineering, or computer science, or equivalent experience.
  • Expertise in Kubernetes and familiarity with related CNCF projects.
  • Experience working with large-scale parallel and distributed accelerator-based systems.
  • Expertise optimizing performance and AI workloads on large-scale systems.
  • Experience with performance modeling and benchmarking at scale.
  • Proficiency in Golang and Python.
  • Experience with the NVIDIA software ecosystem in both training and inference domains.
  • Expertise with at least one public cloud service provider, such as GCP, AWS, Azure, or OCI.

Preferred Qualifications

  • Strong operational experience with a Kubernetes distribution.
  • Experience scaling Kubernetes clusters to ultra-large node and object counts.
  • Demonstrated history of working in the open-source community.
  • Excellent communication and interpersonal abilities.
  • PhD in a relevant area.

Compensation

For Poland, the base salary range is 292,500 PLN–507,000 PLN for Level 4 and 375,000 PLN–650,000 PLN for Level 5. Base salary is determined by location, experience, and the pay of employees in similar positions.

More jobs at Nvidia

Similar jobs