Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
at Nvidia
📍 Germany
PLN 292,500-650,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
AWS
Azure
CI/CD
Distributed Systems @ 7
GCP
GPU
Go
Kubernetes @ 7
Networking @ 7
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Role overview
The DGX Cloud organization at NVIDIA brings together cutting-edge hardware and software innovation to deliver industry-leading accelerated computing for the world’s most adventurous AI workloads. We are looking for an outstanding Senior Systems Software Engineer with deep experience in distributed systems, open-source technologies such as Kubernetes and containers, and a strong background in systems performance and scalability.
In this pivotal role, you will help scale AI infrastructure while optimizing total cost of ownership, driving down cost per token to unlock the next generation of AI innovation and AI factories.
Responsibilities
- Drive end-to-end performance and scale characterization for the NVIDIA DGX Cloud software stack, from Kubernetes control and data planes through NVIDIA components such as GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving, following issues from orchestration down to the metal.
- Collaborate with AI researchers, developers and customers to develop innovative, automated tests that simulate real user workloads using custom-built and leading open-source tools and frameworks.
- Deep dive into performance and scale issues in complex distributed systems, including interactions between Kubernetes and the NVIDIA software stack, to identify and resolve root causes.
- Design and develop monitoring, reporting and analysis tools for performance and scale testing across software, GPU and CPU resources.
- Triage, debug and root cause issues related to operating Kubernetes clusters at ultra-large scale, ensuring reliability and efficiency.
- Build and maintain a high-velocity framework that enables continuous, always-on performance and scale testing via a modern CI/CD pipeline.
- Document research, methodologies and results clearly and concisely, and present findings at internal and external venues, including community conferences such as KubeCon and GTC.
- Engage efficiently with upstream communities—including Kubernetes, CNCF, and NVIDIA open-source projects—to validate performance and scalability of AI workloads early and help shape design and development decisions.
Requirements
- 8+ years of experience in Computer Architecture, Networking, Storage systems, Accelerators, and Bachelors/Masters in Engineering (preferably Electrical Engineering, Computer Engineering, or Computer Science) or equivalent experience.
- Expertise in Kubernetes and familiarity with related CNCF projects.
- Background in working with large scale parallel and distributed accelerator-based systems.
- Expertise optimizing performance and AI workloads on large scale systems.
- Experience with performance modeling and benchmarking at scale.
- Proficiency in Golang/Python.
- Background with the NVIDIA software ecosystem in both training and inference domains.
- Expertise with at least one of public CSP infrastructure (e.g., GCP, AWS, Azure, OCI).
Ways to stand out from the crowd
- Strong operational experience with any one of the Kubernetes distributions.
- Prior experience scaling Kubernetes clusters to ultra-large node and object counts.
- Demonstrated history of working in the open-source community.
- Excellent communication and interpersonal abilities.
- PhD in relevant areas.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Systems Software Engineer, Kubernetes Scale - DGX Cloud
Nvidia · Germany
PLN 176,200-305,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Software Engineer, Accelerated Kubernetes Performance And Scale - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · London, United Kingdom
GBP 325,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year