Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
AWS @ 6
Azure @ 6
CI/CD
Communication @ 6
Distributed Systems @ 7
GCP @ 6
GPU @ 4
Go
Kubernetes @ 7
Networking @ 7
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The DGX Cloud organization at NVIDIA brings together cutting-edge hardware and software innovation to deliver industry-leading accelerated computing for the world's most adventurous AI workloads. The team is dedicated to solving major challenges, driving advancements, and impacting millions of lives worldwide.
The organization is seeking an outstanding Senior Systems Software Engineer with deep experience in distributed systems, open-source technologies such as Kubernetes and containers, and systems performance and scalability. The ideal candidate has broad, end-to-end experience across the stack, from GPU operators and device plugins to distributed inference serving and cloud platforms, along with the technical depth to investigate and resolve real-world problems at scale. This role focuses on scaling AI infrastructure, optimizing total cost of ownership, and reducing cost per token to support the next generation of AI innovation and AI factories.
Responsibilities
- Drive end-to-end performance and scale characterization for the NVIDIA DGX Cloud software stack, from Kubernetes control and data planes through NVIDIA components such as GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving.
- Collaborate with AI researchers, developers, and customers to develop automated tests that simulate real user workloads using custom-built and open-source tools and frameworks.
- Investigate performance and scale issues in complex distributed systems, including interactions between Kubernetes and the NVIDIA software stack, to identify and resolve root causes.
- Design and develop monitoring, reporting, and analysis tools for performance and scale testing across software, GPU, and CPU resources.
- Triage, debug, and identify root causes of issues related to operating Kubernetes clusters at ultra-large scale, ensuring reliability and efficiency.
- Build and maintain a high-velocity framework for continuous, always-on performance and scale testing through a modern CI/CD pipeline.
- Document research, methodologies, and results clearly and concisely, and present findings at internal and external venues, including KubeCon and GTC.
- Engage with upstream communities, including Kubernetes, CNCF, and NVIDIA open-source projects, to validate AI workload performance and scalability and help shape design and development decisions.
Requirements
- 8+ years of experience in computer architecture, networking, storage systems, and accelerators.
- Bachelor's or master's degree in engineering, preferably electrical engineering, computer engineering, or computer science, or equivalent experience.
- Expertise in Kubernetes and familiarity with related CNCF projects.
- Experience working with large-scale parallel and distributed accelerator-based systems.
- Expertise optimizing performance and AI workloads on large-scale systems.
- Experience with performance modeling and benchmarking at scale.
- Proficiency in Golang and Python.
- Experience with the NVIDIA software ecosystem in both training and inference domains.
- Expertise with at least one public cloud service provider, such as GCP, AWS, Azure, or OCI.
Preferred Qualifications
- Strong operational experience with a Kubernetes distribution.
- Experience scaling Kubernetes clusters to ultra-large node and object counts.
- Demonstrated history of working in the open-source community.
- Excellent communication and interpersonal abilities.
- PhD in a relevant area.
Compensation
For Poland, the base salary range is 292,500 PLN–507,000 PLN for Level 4 and 375,000 PLN–650,000 PLN for Level 5. Base salary is determined by location, experience, and the pay of employees in similar positions.