Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
at Nvidia
USD 184,000-356,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
API
AWS @ 6
Azure @ 6
CI/CD
Communication @ 6
Distributed Systems @ 7
GCP @ 6
GPU
Go
Kubernetes @ 3
Networking @ 7
Performance Optimization @ 7
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's DGX Cloud organization delivers accelerated computing infrastructure for large-scale AI workloads. The team is seeking a Senior Systems Software Engineer with deep expertise in distributed systems, Kubernetes, containers, systems performance, and scalability. The role focuses on scaling AI infrastructure, minimizing total cost of ownership, reducing cost per token, and enabling future AI innovation and AI factories.
Responsibilities
- Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack, including control and data planes.
- Analyze NVIDIA components such as GPU Operator, Network Operator, node-feature-discovery, topograph, dra-driver-nvidia-gpu, and nvsentinel, tracking issues from orchestration down to the hardware.
- Design and contribute upstream architectural changes to the Kubernetes control plane and related projects for reliable operation at hyperscale cluster sizes.
- Improve container startup and cold-start latency for low-latency inference scaling across thousands of GPU nodes while minimizing API server pressure and operational fragility.
- Assess, improve, and contribute to open-source projects supporting Kubernetes AI workloads, including Grove and gateway-api-inference-extension.
- Advance the scalability and performance of confidential containers (CoCo) on Kubernetes for encrypted inference workloads.
- Use DSX and related large-scale simulation infrastructure to model AI-factory deployments and validate scalability across thousands of simulated GPUs.
- Collaborate with AI researchers, developers, customers, and upstream communities to design workload tests, including production agent-trace replay.
- Build monitoring and analysis tooling and integrate continuous performance and scale testing into CI/CD workflows.
- Document methods and results and present findings internally and at industry events such as KubeCon and GTC.
- Engage with Kubernetes SIG Scalability, CNCF, and NVIDIA open-source communities.
Requirements
- Bachelor's or Master's degree in Engineering, or equivalent experience; ideally in Electrical Engineering, Computer Engineering, or Computer Science.
- 8+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms.
- Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
- Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads.
- Experience with performance modeling and benchmarking for large-scale systems.
- Proficiency in Golang and/or Python.
- Strong familiarity with the NVIDIA software stack across training and inference.
- Expertise with at least one major public cloud provider, such as AWS, Azure, GCP, or OCI.
Preferred Qualifications
- Strong operational experience with a Kubernetes distribution.
- Experience scaling Kubernetes clusters to ultra-large node and object counts.
- Demonstrated history of contributing to open-source communities.
- Excellent communication and interpersonal abilities.
- PhD or equivalent experience in a relevant area.
Compensation and Benefits
- Base salary range for Level 4: USD 184,000–287,500 per year.
- Base salary range for Level 5: USD 224,000–356,500 per year.
- Eligibility for equity and benefits.
- Hybrid work is preferred, with remote arrangements also available.
NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior DevTech Compute Engineer, Compression and Data Processing
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Platform Engineer, Design Automation
Nvidia · Santa Clara, United States
USD 196,000-368,000 per year
Senior Software Engineer, DGX Cloud Production Engineering
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior DFX Software Engineer - Machine Learning
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Technical Program Manager, AI Infrastructure and Capacity Operations
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale – DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Attestation Services – DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Senior Staff+ Software Engineer, Kubernetes Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year