Systems Software Engineer, Accelerated Kubernetes Performance And Scale - New College Grad 2026
at Nvidia
USD 108,000-195,500 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 5
API
AWS @ 3
Azure @ 3
CI/CD
GCP @ 3
GPU
Go
Kubernetes @ 5
Networking @ 3
OSS @ 5
Performance Optimization @ 3
Python @ 5
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is transforming computer graphics, PC gaming, and accelerated computing, and is now tapping into the potential of AI to define the next era of computing.
The DGX Cloud organization at NVIDIA brings together cutting-edge hardware and software innovation to deliver accelerated computing for AI workloads.
Responsibilities
- Work on end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack (control and data planes), including NVIDIA components such as GPU Operator, Network Operator, node-feature-discovery, topograph, dra-driver-nvidia-gpu, and nvsentinel—tracking issues from orchestration down to the metal.
- Design and contribute upstream architectural changes to the Kubernetes control plane and related projects to enable reliable operation at hyperscale cluster sizes.
- Improve container startup and cold-start latency to enable smooth, low-latency inference scaling on Kubernetes across thousands of GPU nodes, ensuring the AI runtime stack scales without creating API server pressure or operational fragility.
- Assess, improve, and contribute to open-source projects that make Kubernetes an outstanding platform for AI workloads (for example, Grove and gateway-api-inference-extension), composing their architectures with scalability, resilience, and multi-node training/inference in mind.
- Advance scalability and performance of confidential containers (CoCo) on Kubernetes so encrypted inference workloads meet stringent efficiency and latency requirements in production.
- Use DSX and related large-scale simulation infrastructure to model full AI-factory deployments and validate scalability across thousands of simulated GPUs, catching failures that emerge only at scale before hardware arrives.
- Collaborate with AI researchers, developers, customers, and upstream communities to design automated, at-scale workload tests (including replay of production agent traces), build monitoring/analysis tooling, and integrate continuous performance and scale testing into modern CI/CD workflows.
- Document methods and results clearly and present findings internally and at industry events (for example, KubeCon, GTC), while engaging with upstream groups (Kubernetes SIG Scalability, CNCF, and NVIDIA OSS communities) to influence and validate AI workload performance and scalability directions.
Requirements
- Recent graduate of a Bachelor’s, Master’s, or PhD degree in Engineering or equivalent experience, ideally in Electrical, Computer Engineering, or Computer Science.
- Experience in computer architecture, networking, storage systems, and accelerator-based platforms.
- Expertise in Kubernetes and familiarity with the broader CNCF ecosystem.
- Experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads.
- Experience with performance modeling and benchmarking for large-scale systems.
- Proficiency in Golang and/or Python.
- Familiarity with the NVIDIA software stack across training and inference.
- Experience with at least one major public cloud provider (for example, AWS, Azure, GCP, or OCI).
Ways to stand out from the crowd
- Strong operational experience with any one of the Kubernetes distributions.
- Prior experience scaling Kubernetes clusters to ultra-large node and object counts.
- Demonstrated history of working in the open-source community.
- Excellent communication and interpersonal abilities.
Compensation
- Base salary range: 108,000 USD - 178,250 USD for Level 1, and 124,000 USD - 195,500 USD for Level 2.
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Senior Software Engineer, Compute Sanitizer - GPU
Nvidia · United States
USD 184,000-356,500 per year
Senior Backend Platform Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Applied AI Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior System Software Engineer - AV Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Principal Software Engineer, Vehicle Dynamics Simulation - AV
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Similar jobs
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Systems Software Engineer, Accelerated Kubernetes Performance And Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Software Engineer, Infrastructure Security
OpenAI · United States, San Francisco, United States, New York City, United States, Seattle, United States
USD 230,000-385,000 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Systems Software Engineer, Kubernetes Node Lifecycle - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Attestation Services - DGX Cloud
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year