Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
AWS @ 4
Azure @ 4
Communication @ 7
Debugging @ 4
Docker @ 4
GPU
Go @ 4
Grafana @ 6
HPC @ 4
Kibana @ 6
Kubernetes @ 4
Linux @ 4
Machine Learning
Prometheus @ 6
Python @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a systems software engineer to design and build next-generation architecture and software for managing storage services that support GPU design, VLSI, corporate, and AI/ML teams. The role involves building self-service capabilities and highly available infrastructure services supporting NVIDIA users 24/7/365.
Responsibilities
- Define, build, and manage an enterprise software engineering platform for storage infrastructure and services using enterprise appliances, networks, and open-source technologies.
- Design and expand REST APIs used by thousands of engineers for on-demand storage management and workflow support.
- Build and integrate provisioning, metrics, monitoring, and software for storage service management workflows.
- Develop tooling to automate deployment and management of large-scale design-storage environments.
- Automate operational monitoring and alerting and enable self-service consumption of resources.
- Document procedures and practices, perform technology evaluations, and coordinate and track system orders, installations, and deployments.
Requirements
- Bachelor's degree in Computer Science or equivalent experience with 8 or more years of relevant experience; master's degree with 5 or more years of experience; or Ph.D. with 3 years of experience.
- Extensive experience building and owning large-scale, multithreaded, distributed backend systems.
- Experience designing and building REST APIs in Python or Go.
- Experience with containerization and orchestration tools such as Docker and Kubernetes.
- Experience with cloud infrastructure, including AWS, Azure, or Google Cloud.
- Background with telemetry stacks such as Grafana, Prometheus, AlertManager, and Kibana.
- Strong collaboration and communication skills, including the ability to guide and influence others in a dynamic matrix environment.
Preferred Qualifications
- Experience working with open-source software, including building, debugging, patching, and contributing code.
- Experience solving Linux storage-related problems.
- Experience designing, deploying, and managing enterprise NAS solutions such as NetApp and Pure Storage, distributed file systems such as Lustre, and S3 storage.
- Experience with HPC cluster management tools such as Slurm, PBS, or LSF.
Compensation and Benefits
The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 168,000–270,250 for Level 4 and USD 200,000–322,000 for Level 5. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 14, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior Platform Security Engineer – Device Trust, Attestation, and Secure Browser
Nvidia · Santa Clara, United States
USD 196,000-310,500 per year
Senior Staff Software Engineer - Enterprise AI Platform
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
System Software Engineering Intern, GPU - 2027
Nvidia · Poland
PLN 117,800-204,100 per year
Senior Storage Platform Engineer
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Senior Compute Platform Engineer, LSF - EDA Infrastructure
Nvidia · United States
USD 184,000-356,500 per year
Similar jobs
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year