Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Communication @ 7
Distributed Systems @ 6
GPU
Go
Hiring @ 4
JavaScript @ 6
Kubernetes @ 7
LLM @ 4
PostgreSQL
React @ 6
SQL @ 6
Slurm @ 7
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is hiring experienced software engineers to help scale its AI infrastructure. The role focuses on cluster operations, operator development, node health monitoring, GPU resource scheduling, and building reliable infrastructure solutions for AI-based applications.
Responsibilities
- Contribute to the DGX Cloud team responsible for production systems that enable large-scale GPU clusters for various AI workloads.
- Design and develop a massively distributed, scalable platform to identify, diagnose, and remediate non-performant GPU assets.
- Work with teams across NVIDIA to ensure production AI clusters run reliably and consistently at maximum performance.
- Evaluate system failures and improve services through a well-defined incident management process.
- Work across the product stack, including React, Web Components, TypeScript, Golang, PostgreSQL, Temporal, Bazel, and Kubernetes.
Requirements
- Direct experience in a software engineering role within a highly technical organization, with demonstrable impact from previous work.
- Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
- At least 5 years of experience in a similar role and experience with large-scale production systems.
- Knowledge of common software engineering principles, tools, and techniques.
- A bachelor's degree in Computer Science or Engineering, or equivalent experience.
- At least 6 years of full-stack engineering experience.
- At least 3 years of experience building and shipping consumer-facing products.
- Proficiency in React, TypeScript or JavaScript, and Golang.
- Proficiency with a SQL database.
Preferred Qualifications
- Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
- Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, and Base Command Manager.
- Empathy for users, attention to detail, and a passion for creating world-class user experiences.
- Experience with asynchronous workflows and/or event-driven architecture.
- Proven operational excellence in maintaining reliable and performant infrastructure.
- Understanding of responsible LLM usage and the risks of blindly consuming generated output.
Compensation and Benefits
- Base salary range: $184,000–$287,500 USD for Level 4.
- Base salary range: $224,000–$356,500 USD for Level 5.
- Eligibility for equity and benefits.
- Applications will be accepted at least until August 3, 2026.
- This posting is for an existing vacancy.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Software Engineer, Compute Sanitizer - GPU
Nvidia · Canada
CAD 170,000-275,000 per year
Senior Software Engineer, Networking DGX Cloud
Nvidia · United States
USD 200,000-391,000 per year
Senior Customer Program Manager – AI Platform
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Senior System Software Engineer – Data Center Compute Diagnostics
Nvidia · Durham, United States
USD 224,000-356,500 per year
Engineering Manager, AI Compiler Analysis
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Security Engineer, AI Security
Reddit · United States
USD 190,800-267,100 per year
Staff Software Engineer, Identity & Access Management
Reddit · United States
USD 217,000-303,900 per year
Senior Full-Stack Software Engineer – Verification Data and Visualization Platform
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Staff Engineer, Ads Business Manager
Reddit · United States
USD 217,000-303,900 per year
Staff Backend Engineer - Grafana Enterprise
Grafana Labs · United States
USD 175,000-210,000 per year