Senior Full Stack Software Engineer - DGX Cloud

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Communication @ 7 Distributed Systems @ 6 GPU Go Hiring @ 4 JavaScript @ 6 Kubernetes @ 7 LLM @ 4 PostgreSQL React @ 6 SQL @ 6 Slurm @ 7 TypeScript @ 6

Details

NVIDIA is hiring experienced software engineers to help scale its AI infrastructure. The role focuses on cluster operations, operator development, node health monitoring, GPU resource scheduling, and building reliable infrastructure solutions for AI-based applications.

Responsibilities

  • Contribute to the DGX Cloud team responsible for production systems that enable large-scale GPU clusters for various AI workloads.
  • Design and develop a massively distributed, scalable platform to identify, diagnose, and remediate non-performant GPU assets.
  • Work with teams across NVIDIA to ensure production AI clusters run reliably and consistently at maximum performance.
  • Evaluate system failures and improve services through a well-defined incident management process.
  • Work across the product stack, including React, Web Components, TypeScript, Golang, PostgreSQL, Temporal, Bazel, and Kubernetes.

Requirements

  • Direct experience in a software engineering role within a highly technical organization, with demonstrable impact from previous work.
  • Strong communication skills and the ability to work with multifunctional teams, principals, and architects across organizational boundaries and geographies.
  • At least 5 years of experience in a similar role and experience with large-scale production systems.
  • Knowledge of common software engineering principles, tools, and techniques.
  • A bachelor's degree in Computer Science or Engineering, or equivalent experience.
  • At least 6 years of full-stack engineering experience.
  • At least 3 years of experience building and shipping consumer-facing products.
  • Proficiency in React, TypeScript or JavaScript, and Golang.
  • Proficiency with a SQL database.

Preferred Qualifications

  • Technical competency in managing and automating large-scale distributed systems independently of cloud providers.
  • Advanced hands-on experience and deep understanding of cluster management systems such as Kubernetes, Slurm, and Base Command Manager.
  • Empathy for users, attention to detail, and a passion for creating world-class user experiences.
  • Experience with asynchronous workflows and/or event-driven architecture.
  • Proven operational excellence in maintaining reliable and performant infrastructure.
  • Understanding of responsible LLM usage and the risks of blindly consuming generated output.

Compensation and Benefits

  • Base salary range: $184,000–$287,500 USD for Level 4.
  • Base salary range: $224,000–$356,500 USD for Level 5.
  • Eligibility for equity and benefits.
  • Applications will be accepted at least until August 3, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs