Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

at Nvidia
USD 152,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Algorithms @ 4 Communication @ 7 Data Structures @ 4 Distributed Systems GPU Go @ 4 Hiring @ 4 Kubernetes @ 7 Mathematics @ 4 Python @ 4 Slurm @ 7

Details

NVIDIA is hiring experienced software engineers to help scale up its AI Infrastructure. You will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications.

Responsibilities

  • You will be part of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads.
  • Designing and developing a massively distributed scalable platform which would be used to identify, diagnose and remediate non-performant GPU assets.
  • Working with teams across NVIDIA to ensure production AI clusters run reliability and consistently with maximum performance. Evaluating system failures and improving services based on a well-defined incident management process.

Requirements

  • Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work.
  • Highly motivated with strong communication skills; you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies.
  • 5+ years in similar role and experience on large-scale production systems. Experience with common software engineering principles, tools and techniques.
  • You possess a BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience.
  • Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.

Ways to stand out from the crowd

  • Technical competency in managing and automating large-scale distributed systems independent of cloud providers.
  • Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Base Command Manager).
  • Prior experience in asynchronous workflows and/or event driven architecture.
  • Proven operational excellence in maintaining reliable and performant infrastructure.

More jobs at Nvidia

Similar jobs