Senior Distributed Software Engineer, Golang - DGX Cloud

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI Communication @ 7 Debugging @ 4 Deep Learning Distributed Systems @ 7 GPU @ 6 Go Kubernetes @ 7 Machine Learning Software Development @ 6

Details

NVIDIA is seeking a Senior Software Engineer to develop distributed storage services for AI/ML. The role focuses on designing and building reliable, scalable, and efficient storage-as-a-service solutions tailored to AI applications. These services must be deployable anywhere and scale without limitations, supporting NVIDIA's business across graphics drivers, autonomous vehicles, and deep learning frameworks.

Responsibilities

  • Lead the overall architecture and design of a distributed storage service optimized for AI/ML.
  • Develop and maintain robust, scalable distributed Go programs deployed to open-source ecosystems, including Kubernetes.
  • Develop and maintain user-space applications, containers, Go bindings, and CLI tools.
  • Build features that improve availability and reliability for large-scale distributed storage deployments.
  • Collaborate with NVIDIA Research, Computing, Product teams, cross-functional teams, and external customers to deliver cloud services.
  • Automate the distributed storage service end to end, including deployment, management, and monitoring.

Requirements

  • Bachelor's degree in Computer Science or a related field, or equivalent experience.
  • At least 8 years of industry experience.
  • Strong background developing distributed systems with Golang, Kubernetes, and cloud service provider integrations.
  • Proven track record delivering distributed services in a variety of distributed computing environments.
  • Experience implementing storage services and interfaces that provide scalable, high-performance, and reliable solutions.
  • Experience owning product delivery from inception through support.
  • Experience developing and maintaining enterprise software.
  • Experience deploying, managing, and debugging applications in Kubernetes environments.
  • Strong communication and presentation skills.

Preferred Qualifications

  • Experience architecting, building, and deploying distributed services running on large-scale clusters ranging from multi-petabyte to exabyte scale and supporting millions of users.
  • Ownership of all software development and delivery lifecycle stages.
  • Interest in accelerated computing environments and technologies such as GPU Direct Storage, DPU, and RDMA.
  • Experience building and delivering cloud services, particularly distributed systems.

Benefits

  • Equity and benefits are available.
  • NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
  • Applications will be accepted at least until August 30, 2026.
  • NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs