Senior Systems Software Engineer - Infrastructure

at Nvidia
USD 168,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 4 AWS @ 4 Azure @ 4 Communication @ 7 Debugging @ 4 Docker @ 4 GPU Go @ 4 Grafana @ 6 HPC @ 4 Kibana @ 6 Kubernetes @ 4 Linux @ 4 Machine Learning Prometheus @ 6 Python @ 4 Slurm @ 4

Details

NVIDIA is looking for a systems software engineer to design and build next-generation architecture and software for managing storage services that support GPU design, VLSI, corporate, and AI/ML teams. The role involves building self-service capabilities and highly available infrastructure services supporting NVIDIA users 24/7/365.

Responsibilities

  • Define, build, and manage an enterprise software engineering platform for storage infrastructure and services using enterprise appliances, networks, and open-source technologies.
  • Design and expand REST APIs used by thousands of engineers for on-demand storage management and workflow support.
  • Build and integrate provisioning, metrics, monitoring, and software for storage service management workflows.
  • Develop tooling to automate deployment and management of large-scale design-storage environments.
  • Automate operational monitoring and alerting and enable self-service consumption of resources.
  • Document procedures and practices, perform technology evaluations, and coordinate and track system orders, installations, and deployments.

Requirements

  • Bachelor's degree in Computer Science or equivalent experience with 8 or more years of relevant experience; master's degree with 5 or more years of experience; or Ph.D. with 3 years of experience.
  • Extensive experience building and owning large-scale, multithreaded, distributed backend systems.
  • Experience designing and building REST APIs in Python or Go.
  • Experience with containerization and orchestration tools such as Docker and Kubernetes.
  • Experience with cloud infrastructure, including AWS, Azure, or Google Cloud.
  • Background with telemetry stacks such as Grafana, Prometheus, AlertManager, and Kibana.
  • Strong collaboration and communication skills, including the ability to guide and influence others in a dynamic matrix environment.

Preferred Qualifications

  • Experience working with open-source software, including building, debugging, patching, and contributing code.
  • Experience solving Linux storage-related problems.
  • Experience designing, deploying, and managing enterprise NAS solutions such as NetApp and Pure Storage, distributed file systems such as Lustre, and S3 storage.
  • Experience with HPC cluster management tools such as Slurm, PBS, or LSF.

Compensation and Benefits

The base salary is determined by location, experience, and the pay of employees in similar positions. The base salary range is USD 168,000–270,250 for Level 4 and USD 200,000–322,000 for Level 5. The role is also eligible for equity and benefits.

Applications will be accepted at least until September 14, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs