Senior Systems Engineer, Storage - DGX Cloud

at Nvidia
USD 208,000-414,000 per year
SENIOR
✅ Remote

Tech Stack

Algorithms @ 6 Ansible @ 4 ArgoCD @ 4 CI/CD @ 4 Chef @ 4 Data Structures @ 6 Debugging @ 6 Distributed Systems @ 6 Docker @ 4 GPU Git @ 4 Go @ 6 Grafana @ 4 Helm Java @ 6 Kubernetes @ 4 Linux @ 6 Observability @ 4 OpenStack @ 4 Performance Monitoring Prometheus @ 4 Puppet @ 4 Python @ 6 Terraform @ 4

Details

Systems Engineering focuses on building, automating, and operating platforms and tooling that deliver large-scale production systems with high efficiency, reliability, and velocity. The role combines software and systems engineering practices across infrastructure automation, containerized platforms, storage, telemetry, and observability.

NVIDIA's team ensures that internal and external GPU cloud services are deployed reliably, observable end-to-end, and continuously improved through automation. The team uses repeatable CI/CD pipelines, Kubernetes-based deployments, capacity and performance monitoring, self-service tooling, blameless postmortems, proactive failure-mode identification, and iterative improvement.

Responsibilities

  • Design, deploy, and operate Kubernetes solutions for large-scale storage and data platforms, including manifests, Helm charts, and operators.
  • Build tools, services, and automation that improve the lifecycle of storage and data systems, from provisioning and configuration through deployment, scaling, and day-two operations.
  • Develop and operate telemetry and observability for production systems, including metrics, logging, tracing, dashboards, and alerting.
  • Diagnose and resolve complex issues across distributed, containerized infrastructure using strong analytical troubleshooting skills.
  • Collaborate with peers and partner teams to improve services from inception and design through deployment, operation, and refinement.
  • Scale systems sustainably through automation, infrastructure-as-code, and CI/CD.
  • Support services before launch through deployment automation, capacity planning, and launch and readiness reviews.
  • Practice sustainable incident response and postmortems.
  • Participate in an on-call rotation supporting production systems.

Requirements

  • Bachelor's degree or equivalent experience in Computer Science or a related technical field involving coding.
  • 12 or more years of practical experience.
  • Hands-on experience deploying, configuring, and operating production workloads and solutions on Kubernetes.
  • Experience building tools and services for storage, data, or platform infrastructure.
  • Solid software design fundamentals, including algorithms, data structures, and complexity analysis, on large-scale Linux-based systems.
  • Experience building and operating telemetry and observability using tools such as Prometheus, InfluxDB, Grafana, and the Elastic stack.
  • Strong analytical troubleshooting skills with a systematic, root-cause-driven approach.
  • Proficiency in one or more of Python, Go, or Java.
  • Knowledge of infrastructure configuration management and infrastructure-as-code tools such as Ansible, Chef, Puppet, ArgoCD, Git Pipelines, and Terraform.

Preferred Qualifications

  • Customer-first mindset with a focus on customer satisfaction and customer success.
  • Experience with Git, code review, pipelines, and CI/CD.
  • Experience using or operating large private and public cloud systems based on Kubernetes, OpenStack, and Docker.
  • Interest in designing, analyzing, debugging, and fixing large-scale distributed systems.
  • Experience designing storage- or data-focused tooling and automating its operation at scale.
  • Ability to thrive in collaborative environments and adapt to different working styles.

Compensation and Benefits

The base salary range is USD 208,000–333,500 for Level 5 and USD 256,000–414,000 for Level 6. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until July 31, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs