NCX Senior Engineer

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CUDA @ 4 DevOps @ 7 Distributed Systems @ 7 GPU @ 4 Go @ 4 Grafana @ 4 InfiniBand @ 4 Kubernetes @ 7 Linux @ 7 NVLink @ 4 Networking @ 7 Observability @ 7 OpenTelemetry @ 4 Prometheus @ 4 Python @ 4 SRE

Details

NVIDIA is hiring an NCX Senior Engineer passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join the DSX team. This highly technical, hands-on role focuses on helping strategic NVIDIA Cloud Partners build and improve the operational capabilities required to run large-scale NVIDIA-accelerated infrastructure reliably in production.

The role supports partners beyond initial cluster deployment and validation into advanced Day 2 operations, including infrastructure health, observability, lifecycle management, remediation, performance validation, and operational readiness. You will work directly with partner engineering and operations teams across NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations.

Responsibilities

  • Lead NCP Day 2 operational readiness efforts by establishing systems, procedures, automation, and operational methods for NVIDIA-accelerated infrastructure.
  • Develop continuous validation methods for GPU, CPU, storage, and network health across large-scale AI clusters.
  • Help implement telemetry, monitoring, alerting, dashboards, and operational signals across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads.
  • Build automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service while minimizing disruption to customer workloads.
  • Implement scalable GPU fleet lifecycle strategies, including NVIDIA driver and firmware administration, Kubernetes node maintenance, OS patching, configuration management, upgrades, and configuration drift detection.
  • Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable standards.
  • Define health signals, SLOs, metrics, acceptance criteria, and ongoing validation mechanisms for infrastructure reliability and service readiness.
  • Develop reusable tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations for multiple NCP environments.

Requirements

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.
  • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and service-level agreement-driven operations.
  • Experience developing automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
  • Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers.
  • Programming and automation experience using Python, Go, shell scripting, or similar languages.

Preferred Qualifications

  • Experience managing extensive GPU or accelerated-computing infrastructure supporting AI training and inference workloads.
  • Experience with NVIDIA technologies, including DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.
  • Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or large service-provider infrastructure teams.
  • Experience operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency.
  • Knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.
  • Experience translating reference architectures or infrastructure requirements into repeatable production operating models across customer or partner environments.
  • Knowledge of failure modes affecting large distributed AI workloads and the infrastructure features required to support extended training and production inference.

Compensation and Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Salary is determined by location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 27, 2026. NVIDIA is an equal opportunity employer.

More jobs at Nvidia

Similar jobs