Senior Engineer, NCX

at Nvidia
📍 Germany
PLN 292,500-650,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 CUDA @ 4 DevOps @ 7 Distributed Systems @ 7 GPU @ 4 Go @ 4 Grafana @ 4 HPC InfiniBand @ 4 Kubernetes @ 7 Linux @ 7 NVLink @ 4 Networking @ 7 Observability @ 7 OpenTelemetry @ 4 Prometheus @ 4 Python @ 4 SRE

Details

NVIDIA is hiring an NCX Senior Engineer passionate about NVIDIA Cloud Partner (NCP) infrastructure operations to join the DSX team. This highly technical, hands-on role focuses on NVIDIA accelerated computing, cloud infrastructure, distributed systems, and production operations. You will work closely with strategic NVIDIA Cloud Partners to build and improve the operational capabilities required to run large-scale NVIDIA accelerated infrastructure reliably in production.

The role involves guiding partners beyond initial cluster deployment and validation into advanced Day 2 operations, including infrastructure health, observability, lifecycle management, rapid remediation, performance validation, and operational readiness. You will engage directly with partner engineering and operations teams to develop consistent approaches that support NVIDIA workloads and external customer environments.

Responsibilities

  • Lead NCP Day 2 operational readiness efforts by collaborating with NVIDIA Cloud Partners to establish systems, procedures, automation, and operational methods for managing NVIDIA accelerated infrastructure after deployment and activation.
  • Build continuous infrastructure validation methods for GPU, CPU, storage, and network health across large-scale AI clusters, identifying degraded infrastructure before it affects training or inference workloads.
  • Establish observability and operational telemetry across compute, GPU, InfiniBand/RoCE networking, storage, Kubernetes, and AI workloads, including monitoring, alerting, dashboards, and operational signals.
  • Develop automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service while minimizing disruption to customer workloads.
  • Refine fleet lifecycle administration strategies for sizable GPU fleets, including NVIDIA driver and firmware lifecycle management, Kubernetes node maintenance, OS patching, configuration management, upgrades, and configuration drift detection.
  • Translate NVIDIA NCP requirements and reference architectures into production operating practices, validation criteria, runbooks, automation, and measurable operational standards.
  • Define operational health and readiness through health signals, SLOs, key metrics, acceptance criteria, and ongoing validation mechanisms.
  • Build reusable operational frameworks, including tooling, automation, implementation guides, runbooks, operational playbooks, and reference implementations for multiple NCP environments.

Requirements

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
  • 8+ years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • Deep understanding of Kubernetes, containers, cluster scheduling, and the operational lifecycle of large multi-node environments.
  • Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and operations guided by service-level agreements.
  • Experience creating automation for infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
  • Strong networking fundamentals and experience troubleshooting complex distributed systems across compute, network, and storage layers.
  • Programming and automation experience using Python, Go, shell scripting, or similar languages.

Preferred Qualifications

  • Experience managing extensive GPU or accelerated computing infrastructure supporting AI training and inference workloads.
  • Experience with NVIDIA technologies such as DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, Network Operator, or related NVIDIA infrastructure software.
  • Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or large service-provider infrastructures.
  • Experience operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency.
  • Extensive knowledge of infrastructure observability tools, including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines.
  • Experience translating reference architectures or infrastructure requirements into repeatable production operating models across multiple customer or partner environments.
  • Knowledge of failure modes related to large distributed AI workloads and the infrastructure features needed to support extended training and production inference.

NVIDIA is leading developments in Artificial Intelligence, High-Performance Computing, and Visualization. The company is seeking people to help accelerate the next wave of artificial intelligence.

The base salary is determined based on location, experience, and the pay of employees in similar positions. For Poland, the base salary range is 292,500 PLN–507,000 PLN for Level 4 and 375,000 PLN–650,000 PLN for Level 5.

More jobs at Nvidia

Similar jobs