Senior Infrastructure Engineer - Infrastructure Security and Core Services

at Nvidia
USD 208,000-333,500 per year
SENIOR
✅ On-site

Tech Stack

AI Compliance Data Science @ 4 GPU @ 4 Go @ 4 HPC Machine Learning Mathematics @ 4 Networking @ 6 OpenShift @ 4 Python @ 4 Ruby @ 4 Security @ 4 Statistics @ 4

Details

NVIDIA's Contract Manufacturing infrastructure team is seeking an experienced infrastructure engineer focused on infrastructure security, compute services, and storage for manufacturing workloads. The role involves redesigning core services at NVIDIA manufacturing sites, enabling new sites, and rapidly troubleshooting incidents where factory downtime can affect revenue. This is a hands-on engineering position focused on developing and deploying high-speed, resilient, and scalable core services for GPU-accelerated OT and IT environments.

Responsibilities

  • Lead the architecture, design, and deployment of global-scale manufacturing sites and their connectivity to AI factories and offices.
  • Architect and build CPU-based compute, storage, and GPU/HPC clusters.
  • Design high-performance OT and IT networks for NVIDIA and partner connectivity, supporting general compute workloads as well as GPU-dense AI/ML training and inference environments.
  • Partner with systems, operations, supply chain, OS, GPU, storage, and HPC product teams to deliver scalable, highly available network architectures and connectivity solutions.
  • Implement and refine compute, storage, security, telemetry, and performance-engineering practices to detect issues early and improve end-to-end application experience.
  • Manage infrastructure life-cycle activities and develop infrastructure designs for revenue-generating manufacturing sites.
  • Define and enforce security, compliance, and reliability standards for infrastructure supporting mission-critical manufacturing, R&D workloads, and NPI designs.
  • Collaborate with operations, manufacturing partners, and engineering teams to develop NVIDIA-on-NVIDIA reference architectures and best-practice solutions for large-scale compute and AI data designs using NVIDIA products such as Mellanox and GPUs.
  • Troubleshoot DNS, DHCP, and other components of the full connectivity stack.
  • Support test-engineering topologies and infrastructure on the factory floor.

Requirements

  • MS or PhD in Electrical Engineering, Computer Science, Computer Engineering, Artificial Intelligence, Data Science, Mathematics, Statistics, or equivalent experience.
  • 12 or more years of experience building, managing, and supporting large-scale hybrid networks.
  • Experience developing infrastructure automation pipelines using Python, Ruby, Go, or other infrastructure automation languages.
  • Expertise in networking technologies, especially Mellanox.
  • Experience with compute technologies including Dell and OpenShift, cloud compute services, and storage services including Pure and NetApp.
  • Experience designing compute and storage architectures for data centers, offices, manufacturing environments, and labs.
  • Experience building data lakes, caching layers, and cloud-based infrastructure services, and translating changing business and product requirements into new designs.
  • Experience designing test-automation infrastructure in manufacturing environments and strong scripting skills.
  • Experience designing both CPU and GPU workloads.
  • Strong problem-solving abilities and a comprehensive understanding of computer, storage, routing, switching, automation, and fundamental network theory.

Benefits

  • Base salary range of USD 208,000 to USD 333,500 per year, determined based on location, experience, and the pay of employees in similar positions.
  • Eligibility for equity and benefits.
  • NVIDIA is committed to an inclusive work environment and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs