Manager, System Software Engineering - Factory

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 CI/CD Communication @ 7 Compliance Engineering Management @ 6 GPU Git @ 4 Jira @ 4 LLM @ 4 Leadership @ 6 Project Management @ 4 Python @ 4 QA @ 6 Reporting @ 7 Security System Architecture @ 4 Technical Leadership @ 6

Details

NVIDIA's Datacenter System Software team is seeking an Engineering Manager to lead Factory System Software and Diagnostics Integration end to end. The role involves building and leading a global engineering team that delivers embedded code, application programs, and diagnostic updates for factories producing NVIDIA GPU- and DPU-based products, including rack-scale systems such as GB200/GB300 NVL72 and next-generation platforms. The work spans concurrent NPI ramps and sustaining production, in partnership with system architects, firmware developers, SWQA, product engineering, compliance and security teams, program and product management, and ODM/CM manufacturing partners.

Responsibilities

  • Build, lead, mentor, and grow a global factory engineering team across the United States and Taiwan, operating a follow-the-sun coverage model with on-site presence at ODM/CM partner factories.
  • Own hiring, career development, calibration, and succession planning.
  • Define factory-readiness scope and workflows for rack-scale products in coordination with product management, technical architects, and program management.
  • Deliver workflows through the validation matrix and ensure firmware and software releases meet high quality standards and scale reliably.
  • Provide technical leadership for firmware, software, and diagnostics releases reaching factories that build rack-scale systems, including tightly coupled compute and switch trays.
  • Build end-to-end infrastructure and workflows to ensure efficient, high-quality releases.
  • Improve release quality through collaboration with developers, SWQA, and product engineering, using end-to-end CI/CD and quality gates at each product milestone.
  • Publish and track quality indicators regularly and report release progress to collaborators and executives.
  • Own the factory escalation path, including triage SLAs, 24/7 coverage, failure root-cause analysis and deflection, and bonepile burn-down to minimize line-down time during NPI ramps and mass production.
  • Shape the team roadmap and drive automation and AI-assisted validation and triage, including station-readiness automation, firmware-update flows, and log triage.
  • Analyze factory processes, systems, and workflows to identify optimization opportunities, remove bottlenecks, document and publish standard operating procedures, and measure performance against defined targets.

Requirements

  • 10+ years of overall software industry experience, specializing in system software and/or firmware development.
  • 3+ years of engineering management or technical leadership experience, including building and leading geographically distributed teams.
  • BS, MS, or PhD in computer science, computer engineering, electrical engineering, or a related technical field, or equivalent experience.
  • Proven track record of shipping scalable server products through factory ramps, from NPI bring-up to mass production, while collaborating with hardware, firmware, manufacturing, diagnostics, and QA teams.
  • Experience working with ODM/OEM partners to deliver quality servers and solutions for large-scale data centers.
  • Strong written and oral communication skills, including executive-level reporting, along with a strong work ethic and commitment to teamwork.
  • Ability to work and communicate effectively across teams, partners, and time zones.

Preferred Qualifications

  • Experience leading bring-up for sophisticated rack-scale compute architectures such as GB200/GB300 NVL72.
  • Familiarity with manufacturing test flows, including L6/L10/L11/L12 stations, factory test coverage, and MES integration.
  • Hands-on experience with x86/ARM system architecture and coding in C/C++ and Python.
  • Experience with source code management tools such as Git and Perforce, and project management tools such as Jira.
  • Experience integrating AI/LLM tooling into engineering workflows for triage, validation, log analysis, or test generation.
  • Experience establishing follow-the-sun support organizations with measurable response SLAs.

Benefits

The base salary range is USD 224,000–356,500, determined by location, experience, and pay for employees in similar positions. The role also includes eligibility for equity and benefits.

More jobs at Nvidia

Similar jobs