Senior Systems Software Engineer - Fleet Debuggability

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Communication @ 7 Customer Support Debugging @ 7 GPU Git @ 4 Grafana @ 3 Jira @ 4 LLM @ 4 Linux @ 7 Prometheus @ 3 Python @ 7 Reporting @ 7 Rust @ 7 System Architecture @ 4

Details

NVIDIA is looking for a highly motivated Senior Engineer to drive fleet-scale debuggability end to end. As part of the Datacenter System Software team, you will design, architect, and build infrastructure, tooling, and analytics for collecting multi-rack-scale logs. The solution will normalize, correlate, and reason over logs spanning NVIDIA GPUs, CPUs, network products, components, trays, and racks. Logs will be fetched through in-band or out-of-band channels to help triage fleet-level issues seen by customers and shorten the path from raw, noisy log streams to actionable root causes.

Responsibilities

  • Architect, design, and build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time-align logs from heterogeneous sources, including kernel and driver logs, syslog, Redfish event logs, SEL, firmware, and BMC logs, over in-band and out-of-band channels.
  • Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance.
  • Develop debugging and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.
  • Design fleet-scale log collection and analysis while keeping overhead on production compute nodes low.
  • Partner with developers, SWQA, and product engineering on end-to-end logging solutions, event schemas, and contracts between log producers and tooling.
  • Steward the project's open-source release, maintain clean internal and public code paths, review community contributions, and represent the tooling in upstream discussions.
  • Write design documents and own end-to-end delivery from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug-management tools and build realistic end-to-end execution plans with engineers and managers.

Requirements

  • 10+ years of software industry experience specializing in system software and/or firmware development.
  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience.
  • Proven track record of shipping scalable server products or working at fleet-wide scale.
  • Strong written and oral communication skills, including executive-level reporting, along with a strong work ethic and dedication to teamwork.
  • Ability to work and communicate effectively across teams, partners, and time zones.
  • Experience with source control management systems such as Git or Perforce and project-management tools such as Jira.
  • Strong, demonstrable skills in Python or Rust.
  • Deep Linux systems experience, including kernel and driver logs, syslog, journald, and debugging on server platforms.
  • Hands-on experience with out-of-band management and platform interfaces, including BMC, Redfish, IPMI, and SEL, and understanding of in-band versus out-of-band trade-offs.
  • Strong skills in log parsing, normalization, and structured logging, with experience designing schemas and taxonomies for machine-readable events.

Preferred Qualifications

  • Experience leading debuggability solutions on rack-scale compute architectures such as GB200 or GB300 NVL72.
  • Familiarity with log and telemetry analytics stacks such as OpenSearch/ELK, Loki, Prometheus, Grafana, and PagerDuty, as well as time-series databases.
  • Hands-on experience with x86/ARM system architecture and coding in C/C++ or Python.
  • Experience integrating AI/LLM tooling into engineering workflows for triage, validation, log analysis, or test generation.
  • Experience establishing follow-the-sun support organizations with measurable response SLAs.
  • Experience contributing to or maintaining open-source projects, including managing the boundary between internal and public code.

Compensation and Benefits

The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until July 30, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive, equal-opportunity work environment.

More jobs at Nvidia

Similar jobs