Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Communication @ 7
Customer Support
Debugging @ 7
Git @ 4
Jira @ 4
Linux @ 7
Python @ 7
Reporting @ 7
Rust @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Architect, design, build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
- Develop tooling to collect, normalize, and time-align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in-band and out-of-band channels. Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, so triage is repeatable rather than tribal knowledge.
- Develop debug and root-cause tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults. Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
- Partner with all matrixed organizations — developers, SWQA, and product engineering — in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
- Steward the project’s open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
- Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
- Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
- Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.
Requirements
- 10+ years in the software industry with specialization in system software and/or firmware development.
- BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
- Proven track record of shipping scalable server products or fleet-wide experience.
- A self-starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills — including executive-level reporting — strong work ethic, and dedication to teamwork.
- Flexibility to work and communicate effectively across teams, partners, and time zones.
- Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira. Strong, demonstrable skills in Python or RUST.
- Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
- Hands-on experience with out-of-band management and platform interfaces — BMC, Redfish, IPMI, SEL — and an understanding of in-band vs. out-of-band trade-offs.
- Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine-readable events.
Benefits
- Eligible for equity and benefits.
More jobs at Nvidia
Senior System Software Engineer, Interactive World Models
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer – Streaming
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Software Engineer - GeForce Now Low Latency Streaming
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Gpu Architecture Engineer - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Compiler Engineer, Infrastructure - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Similar jobs
Principal Platform Software Engineer - RAS
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Platform Telemetry Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Platform Security Engineer, OpenBMC
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-405,000 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Lead Systems Software Test Engineer – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Host Systems Software Engineer
OpenAI · San Francisco, United States
USD 266,000-445,000 per year
Principal Architect, System Software - Orbital Data Center
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year