Senior System Software Engineer - GPU Power and Performance Management
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Debugging @ 7
GPU @ 4
Networking
Observability @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is redefining accelerated computing and powering the next era of artificial intelligence. Its data-center platforms bring together GPUs, CPUs, networking, systems, and software to solve complex computing problems. The GPU Performance and Power Management Software team is seeking a System Software Engineer to design, develop, and debug production software for performance states (P-states) and power management controllers. The role supports dynamic voltage and frequency scaling (DVFS) and power-management technologies for next-generation data-center GPUs.
The position works at the intersection of hardware architecture, embedded firmware, device drivers, operating systems, and AI workloads. Contributions span the full product lifecycle, including design, pre-silicon development, validation, silicon bring-up, and production support.
Responsibilities
- Architect, build, implement, and debug GPU power and performance management software, with an emphasis on P-state management and DVFS.
- Develop control policies, mechanisms, and real-time software that optimize GPU performance, power, thermals, and energy-efficiency KPIs under demanding data-center workloads.
- Own focus areas and drive features across requirements, architecture, implementation, pre-silicon validation, silicon bring-up, productization, and production support.
- Collaborate with GPU architects and hardware designers to define hardware-software interfaces and shape power-management capabilities for next-generation processors.
- Analyze interactions among workloads, clocks, voltages, power limits, thermals, telemetry, and system-level policies.
- Lead investigations and resolve complex power, performance, stability, and reliability issues spanning firmware, drivers, hardware, and platform software.
- Develop validation strategies and automation for functional correctness, transition latency, performance-per-watt, and robustness across operating conditions.
- Create architecture, interface, and feature specifications for new power and performance management capabilities.
- Partner with distributed teams across architecture, ASIC, firmware, drivers, validation, platform, and data-center systems.
- Influence next-generation system software by building scalable internal architecture frameworks.
Requirements
- BS or MS degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of industry experience developing system software, embedded firmware, device drivers, or other hardware-related production software.
- Strong programming skills in C, including experience developing, debugging, and maintaining complex low-level software.
- Deep understanding of operating-system fundamentals, device-driver architecture, embedded or real-time software, concurrency, interrupt handling, and hardware programming.
- Experience interpreting hardware specifications and developing software interfaces for registers, telemetry, interrupts, and control mechanisms.
- Strong understanding of computer architecture and hardware-software interactions.
- Ability to independently debug complex, cross-layer problems across software and hardware.
- Ability to own technically challenging features or subsystems and drive them to production with cross-functional teams.
- Excellent written and verbal communication and presentation skills.
Preferred Qualifications
- Hands-on experience with GPU, CPU, accelerator, or SoC power and performance management.
- Experience implementing P-state selection, DVFS, clock control, voltage control, power capping, thermal management, throttling, or workload-aware control policies.
- Experience developing real-time feedback controllers, including reasoning about stability, responsiveness, hysteresis, latency, and telemetry noise.
- Experience building validation, observability, and automation infrastructure for low-level software features.
- Understanding of voltage-frequency relationships, transient power, thermal limits, power delivery, and system-level power budgets.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. Applications will be accepted at least until September 26, 2026. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.