Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 4
Debugging @ 7
GPU @ 6
Networking
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is redefining accelerated computing and powering the next era of artificial intelligence. The Data Center Platforms organization brings together GPUs, CPUs, networking, systems, and software to solve complex computing problems.
This role is part of NVIDIA's GPU Performance and Power Management Software team. The engineer will architect production software for performance states (P-states) and Power Management Controllers, including dynamic voltage and frequency scaling (DVFS) and power-management control systems for next-generation data-center GPUs.
The work spans hardware architecture, embedded firmware, device drivers, operating systems, and AI workloads across the entire product lifecycle, from initial architectural design and pre-silicon validation through silicon bring-up, productization, and deployment in large data-center infrastructures.
Responsibilities
- Architect, build, implement, and debug GPU power- and performance-management software, with an emphasis on P-state management and DVFS.
- Develop real-time controllers, policies, and mechanisms that dynamically optimize GPU performance, power, thermals, and energy efficiency under demanding data-center workloads.
- Drive features across the full product lifecycle, including requirements, architecture, implementation, pre-silicon validation, silicon bring-up, productization, and production support.
- Collaborate with GPU architects and hardware designers to define hardware-software interfaces and influence the power-management capabilities of next-generation processors.
- Analyze interactions among workloads, clocks, voltages, power limits, thermals, telemetry, and system-level policies.
- Lead the investigation and resolution of complex power, performance, stability, and reliability issues spanning firmware, drivers, hardware, and platform software.
- Develop validation strategies and automation for functional correctness, transition latency, performance-per-watt, and robustness across operating conditions.
- Build architecture, interface, and composite specifications for new power-management features.
- Partner with geographically distributed architecture, ASIC, firmware, driver, validation, platform, and data-center systems teams.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- At least 12 years of industry experience developing system software, embedded firmware, device drivers, or other hardware-related production software.
- Strong C programming skills, including experience developing, debugging, and maintaining complex low-level software.
- Deep understanding of operating-system fundamentals, device-driver architecture, embedded or real-time software, concurrency, interrupt handling, and hardware programming.
- Experience interpreting hardware specifications and developing software interfaces for registers, telemetry, interrupts, and control mechanisms.
- Strong understanding of computer architecture and hardware-software interactions.
- Ability to independently debug complex problems across multiple software and hardware layers.
- Excellent written and verbal communication skills, with experience working across organizational and geographic boundaries.
Preferred Qualifications
- Hands-on experience with GPU, CPU, accelerator, or SoC power and performance management.
- Experience implementing P-state selection, DVFS, clock control, voltage control, power capping, thermal management, throttling, or workload-aware control policies.
- Experience developing real-time feedback controllers, including reasoning about stability, responsiveness, hysteresis, latency, and noisy telemetry.
- Understanding of voltage-frequency relationships, transient power, thermal limits, power delivery, and system-level power budgets.
Compensation and Benefits
The base salary range is USD 224,000–356,500 for Level 5 and USD 272,000–431,250 for Level 6. Base salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 25, 2026. This posting is for an existing vacancy. NVIDIA is an equal opportunity employer committed to an inclusive work environment.