Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Data Analysis @ 4
Debugging @ 7
LLVM @ 4
Linux @ 7
Profiling @ 4
Python @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is growing a senior engineering team focused on making its compute software stack first-class on NVIDIA CPU platforms. The team turns modern toolchains, build and code-health practices, performance-analysis workflows, and optimization techniques into repeatable improvements across real software components.
The role involves leading cross-stack engineering efforts from ambiguous adoption problems to measurable outcomes. You will work closely with compiler, platform, performance, library, and application teams to validate new capabilities, resolve integration blockers, produce before-and-after evidence, and turn successful approaches into reusable engineering practices. Depending on project needs, the initial work may emphasize toolchain, build, and code-health adoption or profiling, optimization, and evidence-routing workflows.
Responsibilities
- Lead toolchain, build, code-health, and performance-workflow adoption projects across large software components.
- Evaluate and deploy supported GCC and LLVM/Clang toolchains, compiler options, linkers, sysroots, and cross-compilation configurations.
- Integrate modern toolchains and workflows into complex build systems and continuous-integration environments.
- Establish useful Clang diagnostic builds and targeted sanitizer coverage in partnership with component owners.
- Evaluate link-time optimization, profile-guided optimization, AutoFDO, and BOLT on representative software.
- Use profiling, PMU data, flamegraphs, and binary/source analysis to identify actionable performance and code-quality findings.
- Build automation, wrappers, validation scripts, dashboards, and migration helpers that improve adoption and repeatability.
- Measure changes in runtime, code size, build time, launch latency, throughput, quality, or engineering velocity.
- Drive complex blockers to the appropriate compiler, runtime, library, infrastructure, or component owner.
- Document validated approaches as reusable playbooks for other engineering teams.
- Communicate technical results and tradeoffs clearly to engineers and leadership.
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of relevant systems software engineering experience.
- Strong C and C++ development, debugging, and code-review skills.
- Strong Linux systems knowledge and hands-on experience with complex native software stacks.
- Practical experience with GCC or LLVM/Clang, linkers, compiler options, and build systems.
- Experience working in large, multi-component codebases and CI environments.
- Experience with performance profiling, root-cause analysis, and before-and-after validation.
- Ability to lead projects involving substantial technical and organizational ambiguity.
- Excellent written and verbal communication and a record of effective cross-team collaboration.
Preferred Qualifications
- Experience with Arm64 systems, CPU architecture, vectorization, or SVE.
- Experience with LTO, PGO, AutoFDO, BOLT, binary optimization, or code-layout analysis.
- Hands-on experience with Clang diagnostics, AddressSanitizer, or other code-health workflows.
- Familiarity with PMU analysis, perf, flamegraphs, BRBE, SPE, ETM, or similar profiling technologies.
- Experience with cross-compilation, sysroots, large monorepositories, Perforce-scale development, or distributed build systems.
- Experience turning a successful migration or optimization into a maintained workflow used by multiple teams.
- Python or other scripting experience for engineering automation and data analysis.
Compensation and Benefits
- Base salary range: $184,000–$287,500 for Level 4, or $224,000–$356,500 for Level 5. Salary depends on location, experience, and the pay of employees in similar positions.
- Eligible for equity and benefits.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. NVIDIA uses AI tools in its recruiting processes.