Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 7
CUDA @ 7
Communication @ 7
Debugging @ 4
GPU @ 4
JAX @ 3
Leadership @ 6
Machine Learning
Observability @ 4
Profiling @ 4
PyTorch @ 3
Security @ 4
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Help design and ship an always-on, low-overhead GPU profiling service that runs in production, scales across cluster environments, and delivers actionable insights for machine learning workloads. You will lead architecture and hands-on delivery across system software, drivers, and CUDA to make profiling continuously available and reliable.
Responsibilities
- Design the architecture for an always-on profiling service, defining interfaces, data flows, and scalability guarantees for multi-process, multi-GPU, and multi-node systems.
- Drive low-overhead, high-reliability implementations in C/C++, including IPC and shared memory, with bounded CPU and memory budgets.
- Lead end-to-end feature delivery spanning user-mode components, driver and platform layers, and performance counter and trace providers.
- Establish profiling models that integrate with existing ML and AI workflows, such as PyTorch and XLA, to turn low-level signals into actionable insights.
- Set technical direction for an engineering team, mentor engineers, drive technical planning to mitigate architectural risks, and align roadmaps across internal and external partners.
Requirements
- BS or MS degree, or equivalent experience, in Computer Engineering, Computer Science, or a related field.
- 15+ years of system-level C/C++ development experience, including concurrency, memory management, and performance engineering.
- Proven experience designing and shipping production-quality system software or drivers with strict reliability, observability, and performance constraints.
- Demonstrated technical leadership, including defining architecture and success metrics and translating abstract product visions into actionable technical roadmaps with fast-paced, multidisciplinary teams.
- Strong interpersonal, verbal, and written communication skills, with the ability to influence across organizations and build trust with external collaborators.
Preferred Qualifications
- Extensive experience with CPU/GPU profiling and tracing stacks, such as CUPTI, Nsight, performance counters, and event correlation, as well as debugging highly concurrent systems.
- Deep hands-on knowledge of CUDA and GPU architecture, including runtime and driver APIs, CUDA streams and graphs, and kernel behavior.
- Track record of building continuous, always-on, or multi-client profiling systems designed for predictable overhead at scale.
- Hands-on experience tuning ML training and inference loops based on deep profiling analysis, with familiarity with ML ecosystems such as PyTorch and JAX.
- Experience correlating application events with GPU metrics to translate data into actionable performance insights, including bottleneck triage and determining whether workloads are compute-bound or memory-bound.
- Experience with user-mode driver development and integration within platform security and permissions models.
Compensation and Benefits
The base salary range is USD 272,000–431,250 per year. Salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications for this job will be accepted at least until February 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to fostering a diverse work environment.