Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
BI
Debugging @ 4
Distributed Systems @ 4
GPU @ 4
Linux @ 7
Machine Learning
NVLink @ 4
Networking @ 4
Observability @ 4
Profiling @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior GPU Platforms Engineer to join a team working on GPU and SOC systems and platform software for accelerated computing. The role involves developing platform software, observability systems, production applications, and system-level features for NVIDIA SOC products.
Responsibilities
- Own platform software, including kernel device drivers, firmware and BIOS interactions, health monitoring, and system-level features for new GPU and SOC platforms.
- Design and develop high-performance, distributed observability platforms handling high-volume metrics, logs, and traces.
- Build real-time monitoring, logging, alerting, backend pipelines, and analytics platforms optimized for performance at scale.
- Develop production applications, solutions, and features for NVIDIA SOC products.
- Partner with software, driver, product, data center, and cloud teams to launch new GPU platforms, resolve complex system challenges, and improve performance, features, and stability.
- Participate in code reviews and contribute to substantial production codebases using C/C++ and data pipeline and business intelligence skills.
- Understand how AI workloads map onto NVIDIA architectures, including streaming multiprocessors, Tensor Cores, memory hierarchy, NVLink, PCIe, and MIG, and how these features affect latency, throughput, and cost.
- Develop and maintain expertise in operating systems and system software, including Linux internals, device drivers, kernel/user boundaries, concurrency, and performance profiling.
- Apply operating systems and computer architecture expertise to CPU-GPU interactions, PCIe, memory hierarchy, interrupts, and firmware or bootloaders.
Requirements
- BS, MS, PhD, or equivalent experience in Computer Science or a related field.
- Multiple years of experience in systems, platform, embedded, or GPU software roles.
- At least 10 years of experience developing SOC systems and platform software.
- Strong hands-on experience with Linux internals, device drivers, kernel/user boundaries, concurrency, and performance profiling.
- Experience building modern observability architectures involving metrics, logs, traces, distributed systems, business intelligence, and analytics stacks.
- Proficiency developing high-performance, secure production software and backend systems.
- Strong understanding of advanced operating systems, databases, computer architecture, and distributed systems.
- Experience developing AI/ML applications and integrating them with various features and applications.
- Proven ability to design scalable and secure product solutions.
Preferred Qualifications
- Experience programming and debugging SOC platforms across x86 and ARM architectures.
- In-depth knowledge of Linux and/or Windows, networking, and virtualization.
- Demonstrated proficiency in C/C++, extensive production codebases, and code reviews.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. Applications will be accepted at least until October 5, 2026.
More jobs at Nvidia
Legal Operations Engineer
Nvidia · Santa Clara, United States
USD 124,000-195,500 per year
Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
Nvidia · Spain
PLN 292,500-650,000 per year
Staff Site Reliability Operations
Nvidia · Hillsboro, United States
USD 144,000-230,000 per year
Senior Infrastructure Engineer - Connectivity
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Engineering Manager – AI Platform & SRE
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Similar jobs
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Software Engineer, Workload Enablement
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-385,000 per year
Senior System Software Engineer, Software-Defined Networking
Nvidia · United States
USD 224,000-356,500 per year
Principal Site Reliability Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
Software Engineer - Platform Infrastructure (Rust, C++)
SpaceXAI · Palo Alto, United States
USD 180,000-440,000 per year
Principal Software Engineer - Rack-Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year