Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Algorithms @ 6
Cloud Computing @ 7
Communication @ 6
Debugging @ 4
Deep Learning
Distributed Systems @ 8
GPU
Go @ 6
HPC @ 4
MPI @ 6
Microservices
NCCL @ 6
Networking @ 4
Rust @ 6
Security @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking a senior system software engineer to help build scientific computing platform workflows in the cloud. This cloud-based scientific computing platform enables physics-based numerical simulation solvers, AI-based training, inference, and visualization workflows for physical science and engineering problems.
Applications include weather prediction, climate modeling, industrial design, and digital twin simulation across domains such as aerospace, automotive, sports, renewable energy, and biomedical science.
Responsibilities
- Design services and take ownership of the underlying cloud infrastructure for physics-informed and data-driven scientific workflows.
- Design novel algorithms and engage with operations to improve overall system performance across the stack, including deep learning frameworks, numerical solvers, microservices, APIs, and heterogeneous CPU- and GPU-accelerated computing.
- Design, build, deploy, and operate scalable I/O infrastructure for checkpointing, data loading, and pre-processing and post-processing of data.
- Optimize compute, storage, and network architecture for physics- and simulation-driven applications.
Requirements
- BS or MS degree in Computer Science or a related field, or equivalent experience.
- 10+ years of experience building and operating distributed compute- and data-intensive platform-as-a-service systems in the cloud.
- Proven proficiency in a compiled language such as Go, Rust, or C++.
- Strong foundational knowledge of cloud computing, including datacenter architecture, cloud security architecture, virtualization of CPU, memory, and I/O, resource pooling, and elasticity.
- Proven skills in distributed systems and parallel processing, including distributed computation models, topology abstraction, logical time, synchronization, deadlock detection, fault tolerance, failure detection, consensus and agreement protocols, parallel algorithms, shared-memory and distributed-memory architectures, message passing with MPI and NCCL, and cluster scalability and performance.
- Hands-on debugging skills involving processes, threads, deadlocks and synchronization, scheduling, inter-process communication, memory management, file systems, and I/O structures.
- Strong algorithmic thinking and system design skills, including recursion, graphs, trees, stacks, queues, and large-scale loosely coupled distributed system design and operations.
- Self-motivated, with strong interpersonal skills and the ability to work independently with multiple teams with minimal direction.
Preferred Qualifications
- Experience building, deploying, and operating AI platforms on HPC clusters.
- Experience building, deploying, and operating cloud-native systems, including distributed storage, scheduling, and orchestration across compute, storage, and networking.
- Experience configuring and troubleshooting hardware, operating systems, kernels, and compilers for maximum performance.
- Hands-on debugging experience optimizing compute, networking, and I/O frameworks.
- Extensive experience debugging and customizing third-party source code.
Compensation and Benefits
The base salary is determined based on location, experience, and compensation for employees in similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role is also eligible for equity and benefits.
Applications will be accepted at least until September 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to an inclusive work environment.