Senior System Software Engineer - Scientific Computing PaaS

at Nvidia
USD 184,000-356,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API Algorithms @ 6 Cloud Computing @ 7 Communication @ 6 Debugging @ 4 Deep Learning Distributed Systems @ 8 GPU Go @ 6 HPC @ 4 MPI @ 6 Microservices NCCL @ 6 Networking @ 4 Rust @ 6 Security @ 7

Details

We are seeking a senior system software engineer to help build scientific computing platform workflows in the cloud. This cloud-based scientific computing platform enables physics-based numerical simulation solvers, AI-based training, inference, and visualization workflows for physical science and engineering problems.

Applications include weather prediction, climate modeling, industrial design, and digital twin simulation across domains such as aerospace, automotive, sports, renewable energy, and biomedical science.

Responsibilities

  • Design services and take ownership of the underlying cloud infrastructure for physics-informed and data-driven scientific workflows.
  • Design novel algorithms and engage with operations to improve overall system performance across the stack, including deep learning frameworks, numerical solvers, microservices, APIs, and heterogeneous CPU- and GPU-accelerated computing.
  • Design, build, deploy, and operate scalable I/O infrastructure for checkpointing, data loading, and pre-processing and post-processing of data.
  • Optimize compute, storage, and network architecture for physics- and simulation-driven applications.

Requirements

  • BS or MS degree in Computer Science or a related field, or equivalent experience.
  • 10+ years of experience building and operating distributed compute- and data-intensive platform-as-a-service systems in the cloud.
  • Proven proficiency in a compiled language such as Go, Rust, or C++.
  • Strong foundational knowledge of cloud computing, including datacenter architecture, cloud security architecture, virtualization of CPU, memory, and I/O, resource pooling, and elasticity.
  • Proven skills in distributed systems and parallel processing, including distributed computation models, topology abstraction, logical time, synchronization, deadlock detection, fault tolerance, failure detection, consensus and agreement protocols, parallel algorithms, shared-memory and distributed-memory architectures, message passing with MPI and NCCL, and cluster scalability and performance.
  • Hands-on debugging skills involving processes, threads, deadlocks and synchronization, scheduling, inter-process communication, memory management, file systems, and I/O structures.
  • Strong algorithmic thinking and system design skills, including recursion, graphs, trees, stacks, queues, and large-scale loosely coupled distributed system design and operations.
  • Self-motivated, with strong interpersonal skills and the ability to work independently with multiple teams with minimal direction.

Preferred Qualifications

  • Experience building, deploying, and operating AI platforms on HPC clusters.
  • Experience building, deploying, and operating cloud-native systems, including distributed storage, scheduling, and orchestration across compute, storage, and networking.
  • Experience configuring and troubleshooting hardware, operating systems, kernels, and compilers for maximum performance.
  • Hands-on debugging experience optimizing compute, networking, and I/O frameworks.
  • Extensive experience debugging and customizing third-party source code.

Compensation and Benefits

The base salary is determined based on location, experience, and compensation for employees in similar positions. The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role is also eligible for equity and benefits.

Applications will be accepted at least until September 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs