Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Audit
Communication @ 6
Docker @ 4
HPC @ 4
InfiniBand @ 4
Kubernetes @ 4
Leadership @ 6
Linux @ 4
Networking @ 4
OAuth @ 4
Observability @ 6
PKI @ 4
Rust
Security @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for an outstanding Senior Software Engineer to work on its security team focused on securing at scale infrastructure, high-performance computing environments, and AI cluster systems. The role involves designing, implementing, and evolving security systems, policies, and operational practices to protect critical infrastructure and intellectual property.
Responsibilities
- Develop and lead incident response and disaster recovery plans for pre-production clusters.
- Work with infrastructure, networking, storage, OS, firmware, and application teams to harden systems.
- Train peers and users on practical security standard methodologies.
- Build and enforce security controls, systems, and policies for cluster infrastructure of new NVIDIA hardware.
- Identify, assess, and reduce cybersecurity risks; report major risks clearly to leadership.
- Investigate security incidents and drive root-cause analysis.
- Ensure systems meet IT, legal, regulatory, and information security standards.
- Improve security governance, documentation, and audit readiness.
Requirements
- Experience in programming secure computing environments, with strong proficiency in C/C++.
- Linux kernel hardening (SELinux/AppArmor) and observability (eBPF).
- Experience securing large-scale Linux infrastructure.
- Proven understanding of security risks and how to reduce them.
- Demonstrated understanding of incident response and breach handling.
- Clear written and verbal communication with technical and leadership audiences.
- Experience with compute and networking systems security architectures.
- Experience in securing AI agents using sandboxing technologies and AI-based threat detection (e.g. Mythos).
- BS in Computer Science, Engineering, Cybersecurity, or equivalent experience. 8+ yrs of relevant experience.
Ways to Stand Out from the Crowd
- Developing secure software in Rust, prioritizing memory safety.
- Experience with modern authentication and identity frameworks such as OAuth 2.1, OIDC, Kerberos, FIDO2/WebAuthn.
- Experience with Microsoft Active Directory and Entra ID, including cross-realm trusts and identity federation (SCIMv2).
- Experience managing centralized Linux identity (FreeIPA/RHEL IdM/SSSD), including PKI lifecycle management and Host-Based Access Control.
- Experience hardening HPC schedulers and storage, Slurm alongside parallel filesystems like Lustre and NFS.
- Experience securing containerized workloads (Docker, Enroot, Kubernetes).
- Knowledge of high-speed fabric security like InfiniBand PKeys/MKeys.
- Zero Trust, ZTNA, VRFs, VLANs, performance-optimized firewalls.
- Use of advanced vulnerability management and supply chain mitigation (CVSS 4.0, SBOM).
More jobs at Nvidia
HPC Performance Engineer
Nvidia · United States
USD 152,000-241,500 per year
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Dynamo Tools
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Staff+ Software Engineer, Platform
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 405,000-485,000 per year
Staff Forward Deployed Engineer
GitLab · United States
USD 254,000-297,000 per year
Senior GPU and HPC Infrastructure Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Distinguished Engineer, Storage – AI Cloud
Nvidia · Santa Clara, United States
USD 320,000-488,800 per year
Principal Software Engineer - Rack Scale Systems Infrastructure
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer - Datacenter Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Hpc Cluster Engineer
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Perplexity AI · New York City, United States, Ireland, London, United Kingdom, San Francisco, United States
USD 250,000-485,000 per year