Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
Algorithms @ 7
Data Pipelines @ 4
Data Structures @ 7
Distributed Systems @ 7
GPU
Go @ 4
Java @ 4
Kubernetes @ 4
Linux @ 4
Observability @ 6
Python @ 4
Rust @ 4
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. The team is developing next-generation data and storage infrastructure for AI workloads, including storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs.
Responsibilities
- Build storage technologies, client libraries, and filesystem frameworks that help AI workloads access data across object stores, file systems, and hybrid cloud infrastructure.
- Develop high-performance storage paths for training and inference workflows, including data loading, checkpointing, caching, POSIX-style access, and object-store integration.
- Build observability systems that diagnose storage bottlenecks, attribute GPU idle time to I/O behavior, and expose actionable telemetry through production monitoring stacks.
- Improve the performance, scalability, and reliability of storage systems serving massive datasets, deep directory trees, and high-concurrency AI workloads.
- Work closely with internal AI teams, platform teams, SRE, and operations to validate storage behavior against real workloads and production environments.
- Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, performance, and verification.
Requirements
- Bachelor's degree in Computer Science, Information Systems, Computer Engineering, or equivalent experience.
- 5 or more years of software engineering experience.
- Strong foundation in algorithms, data structures, distributed systems, operating systems, and practical software design.
- Experience building performance-sensitive systems, storage, backend, or cloud-native software in languages such as Go, Python, Rust, C/C++, or Java.
- Experience with storage systems, object stores, caching, Linux systems, Kubernetes, or cloud infrastructure.
- Ability to reason about performance, scalability, concurrency, reliability, and operational tradeoffs in production systems.
- Ability to design APIs, document systems, communicate clearly, and break ambiguous infrastructure problems into practical execution plans.
- Curiosity and practical judgment around AI-assisted or agentic engineering workflows, including using clear intent, specifications, acceptance criteria, tests, and verification to guide development.
Preferred Qualifications
- Background with Linux kernel observability, eBPF, tracing, or low-overhead telemetry systems.
- Experience with FUSE, POSIX filesystems, object-store-backed filesystems, or filesystem metadata/indexing.
- Experience optimizing storage performance for AI training, checkpointing, inference, or large-scale data pipelines.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications for this job will be accepted at least until August 1, 2026.
More jobs at Nvidia
Research Engineer, Interactive World Models - New College Grad 2026
Nvidia · Santa Clara, United States
USD 108,000-195,500 per year
Senior Security Engineer, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 170,000-275,000 per year
Systems Software Engineer - AI and Cloud
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · Canada
CAD 245,000-295,000 per year
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Cloud Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Software Engineer - Public Cloud Engineering
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Staff Software Engineer, Infrastructure (Distributed Systems)
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer II
Confluent · United States, Seattle, United States
USD 197,400-232,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year