Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
AWS @ 4
Algorithms @ 7
Azure @ 4
Data Structures @ 7
Distributed Systems @ 7
GCP @ 4
GPU
Go @ 4
Java @ 4
Kubernetes @ 4
Machine Learning @ 4
Observability @ 7
Python @ 4
Rust @ 4
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. The team is developing next-generation data and storage infrastructure for AI, including storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs.
Responsibilities
- Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access.
- Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows.
- Work with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and develop reliable production systems.
- Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale.
- Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification.
Requirements
- Bachelor's degree in Computer Science, Information Systems, Computer Engineering, or equivalent experience.
- 5+ years of software engineering experience.
- Strong foundation in algorithms, data structures, distributed systems, and practical software design.
- Experience building, shipping, and operating backend or cloud-native services using Kubernetes and cloud providers such as AWS, GCP, or Azure.
- Experience with one or more of Go, Python, Rust, C/C++, or Java.
- Ability to design APIs, document systems, reason through tradeoffs, communicate clearly, and break ambiguous problems into practical execution plans.
- Experience working across engineering, product, platform, and operations teams to deliver reliable production software.
- Curiosity and practical judgment around AI-assisted or agentic engineering workflows, including using clear intent, specifications, acceptance criteria, tests, and verification to guide development.
Preferred Qualifications
- Hands-on experience building, scaling, or operating large-scale data, storage, or machine learning infrastructure services.
- Experience solving enterprise-grade data management, governance, analytics, or AI workflow problems with modern data and machine learning infrastructure technologies.
- Strong background in distributed systems, storage systems, cloud infrastructure, performance engineering, observability, or agentic engineering practices.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Compensation
The base salary depends on location, experience, and the pay of employees in similar positions. The base salary range is USD 152,000–241,500 for Level 3 and USD 184,000–287,500 for Level 4. Applications will be accepted at least until August 1, 2026.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Storage Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Software Engineer II
Confluent · Seattle, United States, United States
USD 197,400-232,000 per year
Senior Software Engineer - Public Cloud Engineering
Bloomberg · New York City, United States
USD 160,000-240,000 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year