Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
AWS @ 4
Algorithms @ 7
Azure @ 4
Data Structures @ 7
Distributed Systems @ 7
GCP @ 4
GPU
Go @ 4
Java @ 4
Kubernetes @ 4
Machine Learning
Observability @ 7
Python @ 4
Rust @ 4
SRE
Security
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The NVIDIA DGXC Data Services team builds cloud-native systems, frameworks, and services for managing data across hybrid and multi-cloud infrastructure. They are building the next-generation data and storage infrastructure to solve some of the hardest problems in AI: storage, access, ingestion, governance, observability, and data management for exabyte-scale, high-performance GPU-based training and inference jobs. Their work gives NVIDIA teams the foundational capabilities they need to build, train, deploy, and operate AI products at scale without reinventing critical data infrastructure for every workload.
Responsibilities
- Build cloud-native data and storage services for hybrid and multi-cloud infrastructure, including dataset discovery, ingestion, governance, checkpointing, observability, and low-latency access.
- Develop scalable cloud-native services and APIs that support exabyte-scale, high-performance GPU training and inference workflows.
- Work closely with product managers, internal AI teams, platform teams, and partner engineering teams to understand requirements and turn them into reliable production systems.
- Collaborate with SRE, operations, and support teams to improve service reliability, performance, observability, on-call readiness, and operational scale.
- Use modern software engineering practices, including AI-assisted and agentic development workflows, while maintaining high standards for design, testing, security, and verification.
Requirements
- BS in Computer Science, Information Systems, Computer Engineering, or equivalent experience, with 5+ years of software engineering experience.
- Strong foundation in algorithms, data structures, distributed systems, and practical software design.
- Experience building, shipping, and operating backend or cloud-native services using Kubernetes, cloud providers such as AWS, GCP, or Azure, and languages such as Go, Python, Rust, C/C++, or Java.
- Ability to design APIs, document systems, reason through tradeoffs, communicate clearly, and break ambiguous problems into practical execution plans.
- Experience working across engineering, product, platform, and operations teams to deliver reliable production software.
- Curiosity and practical judgment around AI-assisted or agentic engineering workflows, including using clear intent, specifications, acceptance criteria, tests, and verification to guide development.
Ways to stand out from the crowd
- Hands-on experience building, scaling, or operating large-scale data, storage, or ML infrastructure services.
- Experience solving enterprise-grade data management, governance, analytics, or AI workflow problems with modern data and ML infrastructure technologies.
- Strong background in distributed systems, storage systems, cloud infrastructure, performance engineering, observability, or agentic engineering practices.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.