Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
AWS @ 7
Azure @ 7
CUDA @ 3
Cloud Computing @ 7
Communication @ 6
Data Pipelines @ 4
Distributed Systems @ 4
Docker @ 7
GCP @ 7
GPU
Go @ 6
Grafana
Java @ 6
Kubernetes @ 7
Machine Learning
OpenTelemetry
Prometheus
Python @ 6
Security @ 4
Slurm
Technical Leadership
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Principal Software Engineer to join the DGX Cloud team and build foundational systems for high-performance GPU infrastructure. The role focuses on scalable automation, system integration, and seamless workflows across global cloud operations. As a Principal Engineer, you will provide technical leadership and help shape the platform supporting AI and cloud computing.
Responsibilities
- Lead the development of next-generation APIs, state management, and workflow orchestration systems that automate fleet lifecycle operations at massive scale.
- Drive technical alignment across dependent systems and partner teams to ensure cohesive integration, clear interfaces, and reliable end-to-end workflows.
- Coach and mentor senior engineers while elevating technical standards and guidelines across the organization.
- Maintain a strong focus on customer experience and product requirements, translating technical insight into high-impact business solutions.
- Partner with executive and engineering leadership to codify critical business processes into self-measuring, scalable, and operationally consistent platforms that reduce manual effort.
- Direct integration strategies for technologies including Kubernetes, Slurm, Prometheus, OpenTelemetry, and Grafana.
Requirements
- 16 or more years of progressive industry experience.
- Master's or bachelor's degree, or equivalent experience defining and shipping complex distributed systems.
- Deep hands-on expertise establishing, operating, and scaling services in fast-paced, high-reliability environments.
- Ability to work effectively in ambiguous, fast-paced environments by testing ideas, iterating toward working solutions, and hardening successful approaches into reliable, scalable systems.
- Outstanding proficiency in modern systems programming languages such as Go, Java, or Python.
- Proven experience defining, owning, and evolving the architecture of high-scale distributed systems, including advanced patterns for APIs, control planes, and data pipelines.
- Deep understanding of global cloud infrastructure, including AWS, GCP, and Azure, as well as container ecosystems such as Docker and Kubernetes.
- Ability to drive technical strategy and influence outcomes across organizational boundaries.
- Excellent communication skills, including the ability to explain complex technical concepts, build organizational consensus, and mentor high-performing engineers.
Preferred Qualifications
- Experience leading the development and adoption of organization-wide workflow orchestration systems for petabyte-scale infrastructure.
- Experience working in a Principal or Staff+ capacity and delivering measurable improvements in operational efficiency, reliability, and security across a large engineering organization.
- Familiarity with the operational and deployment aspects of the NVIDIA AI/ML software stack, including CUDA, cuDNN, and containerization.
- Patent contributions or a strong publication record in distributed systems, cloud computing, or infrastructure automation.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package.
The base salary range is USD 272,000–431,250. Applications will be accepted at least until May 3, 2026. NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Senior Applied Research Scientist – AI Native Numerical Methods
Nvidia · United States
USD 192,000-356,500 per year
Senior Applied Research Scientist, Multimodal Foundation Models – Healthcare
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Business Operations Technical Program Manager — DGX Cloud
Nvidia · Santa Clara, United States
USD 200,000-379,500 per year
Senior Applied Research Scientist – GPU Native Numerical Algorithms
Nvidia · United States
USD 192,000-356,500 per year
Senior Software Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer II
Confluent · Seattle, United States, United States
USD 197,400-232,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Backend Engineer - Databases - Loki Query
Grafana Labs · Germany, Spain, Ireland, Sweden, United Kingdom
EUR 97,000-121,000 per year