Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 4
AWS @ 7
Azure @ 7
Communication @ 6
Data Pipelines @ 4
Distributed Systems @ 4
Docker @ 7
GCP @ 7
GPU
Go @ 6
Grafana
Java @ 6
Kubernetes @ 7
Observability
OpenTelemetry
Prometheus
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a Senior Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-performance GPU infrastructure. You will play a critical role in designing scalable automation solutions, integrating diverse systems, and enabling seamless workflows across global cloud operations. NVIDIA is widely recognized as one of the most desirable employers, with some of the most talented people in the world working for us. If you're passionate about building scalable, efficient systems to power cloud operations, we invite you to join our team.
Responsibilities
- Design and develop APIs to orchestrate and integrate operational workflows.
- Build state management and workflow automation systems that streamline infrastructure lifecycle processes.
- Collaborate across teams to codify business processes into scalable, self-measuring systems.
- Develop extensible, schema-driven platforms for reducing manual toil and ensuring operational consistency.
- Drive integrations with container orchestration tools like Kubernetes and observability systems such as Prometheus, OpenTelemetry, Grafana.
- Optimize the reliability and efficiency of cloud operations through automated workflows and telemetry systems.
- Lead and ship impactful technical projects, ensuring quality and scalability at every stage.
Requirements
- 8+ years of industry experience with a Bachelor’s or Master’s degree (or equivalent experience), or 2+ years with a PhD.
- Expertise in designing, building, and operating services in a high reliability environment.
- Proficiency in programming languages such as Go, Java, or Python.
- Strong understanding of cloud infrastructure (AWS, GCP, Azure) and container technologies like Docker and Kubernetes.
- Experience with high-scale distributed systems, including architectural patterns for APIs and data pipelines.
- Outstanding communication and collaboration skills, with a focus on solving complex operational challenges.
- A passion for automating manual processes and driving system efficiency.
Benefits
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
AI and ML Infrastructure Software Engineer, GPU Clusters - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Similar jobs
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer Ii
Confluent · Seattle, United States
USD 197,400-232,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer II
Confluent · Boston, United States, Dallas, United States, Chicago, United States, United States, Portland, United States
USD 176,000-230,000 per year
Tech Lead Manager, Agentic Runtime
Glean · Mountain View, United States, San Francisco, United States
USD 250,000-300,000 per year
Tech Lead Manager, Agentic Runtime
Glean · Mountain View, United States, San Francisco, United States
USD 250,000-300,000 per year