Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
AWS @ 7
Azure @ 7
CUDA @ 3
Communication @ 6
Data Pipelines @ 4
Debugging @ 7
Distributed Systems @ 4
Docker @ 4
GCP @ 7
GPU
Go @ 6
Grafana
Java @ 6
Kubernetes @ 4
Machine Learning
Observability
OpenTelemetry
Prometheus
Python @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a Senior Software Engineer to join NVIDIA's DGX Cloud team and build the foundational systems that drive high-performance GPU infrastructure. You will design scalable automation solutions, integrate diverse systems, and enable seamless workflows across global cloud operations.
Responsibilities
- Design and develop APIs to orchestrate and integrate operational workflows.
- Build state management and workflow automation systems that streamline infrastructure lifecycle processes.
- Collaborate across teams to codify business processes into scalable, self-measuring systems.
- Develop extensible, schema-driven platforms to reduce manual toil and ensure operational consistency.
- Drive integrations with container orchestration tools such as Kubernetes and observability systems including Prometheus, OpenTelemetry, and Grafana.
- Optimize the reliability and efficiency of cloud operations through automated workflows and telemetry systems.
- Lead and ship impactful technical projects, ensuring quality and scalability at every stage.
Requirements
- 8+ years of industry experience with a bachelor's degree or equivalent experience; a master's degree is preferred.
- Expertise in designing, building, and operating services in a high-reliability environment.
- Proficiency in programming languages such as Go, Java, or Python.
- Strong understanding of cloud infrastructure, including AWS, GCP, and Azure.
- Experience with container technologies such as Docker and Kubernetes.
- Experience with high-scale distributed systems, including architectural patterns for APIs and data pipelines.
- Outstanding communication and collaboration skills, with a focus on solving complex operational challenges.
- A passion for automating manual processes and driving system efficiency.
Preferred Qualifications
- A track record of designing workflow orchestration systems for large-scale infrastructure.
- Proven experience reducing operational inefficiencies through automation and integration.
- Strong debugging and problem-solving skills in distributed environments.
- Prior experience or strong familiarity with the operational aspects of the NVIDIA AI/ML software stack, including CUDA, cuDNN, and containerization.
Benefits
- Equity and benefits are provided.
- NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
The base salary range is USD 184,000–287,500 per year. The salary is determined based on location, experience, and the pay of employees in similar positions. Applications will be accepted at least until August 13, 2026.
More jobs at Nvidia
Senior Site Reliability Engineer - Storage
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Director, Global Risk and Compliance
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year
Senior System Software Engineer - CPU SoC Boot Firmware
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Staff Forward-Deployed Engineer, Enterprise AI and Automation
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
DGX Cloud Automation Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · San Francisco, United States, New York City, United States, Seattle, United States
USD 320,000-485,000 per year