Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 7
AWS @ 6
Automated Testing @ 7
Azure @ 6
CI/CD @ 6
Communication @ 7
Docker @ 6
GCP @ 6
Go
Grafana @ 3
IaC
JavaScript @ 6
Kubernetes @ 6
Machine Learning @ 4
Next.js @ 6
Node.js @ 6
Observability @ 3
Prometheus @ 3
Python @ 6
React @ 6
Security
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Senior Full-Stack Software Engineer to join the AI Hub team within the DGX Cloud AI Infrastructure organization. The AI Hub team accelerates AI research by ensuring NVIDIA’s AI infrastructure is used efficiently, transparently, and at scale. The primary goal is to build a unified, self-service “single pane of glass” portal that enables AI researchers to efficiently manage, monitor, and optimize their use of Managed AI research Superclusters.
Responsibilities
- Lead the architecture and delivery of high-scale web products across frontend, backend services, and data layers, with clear availability and latency targets (SLOs/SLAs).
- Own multi-team initiatives end to end: problem discovery, RFCs/design reviews, phased rollouts, and success metrics tied to product and business outcomes.
- Drive reliability, performance, and observability improvements to meet exascale standards.
- Establish engineering standards and reusable platforms/design systems to reduce complexity, support load and long-term tech debt.
- Collaborate with NVIDIA AI Research teams to identify pain points and deliver the next generation user experience that accelerates their work.
- Mentor and sponsor engineers; improve code quality, testing, security, and observability through reviews, pairing, and coaching.
- Stay ahead of AI/ML infrastructure trends and drive adoption of best practices within the team.
Requirements
- 12+ years of software engineering experience delivering production web systems.
- Bachelor’s degree or higher in Computer Science or a related technical field (or equivalent experience).
- Strong cross-functional collaboration skills, including active listening, translating complex use cases into clear technical requirements, and designing data models aligned with business logic and outcomes.
- Deep cloud expertise (AWS, GCP, or Azure), infrastructure as code, containers, and orchestration (Docker, Kubernetes), along with mature CI/CD and safe deployment practices.
- Full-stack depth: modern SPA frameworks (React/Next.js or Vue/Nuxt), JavaScript/TypeScript, and one or more backend languages (Node.js, Python, and/or Golang).
- Familiarity with observability stacks such as OpenSearch, Prometheus, Grafana, or Loki.
- Proficiency in API design (REST), schema evolution, and integration patterns, with a strong commitment to automated testing.
- Experience building machine learning platforms or self-service internal infrastructure tools focused on efficiency, resiliency, and observability.
- Clear written and verbal communication skills, strong problem-solving ability, and a growth mindset.
- Experience leveraging AI-assisted development tools (e.g., Cursor).
Benefits
- NVIDIA provides competitive salaries and a comprehensive benefits package.
- You will also be eligible for equity and benefits.
More jobs at Nvidia
Ncx Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
System Test Engineer
Nvidia · Santa Clara, United States
USD 132,000-253,000 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager, Deep Learning Frameworks
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Senior Software Engineer, CUDA Core Libraries
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Frontend Engineer, NVIDIA Marketplace
Nvidia · Santa Clara, United States
USD 140,000-224,200 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer - HPC
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Full-Stack Swe, Data Acquisition (Foundations)
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Senior Backend Engineer - Databases - Loki Ingest
Grafana Labs · Sweden
SEK 775,000-969,000 per year
Senior Backend Engineer - Databases - Loki Ingest
Grafana Labs · Spain
EUR 83,000-104,000 per year