Senior Software Engineer, Core Infrastructure Services - DGX Cloud
at Nvidia
USD 168,000-322,000 per year
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 4
Ansible @ 4
BGP @ 4
Communication @ 6
Distributed Systems @ 4
FastAPI @ 4
Go @ 7
Grafana @ 4
HPC @ 3
InfiniBand @ 3
Kafka @ 4
Kubernetes @ 4
Linux @ 7
Microservices @ 4
Networking @ 4
OAuth @ 4
Observability @ 4
OpenTelemetry @ 4
Prometheus @ 4
Python @ 7
Redis @ 4
Security @ 4
Terraform @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. The team builds and operates the core infrastructure services that power NVIDIA's DGX Cloud and SuperPod deployments, delivering secure, reliable, and observable platforms at global scale.
Responsibilities
- Build and operate core infrastructure services that power NVIDIA's global AI infrastructure.
- Architect and develop secure, scalable, and highly available cloud-native platform services.
- Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.
- Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.
- Build observability and security capabilities that improve the reliability and resilience of the infrastructure.
- Partner with infrastructure and networking teams to deliver production services at scale.
- Drive operational excellence through automation, monitoring, incident response, and continuous improvement.
Requirements
- Bachelor's degree or equivalent experience, with 8+ years of relevant industry experience.
- Strong proficiency in Python and Go, with experience building production-quality software.
- Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.
- Experience with infrastructure automation using Terraform and Ansible.
- Experience with workflow orchestration using Temporal.
- Experience with distributed systems using databases, Redis, and messaging platforms such as Kafka, NATS, and SQS.
- Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.
- Strong Linux fundamentals.
- Experience with observability technologies including Prometheus, Grafana, OpenTelemetry, and gNMI.
- Experience with networking technologies including BGP, switching, routing, and load balancing.
- Experience with security technologies including VPNs, firewalls, and iptables/nftables.
- Excellent problem-solving, communication, and collaboration skills.
Preferred Qualifications
- Hands-on experience with network infrastructure, including switches, routers, and firewalls.
- Familiarity with InfiniBand, RDMA, and AI/HPC networking.
- Experience with NetBox, Nautobot, or similar network source-of-truth platforms.
- Contributions to open-source software.
- Experience with public cloud platforms.
Compensation and Benefits
The base salary range is USD 168,000–270,250 for Level 4 and USD 200,000–322,000 for Level 5. The base salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until August 8, 2026. This posting is for an existing vacancy. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.
More jobs at Nvidia
Senior Machine Learning Engineer
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior System Test Engineer, Networking
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, Agentic AI
Nvidia · Redmond, United States
USD 152,000-287,500 per year
Director, Technical Program Management
Nvidia · Santa Clara, United States
USD 272,000-425,500 per year
Senior Systems Software Engineer - Infrastructure
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Similar jobs
Principal Network Automation Engineer
Nvidia · Santa Clara, United States
USD 248,000-396,800 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year