Senior Software Engineer, Core Infrastructure Services - DGX Cloud

at Nvidia
USD 168,000-322,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 3 API @ 4 Ansible @ 4 BGP @ 4 Communication @ 6 Distributed Systems @ 4 FastAPI @ 4 Go @ 7 Grafana @ 4 HPC @ 3 InfiniBand @ 3 Kafka @ 4 Kubernetes @ 4 Linux @ 7 Microservices @ 4 Networking @ 4 OAuth @ 4 Observability @ 4 OpenTelemetry @ 4 Prometheus @ 4 Python @ 7 Redis @ 4 Security @ 4 Terraform @ 4 gRPC @ 4

Details

NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. The team builds and operates the core infrastructure services that power NVIDIA's DGX Cloud and SuperPod deployments, delivering secure, reliable, and observable platforms at global scale.

Responsibilities

  • Build and operate core infrastructure services that power NVIDIA's global AI infrastructure.
  • Architect and develop secure, scalable, and highly available cloud-native platform services.
  • Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.
  • Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.
  • Build observability and security capabilities that improve the reliability and resilience of the infrastructure.
  • Partner with infrastructure and networking teams to deliver production services at scale.
  • Drive operational excellence through automation, monitoring, incident response, and continuous improvement.

Requirements

  • Bachelor's degree or equivalent experience, with 8+ years of relevant industry experience.
  • Strong proficiency in Python and Go, with experience building production-quality software.
  • Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.
  • Experience with infrastructure automation using Terraform and Ansible.
  • Experience with workflow orchestration using Temporal.
  • Experience with distributed systems using databases, Redis, and messaging platforms such as Kafka, NATS, and SQS.
  • Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.
  • Strong Linux fundamentals.
  • Experience with observability technologies including Prometheus, Grafana, OpenTelemetry, and gNMI.
  • Experience with networking technologies including BGP, switching, routing, and load balancing.
  • Experience with security technologies including VPNs, firewalls, and iptables/nftables.
  • Excellent problem-solving, communication, and collaboration skills.

Preferred Qualifications

  • Hands-on experience with network infrastructure, including switches, routers, and firewalls.
  • Familiarity with InfiniBand, RDMA, and AI/HPC networking.
  • Experience with NetBox, Nautobot, or similar network source-of-truth platforms.
  • Contributions to open-source software.
  • Experience with public cloud platforms.

Compensation and Benefits

The base salary range is USD 168,000–270,250 for Level 4 and USD 200,000–322,000 for Level 5. The base salary is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until August 8, 2026. This posting is for an existing vacancy. NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment.

More jobs at Nvidia

Similar jobs