Engineering Manager, DGX Cloud Production Engineering

at Nvidia
USD 224,000-356,500 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 Communication @ 7 Distributed Systems @ 4 GPU @ 4 Kubernetes @ 4 Mentoring Networking Observability @ 7 Prioritization @ 7 SRE @ 4 Security

Details

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people.

Today, we’re tapping into the unlimited potential of AI to define the next era of computing, where our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world.

Our DGX Cloud Production Engineering team is looking for an Engineering Manager to lead a dynamic group of engineers in a collaborative environment focused on innovation and world-class engineering.

Responsibilities

  • Leading a team of software and production engineers in building and operating the DGX Cloud infrastructure across NVIDIA Cloud Partner (NCP) and on-prem environments.
  • Driving execution in cluster operations, Kubernetes operability, automation, GitOps, observability, and incident response.
  • Defining team priorities, roadmap, staffing, and operational ownership.
  • Partnering with platform, workload, storage, networking, security, and TPM teams to improve production readiness.
  • Encouraging a healthy on-call and incident review culture focused on learning, ownership, and durable fixes.
  • Mentoring engineers, growing technical leaders, and crafting clear ownership across ambiguous problem spaces.

Requirements

  • 8+ overall years of industry experience, including 2+ years leading or managing engineers.
  • Proven experience in building or operating production infrastructure, cloud platforms, Kubernetes environments, or distributed systems.
  • Strong understanding of reliability engineering, automation, observability, incident response, and operational excellence.
  • Ability to work effectively across teams and influence without direct authority.
  • Clear communication, strong prioritization, and good judgment in fast paced environments.
  • BS/MS in Computer Science or equivalent experience.

Ways to stand out from the crowd

  • Experience leading SRE, production engineering, infrastructure automation, or platform teams.
  • Familiarity with GPU infrastructure, Kubernetes fleet operations, GitOps, BMaaS/VMaaS, managed Kubernetes, or multi-cloud environments.
  • A track record of reducing toil, improving SLOs, and turning operational work into automated systems.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 1, 2026.

More jobs at Nvidia

Similar jobs