Senior DevOps Engineer, Platform Engineering

at Nvidia
USD 176,000-276,000 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 AWS @ 4 Ansible @ 4 Azure @ 4 CI/CD @ 4 Change Management @ 4 DevOps @ 4 Distributed Systems @ 7 GPU @ 4 GitHub @ 4 GitHub Actions @ 4 Grafana @ 4 Helm @ 6 IaC Jenkins @ 4 Kubernetes @ 6 Linux @ 7 Machine Learning Networking @ 7 Observability @ 4 Prometheus @ 4 Python @ 6 Terraform @ 4

Details

NVIDIA is seeking a Senior DevOps Platform Engineer skilled in platform and release engineering to join the Metropolis team. The role focuses on developing, building, and maintaining foundational infrastructure and CI/CD systems that run AI and machine learning video analytics workloads at scale using NVIDIA Data Center GPUs. The engineer will establish reliable release workflows, automation systems, and developer tools to improve efficiency on the Metropolis platform.

Responsibilities

  • Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub Actions, GitLab Actions, and runners for Metropolis software products.
  • Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.
  • Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.
  • Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.
  • Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.
  • Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.
  • Troubleshoot hardware and operating system issues across bare-metal and GPU-accelerated servers to minimize downtime and maintain platform stability.

Requirements

  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • More than 6 years of relevant industry experience.
  • Advanced Python skills for scripting, tooling, and automation.
  • Deep expertise with Kubernetes, Helm, and container orchestration in production environments.
  • Demonstrated experience building and maintaining CI/CD pipelines at scale using Jenkins, GitHub Actions, GitLab Actions, runners, or similar technologies.
  • Strong understanding of Linux systems administration, networking, and distributed systems.
  • Experience with release engineering practices, including semantic versioning, release gating, and change management.
  • Hands-on experience with observability stacks such as Prometheus, Grafana, and ELK.

Preferred Qualifications

  • Experience with GPU infrastructure and AI/ML platform engineering at scale.
  • Experience managing bare-metal and hybrid cloud environments, including AWS, Google Cloud Platform, or Azure.
  • Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms.
  • Experience with GitOps workflows and infrastructure as code using Terraform or Ansible.
  • A track record of driving DevOps culture transformation and improving developer experience.

Benefits

The role includes eligibility for equity and benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.

Applications will be accepted at least until August 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

More jobs at Nvidia

Similar jobs