Senior Devops Engineer, Cloud Simulation Infrastructure

at Nvidia
USD 184,000-287,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 Automated Testing @ 4 CI/CD @ 6 Communication @ 6 Compliance Computer Vision @ 4 DevOps @ 7 Docker @ 6 GPU @ 4 Kubernetes @ 6 Machine Learning Networking Observability Python @ 6 SRE @ 4 gRPC @ 6

Details

We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF). You will be responsible for deploying a robust, multi-GPU pipeline that supports structural validation, AI-driven runtime behavioral testing, and automated asset remediation.

Responsibilities

  • Deployment: Deploy full Isaac Sim runtimes within GPU-aware NVCF containers. Manage container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.
  • Deploy Structural Validation: Deploy services to validate USD structure and compliance without runtime overhead.
  • Deploy Runtime Validation: Architect scalable execution layers to conduct runtime behavior-based testing (e.g., drop/grasp tests). Deploy rule-based systems or AI based systems for automated pass/fail grading.
  • Deploy Automated Remediation: Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results to ensure quality standards.
  • Cloud Infrastructure Ownership: Scale execution from single-workstation validation to massive, multi-GPU cloud environments. Optimize for performance, addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.
  • Artifact & Evidence Pipeline: Automate the generation of verification videos, thumbnails, feature-level reports, and validation metadata. Ensure all assets are traceable and linked to quality gates.
  • Observability & CI/CD: Establish robust CI/CD, cluster verification, and monitoring pipelines. Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.
  • Operational Reliability: Implement atomic update semantics and safe failure handling to ensure validation processes never corrupt the primary asset library.

Requirements

  • BS or MS degree in Computer Science, Computer Engineering, or related field (or equivalent experience).
  • 8+ years of professional experience working on DevOps and/or cloud simulation.
  • Extensive experience in production-grade DevOps, SRE, or Infrastructure Engineering, with a focus on GPU-backed cloud services.
  • Proven expertise in container orchestration (Kubernetes/Docker) and CI/CD pipeline development.
  • Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation.
  • Proficiency in Python and systems scripting for test orchestration and pipeline automation.
  • Strong ability to design and maintain distributed job lifecycle services (submit/poll/fetch/cancel) and handle asynchronous failure states.
  • Ability to diagnose and solve distributed network bottlenecks, including gRPC and function-to-function communication.

Benefits

  • You will also be eligible for equity and benefits.

More jobs at Nvidia

Similar jobs