Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Automated Testing @ 4
CI/CD @ 6
Communication @ 6
Compliance
Computer Vision @ 4
DevOps @ 7
Docker @ 6
GPU @ 4
Kubernetes @ 6
Machine Learning
Networking
Observability
Python @ 6
SRE @ 4
gRPC @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets! This role is critical to our product strategy, enabling us to transition from local, workstation-driven validation to high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF). You will be responsible for deploying a robust, multi-GPU pipeline that supports structural validation, AI-driven runtime behavioral testing, and automated asset remediation.
Responsibilities
- Deployment: Deploy full Isaac Sim runtimes within GPU-aware NVCF containers. Manage container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.
- Deploy Structural Validation: Deploy services to validate USD structure and compliance without runtime overhead.
- Deploy Runtime Validation: Architect scalable execution layers to conduct runtime behavior-based testing (e.g., drop/grasp tests). Deploy rule-based systems or AI based systems for automated pass/fail grading.
- Deploy Automated Remediation: Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results to ensure quality standards.
- Cloud Infrastructure Ownership: Scale execution from single-workstation validation to massive, multi-GPU cloud environments. Optimize for performance, addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.
- Artifact & Evidence Pipeline: Automate the generation of verification videos, thumbnails, feature-level reports, and validation metadata. Ensure all assets are traceable and linked to quality gates.
- Observability & CI/CD: Establish robust CI/CD, cluster verification, and monitoring pipelines. Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.
- Operational Reliability: Implement atomic update semantics and safe failure handling to ensure validation processes never corrupt the primary asset library.
Requirements
- BS or MS degree in Computer Science, Computer Engineering, or related field (or equivalent experience).
- 8+ years of professional experience working on DevOps and/or cloud simulation.
- Extensive experience in production-grade DevOps, SRE, or Infrastructure Engineering, with a focus on GPU-backed cloud services.
- Proven expertise in container orchestration (Kubernetes/Docker) and CI/CD pipeline development.
- Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation.
- Proficiency in Python and systems scripting for test orchestration and pipeline automation.
- Strong ability to design and maintain distributed job lifecycle services (submit/poll/fetch/cancel) and handle asynchronous failure states.
- Ability to diagnose and solve distributed network bottlenecks, including gRPC and function-to-function communication.
Benefits
- You will also be eligible for equity and benefits.
More jobs at Nvidia
HPC Performance Engineer
Nvidia · United States
USD 152,000-241,500 per year
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Dynamo Tools
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Staff Forward Deployed Engineer
GitLab · United States
USD 254,000-297,000 per year
Senior Engineering Manager, Object Storage - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-488,800 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Site Reliability Engineer (In-Office Required)
Nebius · New York City, United States
USD 156,000-262,000 per year
ML Solutions Architect (Early Talent)
Nebius · United States
USD 102,000-126,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year