Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
Automated Testing @ 4
CI/CD @ 6
Communication @ 6
Compliance
Computer Vision @ 4
DevOps @ 7
Docker @ 6
GPU @ 4
Kubernetes @ 6
Machine Learning
Networking
Python @ 6
Robotics @ 6
SRE @ 4
Stress Testing @ 4
gRPC @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are seeking a Senior DevOps / Cloud Simulation Infrastructure Engineer to own the complete end-to-end cloud execution pipeline for SimReady assets. This role enables the transition from local, workstation-driven validation to high-scale, automated cloud validation on NVIDIA Cloud Functions (NVCF). You will be responsible for deploying a robust, multi-GPU pipeline that supports structural validation, AI-driven runtime behavioral testing, and automated asset remediation.
Responsibilities
- Deploy full Isaac Sim runtimes within GPU-aware NVCF containers, including container packaging, GPU initialization, and runtime utilities for physics, sensor, and rendering validation.
- Deploy services to validate USD structure and compliance without runtime overhead.
- Architect scalable execution layers for runtime behavior-based testing, such as drop and grasp tests.
- Deploy rule-based or AI-based systems for automated pass/fail grading.
- Develop an AI-based pipeline that intercepts failures, triggers automated asset fixes, and re-validates results.
- Scale execution from single-workstation validation to massive, multi-GPU cloud environments.
- Optimize performance by addressing function-to-function networking, gRPC bottlenecks, and in-cluster proxy behavior.
- Automate the generation of verification videos, thumbnails, feature-level reports, and validation metadata.
- Ensure that all assets are traceable and linked to quality gates.
- Establish robust CI/CD, cluster verification, and monitoring pipelines.
- Implement logging, metrics, and tracing to ensure services are observable, debuggable, and production-ready.
- Implement atomic update semantics and safe failure handling to ensure validation processes never corrupt the primary asset library.
Requirements
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 8+ years of professional experience working on DevOps and/or cloud simulation.
- Extensive experience in production-grade DevOps, SRE, or Infrastructure Engineering, with a focus on GPU-backed cloud services.
- Proven expertise in container orchestration using Kubernetes and Docker, as well as CI/CD pipeline development.
- Experience with automated testing frameworks, preferably involving AI/ML inference, computer vision, or rule-based validation.
- Proficiency in Python and systems scripting for test orchestration and pipeline automation.
- Strong ability to design and maintain distributed job lifecycle services, including submit, poll, fetch, and cancel operations, and to handle asynchronous failure states.
- Ability to diagnose and solve distributed network bottlenecks, including gRPC and function-to-function communication.
Additional Qualifications
- Direct experience deploying services on NVCF or DGX Cloud.
- Deep familiarity with Isaac Sim, Omniverse, USD, or Sensor RTX workflows.
- Background in robotics simulation, physical AI, or large-scale content creation pipelines.
- Experience building self-healing or automated remediation workflows.
- Experience with cluster verification frameworks, stress testing, and deployment validation at scale.
Benefits
The role includes eligibility for equity and benefits.
Applications for this job will be accepted at least until July 21, 2026. NVIDIA uses AI tools in its recruiting processes and is committed to fostering an inclusive work environment and providing equal employment opportunities.