Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Ansible @ 4
Azure @ 4
CI/CD @ 4
Change Management @ 4
DevOps @ 4
Distributed Systems @ 7
GPU @ 4
GitHub @ 4
GitHub Actions @ 4
Grafana @ 4
Helm @ 6
IaC
Jenkins @ 4
Kubernetes @ 6
Linux @ 7
Machine Learning
Networking @ 7
Observability @ 4
Prometheus @ 4
Python @ 6
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior DevOps Platform Engineer skilled in platform and release engineering to join the Metropolis team. The role focuses on developing, building, and maintaining foundational infrastructure and CI/CD systems that run AI and machine learning video analytics workloads at scale using NVIDIA Data Center GPUs. The engineer will establish reliable release workflows, automation systems, and developer tools to improve efficiency on the Metropolis platform.
Responsibilities
- Compose, build, and maintain scalable CI/CD pipelines using Jenkins, GitHub Actions, GitLab Actions, and runners for Metropolis software products.
- Develop and manage Kubernetes-based platform infrastructure supporting AI/ML workloads on NVIDIA Data Center GPUs.
- Build and implement scaling and performance measurement frameworks within Kubernetes to ensure platform reliability and efficiency under AI/ML workload demands.
- Define and implement release engineering processes, branching strategies, versioning standards, and gating criteria.
- Drive developer efficiency by building and maintaining DevOps MCP servers, tooling, and automation frameworks.
- Own observability and monitoring infrastructure using Prometheus, Grafana, and log aggregation pipelines.
- Troubleshoot hardware and operating system issues across bare-metal and GPU-accelerated servers to minimize downtime and maintain platform stability.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- More than 6 years of relevant industry experience.
- Advanced Python skills for scripting, tooling, and automation.
- Deep expertise with Kubernetes, Helm, and container orchestration in production environments.
- Demonstrated experience building and maintaining CI/CD pipelines at scale using Jenkins, GitHub Actions, GitLab Actions, runners, or similar technologies.
- Strong understanding of Linux systems administration, networking, and distributed systems.
- Experience with release engineering practices, including semantic versioning, release gating, and change management.
- Hands-on experience with observability stacks such as Prometheus, Grafana, and ELK.
Preferred Qualifications
- Experience with GPU infrastructure and AI/ML platform engineering at scale.
- Experience managing bare-metal and hybrid cloud environments, including AWS, Google Cloud Platform, or Azure.
- Familiarity with NVIDIA Metropolis, DeepStream, or similar AI video analytics platforms.
- Experience with GitOps workflows and infrastructure as code using Terraform or Ansible.
- A track record of driving DevOps culture transformation and improving developer experience.
Benefits
The role includes eligibility for equity and benefits. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.
Applications will be accepted at least until August 17, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.