Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Agentic AI @ 4
Ansible @ 4
Automated Testing
Bash @ 6
CI/CD @ 4
Claude Code @ 4
Codex @ 4
Debugging @ 6
Docker @ 4
GPU
Go
Helm @ 4
Jenkins @ 4
Kubernetes @ 4
Mathematics @ 4
Networking @ 4
Python @ 6
Security
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA Cloud Platform Engineering is seeking a Senior Software Engineer to develop and operate infrastructure automation for NVIDIA GPU Cloud services. The team performs more than 250,000 automation runs per week, operates 24/7, and manages thousands of automation projects across multiple operating systems, cloud platforms, and virtualization technologies. You will collaborate with application developers and QA teams to build automation infrastructure for deployment, testing, monitoring, and CI/CD operations.
Responsibilities
- Use configuration management and workflow management tools such as Ansible and StackStorm to deploy and manage colocated, multiplatform server clusters in NVIDIA data centers worldwide.
- Build infrastructure, tools, and automation scripts for image management, switch management, deployment, data analytics, automated testing, logging, monitoring, and microservice alerting.
- Apply agentic AI to develop AI orchestration skills and manage data center operations at scale using NetBox, Foreman, and Mellanox switches.
- Use infrastructure management technologies such as Kubernetes and Docker to support fast and consistent delivery.
- Own microservice infrastructure and provide operational support to application teams, with a focus on automation services, infrastructure and security improvements, and live service troubleshooting.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Mathematics, Physics, or equivalent experience.
- 3 or more years of proven experience.
- Experience with agentic AI, including Codex and Claude Code.
- Proficiency in Python, Bash, Groovy, and Golang.
- Outstanding debugging skills.
- AWS experience with compute, container, and networking services is preferred.
- Experience with Jenkins and Jenkins Pipeline for CI/CD.
- Experience with configuration management tools such as Ansible is highly advantageous.
- Experience with Packer, Terraform, and StackStorm is advantageous.
- Experience with Kubernetes, Docker, and Helm is highly desirable.
Benefits
NVIDIA offers competitive salaries and a comprehensive benefits package. NVIDIA is an AA/EEO/Veterans/Disabled employer and provides reasonable accommodation for individuals with disabilities during the application or interview process and for performing essential job functions and receiving employment benefits.