Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Algorithms @ 6
Ansible @ 4
ArgoCD @ 4
CI/CD @ 4
Chef @ 4
Data Structures @ 6
Debugging @ 6
Distributed Systems @ 6
Docker @ 4
GPU
Git @ 4
Go @ 6
Grafana @ 4
Helm
Java @ 6
Kubernetes @ 4
Linux @ 6
Observability @ 4
OpenStack @ 4
Performance Monitoring
Prometheus @ 4
Puppet @ 4
Python @ 6
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Systems Engineering focuses on building, automating, and operating platforms and tooling that deliver large-scale production systems with high efficiency, reliability, and velocity. The role combines software and systems engineering practices across infrastructure automation, containerized platforms, storage, telemetry, and observability.
NVIDIA's team ensures that internal and external GPU cloud services are deployed reliably, observable end-to-end, and continuously improved through automation. The team uses repeatable CI/CD pipelines, Kubernetes-based deployments, capacity and performance monitoring, self-service tooling, blameless postmortems, proactive failure-mode identification, and iterative improvement.
Responsibilities
- Design, deploy, and operate Kubernetes solutions for large-scale storage and data platforms, including manifests, Helm charts, and operators.
- Build tools, services, and automation that improve the lifecycle of storage and data systems, from provisioning and configuration through deployment, scaling, and day-two operations.
- Develop and operate telemetry and observability for production systems, including metrics, logging, tracing, dashboards, and alerting.
- Diagnose and resolve complex issues across distributed, containerized infrastructure using strong analytical troubleshooting skills.
- Collaborate with peers and partner teams to improve services from inception and design through deployment, operation, and refinement.
- Scale systems sustainably through automation, infrastructure-as-code, and CI/CD.
- Support services before launch through deployment automation, capacity planning, and launch and readiness reviews.
- Practice sustainable incident response and postmortems.
- Participate in an on-call rotation supporting production systems.
Requirements
- Bachelor's degree or equivalent experience in Computer Science or a related technical field involving coding.
- 12 or more years of practical experience.
- Hands-on experience deploying, configuring, and operating production workloads and solutions on Kubernetes.
- Experience building tools and services for storage, data, or platform infrastructure.
- Solid software design fundamentals, including algorithms, data structures, and complexity analysis, on large-scale Linux-based systems.
- Experience building and operating telemetry and observability using tools such as Prometheus, InfluxDB, Grafana, and the Elastic stack.
- Strong analytical troubleshooting skills with a systematic, root-cause-driven approach.
- Proficiency in one or more of Python, Go, or Java.
- Knowledge of infrastructure configuration management and infrastructure-as-code tools such as Ansible, Chef, Puppet, ArgoCD, Git Pipelines, and Terraform.
Preferred Qualifications
- Customer-first mindset with a focus on customer satisfaction and customer success.
- Experience with Git, code review, pipelines, and CI/CD.
- Experience using or operating large private and public cloud systems based on Kubernetes, OpenStack, and Docker.
- Interest in designing, analyzing, debugging, and fixing large-scale distributed systems.
- Experience designing storage- or data-focused tooling and automating its operation at scale.
- Ability to thrive in collaborative environments and adapt to different working styles.
Compensation and Benefits
The base salary range is USD 208,000–333,500 for Level 5 and USD 256,000–414,000 for Level 6. Compensation is determined based on location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 31, 2026. NVIDIA is an equal opportunity employer committed to an inclusive work environment.