Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Ansible @ 4
CI/CD @ 3
CUDA @ 4
Communication @ 6
Data Pipelines
DevOps @ 7
Distributed Systems
GPU @ 7
Go @ 7
Grafana @ 3
IaC
InfiniBand @ 4
Kubernetes @ 7
LLM
Leadership @ 6
Linux @ 7
MLOps @ 3
Machine Learning
Observability @ 3
OpenTelemetry @ 3
Prometheus @ 3
PyTorch @ 7
Python @ 7
SRE
TensorFlow @ 7
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an NCX Senior Engineer to join its DSX team and collaborate closely with strategic customers to implement and enhance advanced AI workloads. The role provides hands-on technical assistance for AI deployments and distributed systems, helping customers achieve efficient performance from NVIDIA's AI platform across varied environments.
Responsibilities
- Build and deploy custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines based on NVIDIA reference architectures.
- Act as the primary technical contact for strategic NVIDIA Cloud Partners, providing remote and on-site support, troubleshooting complex production problems, and guiding partner engineering teams on NVIDIA platform guidelines.
- Deploy and manage AI workloads across DGX Cloud, NCP data centers, and major cloud service provider environments using Kubernetes, containers, and GPU scheduling systems.
- Profile and tune large-scale training and inference workloads on NCP platforms.
- Implement observability and SLO/SLA monitoring, and lead efforts to reduce latency, cost, and operational risk.
- Implement and expand NVIDIA reference architectures on partner platforms.
- Develop integrations with partner control planes and customer environments, ensuring connectivity across APIs, data pipelines, and enterprise software.
- Create implementation guides, runbooks, and post-mortem documentation for running NVIDIA AI workloads at scale on NCP platforms.
- Collaborate with customer and partner engineering teams to investigate complex technical issues and drive them to root cause and resolution.
Requirements
- Bachelor's, master's, or doctoral degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience.
- At least 8 years of experience in customer-facing technical roles such as Solutions Engineering, DevOps, Site Reliability Engineering, or ML Infrastructure Engineering, ideally supporting large-scale cloud or service provider environments.
- Strong expertise in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant or service-provider platforms.
- Experience supporting large-scale AI/ML training and inference workloads, including LLMs, generative models, or recommendation systems, in production or critically important environments.
- Strong programming skills in Python and Go, with hands-on experience using PyTorch or TensorFlow for training and serving.
- Excellent communication and technical presentation skills, with the ability to explain architectures, trade-offs, and recommendations to engineering and leadership audiences.
Preferred Qualifications
- Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, Triton, NIM, InfiniBand, and RoCE.
- Experience collaborating with NVIDIA Cloud Partners, hyperscale cloud service providers, or managed AI cloud platforms.
- Experience implementing NVIDIA reference architectures for AI infrastructure.
- Deep familiarity with MLOps and cloud-native practices, including containerization, CI/CD pipelines, observability stacks such as Prometheus, Grafana, and OpenTelemetry, and GitOps workflows.
- Experience with infrastructure as code using Terraform, Ansible, or similar tools for repeatable deployment and configuration of GPU-accelerated clusters and NCP building blocks.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. Compensation is determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until July 27, 2026. This posting is for an existing vacancy. NVIDIA is an equal opportunity employer.
More jobs at Nvidia
Senior Software Solutions Engineer
Nvidia · Poland
PLN 230,200-487,500 per year
Senior Offensive Security Engineer, Automotive
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Systems Operations and Administrator
Nvidia · Santa Clara, United States
USD 112,000-218,500 per year
Senior Software Engineer, Agent Simulation and Evaluation
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior AI Product Engineer
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Similar jobs
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Software Engineer, Cloud-Native Stack – CSP Engagements
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year