Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 7
Agile
Algorithms @ 6
Communication @ 6
Data Structures @ 6
Debugging @ 7
Distributed Systems @ 7
GPU @ 4
Go @ 7
Kubernetes @ 7
LLM
MLOps @ 4
Machine Learning
Observability @ 7
Python @ 7
Rust @ 7
SRE @ 8
Security @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking to hire a deeply technical, creative, and Senior AI Platform Engineer to build, support, and maintain the next generation of AI-powered enterprise products that improve engineering efficiency, data security, and power product development. This role collaborates with Cloud and AI/ML teams in a multifaceted and agile environment. You will shape the technological future of the organization by ensuring systems are scalable, reliable, and ready for the AI era.
Responsibilities
- Define and lead AI-native infrastructure roadmaps and cross-organizational initiatives.
- Architect and scale LLM/ML infrastructure across cloud-native clusters and on-premises hardware.
- Design and implement observability for infrastructure health and AI model performance.
- Build LLM-aware monitoring and leverage AI to improve incident response and reduce toil.
- Develop automation and tooling to ensure reliability, scalability, and developer self-services.
- Troubleshoot complex distributed systems, including deep Kubernetes and AI/ML scaling challenges.
- Drive AI-assisted engineering practices and mentor engineers to foster an AI-first culture.
- Partner with product engineering and internal business units to translate AI platform capabilities into reliable, scalable solutions that accelerate product development.
Requirements
- 10+ years in cloud, platform, or SRE roles with relevant education or equivalent experience.
- Bachelor’s degree or equivalent experience.
- Strong Python and at least one systems language (C++, Go, or Rust), with proven distributed systems debugging expertise.
- Deep experience building and scaling distributed systems, including Kubernetes and bare-metal infrastructure.
- Strong observability design across infrastructure and AI workloads (metrics, logging, tracing, AI quality signals).
- Hands-on experience operating AI/ML platforms, including MLOps, model serving, and GPU-accelerated environments.
- Experience with infrastructure and application security practices, such as identity/auth, network segmentation, supply chain security, and vulnerability management in cloud-native environments.
- Practical use of AI-assisted development tools and coding agents in daily workflows.
- Solid foundation in data structures, algorithms, and complexity analysis.
- Excellent problem-solving, communication, and collaboration across multiple functions.
Benefits
- Eligible for equity and benefits.
More jobs at Nvidia
HPC Performance Engineer
Nvidia · United States
USD 152,000-241,500 per year
Senior System Software Engineer - Halos Core And Robotics Platform
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer – Dynamo Tools
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Systems Software Engineer - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Perception Engineer, Obstacle Foundation Models - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Senior Cloud Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Germany
PLN 292,500-650,000 per year
Senior Software Engineer, Ai Inference Systems
Nvidia · Toronto, Canada
CAD 170,000-275,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States
USD 179,500-224,300 per year
Senior Storage Software Engineer, DGXC Data Services
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
NCX Engineer, AI Accelerator
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year