Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
AWS @ 4
AWS CloudFront @ 6
Azure @ 4
Cloudflare @ 6
Distributed Systems @ 6
GPU @ 4
Go @ 7
HPC
HTTP @ 7
Kubernetes
Linux @ 7
Machine Learning
Networking @ 6
Observability
Python @ 7
Security @ 6
Terraform
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a Senior Staff Platform Engineer to architect, build, and scale foundational infrastructure for demanding compute and AI/ML workloads. The role spans distributed systems, cloud, networking, content delivery, automation, and reliability engineering, with broad influence across Cloud, Networking, Security, AI/ML, and Developer Infrastructure teams.
Responsibilities
- Architect, build, and scale highly available platform services for AI/ML, distributed compute, and data-intensive workloads.
- Own infrastructure from architecture and design through implementation, production readiness, and large-scale adoption.
- Work across cloud, compute, GPU infrastructure, networking, storage, DNS, load balancing, proxies, traffic management, and content delivery.
- Advance CDN and edge infrastructure, including HTTP caching, origin architecture, TLS, WAF, rate limiting, traffic routing, and global load balancing.
- Drive automation, infrastructure-as-code, and self-service capabilities using technologies such as Python, Go, Kubernetes, and Terraform.
- Use observability, capacity analytics, incident learnings, and performance data to improve reliability, scalability, efficiency, and operational simplicity.
- Collaborate with Cloud, Networking, Security, AI/ML, and infrastructure teams to solve complex cross-domain problems.
- Independently drive architecture and implementation across multiple teams and mentor other engineers.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience.
- 12 or more years of relevant industry experience.
- Proven success architecting, building, and operating large-scale distributed platforms or infrastructure systems in production.
- Strong technical depth in several areas, including cloud infrastructure, distributed systems, networking, compute, storage, platform engineering, or content delivery.
- Deep knowledge of Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, proxies, load balancing, availability, scalability, and fault-tolerant system design.
- Strong programming and automation skills with Python, Go, or similar languages, along with hands-on experience with infrastructure-as-code and orchestration.
- Experience with AWS, Azure, or Google Cloud Platform.
- Ability to troubleshoot complex systems across application, operating system, network, and infrastructure layers.
- Ability to communicate effectively in complex situations and mentor other engineers.
Preferred Qualifications
- Experience building platforms for AI/ML training, inference, model serving, GPU-accelerated workloads, distributed compute, or high-performance computing.
- Deep expertise with CDN and edge platforms such as Akamai, AWS CloudFront, Fastly, or Cloudflare, including caching, origin design, WAF, DNS, TLS, and global traffic management.
- Experience developing self-service platform capabilities that enable engineering teams to consume infrastructure reliably and at scale.
- Proven use of SLIs, SLOs, error budgets, capacity analytics, and reliability metrics to deliver measurable improvements.
- Experience distributing models, datasets, containers, software artifacts, or other large objects across globally distributed environments.
Benefits
- Competitive salary.
- Comprehensive benefits package.
- Equity eligibility.
NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment. Applications will be accepted at least until August 23, 2026.
More jobs at Nvidia
Senior Compute Platform Engineer, LSF
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Deep Learning Engineer, Accuracy Evaluation
Nvidia · Poland
PLN 375,000-650,000 per year
ML and Agentic Systems Engineer
Nvidia · Santa Clara, United States
USD 224,000-431,200 per year
Senior Product Architect, K8s-Based AI Infrastructure
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior System Software Engineer - SoC Power
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Similar jobs
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Staff+ Software Engineer, Infrastructure (Distributed Systems)
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 320,000-485,000 per year
Senior Staff+ Software Engineer, Node Infra
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 405,000-485,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior DevOps Engineer, Platform Engineering
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Software Engineer, AI Inference Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Software Engineer, Infrastructure (Distributed Systems)
Anthropic · London, United Kingdom
GBP 325,000-390,000 per year