Senior Staff Platform Engineer

at Nvidia
USD 200,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 AWS @ 4 AWS CloudFront @ 6 Azure @ 4 Cloudflare @ 6 Distributed Systems @ 6 GPU @ 4 Go @ 7 HPC HTTP @ 7 Kubernetes Linux @ 7 Machine Learning Networking @ 6 Observability Python @ 7 Security @ 6 Terraform

Details

NVIDIA is seeking a Senior Staff Platform Engineer to architect, build, and scale foundational infrastructure for demanding compute and AI/ML workloads. The role spans distributed systems, cloud, networking, content delivery, automation, and reliability engineering, with broad influence across Cloud, Networking, Security, AI/ML, and Developer Infrastructure teams.

Responsibilities

  • Architect, build, and scale highly available platform services for AI/ML, distributed compute, and data-intensive workloads.
  • Own infrastructure from architecture and design through implementation, production readiness, and large-scale adoption.
  • Work across cloud, compute, GPU infrastructure, networking, storage, DNS, load balancing, proxies, traffic management, and content delivery.
  • Advance CDN and edge infrastructure, including HTTP caching, origin architecture, TLS, WAF, rate limiting, traffic routing, and global load balancing.
  • Drive automation, infrastructure-as-code, and self-service capabilities using technologies such as Python, Go, Kubernetes, and Terraform.
  • Use observability, capacity analytics, incident learnings, and performance data to improve reliability, scalability, efficiency, and operational simplicity.
  • Collaborate with Cloud, Networking, Security, AI/ML, and infrastructure teams to solve complex cross-domain problems.
  • Independently drive architecture and implementation across multiple teams and mentor other engineers.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience.
  • 12 or more years of relevant industry experience.
  • Proven success architecting, building, and operating large-scale distributed platforms or infrastructure systems in production.
  • Strong technical depth in several areas, including cloud infrastructure, distributed systems, networking, compute, storage, platform engineering, or content delivery.
  • Deep knowledge of Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, proxies, load balancing, availability, scalability, and fault-tolerant system design.
  • Strong programming and automation skills with Python, Go, or similar languages, along with hands-on experience with infrastructure-as-code and orchestration.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Ability to troubleshoot complex systems across application, operating system, network, and infrastructure layers.
  • Ability to communicate effectively in complex situations and mentor other engineers.

Preferred Qualifications

  • Experience building platforms for AI/ML training, inference, model serving, GPU-accelerated workloads, distributed compute, or high-performance computing.
  • Deep expertise with CDN and edge platforms such as Akamai, AWS CloudFront, Fastly, or Cloudflare, including caching, origin design, WAF, DNS, TLS, and global traffic management.
  • Experience developing self-service platform capabilities that enable engineering teams to consume infrastructure reliably and at scale.
  • Proven use of SLIs, SLOs, error budgets, capacity analytics, and reliability metrics to deliver measurable improvements.
  • Experience distributing models, datasets, containers, software artifacts, or other large objects across globally distributed environments.

Benefits

  • Competitive salary.
  • Comprehensive benefits package.
  • Equity eligibility.

NVIDIA is an equal opportunity employer committed to fostering an inclusive work environment. Applications will be accepted at least until August 23, 2026.

More jobs at Nvidia

Similar jobs