Principal Engineer, Cloud Site Reliability Engineering

at Nvidia
USD 272,000-431,200 per year
SENIOR
✅ On-site

Tech Stack

AI @ 8 API @ 4 Algorithms @ 6 Android CI/CD Cassandra @ 4 Ceph @ 6 Chef @ 6 Deep Learning @ 6 Distributed Systems @ 4 Docker @ 8 ElasticSearch @ 4 GPU Git @ 6 Hadoop @ 6 Java @ 7 Kafka @ 6 Kubernetes @ 6 Linux Machine Learning @ 6 MongoDB @ 4 MySQL @ 4 NoSQL @ 4 OpenStack @ 6 Puppet @ 6 Python @ 7 SQL @ 4 SRE Software Development @ 8

Details

NVIDIA is seeking a Cloud Site Reliability Engineering Architect to join IPP's Cloud Infrastructure Team. The global team supports infrastructure needs across NVIDIA, including Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence, and Autonomous Vehicles. Its cloud services run nearly half a million automated jobs per day on thousands of servers, supporting software engineers worldwide across Windows, Linux, and Android systems, as well as NVIDIA GPUs and Tegra processors. The team delivers unified CI/CD solutions and cloud-based software development infrastructure.

Responsibilities

  • Serve as an SRE Architect on the GPU Private Cloud team, supporting interactive development, centralized CI/CD, and QA testing for NVIDIA employees globally.
  • Evaluate, identify, and develop software solutions to optimize critical software development workflows across NVIDIA organizations.
  • Architect, implement, and support end-to-end CI/CD systems using open-source and NVIDIA proprietary software.
  • Onboard internal NVIDIA development teams to private cloud infrastructure by discovering use cases and identifying available cloud solutions.
  • Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.
  • Lead software development projects and provide technical direction to engineering teams.
  • Identify and resolve problems within software systems.
  • Develop and implement critical metrics using analytics methods and dashboards.

Requirements

  • Bachelor's or master's degree in Electrical Engineering, Computer Science, or a relevant field, or equivalent experience.
  • 15+ years of systems software development experience, including at least one year developing or exploring AI.
  • Experience maintaining cloud infrastructure and highly available production environments.
  • Strong programming and software development skills in Java, Python, and shell scripting.
  • Good understanding of distributed systems and REST APIs.
  • Experience with SQL and NoSQL database systems such as MySQL, Cassandra, MongoDB, or Elasticsearch.
  • Excellent knowledge of and experience with Docker containers and virtual machines.
  • Background in cloud technologies including OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, and Kafka.
  • Ability to work effectively across organizational boundaries in a multinational, multi-time-zone corporate environment.

Preferred Qualifications

  • Depth in artificial intelligence, machine learning, and deep learning algorithms and techniques.
  • Strong collaborative and interpersonal skills, with a record of guiding and influencing others in dynamic environments.
  • Experience developing large-scale software systems using modular architecture under real-time performance requirements.
  • Experience designing high-performance, scalable software systems with a focus on hardware cost optimization.

Compensation and Benefits

  • Base salary range: $272,000–$431,250 USD per year.
  • Eligible for equity and benefits.
  • Applications will be accepted at least until August 9, 2026.
  • NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs