Senior Site Reliability Engineer - Cloud

at Nvidia
USD 168,000-264,500 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 AWS @ 4 Communication @ 6 GEO GPU @ 4 GenAI Generative AI @ 4 Kubernetes @ 7 LLM @ 4 Machine Learning Marketing Python @ 7 SEO SRE @ 4 Vector Databases @ 4

Details

NVIDIA's Digital Marketing Organization is seeking a Senior Site Reliability Engineer to help keep its Digital Marketing Services reliable, fast, and efficient. The role supports production-grade applications, CDN and cloud infrastructure, and AI/ML services.

Responsibilities

  • Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.
  • Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.
  • Provide on-call support for production applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.
  • Author, test, and activate shared and non-shared Cloudlet Policies through the Akamai Cloudlets Policy Manager.
  • Maintain custom match criteria, including Geo, Device Characteristics, RegEx, and Query Strings, to support origin offload and intelligent content delivery.
  • Identify and resolve user-reported problems across the Digital Marketing Organization ecosystem.
  • Onboard applications, AI/ML services, and model endpoints on AWS infrastructure.
  • Implement monitors, alerts, and standard operating procedures for early detection and accurate response to service-impacting issues, including model drift and inference latency.
  • Participate in incident management, including issue recognition, triage, partner communication, impact containment, service restoration, and post-incident follow-up.

Requirements

  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 8 or more years of experience supporting technical operations in a live-site production environment, with an interest in CDN automation, tooling, and infrastructure for AI applications.
  • Strong knowledge of the Kubernetes platform, deployments, and cloud-native automation.
  • Strong problem-solving and root-cause analysis skills, with a focus on optimization and efficiency.
  • Advanced scripting and development experience with Python, including full automation of operational steps.
  • SRE on-call experience is required.

Preferred Qualifications

  • Strong Akamai CDN support skills and an understanding of edge computing and edge AI.
  • Experience with AWS and Kubernetes as a platform.
  • Hands-on experience deploying and scaling generative AI or LLM applications, integrating vector databases, or managing GPU-accelerated infrastructure.
  • Excellent communication, presentation, and analytical skills, including the ability to explain infrastructure and AI concepts to varied audiences.

Benefits

The position offers a competitive salary, equity, and benefits package. NVIDIA is an equal opportunity employer committed to an inclusive work environment.

More jobs at Nvidia

Similar jobs