Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Communication @ 6
GEO
GPU @ 4
GenAI
Generative AI @ 4
Kubernetes @ 7
LLM @ 4
Machine Learning
Marketing
Python @ 7
SEO
SRE @ 4
Vector Databases @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Digital Marketing Organization is seeking a Senior Site Reliability Engineer to help keep its Digital Marketing Services reliable, fast, and efficient. The role supports production-grade applications, CDN and cloud infrastructure, and AI/ML services.
Responsibilities
- Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.
- Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.
- Provide on-call support for production applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.
- Author, test, and activate shared and non-shared Cloudlet Policies through the Akamai Cloudlets Policy Manager.
- Maintain custom match criteria, including Geo, Device Characteristics, RegEx, and Query Strings, to support origin offload and intelligent content delivery.
- Identify and resolve user-reported problems across the Digital Marketing Organization ecosystem.
- Onboard applications, AI/ML services, and model endpoints on AWS infrastructure.
- Implement monitors, alerts, and standard operating procedures for early detection and accurate response to service-impacting issues, including model drift and inference latency.
- Participate in incident management, including issue recognition, triage, partner communication, impact containment, service restoration, and post-incident follow-up.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 8 or more years of experience supporting technical operations in a live-site production environment, with an interest in CDN automation, tooling, and infrastructure for AI applications.
- Strong knowledge of the Kubernetes platform, deployments, and cloud-native automation.
- Strong problem-solving and root-cause analysis skills, with a focus on optimization and efficiency.
- Advanced scripting and development experience with Python, including full automation of operational steps.
- SRE on-call experience is required.
Preferred Qualifications
- Strong Akamai CDN support skills and an understanding of edge computing and edge AI.
- Experience with AWS and Kubernetes as a platform.
- Hands-on experience deploying and scaling generative AI or LLM applications, integrating vector databases, or managing GPU-accelerated infrastructure.
- Excellent communication, presentation, and analytical skills, including the ability to explain infrastructure and AI concepts to varied audiences.
Benefits
The position offers a competitive salary, equity, and benefits package. NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Senior MLOps Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Technical Program Manager - Autonomous Vehicles
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Nvidia · Santa Clara, United States
USD 92,000-184,000 per year
Senior Data Analyst - Automotive
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
ML Solution Architect (Early Talent)
Nebius · United States
USD 102-126 per hour
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
Senior Technical Marketing Engineer, Enterprise AI Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year
Senior Staff Machine Learning Engineer, GenAI Platform
Reddit · United States
USD 292,500-409,500 per year
Senior Software Engineer, GenAI Platform
Reddit · United States
USD 190,800-267,100 per year
Senior ML Solutions Architect - Token Factory
Nebius · United States
USD 210,000-260,000 per year