Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 4
Communication @ 6
GEO
GPU @ 4
GenAI
Generative AI @ 4
Kubernetes @ 7
LLM @ 4
Machine Learning
Marketing
Python @ 7
SEO
SRE @ 4
Vector Databases @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's Digital Marketing Organization is seeking a Senior Site Reliability Engineer to help keep its Digital Marketing Services reliable, fast, and efficient. The role supports production-grade applications, CDN and cloud infrastructure, and AI/ML services.
Responsibilities
- Build and deploy large-scale dynamic URL redirects using Akamai Edge Redirector Cloudlets for promotional efforts and site migrations.
- Configure Akamai Forward Rewrite Cloudlets to map inbound requests to SEO-friendly paths.
- Provide on-call support for production applications, responding to incidents, prioritizing issues, and driving resolution across deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.
- Author, test, and activate shared and non-shared Cloudlet Policies through the Akamai Cloudlets Policy Manager.
- Maintain custom match criteria, including Geo, Device Characteristics, RegEx, and Query Strings, to support origin offload and intelligent content delivery.
- Identify and resolve user-reported problems across the Digital Marketing Organization ecosystem.
- Onboard applications, AI/ML services, and model endpoints on AWS infrastructure.
- Implement monitors, alerts, and standard operating procedures for early detection and accurate response to service-impacting issues, including model drift and inference latency.
- Participate in incident management, including issue recognition, triage, partner communication, impact containment, service restoration, and post-incident follow-up.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 8 or more years of experience supporting technical operations in a live-site production environment, with an interest in CDN automation, tooling, and infrastructure for AI applications.
- Strong knowledge of the Kubernetes platform, deployments, and cloud-native automation.
- Strong problem-solving and root-cause analysis skills, with a focus on optimization and efficiency.
- Advanced scripting and development experience with Python, including full automation of operational steps.
- SRE on-call experience is required.
Preferred Qualifications
- Strong Akamai CDN support skills and an understanding of edge computing and edge AI.
- Experience with AWS and Kubernetes as a platform.
- Hands-on experience deploying and scaling generative AI or LLM applications, integrating vector databases, or managing GPU-accelerated infrastructure.
- Excellent communication, presentation, and analytical skills, including the ability to explain infrastructure and AI concepts to varied audiences.
Benefits
The position offers a competitive salary, equity, and benefits package. NVIDIA is an equal opportunity employer committed to an inclusive work environment.
More jobs at Nvidia
User Interface - User Experience Designer
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior Application Engineer, HPC and AI for Physics
Nvidia · United States
USD 140,000-270,200 per year
Senior QA Software Engineer, Networking
Nvidia · Warsaw, Poland
PLN 157,500-357,500 per year
Senior System Software Engineer
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Similar jobs
ML Solution Architect (Early Talent)
Nebius · United States
USD 102-126 per hour
Senior Architect, Agentic AI for Marketing
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Technical Marketing Engineer, Enterprise AI Software
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Forward Deployed Engineer, Ecosystem
Nebius · United States
USD 208,800-261,000 per year
Senior Staff Machine Learning Engineer, GenAI Platform
Reddit · United States
USD 292,500-409,500 per year
Senior ML Solutions Architect - Token Factory
Nebius · United States
USD 210,000-260,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year