Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
AWS @ 6
Agentic AI @ 4
Azure @ 6
CI/CD @ 6
Communication @ 6
Data Pipelines @ 6
Datadog @ 4
DevOps @ 8
GCP @ 6
Go @ 6
Grafana @ 4
Hiring @ 6
IaC
Kubernetes @ 6
LLM
Leadership @ 6
Observability @ 4
Prometheus @ 4
Python @ 6
RAG
SRE @ 8
ServiceNow
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking a hands-on technical leader to build and lead a high-performance engineering organization that architects, delivers, and operates production-grade software systems at global scale. The role focuses on transforming enterprise IT operations from manual, reactive workflows into fully automated, AI-driven platforms that scale with NVIDIA's growth.
Responsibilities
- Architect and ship agentic AI systems using LLM-based agents, tool calling, RAG, and orchestration frameworks for production-grade AI-assisted operations across employee support, endpoint services, and IT support operations.
- Design and deploy autonomous AI agents that execute complex, multi-step enterprise workflows, including approvals, vendor handoffs, cross-system data reconciliation, and exception handling with human-in-the-loop controls.
- Engineer integration and automation platforms spanning ServiceNow, ERP and procurement systems, and endpoint-management platforms.
- Own the full-stack infrastructure, data pipelines, APIs, and user-facing applications.
- Set engineering standards through hands-on technical leadership, production coding, rigorous code reviews, and system design for critical components.
- Recruit, develop, and retain engineering talent while building a culture of engineering excellence, ownership, and continuous delivery.
- Define and execute a multi-quarter technical roadmap for automation and agentic operations across enterprise IT, tied to measurable business outcomes such as cost reduction, throughput, SLA improvement, and headcount avoidance.
- Drive project prioritization, milestone tracking, capacity planning, and on-time delivery.
- Own talent strategy, including hiring pipelines, performance calibration, and career development.
Requirements
- Bachelor's or Master's degree in a related field, or equivalent experience.
- 10+ years of hands-on software engineering experience, with deep expertise in at least one of Infrastructure, SRE, DevOps, or Production Engineering.
- 5+ years leading engineering teams, including hiring, developing, and managing IT engineers.
- Experience building engineering teams from zero and scaling them in a high-growth, high-ambiguity environment.
- Deep expertise designing and shipping production software systems, including integrations, automation platforms, and data pipelines for complex enterprise operations at scale.
- Experience modernizing enterprise IT operations platforms, such as asset management, endpoint services, IT supply chain, and infrastructure operations.
- Experience deploying agentic AI into production, including multi-step autonomous execution, human-in-the-loop safeguards, exception handling, governance frameworks, and measurable business outcomes.
- Production-grade proficiency with infrastructure as code, CI/CD, Kubernetes, and at least one cloud platform: AWS, GCP, or Azure.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, PagerDuty, or similar.
- Fluency in Python, Go, or equivalent programming languages, with the ability to architect, write, and review production-quality code.
- Executive-level communication skills and the ability to influence technical direction across engineering, product, and senior leadership.
- Ability to translate complex technical capabilities into quantifiable business value and present to VP and C-level audiences.
Compensation and Benefits
The base salary range is $248,000-$391,000 USD, determined by location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until September 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Principal ML Solutions Architect - Token Factory
Nebius · United States
USD 208,000-261,000 per year
Senior Full-Stack Lead Engineer
Nvidia · Santa Clara, United States
USD 224,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer, DGX Cloud Orchestration
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year