Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Communication @ 6
Debugging @ 4
Deep Learning @ 4
Distributed Systems @ 4
Engineering Management @ 7
GPU
GenAI @ 4
LLM @ 7
Leadership @ 7
Machine Learning
Observability @ 4
Performance Analysis @ 7
Profiling @ 4
Technical Leadership @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale. In this role, you will lead a team building an agentic platform that provides visibility into model behavior, inference performance, reliability, and cost across large-scale GenAI workloads.
This agentic platform will capture, correlate, and analyze logs, traces, metrics, and performance signals across large scale LLM and VLM deployments. It will help engineers understand model-serving behavior, find regressions, optimize latency and throughput, and improve the reliability of GenAI systems in production.
You will work across the NVIDIA AI software stack with teams focused on inference serving, model optimization, distributed systems, GPU performance, and production operations. This is a highly cross-functional role for someone who understands deep learning systems, observability, and large-scale software platforms, and who is excited about building agentic workflows that help teams reason over complex telemetry and performance data.
Responsibilities
- Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
- Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
- Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
- Collaborate with internal customers and business units to align priorities and deliver production grade platform capabilities.
Requirements
- BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
- 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
- Experience leading software engineering teams building large-scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
- Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
- Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
- Strong programming, debugging, performance analysis, and test design skills.
- Ability to work across organizations and align technical priorities with product and business goals.
- Excellent communication and collaboration skills.
Benefits
You will also be eligible for equity and benefits.