Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Agentic Systems @ 4
ClickHouse @ 4
Communication @ 6
Debugging @ 4
Deep Learning @ 4
Distributed Systems @ 4
Engineering Management @ 7
GPU @ 4
GenAI @ 4
Grafana @ 4
LLM @ 4
Leadership @ 7
Machine Learning
Observability @ 4
OpenTelemetry @ 4
Performance Analysis @ 7
Profiling @ 4
Prometheus @ 4
Technical Leadership @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is seeking an Engineering Manager to lead the development of an agentic platform for observing, debugging, and optimizing GenAI models deployed at scale. The platform will provide visibility into model behavior, inference performance, reliability, and cost across large-scale GenAI workloads. It will capture, correlate, and analyze logs, traces, metrics, and performance signals across large-scale LLM and VLM deployments to help engineers understand model-serving behavior, find regressions, optimize latency and throughput, and improve the reliability of GenAI systems in production.
You will work across the NVIDIA AI software stack with teams focused on inference serving, model optimization, distributed systems, GPU performance, and production operations. This is a highly cross-functional role for someone who understands deep learning systems, observability, and large-scale software platforms, and who is excited about building agentic workflows that help teams reason over complex telemetry and performance data.
Responsibilities
- Lead, mentor, and grow a team building an agentic platform for monitoring and improving large-scale LLMs and VLMs in production.
- Build systems that collect, correlate, and analyze telemetry across inference servers, GPUs, schedulers, model runtimes, and customer-facing APIs.
- Develop agentic workflows that help engineers identify root causes, explain regressions, and recommend performance optimizations.
- Collaborate with internal customers and business units to align priorities and deliver production-grade platform capabilities.
Requirements
- BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
- 8+ years of relevant software engineering experience, including 3+ years in engineering management or technical leadership.
- Experience leading software engineering teams building large-scale distributed systems, observability platforms, ML infrastructure, or production AI systems.
- Strong understanding of LLM/VLM inference systems, deployment patterns, and production performance challenges.
- Experience with logs, metrics, traces, profiling, alerting, dashboards, or incident/debugging workflows.
- Strong programming, debugging, performance analysis, and test design skills.
- Ability to work across organizations and align technical priorities with product and business goals.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Background in GPU performance analysis, distributed inference, model serving optimization, or reliability engineering.
- Experience building observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure.
- Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, Elastic, or similar observability tools.
- Experience building agentic systems that reason over logs, traces, performance data, incidents, or operational workflows.
- Hands-on experience with production GenAI serving systems and metrics such as TTFT, TPOT, throughput, queueing delay, GPU utilization, KV cache pressure, error rates, and cost per token.
Compensation and Benefits
The base salary range is USD 224,000–356,500 for Level 3 and USD 272,000–431,250 for Level 4. Base salary will be determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications for this job will be accepted at least until July 20, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. The company does not discriminate in hiring or promotion practices on the basis of protected characteristics.