Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 6
Debugging @ 4
Go @ 6
LLM @ 4
Leadership @ 7
Mathematics @ 4
Networking @ 4
Observability @ 4
OpenTelemetry @ 4
Python @ 6
Rust @ 6
TypeScript @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building foundational technology for autonomous software systems, including open-source agent harnesses, sandboxed execution, model-provider infrastructure, and agent benchmarks. The role focuses on turning agent-evaluation evidence into improvements in quality, reliability, cost, and latency across models.
Responsibilities
- Define and evolve the technical approach for evaluating and improving open agent harnesses across frameworks, models, and benchmark environments.
- Build and operate scalable, reproducible evaluation systems spanning CI, scheduled compute, sandboxed execution, artifact provenance, trace capture, and analysis.
- Diagnose agent failures using traces, logs, tool results, and benchmark artifacts, distinguishing harness, evaluator, environment, inference, and model causes.
- Design controlled experiments and promotion criteria that distinguish genuine harness improvements from noise, benchmark artifacts, or model-specific gains.
- Ship focused harness and runtime improvements, including upstream open-source contributions, detectors, regression tests, and maintainable documentation.
- Partner with Relay, Hermes, model, infrastructure, and research teams to turn recurring optimization needs into reusable runtime, trace, and evaluation capabilities.
Requirements
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Applied Mathematics, or a related field, or equivalent experience.
- 8+ years of hands-on software-engineering experience, with demonstrated technical ownership of production systems and architecture design leadership.
- Expert Python skills and working proficiency in Rust, Go, C++, or TypeScript, with the ability to debug systems across language and process boundaries.
- Experience building or extending LLM agents, coding agents, tool-use loops, model-provider integrations, developer tools, or evaluation infrastructure.
- Solid understanding of asynchronous execution, subprocesses, containers, sandboxing, callbacks, retries, networking, and distributed compute.
- Experience crafting reproducible experiments, benchmark methodology, performance investigations, regression gates, or reliability analysis under nondeterministic conditions.
- A record of shipping and maintaining open-source or developer-facing software with sound testing, API development, documentation, and code reviews.
Preferred Qualifications
- Meaningful contributions to open-source coding agents, harnesses, agent runtimes, evaluation frameworks, developer tools, or observability projects.
- Published research, technical writing, patents, or substantive open-source work in autonomous software engineering, agent evaluation, reliability, tool use, or inference efficiency.
- Experience with SWE-bench, Terminal-Bench, AgentBench, or comparable public agent benchmarks.
- Experience shipping improvements to context management, tool interfaces, retries, sandboxed repository execution, long-running agent loops, or evaluation environments.
- Experience with trace-analysis and observability systems such as OpenTelemetry, OpenInference, structured event pipelines, exporters, or debugging tools.
Compensation and Benefits
The base salary range is USD 184,000–287,500 for Level 4 and USD 224,000–356,500 for Level 5. The role also includes eligibility for equity and benefits. NVIDIA provides a comprehensive benefits package.
Applications will be accepted at least until September 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.
More jobs at Nvidia
Senior NPI Program Manager
Nvidia · Santa Clara, United States
USD 168,000-258,800 per year
GPU PCIe and Boot Architect - New College Grad 2026
Nvidia · Santa Clara, United States
USD 124,000-241,500 per year
Senior AI Engineer, High Performance AI
Nvidia · Santa Clara, United States
USD 152,000-241,500 per year
Senior Salesforce CPQ Developer
Nvidia · Santa Clara, United States
USD 176,000-276,000 per year
Senior Technical Program Manager - LLM Safety
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year
Similar jobs
Principal Software Engineer — Agentic AI Applications and Foundations
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Software Engineer, Build Systems / CI
OpenAI · New York City, United States, San Francisco, United States, Seattle, United States
USD 185,000-490,000 per year
Systems Generalist, GPT Infrastructure
OpenAI · San Francisco, United States, Seattle, United States
USD 293,000-445,000 per year
Principal Software Engineer – Infrastructure
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
Senior Software Engineer, Fleet Intelligence Agent Systems
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Software Engineer, API Enterprise Controls
OpenAI · San Francisco, United States
USD 293,000-385,000 per year
Senior Software Engineer, AIOps and Observability
Nvidia · Santa Clara, United States
USD 200,000-322,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year