Senior Software Engineer, Evaluation Flywheel — Autonomous Vehicles
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Communication @ 6
Data Engineering @ 7
Data Pipelines @ 7
LLM
Leadership @ 6
Machine Learning @ 8
Python @ 7
Robotics @ 4
Software Development @ 8
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is building the future of autonomous driving, and evaluation is how the organization measures whether the driving system is improving. The AV Evaluation team owns the metrics, golden datasets, and closed-loop evaluation workflows that determine what ships in NVIDIA's self-driving stack. The team is applying frontier vision-language models and agentic techniques to evaluation at scale.
This role owns the evaluation flywheel end to end, including its architecture, quality, adoption, implementation, and technical direction.
Responsibilities
- Own the evaluation flywheel's strategy and architecture, including how road and simulation driving data becomes curated golden datasets, how metrics are measured against them using precision and recall, and how results earn lasting trust with dependent teams.
- Set standards for evaluation quality, including golden dataset curation, versioning, health, metric performance measurement, and release processes.
- Build tooling that enables metric developers to iterate quickly, including self-service dataset pipelines, metric performance measurement, and quality reporting.
- Partner with senior engineers and leaders across test engineering, behavior planning, and infrastructure to set expectations, resolve trade-offs, and represent evaluation quality in cross-team decisions.
- Work directly with AI model developers to improve evaluation iteration speed, including the use of learned and vision-language-model-based evaluation.
Requirements
- Track record of independent execution and technical leadership, including identifying high-leverage problems, driving initiatives across team boundaries, and delivering without waiting to be asked.
- Clear and proactive communication with engineers and senior leaders in a fast-paced environment.
- Bachelor's or master's degree in Computer Science, Robotics, or a related field, or equivalent experience.
- At least 12 years of software development experience, including significant experience in autonomous vehicles, robotics, or large-scale machine learning systems.
- Deep experience evaluating machine learning or robotic systems, including metric design, ground-truth and golden dataset curation, precision/recall methodology, and the supporting data pipelines.
- Strong Python and data engineering skills for production-scale pipelines.
Preferred Qualifications
- Experience building an evaluation flywheel involving dataset curation, metric measurement, and developer tooling, with demonstrated impact on model development velocity.
- Experience with closed-loop simulation evaluation for autonomous driving and related realism considerations.
- Experience applying large language models or vision-language models to evaluation, or productionizing the infrastructure behind them.
- History of earning trust for metrics across skeptical partner teams.
Compensation and Benefits
The base salary range is $224,000–$356,500 USD, determined by location, experience, and the pay of employees in similar positions. The role is also eligible for equity and benefits.
Applications will be accepted at least until July 31, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.