Evaluation And ML Systems Engineer, AI Safety And Security Engineering
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Data Pipelines @ 6
LLM @ 4
Machine Learning
Python @ 6
Security @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's AI Safety & Security Engineering team builds and evaluates AI-powered tooling that helps find, validate, and patch software vulnerabilities. The team believes open-weight models are foundational to American AI leadership and cybersecurity, and that trust in AI grows through participation, transparency, and broad scientific scrutiny.
The Evaluation/ML-Systems Engineer will own how the program measures progress, separating real capability from anecdote. The role is responsible for creating benchmarks, runs, and reproducible code so that conclusions can be traced to reliable evidence. The engineer will design metrics, choose baselines, write fair comparison protocols, and build infrastructure that connects every result to the run that produced it.
The role involves partnering with security researchers and platform engineers to design experiments that can be trusted and rerun. The engineer will also help automate measurement processes and support internal review of results and conclusions.
Responsibilities
- Build benchmarking and reproducibility systems for evaluation infrastructure.
- Define metrics and protocols used to measure ML systems.
- Map every result to the code and runs that produced it.
- Keep findings reviewable and conclusions traceable.
- Design benchmarks, metrics, and statistically sound comparisons for ML systems.
- Build shared Python infrastructure, including experiment tracking and data pipelines.
- Partner with security researchers and platform engineers to design reproducible experiments.
- Automate routine measurement tasks.
Requirements
- Bachelor's degree or equivalent experience.
- 5+ years of experience in ML engineering or evaluation.
- Experience designing benchmarks, metrics, and statistically sound comparisons for ML systems.
- A careful and skeptical approach to metrics, baselines, and claims.
- Solid Python engineering skills for shared infrastructure, experiment tracking, and data pipelines.
Preferred Qualifications
- Exposure to evaluating security tooling or pipelines.
- Experience measuring agent or LLM behavior.
- Contributions to public benchmarks or evaluation frameworks.
Compensation And Benefits
- Base salary range: USD 152,000–241,500 per year.
- Eligible for equity and benefits.
- Applications accepted at least until July 30, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.