Definition
An evaluation criterion is one property of an AI system’s output or behavior that an evaluator should assess. It answers “what are we checking?” Examples include whether an answer follows the user’s instruction or whether its claims are supported by the supplied documents.
A scoring scale defines the available ratings. A scoring rubric explains what earns each rating for that criterion. The same criterion might be checked as pass or fail in one evaluation and scored on several levels in another.
Simple example
A customer asks whether a damaged item can be replaced after 20 days. The current policy allows replacements within 14 days, but the support assistant says the customer has 30 days. One criterion is: “Does the answer state the replacement window correctly under the current policy?”
The policy provides evidence for the check. A reviewer could mark the answer pass or fail, or use a rubric that distinguishes a wrong window from an answer that omits the window. Those are scoring choices. The property being assessed stays the same.
Why it matters
A single “answer quality” score can hide a policy error behind clear, polite writing. A separate criterion for policy accuracy makes the error visible. It directs evaluators to check the current policy. If teams apply the criterion consistently to comparable test cases, they can spot a drop in policy accuracy after a prompt or retrieval change.
One important nuance
“Accuracy” is often too broad to guide a review. Specify which claim or behavior matters and what source defines correctness. An answer can faithfully repeat a supplied policy that is out of date. “Faithful to the supplied document” and “correct under the current policy” are different criteria, even when they agree on most test cases.