Definition
An evaluation guideline tells a human reviewer or automated evaluator how to apply a criterion to a case. It specifies what material to inspect, what counts as evidence, and how to handle cases the criterion alone does not settle.
The criterion names the property to check. A scoring rubric describes the available ratings and what earns them. The guideline supplies the instructions for making that judgment in practice. It may sit alongside a rubric or be included in one, but it should not quietly introduce a new criterion.
Simple example
A team evaluates whether a support assistant applies the replacement policy correctly. The current policy allows replacements within 14 days. The guideline tells reviewers to use the policy version in effect when the customer asked and compare the purchase date with the request date. For a request governed by the 14-day policy, an answer claiming that a purchase made 20 days earlier qualifies is incorrect. For a different customer, the purchase date is missing. If that prevents a defensible judgment, reviewers use the evaluation’s predefined “insufficient information” outcome rather than assume the customer qualifies.
The same instructions can go into a model judge’s prompt. That does not guarantee the judge will follow them, so the team checks its decisions against reviewed examples.
Why it matters
Two evaluators may use different policy versions or make different assumptions about missing dates. Their scores then mix the assistant’s behavior with their own choices. A written guideline states which policy to use and when to withhold a judgment.
Guidelines also make evaluation runs easier to compare. When the instructions change, record the revision with the results. A score shift may come from the guideline rather than the system under test.
One important nuance
More detailed instructions do not fix a badly chosen criterion or missing evidence. If scoring a case would require guessing at facts, the evaluation may need a predefined “insufficient information” outcome or better source material. Review borderline examples with the people who will use the results, and revise instructions that produce inconsistent judgments.