Definition

A scoring rubric defines the criteria and score levels used to evaluate an output. It explains what each criterion means, what counts at each level, and may specify which evidence the evaluator should consider.

A rubric can guide human reviewers or model judges. It structures judgment but does not perform the measurement by itself.

Simple example

A grounded-answer rubric scores faithfulness from zero to two. Zero means the core answer is unsupported by the supplied evidence or contains a material claim contradicted by it. One means the core answer is supported but includes a material claim without evidence. Two means every material claim is supported by an identified passage.

Completeness and writing clarity get their own scores, so a faithfulness failure stays visible. If those scores are combined, a minimum faithfulness threshold can stop strong writing from offsetting an unsupported claim.

Why it matters

Words like “good”, “helpful”, and “accurate” leave room for different interpretations. A rubric makes expectations explicit. Clear criteria can help evaluators agree and make low scores easier to trace to specific problems.

In a repeatable evaluation system, treat the rubric as versioned configuration. If its wording or thresholds change, scores may no longer be directly comparable with earlier runs.

One important nuance

Detailed score levels cannot make an irrelevant criterion useful, and ambiguous criteria still need clearer definitions. Evaluators may still disagree about borderline cases, and a numeric scale can suggest more precision than the evidence supports. Check the rubric against reviewed cases, measure agreement, and inspect score distributions. Revise definitions that evaluators repeatedly interpret differently. Keep hard safety or contract checks separate when a high average must not hide their failures.