AI Concepts category
Evaluation
AI-as-a-Judge
AI-as-a-Judge uses a model to evaluate outputs against defined criteria, references, or competing responses.
Comparative Evaluation
Comparative evaluation judges outputs or systems directly against one another rather than assessing each independently.
Evaluation Harness
An evaluation harness runs cases against an AI system in a repeatable way and records outputs, measurements, and run configuration.
Factual Consistency
Factual consistency assesses whether an output's claims are compatible with the reference information being evaluated.
Faithfulness
Faithfulness measures whether an answer's claims are supported by the evidence supplied to the model.
Reference-Based Evaluation
Reference-based evaluation compares a system output with known answers, evidence, or other designated reference data.
Scoring Rubric
A scoring rubric defines criteria and score levels that guide evaluation of an output.