Version prompts, models, datasets, and evaluation results together
Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.
Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.