Version prompts, models, datasets, and evaluation results together
Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.
Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.
Treat prompt changes like code changes: measure the behavior before deciding whether the edit helped.
Use a small golden dataset to catch prompt regressions, compare changes against a baseline, and validate model updates before users do.