Version prompts, models, datasets, and evaluation results together

Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.

August 27, 2026 · 4 min · Lukas Walter