Version prompts, models, datasets, and evaluation results together

Keep every evaluation result tied to the exact prompt, model identities, dataset, evaluator, and configuration that produced it.

August 27, 2026 · 4 min · Lukas Walter

Use evals before changing prompts

Treat prompt changes like code changes: measure the behavior before deciding whether the edit helped.

July 9, 2026 · 2 min · Lukas Walter

Stop Guessing – Use Golden Datasets for Prompt Evals

Use a small golden dataset to catch prompt regressions, compare changes against a baseline, and validate model updates before users do.

March 25, 2026 · 2 min · Lukas Walter