Use evals before changing prompts
Treat prompt changes like code changes: measure the behavior before deciding whether the edit helped.
Treat prompt changes like code changes: measure the behavior before deciding whether the edit helped.
Use a small golden dataset to catch prompt regressions, compare changes against a baseline, and validate model updates before users do.