Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/02-measure/before-and-after before running the commands below. Browse this recipe on GitHub.
Run the old prompt and the new one on the same questions, on a model on your own machine, and get one verdict: the gain, its 95% range over tasks, and PASS only when the range clears zero. Free: no key, no credits. What you will learn: how to turn “the new prompt feels better” into a paired number with an interval, why both prompts draw on the same seeds, and what an honest null looks like. You need nothing for the offline run; Ollama for a real model. Seconds offline; an external case study ran this shape on qwen3:4b-instruct through Ollama on one laptop in 22 minutes (gentlyventures.com).

Run it

The command line takes your own files:
tasks.jsonl is one {"id", "prompt", "reference"} per line. --reward is any wai.verify class (Numeric, ExactMatch, Includes, MathEqual) or module:function.

What you get

The number that matters is the gain, +0.306, with its 95% range +0.253..+0.358: the range clears zero and the gain clears the eval’s own re-run noise (0.035, measured from the three runs a side), so PASS. The compare block is wai.compare printing itself; nothing in this check computes a statistic of its own. With --same the gain is +0.000 and the verdict is NO DIFFERENCE, which is what a check that cannot be fooled by its own noise has to say. Each arm got draw j of a question on the same seed, so the sampler’s luck is shared and the paired difference is the prompt’s. Why three runs: one run per side is one draw of the eval, and wai.compare reads a gain on one draw as INCONCLUSIVE, because it cannot tell the prompt from the luck of that run. Three is the fewest re-runs a spread can be read from. --runs 1 is a third of the time and says so. The case study’s 22 minutes was one run; budget about three times that for the default on the same laptop.

Next

Swap in your prompts and your test set, then wai compare --model ollama:<your-model> .... The rows are on the report (report.before_rows, report.after_rows) for wai.pass_at, wai.select or wai.harness.attribute.
Last modified on October 1, 2026