The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/02-measure/before-and-after before running the commands below. Browse this recipe on GitHub.qwen3:4b-instruct through Ollama on one
laptop in 22 minutes
(gentlyventures.com).
Run it
tasks.jsonl is one {"id", "prompt", "reference"} per line. --reward
is any wai.verify class (Numeric, ExactMatch, Includes,
MathEqual) or module:function.
What you get
wai.compare printing itself; nothing in this check computes a statistic
of its own. With --same the gain is +0.000 and the verdict is
NO DIFFERENCE, which is what a check that cannot be fooled by its own
noise has to say. Each arm got draw j of a question on the same seed,
so the sampler’s luck is shared and the paired difference is the
prompt’s.
Why three runs: one run per side is one draw of the eval, and
wai.compare reads a gain on one draw as INCONCLUSIVE, because it cannot
tell the prompt from the luck of that run. Three is the fewest re-runs a
spread can be read from. --runs 1 is a third of the time and says so.
The case study’s 22 minutes was one run; budget about three times that
for the default on the same laptop.
Next
Swap in your prompts and your test set, thenwai compare --model ollama:<your-model> .... The rows are on the report
(report.before_rows, report.after_rows) for wai.pass_at,
wai.select or wai.harness.attribute.