wai.compare runs, so you can see
the statistics hold instead of taking our word for it.
Who it is for: anyone about to trust a “moved” or “no change” from
this library.
Needs: nothing. No key, no network, no GPU.
What it does
Every verdict this library prints rests on one comparison: per-case differences between two runs, a bootstrap interval around their mean, and “moved” only when that interval leaves out zero. The question is whether that rule does what it claims: say “moved” by accident 5% of the time, cover the true change 95% of the time, and catch a real gain as often as the sizing inholdout_size promised.
You can answer that without trusting us. Script two agents whose chance of
passing every case is set by hand, run the comparison hundreds of times,
and count. That is what wai self-check does. It runs the same function
wai.compare calls (stats.compare_runs), and the first trial of every
check also goes through wai.compare itself and has to give the identical
answer, so the check is on the code your verdict comes from, not on a copy.
The idea is not ours. An external case study graded the grader by hand on
whileai 0.126 with scripted stand-in agents
(Gently Ventures); this
command packages their checks at their sizes.
The checks
Each agent tries each case once per arm. Each trial draws fresh cases and
fresh outcomes from a seed that derives from
seed, the check and the trial
number, so a report reproduces on any machine.
Run it
Expected numbers
wai self-check --full, seed 0, 400 trials per check, whileai 0.126 plus
this command, beside what the case study measured by hand on 0.126 with
300 to 400 trials.
All eight are within three Monte Carlo errors. Two readings worth keeping:
- At 10 cases the 95% range is about a 92% range. 91.8% sits 2.9 errors under target, just inside the flag, and matches the case study. A percentile bootstrap undercovers on small samples (Efron and Tibshirani 1993, chapter 13). From 20 cases it is on target. Below 20 cases, read a “moved” as weaker than its 95% says.
- Power is low at these sizes, and the formula says so. A real
+10-point gain on 30 cases is caught about one time in ten. Run
holdout_sizebefore the eval, not after the verdict.
wai.decontaminate finds
near-duplicate cases before that happens.
Reading a flag
The Monte Carlo error issqrt(t * (1 - t) / trials) at the target t:
how far the measured rate moves from trial count alone. A check is flagged
when it sits more than three of those on its bad side: above target for a
false alarm, below for coverage, either way for power, where the target is
a prediction. Three errors keeps a correct check from flagging by its own
noise about one time in a thousand. A flag at 40 trials says re-run with
--full; a flag at 400 is a finding about the statistics, and we want the
issue.
wai self-check exits 1 when any check is flagged, so it can sit in CI.
Defaults
Every number has a name inwhileai/simulations/defaults.py, with the
reason beside it, and the call takes the ones a user moves.
wai.self_check(trials=, seed=, n_boot=, alpha=, flag_at=) takes the
trial count, the seed and the flag rule from this table. n_boot
(BOOTSTRAP_DRAWS, 2000) and alpha (ALPHA, 0.05) go to the comparison
exactly as a wai.compare call passes them, so you can check a setting you
plan to use.