Skip to main content
What you get: a measured false-alarm rate, interval coverage and detection power for the comparison wai.compare runs, so you can see the statistics hold instead of taking our word for it. Who it is for: anyone about to trust a “moved” or “no change” from this library. Needs: nothing. No key, no network, no GPU.

What it does

Every verdict this library prints rests on one comparison: per-case differences between two runs, a bootstrap interval around their mean, and “moved” only when that interval leaves out zero. The question is whether that rule does what it claims: say “moved” by accident 5% of the time, cover the true change 95% of the time, and catch a real gain as often as the sizing in holdout_size promised. You can answer that without trusting us. Script two agents whose chance of passing every case is set by hand, run the comparison hundreds of times, and count. That is what wai self-check does. It runs the same function wai.compare calls (stats.compare_runs), and the first trial of every check also goes through wai.compare itself and has to give the identical answer, so the check is on the code your verdict comes from, not on a copy. The idea is not ours. An external case study graded the grader by hand on whileai 0.126 with scripted stand-in agents (Gently Ventures); this command packages their checks at their sizes.

The checks

Each agent tries each case once per arm. Each trial draws fresh cases and fresh outcomes from a seed that derives from seed, the check and the trial number, so a report reproduces on any machine.

Run it

The 40-trial run is a smoke test: its Monte Carlo error is 3 to 8 points, enough to catch a broken interval, not a small miss. Read the full run for the numbers.

Expected numbers

wai self-check --full, seed 0, 400 trials per check, whileai 0.126 plus this command, beside what the case study measured by hand on 0.126 with 300 to 400 trials. All eight are within three Monte Carlo errors. Two readings worth keeping:
  • At 10 cases the 95% range is about a 92% range. 91.8% sits 2.9 errors under target, just inside the flag, and matches the case study. A percentile bootstrap undercovers on small samples (Efron and Tibshirani 1993, chapter 13). From 20 cases it is on target. Below 20 cases, read a “moved” as weaker than its 95% says.
  • Power is low at these sizes, and the formula says so. A real +10-point gain on 30 cases is caught about one time in ten. Run holdout_size before the eval, not after the verdict.
Near-duplicates do not inflate false alarms here because each restated case is a fresh try. If your eval copies a case’s outcome rather than re-running it (a cached reply graded twice), the copies are not independent, the interval is too narrow, and false alarms rise. wai.decontaminate finds near-duplicate cases before that happens.

Reading a flag

The Monte Carlo error is sqrt(t * (1 - t) / trials) at the target t: how far the measured rate moves from trial count alone. A check is flagged when it sits more than three of those on its bad side: above target for a false alarm, below for coverage, either way for power, where the target is a prediction. Three errors keeps a correct check from flagging by its own noise about one time in a thousand. A flag at 40 trials says re-run with --full; a flag at 400 is a finding about the statistics, and we want the issue. wai self-check exits 1 when any check is flagged, so it can sit in CI.

Defaults

Every number has a name in whileai/simulations/defaults.py, with the reason beside it, and the call takes the ones a user moves. wai.self_check(trials=, seed=, n_boot=, alpha=, flag_at=) takes the trial count, the seed and the flag rule from this table. n_boot (BOOTSTRAP_DRAWS, 2000) and alpha (ALPHA, 0.05) go to the comparison exactly as a wai.compare call passes them, so you can check a setting you plan to use.
Last modified on October 1, 2026