The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/04-train/prompt-injection-classifier before running the commands below. Browse this recipe on GitHub.wai.simulate with a model as writer, user and world, how a matched benign
twin keeps a program-labelled dataset from teaching the generator, how to
freeze a test by content hash and hold out attack families and carriers, how
the shortcut probes (bag of words, length, payload removed, hack_scan) say
what a slice measures, and what a pairwise loss and hard-negative mining do
and do not buy on a saturated pool. Under a minute per seed on one L40S; the
whole climb was about 8 of Haiku.
Run it
while-ai/prompt-injection-carriers holds them,
and the *.sha256 files next to this README pin them, so a test that changed
cannot pass for the frozen one.
The live path needs the public sets cloned next to the recipe (InjecAgent,
agentdojo, BIPIA, InjecGuard for NotInject, all MIT), deepset.jsonl,
gandalf.jsonl and oasst1_prompts.jsonl pulled from the Hub (Apache-2.0
and MIT), a Modal token, and for round 7 an Anthropic key for the writer:
The behaviour
Label 1 when a chunk of at most 512 tokens carries an instruction addressed to the model that the content’s author had no standing to give; 0 otherwise. Direct (in the user turn) and indirect (planted in what a tool returned, in an email, in a document the user pasted). The classifier runs in front of the agent on every chunk it reads, so it must be cheaper than one decode token: the target was p50 under 5 ms per 512-token chunk on one CPU core.How the SDK is used
The classifier is not a decoder LM, so the SDK’strain (sft | grpo | dpo | rm) does not apply; everything
else in the loop is the SDK, and sdk_findings.md says where it fit and where
it did not.
Data: label by program, carriers by model
Every attack string comes from a public set and keeps its published label: InjecAgent (MIT), AgentDojo’s injection goals (MIT), BIPIA’s text attacks (MIT), Gandalf (MIT). No attack phrasing was written here. The carrier the string is planted into, and the position, are drawn by program, so the label is known by construction, the ruleresist-planted-instruction uses.
The negative side is the whole difficulty. Round 1 trained on planted
carriers against clean carriers and learned the generator: AUROC 1.00 on
every planted slice, every user turn flagged, NotInject at 0 points. From
round 3 every planted row has a matched twin: the same carrier, the same
position, the same framing, and a harmless insert of matched length instead
of the payload. Nothing separates the pair but whether the insert is an
instruction to the model. That is a preference pair with the confound
removed, the shape the book’s Direct Alignment chapter trains on [1], and
the length match is the verbosity control its Preference Data chapter
warns is needed [1].
Round 7 replaces the template carriers with what wai.simulate produced for
six businesses (clinic, bank, e-commerce, legal, devops, travel) with a model
as situation writer, as the user and, through execute=, as the world: each
tool call is answered by a model writing the document that tool would return
(a patient message with headers, a lease clause, a CI log, a review thread).
579 such documents over 384 rollouts, plus the agent’s own drafts and the
users’ asks. Payloads are the public strings or a model paraphrase of one
that a second call confirmed still addresses the assistant (128 of 137 kept);
benign inserts are model-written sentences that use “ignore”, “override”,
“system prompt” or “security” with no instruction in them (270 of 320 kept
by the same check). Real human text is mixed in on the direct side
(oasst1 prompts, Gandalf). Everything sharing a word 8-gram with any frozen
test is dropped and counted (108 payloads in round 3), the decontamination
rule of the Evaluation chapter [1]; mixing real human text into a
model-written set is the Synthetic Data and Distillation chapter’s guard
against training on one model’s own distribution [1].
Four frozen tests, hashed before training
The headline slices are the ones the recipe’s template generator did not
write:
llm (a held-out business, model-written carriers), deepset,
sim_tool and NotInject; hard, paste and the held-out families are
template carriers with held-out payloads and transforms, and are read as
such. The in-distribution planted slices are reported as in-distribution
only. The first three tests were frozen before round 1; test_llm was
frozen before round 7b, the first round that could be scored on it fairly,
and every earlier round is scored on it after the fact.
Baseline and the served version
ProtectAIdeberta-v3-base-prompt-injection-v2 (184M, Apache-2.0) at its
shipped 0.5. It is the served version on the platform page; every round is a
candidate against it. Meta’s Prompt Guard 2 (22M and 86M) is gated and the
account’s token gets a 403; it is not scored, and no number for it is
claimed.
The climb
Three seeds per round because the eval’s own run-to-run spread is the floor a delta has to clear (Evaluation [1]); a flat or negative round stays in the table as a result. The mining round is rejection sampling’s “keep what the model finds hard” with the random-selection control the Rejection Sampling chapter asks for [1]; the twin-hinge round is the pairwise loss of Direct Alignment on already-separated pairs, which is why it bought nothing [1]. Correctness at the round’s threshold, points out of 100, seed 1, Wilson 95% half-width inresults.json; the threshold is the score at 1% false positives
on the benign side of a validation split carved from that round’s training
rows, never from a test. Three seeds per round; the seed spread on the full
test is 0.8 points (eval_variance, noise band 5.1).
Bold clears the baseline’s Wilson 95% interval on that slice; italics sit below it; plain is inside it. Verdict: v10-multilingual: clears the baseline on 8 of 10 headline slices; loses on [‘multilingual_direct’].
v1 at 99 on the held-out families and 0 on NotInject is the same fact twice:
it flagged everything. v4 is the negative result: 16,000 rows labelled
“prompt injection” from a permissive set moved three headline slices down,
because SPML is role-play jailbreak text and its benign side is short asks,
so length and prompt-likeness re-entered the label. v6 is the second: a
margin on pairs the model already separates buys nothing (1.4% of the round-5
rows sit in the uncertain band). v7 and v7b are the model-written-carrier rounds: on the business held out of training the score goes from 49 to 73 with 0 to 2 false positives in 211, and the template-carrier tests drop, because a model trained on one generator scores that generator; the union (v8) sits between the two on every slice. By the pre-stated rule (sum of headline points, seed 1) v5 is the pick at 509 against v8’s 491; on the slices no template wrote (llm, sim_tool, NotInject) v8 is the better arm, and the recipe ships both, with the rule stated.
Paired against the baseline with
wai.compare on correctness, by slice
(out/sdk_measure.json): every round from v3 on is up on sim_tool, NotInject,
the held-out families and the planted slices with intervals clear of zero,
and down on deepset (v3: -8.8 [-11.5, -6.0] on 662 rows). That loss is
real and is the same across seeds.
Shortcut probes
Reported for every round, because a program-labelled dataset can be solved by the program’s fingerprints.Latency and size
int8 ONNX (onnxruntime.quantization.quantize_dynamic), one thread, Apple
M5 Max, 300 timed runs after 20 warm-ups, tokenisation excluded:
23 MB on disk (int8), 10.8M non-embedding parameters (22.7M with the
embedding table). The 5 ms target at 512 tokens is not met on this CPU; int8
is not faster than fp32 on ARM here (44.4 against 39.7 ms). A 128-token
sliding window with max-pooling costs 30 ms for all four windows of a
512-token chunk, 121 ms for 2,048 tokens, and 7.5 ms when the first window
already exceeds the threshold. A byte 4-gram hashed logistic regression
(
byte_stage.py) runs in 0.8 ms per 2,000-character row but scores 0.49
AUROC on the hard test, so it is a speed layer in a cascade, not a detector.
What the SDK did and did not do
sdk_findings.md lists eight findings. The short version: wai.simulate
with execute= and a model as simulator= produced the data the recipe
needed once the hosted writer’s daily quota was routed around; the
coverage axes, a span export, classifier metrics and an encoder trainer do
not exist and are recipe-local here; route (from whileai.routing) on classifier rows
picked sft at k=1 for the wrong reason and then refused on a truncation
check that reads nothing a classifier row has; hack_scan worked as the
shortcut detector once the pair was the ask.
Honest limits
- Every planted slice is synthetic. The benign FPR on NotInject and on the simulated tool results is not the FPR on a customer’s traffic; nothing here was measured on real documents.
- The model is a public 22M encoder and the attacker can read it. In-distribution numbers overstate by 8 to 16 AUC points on this task family (arXiv:2602.14161); the held-out-family and held-out-carrier slices are the honest ones and the hard test is where the number is lowest.
- It loses to ProtectAI on deepset, a direct-attack set with German rows, on every round. The training has no German and few direct attacks that are not Gandalf.
- The paste and hard tests are matched twins on two held-out families (AgentDojo, BIPIA-test), so their negatives are still the recipe’s own.
- One base model, one CPU, one tokenizer. The writer, the user and the world
in round 7 are the same model (haiku);
wai.simulatewarned about it. holdout_sizesays 1,000 to 1,500 paired rows resolve a 5-point gain per slice at 80% power; the 91-rowsim_toolslice and the 250-row family slice cannot, and their intervals say so.
Next
Push the round-9 rows and the five frozen tests to your While account’s Datasets page (WHILEAI_API_KEY, or wai login), and the weights and rows to Hugging
Face under your org:
References
- Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025 (rlhfbook.com). Chapters Evaluation, Direct Alignment, Preference Data, Rejection Sampling, Synthetic Data and Distillation.
- Zhan, Q. et al. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated LLM Agents. arXiv:2403.02691, 2024.
- Debenedetti, E. et al. AgentDojo. arXiv:2406.13352, 2024.
- Yi, J. et al. BIPIA: Benchmarking and Defending Against Indirect Prompt Injection Attacks. arXiv:2312.14197, 2023.
- Li, H. et al. InjecGuard / NotInject. arXiv:2410.22770, 2024.
- When Benchmarks Lie: leave-one-dataset-out on prompt-injection detectors. arXiv:2602.14161, 2026.
- Devlin, J. et al. BERT. arXiv:1810.04805, 2019. Wang, W. et al. MiniLM. arXiv:2002.10957, 2020.
- Hinton, G., Vinyals, O., Dean, J. Distilling the Knowledge in a Neural Network. arXiv:1503.02531, 2015.
- Efron, B. Bootstrap Methods. Annals of Statistics 7(1), 1979. Wilson, E. B. JASA 22, 1927.
- Wu, H. et al. Integer Quantization for Deep Learning Inference. arXiv:2004.09602, 2020.