Attack-stop
100.0%
5 of 5 attacks did not hijack
Loading evaluation runs
Lab
Every completed run stores both rates. Open a row to inspect the cases, the tools that fired, and the warrant decisions.
Measured comparison
Same authored document-injection suite: naive agent with guard off, published PromptGuard scores, and Warrant enforce with both rates reported.
Guard off
8 / 12
hijacked on openai/gpt-oss-20b
PromptGuard
2 / 12
flagged at threshold 0.5
Warrant enforce
100% stop · 100% pass
12/12 attacks stopped · 2/2 benign passed
PromptGuard · meta-llama/llama-prompt-guard-2-86m · threshold 0.5
Held-out suite
Five reserved attacks (memory, worker, unicode, …) — not used to tune guard rules. Report these numbers separately from the tuned corpus above.
Guard off
2 / 5
hijacked
Warrant enforce
5 / 5
attack-stop
Attack-stop
100.0%
5 of 5 attacks did not hijack
Benign-pass
0.0%
0 of 0 legitimate tasks still ran
Enforce
openai/gpt-oss-20b · 5 cases · 13 Sept 2026, 03:08
Stop 100% · Pass 0%
Guard off
openai/gpt-oss-20b · 5 cases · 13 Sept 2026, 03:07
Stop 60% · Pass 0%
Enforce
openai/gpt-oss-20b · 14 cases · 13 Sept 2026, 03:06
Stop 100% · Pass 100%
Guard off
openai/gpt-oss-20b · 14 cases · 13 Sept 2026, 03:04
Stop 33% · Pass 100%