Forgeable Confirmation in Automated Computer Security Testing: Deterministic Rules versus AI Judges Akihisa Fujiyama1,2 , Niwase Shamim1,* 1
arXiv:2609.24200v1 [cs.CR] 21 Sep 2026
2
OmiCore Inc., Fukuoka, Japan
Graduate School of Engineering, Kyushu University, Fukuoka, Japan *
Corresponding author: Niwase Shamim ([email protected])
Abstract AI is increasingly used to automate computer security testing, and the tools must decide for themselves whether an attack succeeded. A finding that a deterministic rule confirms by observation is reported as fact, whereas one that an LLM judges exploitable is treated as an opinion. We ask whether the system under test can forge that confirmation. In offline security testing of a four-stage AI-assisted pipeline, nine of its fifteen confirmation mechanisms are forgeable, and forgeability is predicted entirely by whether the decision reads attacker-controlled data. We formalise this as an auditable attack surface and test it prospectively: on sixteen held-out mechanisms, predictions fixed before any attack separated forgeable from unforgeable mechanisms exactly (Fisher p = 5.5 × 10−4 ; 13 of 16 under the originally specified adversary), and across 12,203 mechanisms in public scanner templates the prediction was 99.9 % accurate. Deterministic rules proved cheaper to forge than eight open-weight LLM judges, failing at 2 % of attacker-controlled response content against a median of 50 %. No implementation of one check was both robust and precise, and routing between a rule and an AI judge raised forgery to 99 %. Moving the decisive evidence to a channel the attacker cannot write cuts attack success from 97 % to 0 %, and an escalate verdict recovers the sensitivity this costs. The protection fails when the scanned host is itself the adversary. The results bear on AI security agents and on benchmarks that score success by string matching.
Keywords: computer security; AI-assisted security testing; offline security testing; LLM-as-a-judge; forgeable confirmation; auditable attack surface; adversarial evaluation.
1
Introduction
Security-operations analysts triage large volumes of machine-generated findings; practitioners report that most are false and that manual validation dominates their workload1 . The problem predates security operations centres. Developers abandon static-analysis tools with high false-positive rates2,3 , tool builders treat the false-positive budget as the primary design constraint4 , security-focused static analysis shows the same pattern5 , and black-box web scanners have long missed entire vulnerability classes while emitting findings that a human must adjudicate6,7 . Verifier-gated pipelines address this by attaching a different kind of evidence to some findings. A finding may carry a model's judgement, or it may carry an observation: for example, a forged token returned the same protected resource, and the two response bodies are hash-identical. Pipelines label such findings accordingly (confirmed by observation rather than 84 % likely). When queues are ordered, bug-bounty reports filed or deployment gates opened on that label, an attacker who can manufacture it also manufactures the priority. The question is becoming more pressing. Language-model agents exploit real one-day vulnerabilities8 , automate penetration-testing workflows9 , and are evaluated on capture-the-flag and real-world exploitation benchmarks10,11 . Each of these systems decides whether an exploit succeeded by reading bytes that the target produced. 1
We study a four-stage pipeline (generate → judge → confirm → measure) in which the confirmation stage is the verdict authority, and ask whether its confirmations can be forged. Our contributions are: (i) the auditable attack surface, a criterion that predicts from source code which confirmation mechanisms are forgeable, derived on one pipeline, tested prospectively on held-out modules and applied to a third-party corpus; (ii) measurements of attacker cost for deterministic rules and LLM judges, both as a robustness curve over attacker budgets and as an exact minimum payload length; (iii) evidence from six independent implementations, and from rule–judge composition, that the weakness lies in the task definition rather than in any one implementation; and (iv) a defence, its measured cost, and the adversary model under which it holds.
2
Background
Attack surface of a decision. Attack-surface measurement for whole systems is well established: Manadhata and Wing formalise attackability along method, data and channel dimensions12 , and a systematic review documents the many definitions in use13 . We apply the idea to a single decision procedure, so that two verifiers can be compared by what an attacker must supply rather than by their stated confidence. To our knowledge, the attack surface of the check that decides whether an attack succeeded has not been measured. The difficulty of trusting a tool to certify its own output is long recognised14 ; in software supply chains the current response is to attach verifiable provenance to each step15 . LLM-as-a-judge. Using one model to score another's output is standard practice with documented limitations: position, verbosity and self-enhancement biases16 , sensitivity to answer order17 , preference for a model's own generations18 , and a broader design space19 . Two common mitigations are used in our pipeline: repeated sampling of one judge20 and a panel of smaller judges21 . Judges can be flipped by a single token, with false-positive rates up to 80 %22 , and such flips can be found by optimisation23 . We transfer this attack to a security setting: the transfer to our judges is weak (at most 20 % of cases flipped, and three of five judges never flip), whereas the deterministic marker rule of the same pipeline flips on 100 % (Fig. S1d). Prompt injection. Indirect prompt injection compromises LLM-integrated applications through retrieved content24,25 ; attacks and defences have been formalised and benchmarked26 , including for tool-using agents27,28 . These attacks exploit a system that does not separate data from instructions. A rule that searches a response body for marker strings has the analogous flaw without any model: it does not separate data from verdict. Consistent with this, the prompt-injection payload family wins no case against any deterministic arm in our search, while a three-character literal wins most. Adversarial machine learning. Evasion of learned security classifiers is well studied29 . Two methodological lessons apply here: defences must be evaluated against adaptive attackers30,31 , and attacks must be realisable in the problem space32 . Evaluation bias is a known hazard in security machine learning33,34 . Our budget sweep applies the same discipline to rules: every payload is a byte sequence that a server could return, and a deletion-only control separates forgery from loss of evidence. Ground truth. Fuzzing evaluations are sensitive to methodological choices35 , which motivated ground-truth benchmarks with injected bugs36 . Vulnerability datasets carry 20–71 % inaccurate labels37 , performance on curated data does not transfer to realistic data38,39 , and the effect persists for code language models40 ; guidelines for sound experimental design are long established41 . The problem applies to this study directly, because our positive class is labelled by the mechanisms under examination; we discuss this in the Limitations. Exploitation, reliance and measurement. Automatic exploit generation includes a verification step whose signal is not attacker-authored: a segmentation fault is observed by the operating system, not asserted by the target42,43 . LLM-driven offensive tools instead verify success from application responses8,9 , on benchmarks whose success criteria are string matches10,11 . How operators use automated verdicts is well studied: automation is misused and disused in predictable ways44 , explanations increase acceptance regardless of correctness45 , cognitive forcing reduces overreliance46 , reliance is measured with appropriateness constructs47,48 , and hedged advice is followed less than unhedged advice of equal accuracy49 . The label confirmed by observation is therefore consequential. We measure the machine rather than the operator (Fig. S3). Agreement is reported chance-corrected50–52 . The relevant failure mode is specification gaming, in 2
which a proxy diverges from its goal under optimisation53–56 . Vulnerability prioritisation systems already condition on exploitation likelihood57,58 , so a forgeable confirmed flag propagates into them.
3
Results
3.1
Forgeability tracks what a decision reads, not whether it is deterministic
Figure 1. The reachability gate at three scales. (a) The fifteen deterministic confirmation mechanisms of the pipeline, ordered by outcome and then by | S |. Bar length is the number of literal constants the decision consults; colour is the audit result (red, forged; green, unreachable, G = 0; blue, reachable but not forged, which is empty here). Annotations give forged over attempted attacks. Bar length and outcome are nearly unrelated: ede·cross consults fourteen constants and cannot be attacked, whereas timing consults one and is forged. (b) The gate as a prediction. Rows are the gate's call and columns the attack outcome. Left: the 15 mechanisms on which the criterion was defined (Fisher exact p = 2 × 10−4 ). Centre: 16 held-out mechanisms from six modules, including an Active Directory scanner, predicted from source and hash-pinned before any attack (p = 5.5 × 10−4 ). Right: 10,925 third-party mechanisms, with forgeries constructed and evaluated automatically (accuracy 0.999). The bottom-right cell holds by definition; the informative cells are the top row, where 15 of 10,340 predicted-attackable mechanisms resisted, and the bottom-left cell (predicted 3
unreachable but forged), which is zero at all three scales. (c) Decision channels of the 12,203 third-party mechanisms; a mechanism reading two channels is counted twice. Orange marks the two out-of-band channels, which the scanned host nonetheless receives. (d) Top: 95.8 % of mechanisms read only channels authored by the scanned host. Bottom: of the 508 that consult an out-of-band channel, 350 combine it with a host-written channel, the structure of the hardened policy in eq. (7), and 158 rely on the out-of-band anchor alone. We audited all fifteen deterministic confirmation mechanisms of the pipeline (Fig. 1a, Table 1). Nine are forgeable and six are unreachable, and the number of constants a mechanism consults does not determine the outcome: a mechanism consulting fourteen constants cannot be attacked, whereas one consulting a single threshold is forged (Fig. 1a). Every forgeable mechanism decides by reading text or structure the attacker wrote; every unreachable one relies on a tester-generated secret, an out-of-band callback or a server-side read-back. Table 1. Deterministic confirmation mechanisms of the pipeline, audited for forgeability. R = 1 when attacker-controlled input reaches the decision; K = 1 when the decisive value is attacker-knowable; AAS = |S| when R·K = 1 and 0 otherwise. Of the 15 mechanisms, 9 are forgeable and 6 are unreachable.
Note. Overall attack success 61%. Every mechanism passed its clean-accuracy check before it was attacked. Held-out prediction. A criterion that separates the cases it was defined on may simply fit them, so we tested it prospectively (Table 2): for sixteen mechanisms from six modules not used in its derivation, predictions were read from source and hash-pinned before any attack was written. All twelve mechanisms predicted forgeable were forged and all four predicted unreachable resisted (p = 5.5 × 10−4 ; Fig. 1b, centre), compared with p = 2 × 10−4 on the fifteen mechanisms from which the criterion was derived (Fig. 1b, left). The adversary model was widened after one misprediction, to allow an injected line break, and the wider model was then applied to all seven anchors of that module. Under the original, narrower model the accuracy is 81.2 % (p = 0.019); we report both. 4
Table 2. Held-out prediction test of the reachability gate. 16 confirmation mechanisms from 6 modules not used to derive the criterion, including an Active Directory scanner. Predictions were read from source and hash-pinned (PREDICTION_SHA = a41c7a1fec100b9f) before any attack was written.
Note. Wide adversary: 12 true positives, 0 false negatives, 0 false positives, 4 true negatives — accuracy 100%, Fisher exact p = 0.000549. Narrower adversary (no injected line break): accuracy 81.2%, p = 0.019231. The wider adversary model was adopted after one misprediction and then applied to all seven AD anchors; both scorings are therefore reported. Deletion alone flips 0 of 16. Third-party corpus. The six re-implementations reported below address the concern that the result is specific to one codebase, but not that it is specific to one set of authors. We therefore applied the gate to 11,137 public scanner templates59 containing 12,203 confirmation mechanisms written by 1,611 contributors, scoring each automatically from its matcher specification (Fig. 1b right, 1c, 1d; Table S7). Before any attack, 121 mechanisms were excluded because they already fired on a benign placeholder (the clean-accuracy control of eq. (3)); 1,102 offered no writable matcher and 55 could not be evaluated from their specification. Of the remaining 10,925, 10,325 were forged, 15 resisted and 585 were out of the attacker's reach; no mechanism predicted unreachable was forged (accuracy 0.999; Fig. 1b, right). Response body and status code dominate the channels these mechanisms read (Fig. 1c), and 95.8 % of mechanisms read only channels authored by the scanned host; of the 508 that consult an out-of-band channel, 350 combine it with a host-written channel, the structure of the hardened policy of eq. (7), and 158 rely on the out-of-band anchor alone (Fig. 1d). Constant count does not predict cost. | S | is not associated with the minimum attacker budget b* across the six implementations (ρ = −0.563, p = 0.33) or the nine held-out mechanisms for which both are measured (ρ = −0.183, exact p = 0.638), whereas over 400 third-party mechanisms the association is positive (ρ = +0.483, p = 8.2 × 10−25 ). Because the direction of the association changes between populations, | S | is not a cost model, and we report G and | S | separately.
5
3.2
Rules and judges both fail abruptly, 25× apart in attacker budget
Figure 2. Attacker budget required to forge a confirmation. (a) Verifier Robustness Curve: FA (α) of eq. (3) over 147 clean-negative findings and 13 attacker budgets. Red, reachable deterministic rule; blue, six-judge ensemble; green, hardened rule (dashed, zero throughout); the shaded band spans the eight model arms. Dotted verticals mark the breakpoints α* of eq. (4): 2 % for the rule and a median of 50 % for the judges, roughly 6 and 157 characters at the median body length. (b) Minimum budget b* of eq. (6) for six independent implementations of one check, found by generate-then-minimise search over 161 attackable cases. Bars give the median in characters; annotations give the share of cases flipped. Substring and word-boundary regex implementations fall at three characters (otp, ssn, cvv), a JSON-key check at 11, key-plus-value at 15 and both shape detectors at 22. (c) The same search against four judges over twelve cases each; dots mark the median b* and bars the range. Judges require a median of 20 to 79 characters, seven to twenty-six times the cheapest rule, and are flipped on 8–9 of 12 cases. (d) b* for 400 sampled third-party mechanisms, grouped by what the matcher reads; the dashed line is the corpus median (29 bytes) and the minimum is 3. A single forgery rate cannot distinguish a verifier that degrades gradually from one that collapses, and the two call for different defences; we therefore sweep the attacker's budget (Fig. 2a). Both kinds of decider behave as threshold devices: on the common pool, three of the eight model arms show the same abrupt (cliff) failure as the rule. What separates them is how much of the response the attacker must control before the threshold is crossed, a quantity a defender can budget against: the deterministic rule breaks at α* = 2% of the response body and the judges at a median of 50 %, roughly 6 and 157 characters at the median body length (Fig. 2a).
6
Because the curve is a bound for a fixed payload ladder, we also measured the exact minimum budget b* by adaptive search. Across six independent implementations of one check, b* rises from three characters for substring and word-boundary matching to 11 for a JSON-key check, 15 for key-plus-value and 22 for both shape detectors (Fig. 2b). Four judges require a median of 20 to 79 characters, seven to twenty-six times the cheapest rule, yet are still flipped on 8–9 of 12 cases (Fig. 2c). In 400 third-party mechanisms the median is 29 bytes and the minimum 3 (Fig. 2d), so forgery outside this pipeline is similarly cheap. The label asserting the strongest warrant is thus attached to the decision that is cheapest to forge. Per-arm shape, area, breakpoint and clean false-confirmation rate are given in Table S1, and the minimum budgets for every implementation, judge and third-party matcher class in Table S2.
3.3
The result holds across model families; ensemble statistics depend on the roster
Figure 3. Generality across model families, and variance due to ensemble composition. (a) AU F C of eq. (4) for the eight model arms on the 147-case pool, coloured by the shape verdict of eq. (5) (blue, flat; orange, gradual; red, cliff). The vertical line marks the reachable deterministic rule (0.980); the nearest model arm is 0.607 below it, and the hardened rule scores 0.000. (b) AUFC against parameter count, labelled by family, with the 6.7–9B band shaded (ρ = −0.193, exact p = 0.654). Within that band, five arms from five families split two flat against three degrading. (c) Fleiss κ over the 39 observed findings (blue) and sensitivity on the 31 findings judged by every arm (orange) against the number of judges, computed on the two six-judge rosters described in Methods. Dots are medians over rosters of each size and bars the range. At
7
three judges, κ ranges from +0.545 to +0.883 with only membership varying; median sensitivity does not increase with size (best single judge 97 %, six-judge fleet 55 %). Values are listed in Tables S6 and S8. (d) Two sources of verdict variation on the same 34 cases. Left: repeated queries to the same judge (ten times) or the same fleet (twice) change 1.8 % and 0.3 % of verdicts. Right: a differently composed fleet changes 4.5 % of (roster, case) verdicts at three judges and 2.5 % at five; at seven judges only one roster exists. To test generality, we evaluated eight model arms from six families on a common pool. Three arms fail as cliffs, but the most vulnerable still has an area under the forgeability curve 0.607 below that of the deterministic rule (0.980), and the hardened rule scores 0.000 (Fig. 3a). Robustness does not track parameter count (ρ = −0.193, exact p = 0.654; Fig. 3b); within the 6.7–9B band, a small model from a robust family stays flat while a small model from a less robust family degrades. Scales above 16B and sampled decoding were not evaluated on this pool. Judges are highly self-consistent: across 2,380 inferences (seven judges, ten repetitions of identical evidence), per-judge self-consistency is 94–100 % and fleet self-consistency 99.7 %, so disagreement between judges cannot be attributed to sampling noise. Fleet composition, however, changes verdicts. With three judges, Fleiss κ ranges from +0.545 to +0.883 depending only on which judges are seated, and median sensitivity does not increase with fleet size: the best single judge confirms 97 % of observed positives and the six-judge fleet 55 % (Fig. 3c). Replacing members changes the verdict on 6 of 34 cases, whereas repeating the same judge or fleet changes 1.8 % and 0.3 % of verdicts (Fig. 3d), a source of variance that repeatability checks cannot detect because the evidence is unchanged. Ensemble statistics for every fleet size and roster are listed in Table S6. Within the six-judge agreement fleet, agreement is substantial (Fleiss κ = +0.683, Krippendorff α = +0.685) and bimodal, with 74 % of findings at the extremes of the vote (Fig. S1a). It is not organised by training lineage: pairs sharing a lineage agree only +0.040 more than other pairs (exact permutation p = 0.30; Fig. S1b), whereas removing a single judge changes fleet κ by up to +0.066 (Fig. S1c). Per-judge centrality, leave-one-out effect, self-consistency and cost are given in Table S8.
8
3.4
No implementation of one check is both attack-resistant and precise
Figure 4. Six implementations of one check, and their composition with a judge ensemble. (a) Attack success for six independent implementations of confirm_unauth (rows) under six conditions (columns): the share of 161 attackable cases flipped. The clean control column is non-zero because the two value-shape implementations fire on 79 % and 66 % of unmodified evidence. Text-matching implementations fall to a single English sentence, structure-parsing ones to a single injected JSON field, and all six to a high-entropy JSON value. (b) The same implementations on three rates, with b* in the row label: clean false-confirmation rate (red), share flipped by one English sentence (orange) and by one JSON field (blue). The two implementations that resist prose are the two with high clean false-confirmation rates. (c) The status-inversion branch over 200 instances: six implementations (rows) under activation conditions ranging from clean evidence to three attacker-authored fields (columns). The shipped rule returns inconclusive on all unmodified instances through both branches; one authored status code activates its three-way conjunction on 58 % of the pool, and three fields on 100 %. The anchored and read-back implementations never activate. (d) The rule-then-ensemble hybrid against its components, on the 31 observed positives with verdicts from every arm. Red, attack success; blue, sensitivity. The hybrid confirms 100 % and is forged on 99 %, more than the deterministic rule (97 %) or the ensemble (6 %), because it confirms when either component confirms. The SQL-injection confirmer returns inconclusive on every real finding in this corpus, which consists of authorisation evidence, so forging it would only demonstrate activation of a dormant rule; this is why the marker rule flips on 100 % of cases in our initial master-key test while the judges flip on at most 20 % (Fig. S1d). We therefore attacked a mechanism that fires in practice, confirm_unauth, and implemented it
9
six independent ways. It fires on 39 of 200 instances and declines on the other 161, which form the attack set NA . The full grid is given in Table S3 and the status-inversion branch in Table S4. Appending the sentence "no password is returned to the client", which tells a human reader that the endpoint is safe, flips both text-matching implementations on 88 % of cases (Fig. 4a). Because a benign API could return such a sentence, the same defect also produces false positives in normal operation. Parsing resists prose but falls to one injected field (80 %), and every implementation falls to a high-entropy JSON value (Fig. 4a). Shape detection resists prose, but fires on two-thirds of unmodified traffic, because response bodies contain UUIDs and digests with the length, character mix and entropy of a credential; explicitly excluding those formats reduces the clean firing rate only from 79 % to 66 %. The two implementations that resist prose are therefore exactly the two with high clean false-confirmation rates, and none is low on all three rates (Fig. 4b). The weakness therefore lies in the task definition: detecting unauthenticated exposure of sensitive data from a response body requires a decision made from attacker-authored data, whatever the implementation. The same holds for the status-inversion branch of the SQL-injection confirmer. The shipped rule returns inconclusive on all 200 unmodified instances, but one attacker-authored status code satisfies its three-way conjunction on 58 % of the pool and three authored fields on 100 %, whereas the anchored and read-back implementations cannot be activated at all (Fig. 4c; Table S4). A conjunction costs the attacker only the conjuncts that are not already true. Composition does not help. The rule-then-ensemble hybrid confirms 100 % of observed positives and is forged on 99 % of attacked evidence, more than either the deterministic rule (97 %) or the ensemble (6 %), because it confirms whenever either component does (Fig. 4d). Hardening the deterministic component does not repair the hybrid, because the hardened rule confirms none of these findings and the hybrid then reduces to the ensemble. Only conjunctive routing improves robustness, and on this corpus it confirms nothing. End-to-end demonstration. To confirm that the offline results transfer to a live system, we served attacker-authored responses from an endpoint on the loopback interface, fetched them with the pipeline's own HTTP client over TCP, and applied the pipeline's own confirmation rule. Three of five responses, none of which disclosed any data, were confirmed as vulnerabilities and would have reached an operator labelled confirmed by observation.
10
3.5
Moving the decisive value off the attacker's channel closes the attack, at a cost
Figure 5. The defence, its cost, and the adversary against which it fails. (a) Seven confirmation policies over 1,932 attacked negatives. Red, attack success; blue, sensitivity on five constructed positives; green, sensitivity on the 39 observed findings. anchored_or_conj2 (eq. (7)) reduces attack success from 97 % to 0 % with unchanged repeatability and cost. On constructed positives its sensitivity is unchanged; on observed findings it falls from 100 % to 18 %, because each observed finding fires a single mechanism and the policy requires corroboration. (b) Outcome shares of the observed positives under each policy: confirmed (blue), escalated to an analyst (green) and missed (grey); filled circles mark ternary and open circles binary policies, with attack success and escalation rate under attack annotated. escalate_conj (eq. (8)) makes the same confirmation decisions as the hardened binary policy and escalates the remainder: 100 % total recall, 0 % missed and 0 % attack success, with escalation of 97 % of attacked evidence and 0 % of clean evidence. (c) Five policies on the four higher-is-better axes defined in Methods, shown as parallel coordinates because three policies coincide on any two axes. Coloured lines are the three frontier policies; the grey dashed line is the two dominated policies, which coincide. binary_orig is on the frontier only through its unattacked sensitivity, and is forged by a three-character payload. (d) The third-party corpus under two adversaries. Bottom: a third party that controls responses but not callbacks. Top: the scanned host, which receives the
11
interactsh60 callback URL in the payload and the per-run value in the request; a further 569 mechanisms fall and none remains out of reach. The criterion identifies the remedy. A verifier is forgeable to the extent that its decision reads attackercontrolled data, so hardening means moving the decisive value to a channel the attacker cannot write (eq. (7)), not adding determinism. The pipeline already contains such anchors, namely the mechanisms with G = 0: a per-run nonce reflected back, an out-of-band canary callback, and a sentinel value that the server persisted. Where a vulnerability class has such an anchor, it should be used; where it has none, confirmation should require corroboration or be referred to a human. The policy anchored_or_conj2 (eq. (7)) reduces attack success over 1,932 attacked negatives from 97 % to 0 % with unchanged repeatability and cost (Fig. 5a). Its sensitivity is unchanged on five constructed positives but falls from 100 % to 18 % on the 39 observed findings, because each observed finding fires a single mechanism and the policy requires corroboration (Fig. 5a). All policies are scored in Table S5, and Table 3 compares the resulting systems on six axes. Table 3. Verifier systems compared on six axes. Sensitivity is measured on the 31 findings the pipeline labelled positive. Attack success is the mean false-confirmation rate over the 12 non-zero attacker budgets of the α-sweep (147 cases); it is therefore not the AUFC of Table S1.
Note. On five hand-constructed positives (one per signal type) the original and hardened rules both score 100%; on the observed positives they score 100% and 18%, because each observed finding fires a single mechanism and the hardened policy requires corroboration. Auditability is the number of constants the decision consults: for the original rule, the marker list of Table 1; for the hardened policy, none on its anchored path and the reachable mechanisms' constants on its corroboration path (Table S5). Model arms have no enumerable list. The cost in sensitivity is substantial. A conjunctive policy is affordable only where findings produce corroborating signals; where they produce a single signal, the binary choice is between an attackable confirmation and an inconclusive verdict referred to a human. A third output (eq. (8)), which escalates such findings together with the probe that would resolve them, keeps attack success at 0 % with 100 % total recall and no missed positives; it escalates 97 % of attacked evidence and none of the clean evidence (Fig. 5b). Escalations then occur only under attack, which is the appropriate failure direction for a triage system. On the four scoring axes defined in Methods, three policies form the frontier (Fig. 5c): the hardened binary policy, its ternary counterpart, and the original rule, which is on the frontier only through its unattacked sensitivity and is forged by a three-character payload. These results assume a third-party adversary. The scanned host itself receives the out-of-band callback URL in the payload and the per-run value in the request; under that adversary a further 569 third-party mechanisms are forged and none remains out of reach (Fig. 5d).
12
4
Discussion
The distinction between judged and confirmed findings is real but does not follow the stochastic/deterministic divide. What matters is whether the decisive value is on a channel the attacker writes; by that criterion the rule and the judge fall on the same side and differ in threshold rather than in kind. Three practical recommendations follow. Audit the channel, not the paradigm. Whether a decision reads attacker-controlled data can be determined from source code before any attack is attempted: the gate predicted all 16 held-out mechanisms, and none of the 10,925 scored third-party mechanisms was predicted unreachable and then forged (Fig. 1b). The number of constants a decision consults is not a substitute: it neither separates forgeable from unforgeable mechanisms (Fig. 1a) nor orders their cost consistently (Fig. 2b,d). This makes the check suitable for routine code review. Report the roster with the number. Every ensemble statistic in this study changed with fleet composition on identical evidence (Fig. 3c,d), and a single judge moved fleet agreement more than training lineage did (Fig. S1b,c); repeatability checks detect neither effect. Agreement, sensitivity or confidence reported for a panel of models should therefore be accompanied by the panel's exact membership. Prefer a ternary verdict. Every hardening we measured traded sensitivity for robustness until an escalate verdict was allowed (Fig. 5a,b). Neither a different implementation (Fig. 4a,b) nor composition with a judge ensemble (Fig. 4d) removed the trade-off. With it, the trade-off disappears, at the cost of analyst attention only when an attack is under way. The findings also bear on evaluation. Benchmarks for offensive and defensive agents often decide success by matching a string in the output10,11 . Such criteria are confirmation mechanisms with G = 1: they can be satisfied without performing the task, and an agent optimised against them is optimised against a proxy that can diverge from the intended goal53,55 . Our judges needed only 20 to 79 characters to be flipped (Fig. 2c), and a string-matching criterion only three (Fig. 2b). A benchmark whose success signal the system under test can author does not measure capability reliably.
4.1
Limitations
One pipeline and evidence distribution. The evidence comprises 200 instances from five target families, 84 % from one application. The six re-implementations and the third-party corpus show that the weakness is not specific to one codebase, but all reported rates are properties of this distribution, and the third-party corpus was scored from matcher specifications rather than executed against live hosts. Open-weight judges only. The eight model arms are 6.7–16B open-weight models, run locally with greedy decoding. We make no claim about hosted frontier models, to which hard cases are most likely to be sent, or, on this pool, about larger models or sampled decoding. Evaluating hosted judges is the most important extension. Mechanism-derived labels. The positive class is labelled by the pipeline's own mechanisms, which this paper shows to be forgeable. These mechanisms return inconclusive rather than guess, which supports but does not establish their labels. Fifty blinded annotation packets, stratified and stripped of every field stating a pipeline conclusion, have been prepared for two expert annotators; of the four steps of the labelling procedure, only the first is complete (Fig. S2a); a machine-annotated pilot on the same packets, reported as a pilot rather than ground truth, finds that the implementations most resistant to forgery agree least with a blinded reader (73 % agreement for substring matching against 34 % and 23 % for the two shape detectors; Fig. S2b,c). No operator data. We measure the machine, not the analyst. We have designed an operator study in which analysts see the same finding under labels ranging from a hedge to an assertion of an observed state change (Fig. S3a). It cannot yet be run: the run records contain one matched pair against a pre-registered minimum of eight (Fig. S3b), and every candidate stimulus for two of the three studies lies on a publicly taught target, so provenance cannot be separated from recognition (Fig. S3c). Power simulations indicate 24–48 participants for the binary contrast at effect sizes of 0.4–0.8 log-odds (Fig. S3d), but show that a two-level design cannot distinguish a gradual effect from a step (Fig. S3e). Design targets are listed in Table S9. 13
Bounds. Each point on a robustness curve is a lower bound on attack success, because the attacker draws from a fixed payload list; each b* is an upper bound on the minimum attacker cost. Where the two disagree, b* is the tighter statement. Adversary dependence. The unreachable mechanisms and the hardened policy rely on anchors the attacker cannot reach, and reachability depends on the adversary. A scanned host receives the out-of-band callback URL in the payload and the per-run value in the request; against that adversary none of the third-party mechanisms remains out of reach (Fig. 5d). Bug-bounty triage bots and agentic scanners operate under this model. All claims of unforgeability in this paper hold against a third party, not against a hostile target. Live demonstration. The end-to-end test shows that the pipeline confirms attacker-authored responses fetched over HTTP; it does not show how often an attacker attains that position.
5
Ethics and disclosure
No third-party system was scanned, contacted or attacked. The live demonstration used a loopback endpoint serving invented data. The third-party corpus is a public, openly licensed collection of scanner templates59 ; forgeries were constructed against matcher specifications, never against hosts, and no template author is named in connection with a defect. The vulnerabilities in the evidence corpus are in deliberately vulnerable applications maintained for security education. The weakness is a design property of a class of confirmation mechanisms rather than a flaw in a single product; because the attack requires only a few characters and no tooling, we publish it together with the criterion and the mitigation.
6
Methods
6.1
Pipeline and evidence
The pipeline is a four-stage API-security pipeline (generate → judge → confirm → measure) in which the third stage is the verdict authority: a finding is accepted when a deterministic mechanism observes the target enter a forbidden state. The evidence consists of 457 distinct instances from live scans of five target families; analyses use a 200-instance stratified pool and, for the robustness curves, the 147 instances on which every arm is a clean negative. 84 % of the pool comes from one application. The observed positive class is the 39 findings that the pipeline's own mechanisms confirm on unmodified evidence, and policies are scored against all 39. Comparisons involving model arms (Fig. 3c, Fig. 4d, Table 3) use the 31 of these that carry a verdict from every arm; the remaining 8 have no model verdict and are excluded rather than imputed.
6.2
Judges
Eleven open-weight models are used across four rosters, and each ensemble statistic is reported with its roster. The robustness curves use eight arms from six families (qwen, deepseek, llama, gemma, mistral, phi) at 6.7–16B. Agreement, the lineage contrast, the leave-one-out analysis and the master-key grid use a six-judge agreement fleet: qwen2.5:14b, qwen2.5:32b, qwen2.5-coder:32b, deepseek-r1:32b, phi4:14b and deepseek-coder-v2:16b. The sensitivity sweep uses a six-judge sweep fleet in which llama3.1:8b, mistral:7b and gemma2:9b replace the three models above 16B, and the repeatability study uses these six plus deepseek-r1:32b. All models are greedy-decoded and served locally through Ollama on one RTX 5090; no hosted model is used.
6.3
Threat model
The attacker controls the body an endpoint returns and may choose its status code and headers. This capability arises from server-side injection, a compromised upstream service, an attacker-registered tenant on a multi-tenant API, or server-side request forgery that directs a scanner to an attacker-owned host. The attacker does not control values the tester generates per run, out-of-band callbacks to tester infrastructure, or what the real server persisted and returns on an independent read-back. The Limitations discuss the case in which the last assumption fails. 14
6.4
The auditable attack surface
Let D be a decision procedure returning a confirmation verdict, S(D) the set of literal constants it consults (marker lists, status sets, regular expressions, thresholds), and let R(D) = ⊮[attacker-controlled input reaches D], K(D) = ⊮[the decisive value is attacker-knowable].
(1)
The auditable attack surface is the constant count gated by reachability: AAS(D) =
| S(D) |, 0,
R(D) · K(D) = 1, otherwise.
(2)
AAS is not a probability. It counts what an auditor must read to determine how a decision can be driven, and is zero when no attacker-reachable input affects it. An LLM judge has R = K = 1 and an unbounded, non-enumerable | S |. We call G(D) = R(D)K(D) the reachability gate and report it separately from | S |, because the gate predicts forgeability across populations and the magnitude does not. A reachable mechanism with no literal constant (G = 1, | S |= 0; one case in Table 2) has AAS = 0 but is classified by G, which is the quantity used for prediction.
6.5
Attacker budget and the robustness curve
Let x be an evidence instance with response body b(x), let x ⊕ p denote x with payload p composed into that body, and let P be the payload families available to the attacker. For an arm A, let NA = {x : A(x) = decline} be the instances on which it is correct before the attack. The budget α ∈ [0, 1] is the fraction of the body the attacker may author, and the false-confirmation rate is FA (α) =
X 1 ⊮[ ∃ p ∈ P : | p |≤ α | b(x) |, A(x ⊕ p) = confirm ]. | NA |
(3)
x∈NA
Restricting to NA is the clean-accuracy control: no forgery is scored against a detector that was already wrong. Over a grid α1 < . . . < αm we summarise the curve by its normalised area and breakpoint, Rα AU F C(A) = αm1−α1 α1m FA (α) dα, * = min{α : FA (α) ≥ FAmin + 12 ρA }, αA
(4)
with rise ρA = FAmax − FAmin and the integral evaluated by the trapezoid rule. A shape verdict is issued only when the rise exceeds a noise floor that scales with the sample size behind each point, q εA = max 0.05, 2 F̄A 1 − F̄A /n, 2/n ,
(5)
so that at small n a change in one or two findings is not read as degradation. Where ρA > εA , the arm is a cliff if its 10–90 % transition width WA = (α90 − α10 ) / (αm − α1 ) satisfies WA ≤ 0.25 and gradual otherwise; when the transition falls within a single wide grid interval, the estimator returns unresolved. The grid is dense below α = 10% and refined at 30, 35, 40, 45 and 60 %, so that shape verdicts are not artefacts of grid resolution. Because the curve is a bound for a laddered adversary, we also compute the exact minimum budget in characters by generate-then-minimise search: b*A (x) = min{| p |: p ∈ P, A(x ⊕ p) = confirm}, reported per arm as the median over NA . Each b* is an upper bound on the true minimum.
15
(6)
6.6
Confirmation policies and scoring
Let M = U ∪˙ R partition the mechanisms into unreachable (G = 0) and reachable (G = 1) ones, and let m(x) ∈ {0, 1} indicate whether m fires on x. The hardened binary policy is W Πhard (x) =P confirm ⇔ m∈U m(x) ∨ m∈R m(x) ≥ 2 ,
(7)
and its ternary counterpart escalates what neither branch decides: confirm, Πesc (x) = escalate, decline,
Πhard (x) = confirm, ∃m ∈ M : m(x) = 1, otherwise.
(8)
Over an observed positive class P , an attacked negative set N † and a clean negative set N , we score four higher-is-better axes: robustness 1 − PrN † [confirm], confirmed sensitivity PrP [confirm], total recall PrP [confirm ∨ escalate], and attention efficiency 1 − PrN † [escalate]. A policy Π is dominated if some Π′ is at least as good on all four axes and strictly better on one. Both sensitivity axes are needed, because confirmed sensitivity alone treats an escalated positive as a missed one.
6.7
Controls
We applied three controls. Clean accuracy: eq. (3) restricts every attack to NA . Deletion only: truncating evidence without adding content flips 0 of 161 cases on every arm, so the curves measure forgery rather than loss of evidence. Dormancy: a rule that returns inconclusive on every real instance is trivially flipped, so the six-implementation study attacks a mechanism that fires on 39 of 200 instances.
6.8
Statistics and pre-registration
We use exact permutation tests and Fisher's exact test, because the panels are small; where a design has a smallest attainable p, it is reported. Spearman correlations over the 400-mechanism corpus sample use the t approximation. For the held-out test, every prediction was read from source and hashed before any attack was written (PREDICTION_SHA = a41c7a1fec100b9f).
6.9
Reproducibility and compute
All judging ran locally through Ollama on one RTX 5090, with no external API, for approximately 16 GPU-hours in total. Deterministic audits are exhaustive, with no sampling. Every figure and table is generated programmatically from the machine-written record of the analysis that produced it.
6.10
Author contributions
A.F. and N.S. designed the study. A.F. built the pipeline, the attack apparatus and the analysis, and ran the experiments. N.S. designed the audit criterion and the held-out prediction protocol. Both authors analysed the results and wrote the manuscript.
6.11
Data and code availability
The analysis records from which every figure and table is generated, the vector figures and the bibliography are provided with this preprint. The evidence corpus is drawn from deliberately vulnerable applications maintained for security education; the third-party template corpus is public and openly licensed59 and is not redistributed. The code used in this study is available from the corresponding author on reasonable request. Because parts of it implement working attacks on security verification mechanisms, access to those components may be limited, or granted after review of the intended use. No human-subject data were collected.
16
6.12
Competing interests
N.S. is employed by OmiCore Inc., which develops the security pipeline audited here. A.F. is a graduate student at Kyushu University and contributes to OmiCore Inc. as an unpaid volunteer. All defects found in that pipeline are reported, together with the mitigation.
7
References
1. B. A. Alahmadi, L. Axon and I. Martinovic. 99% False Positives: A Qualitative Study of SOC Analysts' Perspectives on Security Alarms. 31st USENIX Security Symposium, 2022. 2. B. Johnson, Y. Song, E. Murphy-Hill and R. Bowdidge. Why Don't Software Developers Use Static Analysis Tools to Find Bugs?. ICSE, 2013. doi:10.1109/ICSE.2013.6606613. 3. M. Christakis and C. Bird. What Developers Want and Need from Program Analysis: An Empirical Study. ASE, 2016. doi:10.1145/2970276.2970347. 4. C. Sadowski, E. Aftandilian, A. Eagle, L. Miller-Cushon and C. Jaspan. Lessons from Building Static Analysis Tools at Google. Communications of the ACM 61(4):58–66, 2018. doi:10.1145/3188720. 5. J. Smith, L. N. Q. Do and E. Murphy-Hill. Why Can't Johnny Fix Vulnerabilities: A Usability Evaluation of Static Analysis Tools for Security. SOUPS, 2020. 6. A. Doupé, M. Cova and G. Vigna. Why Johnny Can't Pentest: An Analysis of Black-box Web Vulnerability Scanners. DIMVA, 2010. doi:10.1007/978-3-642-14215-4_7. 7. J. Bau, E. Bursztein, D. Gupta and J. Mitchell. State of the Art: Automated Black-Box Web Application Vulnerability Testing. IEEE Symposium on Security and Privacy, 2010. doi:10.1109/SP.2010.27. 8. R. Fang, R. Bindu, A. Gupta and D. Kang. LLM Agents can Autonomously Exploit One-day Vulnerabilities. arXiv preprint, 2024. arXiv:2404.08144. 9. G. Deng, Y. Liu, V. Mayoral-Vilches, P. Liu, Y. Li, Y. Xu, T. Zhang, Y. Liu, M. Pinzger and S. Rass. PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing. 33rd USENIX Security Symposium, 2024. arXiv:2308.06782. The arXiv preprint carries an earlier title; the version cited is the USENIX Security 2024 paper. 10. A. K. Zhang, N. Perry, R. Dulepet, J. Ji, C. Menders, J. W. Lin, E. Jones, G. Hussein, S. Liu, D. Jasper, P. Peetathawatchai, A. Glenn, V. Sivashankar, D. Zamoshchin, L. Glikbarg, D. Askaryar, M. Yang, T. Zhang, R. Alluri, N. Tran, R. Sangpisit, P. Yiorkadjis, K. Osele, G. Raghupathi, D. Boneh, D. E. Ho and P. Liang. Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models. ICLR, 2025. arXiv:2408.08926. 11. M. Shao, S. Jancheska, M. Udeshi, B. Dolan-Gavitt, H. Xi, K. Milner, B. Chen, M. Yin, S. Garg, P. Krishnamurthy, F. Khorrami, R. Karri and M. Shafique. NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. NeurIPS Datasets and Benchmarks Track, 2024. arXiv:2406.05590. 12. P. K. Manadhata and J. M. Wing. An Attack Surface Metric. IEEE Transactions on Software Engineering 37(3):371–386, 2011. doi:10.1109/TSE.2010.60. 13. N. Munaiah and A. Meneely. Attack Surface Definitions: A Systematic Literature Review. Information and Software Technology 104:94–103, 2019. doi:10.1016/j.infsof.2018.07.008. 14. K. Thompson. Reflections on Trusting Trust. Communications of the ACM 27(8):761–763, 1984. doi:10.1145/358198.358210. 15. S. Torres-Arias, H. Afzali, T. K. Kuppusamy, R. Curtmola and J. Cappos. in-toto: Providing farm-totable guarantees for bits and bytes. 28th USENIX Security Symposium, 2019. 16. L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez and I. Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks Track, 2023. arXiv:2306.05685. 17. P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu and Z. Sui. Large Language Models are not Fair Evaluators. ACL, 2024. arXiv:2305.17926. 18. A. Panickssery, S. R. Bowman and S. Feng. LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS, 2024. arXiv:2404.13076. 19. D. Li, B. Jiang, L. Huang, A. Beigi, C. Zhao, Z. Tan, A. Bhattacharjee, Y. Jiang, C. Chen, T. Wu, K. Shu, L. Cheng and H. Liu. From Generation to Judgment: Opportunities and Challenges of 17
LLM-as-a-judge. arXiv preprint, 2024. arXiv:2411.16594. 20. X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery and D. Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR, 2023. arXiv:2203.11171. 21. P. Verga, S. Hofstatter, S. Althammer, Y. Su, A. Piktus, A. Arkhangorodsky, M. Xu, N. White and P. Lewis. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv preprint, 2024. arXiv:2404.18796. 22. Y. Zhao, H. Liu, D. Yu, S. Y. Kung, H. Mi and D. Yu. One Token to Fool LLM-as-a-Judge. arXiv preprint, 2025. arXiv:2507.08794. 23. J. Shi, Z. Yuan, Y. Liu, Y. Huang, P. Zhou, L. Sun and N. Z. Gong. Optimization-based Prompt Injection Attack to LLM-as-a-Judge. ACM CCS, 2024. arXiv:2403.17710. 24. K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz and M. Fritz. Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. ACM AISec, 2023. arXiv:2302.12173. 25. F. Perez and I. Ribeiro. Ignore Previous Prompt: Attack Techniques For Language Models. NeurIPS ML Safety Workshop, 2022. arXiv:2211.09527. 26. Y. Liu, Y. Jia, R. Geng, J. Jia and N. Z. Gong. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. 33rd USENIX Security Symposium, 2024. arXiv:2310.12815. 27. E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer and F. Tramèr. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. NeurIPS Datasets and Benchmarks Track, 2024. arXiv:2406.13352. 28. Q. Zhan, Z. Liang, Z. Ying and D. Kang. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. Findings of ACL, 2024. arXiv:2403.02691. 29. B. Biggio and F. Roli. Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning. Pattern Recognition 84:317–331, 2018. arXiv:1712.03141. 30. N. Carlini and D. Wagner. Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods. ACM AISec, 2017. arXiv:1705.07263. 31. F. Tramèr, N. Carlini, W. Brendel and A. Madry. On Adaptive Attacks to Adversarial Example Defenses. NeurIPS, 2020. arXiv:2002.08347. 32. F. Pierazzi, F. Pendlebury, J. Cortellazzi and L. Cavallaro. Intriguing Properties of Adversarial ML Attacks in the Problem Space. IEEE Symposium on Security and Privacy, 2020. arXiv:1911.02142. 33. F. Pendlebury, F. Pierazzi, R. Jordaney, J. Kinder and L. Cavallaro. TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time. 28th USENIX Security Symposium, 2019. arXiv:1807.07838. 34. D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro and K. Rieck. Dos and Don'ts of Machine Learning in Computer Security. 31st USENIX Security Symposium, 2022. arXiv:2010.09470. 35. G. Klees, A. Ruef, B. Cooper, S. Wei and M. Hicks. Evaluating Fuzz Testing. ACM CCS, 2018. arXiv:1808.09700. 36. A. Hazimeh, A. Herrera and M. Payer. Magma: A Ground-Truth Fuzzing Benchmark. ACM SIGMETRICS, 2021. arXiv:2009.01120. 37. R. Croft, M. A. Babar and M. M. Kholoosi. Data Quality for Software Vulnerability Datasets. ICSE, 2023. arXiv:2301.05456. 38. S. Chakraborty, R. Krishna, Y. Ding and B. Ray. Deep Learning based Vulnerability Detection: Are We There Yet?. IEEE Transactions on Software Engineering, 2021. arXiv:2009.07235. 39. B. Steenhoek, M. M. Rahman, R. Jiles and W. Le. An Empirical Study of Deep Learning Models for Vulnerability Detection. ICSE, 2023. arXiv:2212.08109. 40. Y. Ding, Y. Fu, O. Ibrahim, C. Sitawarin, X. Chen, B. Alomair, D. Wagner, B. Ray and Y. Chen. Vulnerability Detection with Code Language Models: How Far Are We?. ICSE, 2025. arXiv:2403.18624. 41. C. Rossow, C. J. Dietrich, C. Grier, C. Kreibich, V. Paxson, N. Pohlmann, H. Bos and M. van Steen. Prudent Practices for Designing Malware Experiments: Status Quo and Outlook. IEEE Symposium on Security and Privacy, 2012. doi:10.1109/SP.2012.14. 42. T. Avgerinos, S. K. Cha, B. L. T. Hao and D. Brumley. AEG: Automatic Exploit Generation. NDSS, 2011. 43. Y. Shoshitaishvili, R. Wang, C. Salls, N. Stephens, M. Polino, A. Dutcher, J. Grosen, S. Feng, C. Hauser,
18
C. Kruegel and G. Vigna. SoK: (State of) The Art of War: Offensive Techniques in Binary Analysis. IEEE Symposium on Security and Privacy, 2016. doi:10.1109/SP.2016.17. 44. R. Parasuraman and V. Riley. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors 39(2):230–253, 1997. doi:10.1518/001872097778543886. 45. G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro and D. S. Weld. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. ACM CHI, 2021. arXiv:2006.14779. 46. Z. Buçinca, M. B. Malaya and K. Z. Gajos. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proc. ACM Human-Computer Interaction 5(CSCW1), 2021. arXiv:2102.09692. 47. M. Schemmer, N. Kuehl, C. Benz, A. Bartos and G. Satzger. Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. ACM IUI, 2023. arXiv:2302.02187. 48. H. Vasconcelos, M. Jörke, M. Grunde-McLaughlin, T. Gerstenberg, M. S. Bernstein and R. Krishna. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. Proc. ACM HumanComputer Interaction 7(CSCW1), 2023. arXiv:2212.06823. 49. K. Zhou, J. Hwang, X. Ren and M. Sap. Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty. ACL, 2024. arXiv:2401.06730. 50. J. L. Fleiss. Measuring Nominal Scale Agreement Among Many Raters. Psychological Bulletin 76(5):378– 382, 1971. doi:10.1037/h0031619. 51. J. R. Landis and G. G. Koch. The Measurement of Observer Agreement for Categorical Data. Biometrics 33(1):159–174, 1977. doi:10.2307/2529310. 52. K. Krippendorff. Reliability in Content Analysis: Some Common Misconceptions and Recommendations. Human Communication Research 30(3):411–433, 2004. doi:10.1111/j.1468-2958.2004.tb00738.x. 53. D. Manheim and S. Garrabrant. Categorizing Variants of Goodhart's Law. arXiv preprint, 2018. arXiv:1803.04585. 54. D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman and D. Mané. Concrete Problems in AI Safety. arXiv preprint, 2016. arXiv:1606.06565. 55. J. Skalse, N. H. R. Howe, D. Krasheninnikov and D. Krueger. Defining and Characterizing Reward Hacking. NeurIPS, 2022. arXiv:2209.13085. 56. A. Pan, K. Bhatia and J. Steinhardt. The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. ICLR, 2022. arXiv:2201.03544. 57. J. Jacobs, S. Romanosky, B. Edwards, M. Roytman and I. Adjerid. Exploit Prediction Scoring System (EPSS). Digital Threats: Research and Practice 2(3), 2021. arXiv:1908.04856. 58. A. D. Householder, J. Chrabaszcz, T. Novelly, D. Warren and J. M. Spring. Historical Analysis of Exploit Availability Timelines. USENIX CSET Workshop, 2020. 59. ProjectDiscovery. nuclei-templates: Community-curated detection templates for the Nuclei scanner. GitHub repository, https://github.com/projectdiscovery/nuclei-templates (MIT licence; snapshot of 11 August 2026), 2026. 60. ProjectDiscovery. interactsh: An out-of-band interaction gathering server and client library. GitHub repository, https://github.com/projectdiscovery/interactsh, 2026.
8
Supplementary material
Supplementary Figures S1–S3 and Tables S1–S9 follow. Null results are reported alongside the smallest attainable p of their design where one exists.
19
Figure S1, related to Figure 3. Agreement between model judges. (a) Distribution of the fleet vote over 39 findings and six judges: the x-axis is the number of judges calling a finding a true positive and the bar height the number of findings. Agreement is substantial (Fleiss κ = +0.683, Krippendorff α = +0.685) and bimodal: 74 % of findings lie at the extremes of the vote (red) and the middle is nearly empty. This is why majority voting suppresses individual wobble in Figure 3d, since a member must both wobble and be pivotal. (b) Pairwise κ for all 15 judge pairs, split by whether the pair shares a training lineage; dots are pairs and horizontal lines group means. The difference, +0.040 (exact permutation p = 0.30, 60 relabellings), is not significant. (c) Change in fleet κ when each judge is removed, annotated with that judge's mean pairwise agreement with the others; red bars mark judges whose removal increases agreement. Removing one judge changes fleet κ by +0.066, more than the lineage difference in (b). Self-consistency does not predict agreement (ρ = +0.154, p = 0.77). (d) Master-key grid: five judges and the deterministic rule (rows) under seven conditions (columns); cells give the share of ten cases confirmed as true positives. Transfer to judges is weak (the most affected judge flips on 20 % and three of five never flip), whereas the deterministic marker rule flips on 100 %. This rule is the SQL-injection confirmer, which is dormant on this authorisation evidence; Figure 4 therefore attacks a mechanism that fires.
20
Figure S2, related to Figure 4 and the Limitations. Independent labels. All accuracies in the main figures are measured against labels produced by the pipeline itself. (a) The independent-labelling procedure in four steps: blinded packets drawn, annotations collected, inter-annotator κ computed, arms re-scored. Only the first step is complete: fifty findings were drawn from the pool of 200 with a fixed seed, with five conclusion-bearing fields withheld, and the collection workbook validates each return before analysis. Two expert annotators are required for the remaining steps. (b) A machine-annotation pilot, which is not ground truth, since an LLM cannot serve as the reference in a study of the reliability of machine verdicts. It measures agreement between each implementation and a blinded reader of the same bytes over the 44 decidable packets. v1_substring, which has 100 % sensitivity against the pipeline's own labels by construction, agrees with the reader on 73 %; the two shape detectors agree on 34 % and 23 % (red, below chance-level usefulness). (c) The six implementations by resistance to a one-sentence prose payload (x) and agreement with the blinded reader (y); marker area is proportional to the clean false-positive rate. Resistance to forgery and agreement with the reader are inversely related. Six packets are undecidable because redaction removes the deciding information; in three cross-object reads the attack and baseline requests become identical once the token is masked. This packet-builder defect was detected by the pilot before expert annotation.
21
Figure S3, related to the Limitations. Design of the planned operator study. (a) The manipulation: a five-level scale of asserted warrant, from a hedge (possibly exploitable, orange) through probability statements (blue) to assertions of an observed state change (green). The quantities of interest are the acceptance gap between deterministic and model-judged labels, ∆det = P (accept | det) − P (accept | model), the corresponding accuracy gap ∆just , and their difference ∆excess = ∆det − ∆just , i.e. acceptance beyond what accuracy justifies. (b) Supply of matched pairs, in which the same finding was both model-judged and deterministically confirmed: the run records contain one, against a pre-registered minimum of eight (red dashed) and a target of twelve (dotted). A further 91 pairs (grey), in which a detector's finding was kept without running the judge, do not involve an observed state change and are not eligible. (c) Stimulus exposure by study. Red, share of candidate stimuli on a publicly taught target; blue, share of trials retained at an assumed 40 % recognition rate, annotated with the resulting recruitment multiplier. All candidates for S1 and S2 are exposed, so provenance cannot be separated from recognition. (d) Participants required for 80 % power on the binary contrast, by effect size, from 2,000 simulations per cell with participant and item random effects. (e) Participants required for 80 % power on the shape of the effect (gradient, cliff or mixed), by design (rows) and true regime (columns); a dash marks a cell the design cannot identify at any sample size. With two levels, the cliff indicator is an affine function of the manipulation, so the binary contrast cannot distinguish a gradient from a step. No human-subject data have been collected, and no effect size is assumed. Design targets are listed in Table S9.
22
9
Supplementary tables
Table S1. Failure geometry, arm by arm. All arms measured on the same 147-case clean-negative pool over 13 attacker budgets. AUFC is the area under the forgeability curve; α∗ is the budget at which the arm breaks.
Note. ρ(parameters, AUFC) = -0.193, exact p = 0.654. The within-scale control is the 5 arms in the 6.7–9B band, drawn from 5 families. Separation margin from the deterministic rule: 0.607. Four arms evaluated only on a 10-case pilot pool are excluded, because areas over 10 and 147 cases are not comparable. Table S2. Minimum attacker budget b*, by adaptive search. Generate-then-minimise search over seven payload families, not a fixed ladder. Every b* is an upper bound: a cleverer attacker pays no more. Adversary model: append.
Note. Each deterministic arm is searched only on its own clean negatives. The deletion-only control (truncation without added content) flips no case. The prompt-injection payload family wins no case against
23
any deterministic arm. Table S3. Six independent implementations of one check under six attack conditions. confirm_unauth fires on 39 of 200 evidence instances and declines on the remaining 161, which form the attack set. Each cell is the share of attackable cases flipped; the clean column is the unattacked control.
Note. Worst condition: json_real. The two shape detectors resist prose payloads but already fire on most unmodified evidence (clean column). Table S4. Activation of the status-inversion branch, six implementations. 200 instances. The shipped rule returns inconclusive on all of them through both branches; the table therefore reports activation, i.e. how many attacker-authored fields are needed to satisfy the conjunction.
Note. A conjunction costs the attacker only the conjuncts that are not already true. The anchored and read-back implementations cannot be activated. Table S5. Confirmation policies, binary and ternary. Measured on 5 constructed positives, 39 observed positives, 27 clean negatives and 1,932 attacked negatives. The two sensitivity columns use different denominators.
24
Note. Frontier membership uses both sensitivity axes; scoring confirmed sensitivity alone would treat an escalated positive as a missed one. Escalation on clean evidence is 0 % for every ternary policy. Table S6. Ensemble statistics by fleet size and roster. κ computed over 39 findings, sensitivity over 31 observed positives. Instability is the share of (roster, case) pairs whose verdict differs from the full fleet's on identical evidence.
Note. The full fleets give κ = +0.683 (6 judges) and sensitivity 54.8% (6 judges); a different roster of the same size changes both. The 0 % instability at seven judges reflects a single possible roster. The same fleet queried twice is 99.7% self-consistent over 2,380 inferences. Table S7. The reachability gate applied to a third-party template corpus. Each mechanism is scored automatically from the template's matcher specification (nuclei-templates). The two blocks differ only in the adversary model.
25
Note. Under the third-party adversary, the 585 unreachable mechanisms resist by definition, since the attacker cannot write the channel they read. The scanned host, by contrast, receives the out-of-band callback URL in the payload and the per-run value in the request, so under the hostile-host adversary no mechanism is out of reach. Table S8. Per-judge statistics of the agreement fleet. Fleet Fleiss κ = +0.683 and Krippendorff α = +0.685 over 39 findings and 6 judges. 74% of findings sit at an extreme of the vote. Centrality is a judge's mean pairwise κ against the rest of the fleet; ∆κ is what removing it does to the fleet.
Note. Within- minus across-lineage κ is +0.040 (exact permutation p = 0.30); self-consistency does not predict agreement (ρ = +0.154, p = 0.77). Self-consistency is from 2,380 repeated inferences; a dash marks a judge not in the repeatability study. Table S9. Planned operator study: supply, stimuli and power. No human-subject data have been collected. Values are design targets and simulated sample-size requirements, not results.
26
Note. The study is blocked by supply: the pre-registered matched-pair floor is not met, and every candidate stimulus lies on a publicly taught target, so provenance cannot be separated from recognition without a non-public target.
27