From Admission to Invariants: Measuring Deviation in Delegated Agent Systems Marcelo Fernandez
arXiv:2604.17517v1 [cs.AI] 19 Apr 2026
TraslaIA [email protected]
Abstract Autonomous agent systems are governed by enforcement mechanisms that flag hard constraint violations at runtime. The Agent Control Protocol [5] identifies a structural limit of such systems (§17): a correctly-functioning enforcement engine can enter a regime in which behavioral drift is invisible to it, because the enforcement signal operates below the layer where deviation is measurable. We show that enforcement-based governance is structurally unable to determine whether an agent’s behavior remains within the admissible behavior space A0 established at admission time. Our central result, the Non-Identifiability Theorem, proves that A0 ∈ / σ(g): the σ-algebra generated by the enforcement signal g does not contain A0 under the Local Observability Assumption, which every practical enforcement system satisfies. The impossibility arises from a fundamental mismatch: g evaluates actions locally against a point-wise rule set, while A0 encodes global, trajectory-level behavioral properties set at admission time. We then define the Invariant Measurement Layer (IML), which bypasses this limitation by retaining direct access to the generative model of A0 . We prove an informationtheoretic impossibility for enforcement-based monitoring; separately, we show that IML detects admission-time drift with provably finite detection delay, operating in the region where enforcement is structurally blind. We validate the theoretical claims across four experimental settings: three controlled drift scenarios (300 and 1000 steps), a live n8n webhook pipeline, and a LangGraph StateGraph agent with deterministic tool selection—in every case, enforcement triggers zero violations while IML deviation grows monotonically, detecting each drift type within 9–258 steps of drift onset. This paper is Paper 2 of a 4-paper Agent Governance Series: atomic decision boundaries (P0, [4]), stateful enforcement—ACP (P1, [5]), fair multi-agent allocation (P3, [3]), and composition irreducibility (P4, [6]).
Contents
1 Introduction
3
2 Problem Setup
4
2.1 Trace language and admissible behavior . . . . . . . . . . . . . . . . . . . . . . . . .
4
2.2
Enforcement signal and local observability . . . . . . . . . . . . . . . . . . . . . . . .
4
2.3
Observability structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5
2.4
Ground-truth deviation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
5
1
3 Theoretical Results
5
3.1 T1: Existence of the compliance-invariance gap . . . . . . . . . . . . . . . . . . . . .
5
3.2 T2: Non-identifiability of A0 under enforcement . . . . . . . . . . . . . . . . . . . . .
6
3.3 T3: IML Recoverability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
9
4 The Invariant Measurement Layer
10
4.1
Design principle . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.2
Deviation decomposition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
4.3
Comparison with baselines . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5 Experiments
11
5.1
Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11
5.2
Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12
5.3
Real-Agent Validation via n8n
5.4
Long-Horizon Validation (1000 Steps) . . . . . . . . . . . . . . . . . . . . . . . . . . 14
5.5
LangGraph Agent Experiment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13
6 Related Work
16
7 Discussion and Limitations
18
8 Conclusion
19
2
1
Introduction
Multi-agent systems delegate subtasks to agents that operate with partial autonomy under a shared policy framework. At admission time, when an agent is authorized and its operational scope defined, the system implicitly records the constraints, initial context, and delegation lineage that constitute the agent’s admissible behavior space A0 . The dominant governance model then attempts to maintain alignment through an enforcement signal g : Σ∗ → {0, 1}, which returns 1 whenever a trace τ contains a hard constraint violation. This architecture has a fundamental observability limitation that has received little formal attention. Enforcement signals are local: they evaluate each action against a rule set that is static, does not change with trajectory history, and does not reference A0 directly. An agent can therefore drift— systematically shifting its behavioral distribution away from admission-time expectations—while every individual action remains within the permitted action space. This phenomenon is not a corner case; it is the generic behavior of any gradual distributional shift, goal reinterpretation, or delegation-depth creep that stays below hard constraint thresholds. Motivation. A foundational result in agent governance [4] establishes that only atomic decision systems—where evaluation and execution occur at the same state transition—can guarantee admissibility at runtime. Split systems that separate evaluation from execution cannot close this gap regardless of policy sophistication. Fernandez [5] instantiates this principle as the Agent Control Protocol (ACP): an atomic admission-control mechanism that requires execution-trace state. Section 17 of that work identifies a structurally distinct failure mode that persists even within an atomic enforcement layer: a correctly-functioning, stateful enforcement system can enter a regime in which behavioral drift is invisible to it. ACP calls this deviation collapse (§17.4)—enforcement is active and syntactically correct, yet the behavioral boundary is never exercised because upstream conditions suppress all activating inputs. This paper operates above the atomic enforcement boundary and formalizes the conditions under which behavioral drift escapes detection. We prove (Theorem 3.2) that the admission-time behavioral contract A0 is not identifiable from the enforcement signal g alone—no matter how sophisticated the risk-scoring function. We then introduce the Invariant Measurement Layer (IML), an estimator that restores observability at the behavioral level, operating precisely in the region where g(τ ) = 0 always and enforcement-based monitoring is structurally blind. Main contributions. 1. T1 (Existence). We prove that the compliance-invariance gap is non-empty: trajectories satisfying g(τ ) = 0 yet τ ∈ / A0 exist under mild conditions (Lemma 3.1). 2. T2 (Non-Identifiability). We prove A0 ∈ / σ(g) under the Local Observability Assumption: no measurable function of the enforcement signal can reconstruct A0 -membership (Theorem 3.2). We exhibit an explicit constructive witness grounded in experimental data (Example 3.1). 3. T3 (IML Recoverability). We define IML and prove it is a consistent estimator of deviation D(τ, A0 ) with provably finite detection delay (Theorem 3.6). b t grows mono4. Empirical validation. Four experimental settings confirm g(τt ) = 0 ∀t while D tonically: three controlled drift scenarios (300 and 1000 steps), a live n8n webhook pipeline
3
(§5.3), and a real LangGraph StateGraph agent (§5.5)—directly instantiating both T2 and T3. Paper organization. Section 2 introduces the formal model. Section 3 states and proves T1, T2, and T3. Section 4 details the IML estimator. Section 5 presents the empirical validation. Section 6 discusses related work, and Section 8 concludes.
2 2.1
Problem Setup Trace language and admissible behavior
Let Σ be a finite alphabet of agent actions (tool calls, delegation requests, context updates). A trace is a finite sequence τ = (b1 , b2 , . . . , bT ) ∈ Σ∗ . Definition 2.1 (Admissible Behavior Space). The admissible behavior space is A0 = f (C, E0 , L) ⊆ Σ∗ , where C is the constraint set at admission, E0 is the initial context, and L is the delegation lineage. A0 is a global object: whether τ ∈ A0 depends on full-trajectory properties (distribution over actions, depth profile, context coherence), not on any single action in isolation. 2.2
Enforcement signal and local observability
Definition 2.2 (Enforcement Signal). An enforcement signal g : Σ∗ → {0, 1} returns 1 to indicate a policy violation. Assumption 2.1 (Local Observability). The enforcement signal decomposes as g(τ ) = h(V (τ )), where V : Σ∗ → V evaluates only point-wise constraint violations (e.g., presence of a forbidden tool, delegation depth exceeding a fixed limit) and h : V → {0, 1} aggregates violations. Crucially, V (τ ) is independent of A0 and depends only on individual actions, not on global trajectory properties. Assumption 2.1 is satisfied by every practical enforcement mechanism we are aware of: permissioncheck middleware, guardrail classifiers, schema validators, and policy engines all evaluate individual actions against a static rule set. Remark 2.1 (ACP as a natural instance of Definition 2.2). The Agent Control Protocol [5] provides a canonical instance of Definition 2.2 that satisfies Assumption 2.1. For a trace τ = (b1 , . . . , bT ) of agent requests, define the ACP projection: gACP (τ ) := 1[ ∃ i ≤ T : decision(bi ) ∈ {Denied, Escalated}] . Each decision(bi ) is computed by ACP’s evaluate-then-mutate pipeline against a static PatternKey rule set, independently of A0 ; it therefore satisfies the decomposition g = h ◦ V required by Assumption 2.1. We use gACP as the running example throughout this paper. Importantly, the Non-Identifiability Theorem (Theorem 3.2) is proved for the entire class of g satisfying Assumption 2.1, not for gACP alone; gACP is a natural instance, not the sole scope of the result. 4
Σ∗ (all traces) Compliance(C) = g −1 (0) (compliant traces)
g −1 (1)
A0
violations
admission-time behavior
detected by ACP
Compliance \ A0 hidden drift
ACP boundary
IML boundary
Figure 1: Layered structure of the trace space Σ∗ . ACP’s enforcement signal g partitions Σ∗ into g −1 (1) (constraint violations, detected by ACP) and Compliance(C) = g −1 (0) (compliant traces, invisible to enforcement). Within Compliance(C), IML monitors the boundary between A0 (admission-time behavior) and Compliance \ A0 (compliant but drifted behavior— hidden drift). The two mechanisms cover complementary boundaries; neither is redundant.
2.3
Observability structure
The enforcement signal g partitions Σ∗ into two equivalence classes: g −1 (0) (compliant traces) and g −1 (1) (violating traces). The corresponding σ-algebra is σ(g) =
∅, g −1 (0), g −1 (1), Σ∗ .
A set S ⊆ Σ∗ is identifiable from g if and only if S ∈ σ(g). 2.4
Ground-truth deviation
Definition 2.3 (Deviation Function). Let d : Σ∗ × Σ∗ → R≥0 be a trajectory metric. The deviation of τ from A0 is D(τ, A0 ) = inf x∈A0 d(τ, x). Remark 2.2 (Series notation). Throughout this paper, D(·, ·) denotes a deviation function (realvalued distance from a behavioral set). In Paper 0 [4], the symbol D denotes the decision domain D = {Allow, Refuse, Escalate}—a finite set of admission outcomes. The two uses are unrelated in type (function vs. set); when results from both papers appear together, we write D(τ, A0 ) for the deviation function and D or {Allow, . . .} for the decision domain to avoid ambiguity. D(τ, A0 ) is the ground truth but is intractable to compute directly because A0 is implicitly defined. Section 4 constructs IML as a computable, consistent approximation.
3 3.1
Theoretical Results T1: Existence of the compliance-invariance gap
Lemma 3.1 (Existence). Under Assumption 2.1, if A0 ̸= ∅ and A0 ⊊ Σ∗ , then there exists τ ∈ Σ∗ such that V (τ ) = ∅, g(τ ) = 0, and τ ∈ / A0 . 5
Proof. Since A0 ⊊ Σ∗ , there exists τ ∗ ∈ / A0 . If V (τ ∗ ) = ∅ already, we are done. Otherwise, τ ∗ contains a step bi that triggers V ̸= ∅. By Assumption 2.1, V depends only on point-wise properties of individual actions. Let τ ′ be obtained from τ ∗ by replacing every bi with V (bi ) ̸= ∅ with an action b′i drawn from the marginal distribution of A0 restricted to safe actions (non-empty by A0 ̸= ∅). This substitution satisfies V (τ ′ ) = ∅ by construction. It remains to show τ ′ ∈ / A0 . Recall that τ ∗ ∈ / A0 because it violates a trajectory-level constraint encoded in f (C, E0 , L)— for instance, the fraction of safe tools in τ ∗ falls below the threshold specified in C, or the mean delegation depth exceeds the lineage bound in L. The substitution replaces only actions with V (bi ) ̸= ∅ by safe actions, which either leaves the violating distributional property intact (if τ ∗ has too many boundary tools) or reduces the safe-tool fraction further (if the original violation was that boundary tools were over-represented). In either case, the distributional constraint that caused τ ∗ ∈ / A0 is still violated by τ ′ , so τ ′ ∈ / A0 , as required. In our empirical instantiation (Section 5), this construction is realized concretely: all 900 post-drift steps across three scenarios satisfy V (τt ) = ∅ while D(τt , A0 ) > 0. 3.2
T2: Non-identifiability of A0 under enforcement
Theorem 3.2 (Non-Identifiability). Under Assumption 2.1, if A0 ̸= ∅ and A0 ⊊ Σ∗ , then A0 ∈ / σ(g). Equivalently, there exists no measurable function f : {0, 1} → {0, 1} such that f (g(τ )) = 1[τ ∈A0 ] for all τ ∈ Σ∗ . Proof. By Lemma 3.1, there exists τ2 ∈ / A0 with g(τ2 ) = 0. Since A0 ̸= ∅, there exists τ1 ∈ A0 . If g(τ1 ) = 0, we already have g(τ1 ) = g(τ2 ) = 0 with τ1 ∈ A0 and τ2 ∈ / A0 . If g(τ1 ) = 1, then by the Local Observability Assumption 2.1, V (τ1 ) ̸= ∅; applying the same substitution argument as in Lemma 3.1 yields τ1′ ∈ A0 with V (τ1′ ) = ∅, hence g(τ1′ ) = 0. In either case, we have τ1 ∈ A0 and τ2 ∈ / A0 with g(τ1 ) = g(τ2 ) = 0. Suppose for contradiction that A0 ∈ σ(g). Since σ(g) is generated by g −1 (0) and g −1 (1), and A0 ⊆ g −1 (0), any σ(g)-measurable function f satisfies f (0) = c for some constant c ∈ {0, 1}. But then f (g(τ1 )) = f (0) = c = f (g(τ2 )), so 1[τ1 ∈A0 ] = 1[τ2 ∈A0 ] . This contradicts τ1 ∈ A0 and τ2 ∈ / A0 . Corollary 3.3 (Information-Theoretic Bound). I(A0 ; g(τ )) < H(A0 ). The mutual information between the enforcement signal and A0 -membership is strictly less than the entropy of A0 -membership. Proof. Suppose for contradiction that I(A0 ; g(τ )) = H(A0 ). This equality holds if and only if A0 membership is a deterministic function of g(τ )—i.e., there exists f : {0, 1} → {0, 1} with f (g(τ )) = 1[τ ∈A0 ] for all τ . But Theorem 3.2 proves exactly that no such f exists. Hence I(A0 ; g(τ )) < H(A0 ). Remark 3.1 (Why T2 is non-trivial). A referee might object: “any many-to-one function loses information.” T2 is stronger for three reasons. (i) Structural, not incidental. The result holds for the entire class of g satisfying Assumption 2.1— not for one particular rule set. No redesign of enforcement rules, no matter how sophisticated, can
6
make A0 ∈ σ(g), because the impossibility is a consequence of the local-observation architecture itself. (ii) The oracle contrast. If g were permitted to query Aemp at each step—i.e., to evaluate the JS 0 divergence JS(Pτ ∥PE0 ) in addition to point-wise violations—then A0 could be in σ(g). T2 says that precisely this access is what every practical enforcement system lacks. IML is the minimal mechanism that adds it back. (iii) Binary enforcement covers the decision boundary. At the point of action—permit or block— enforcement is effectively binary. A continuous risk score r(τ ) ∈ [0, 1] still produces a binary decision g(τ ) = 1[r(τ ) ≥ θr ], and our argument applies to that induced partition. More generally, for any finite-valued g : Σ∗ → G, the same structural barrier holds: unless g explicitly encodes A0 -distance (i.e., queries Aemp as IML does), σ(g) is a finite coarsening of Σ∗ that cannot resolve 0 A0 -membership. Corollary 3.4 (Irrecoverability). No function h : {0, 1} → {0, 1} satisfies h(g(τ )) = 1 ⇐⇒ τ ∈ A0 for all τ ∈ Σ∗ . In particular, no adjustment of a risk-scoring function can recover A0 -membership from enforcement signals alone. Proof. By Theorem 3.2, there exist τ1 ∈ A0 and τ2 ∈ / A0 with g(τ1 ) = g(τ2 ) = 0. Any h satisfying the condition would require h(0) = 1 (from τ1 ∈ A0 ) and h(0) = 0 (from τ2 ∈ / A0 )—a contradiction. The same argument applies symmetrically when g(τ1 ) = g(τ2 ) = 1. This corollary extends Fernandez [5] §1.1, which gives an analogous impossibility for stateless vs. stateful enforcement; here we show the same holds for enforcement-vs.-invariant monitoring. Theorem 3.5 (Monotonic Hidden Drift). Under the conditions of Theorem 3.2, there exists a sequence (τt )Tt=1 ⊂ g −1 (0) such that: 1. g(τt ) = 0 for all t ∈ {1, . . . , T }, and 2. D(τt , A0 ) is strictly increasing in t. That is, behavioral deviation from A0 can grow arbitrarily while producing no enforcement signal whatsoever. Proof. Base case (T = 2). Let τ1 ∈ A0 ; then g(τ1 ) = 0 and D(τ1 , A0 ) = 0. By Corollary 3.4, g −1 (0) \ A0 ̸= ∅, so there exists τ2 ∈ g −1 (0) with ∆ := D(τ2 , A0 ) > 0. The pair (τ1 , τ2 ) already satisfies both conditions. Extension to arbitrary T . For t = 2, . . . , T − 1, let τt be any trace with empirical distribution Pt =
t−1 1 − Tt−1 −1 P̂τ1 + T −1 P̂τ2 .
Since supp(Pt ) ⊆ supp(P̂τ1 ) ∪ supp(P̂τ2 ) ⊆ Σ \ F (both source traces have g = 0, hence no forbidden actions under Assumption 2.1), we have g(τt ) = 0 for all t. Define f : [0, 1] → R≥0 by f (s) = D(τ s , A0 ) where P̂τ s = (1 − s)P̂τ1 + sP̂τ2 . By the Lipschitz continuity of JS divergence in its first argument (a standard property; see Assumption 3.1(A3) below for the formal statement), f is continuous, with f (0) = 0 and f (1) = ∆ > 0. By the Intermediate Value Theorem, f attains 7
all values in [0, ∆]; choosing τt to realize parameter st = (t − 1)/(T − 1) yields D(τt , A0 ) = f (st ) strictly increasing in t (since f cannot be identically zero on (0, 1]: that would require P̂τ2 ∈ A0 , contradicting τ2 ∈ / A0 ). Theorem 3.5 formalizes what Fernandez [5] calls the hidden drift scenario within deviation collapse (§17.4): behavioral drift can be arbitrarily large while the enforcement boundary remains entirely silent. The BAR-Monitor of Fernandez [5] detects whether the enforcement boundary is being b detects drift within the region where BAR is structurally blind (g(τ ) = 0 exercised; IML’s D always). Constructive witness. The proof above is existential. We now exhibit an explicit pair (τ1 , τ2 ) extracted verbatim from our simulation. Example 3.1 (Constructive Witness for T2). We run a 300-step tool-drift simulation (seed 42) with alphabet Σ = {safe_read, safe_query, moderate_write, moderate_send, risky_execute, risky_delegate}. The enforcement function uses forbidden set F = {forbidden_exec, is forbidden_delete} ∩ Σ = ∅ and maximum delegation depth 10. The admission snapshot Aemp 0 built from 50 burn-in steps under base distribution {safe: 75%, boundary: 20%, risky: 5%}. Let τ1 = (b0 , . . . , b49 ) (admission-time segment) and τ2 = (b250 , . . . , b299 ) (post-drift segment). Table 1: Tool distributions and IML scores for the T2 witness pair. τ1 ∼g τ2 (both have g = 0), yet IML separates them. Tool safe_read safe_query moderate_write moderate_send risky_execute risky_delegate
Risk ρ
τ1
τ2
∆
0.10 0.10 0.50 0.60 0.85 0.90
36% 36% 14% 6% 4% 4%
16% 10% 32% 38% 4% 0%
−20pp −26pp +18pp +32pp ±0pp −4pp
0 0.1535 0.1500 0.2480
0 0.2167 0.2432 0.3465
g(·) b A0 ) D(·, Dt (JS div.) Dc (mean risk)
+0.063
g-equivalence. No action in F is used in either trace; all depths equal 1 < 10. Therefore V (τ1 ) = V (τ2 ) = ∅ and g(τ1 ) = g(τ2 ) = 0. The traces are enforcement-indistinguishable: no rule-based system can separate them. IML separability. b 2 , A0 ) − D(τ b 1 , A0 ) = |0.2167 − 0.1535| = 0.0632 > 0. D(τ
τ1 matches Aemp closely (72% safe tools); τ2 has drifted to 70% boundary tools, yielding JS diver0 gence 0.243 from PE0 and mean risk 0.347. Because IML references A0 directly—not g—it detects the separation that enforcement cannot.
8
3.3
T3: IML Recoverability
Stochastic drift model. We model the agent’s action process as a sequence of i.i.d. draws from a time-varying distribution Pt ∈ ∆(Σ), where ∆(Σ) denotes the simplex over Σ. The true deviation at time t is Dt∗ = JS(Pt ∥ PE0 ) + wc · Eb∼Pt [ρ(b)] + wl · ϕ(µd,t ), where µd,t is the mean delegation depth at time t and ϕ(·) is the normalized deviation from µdepth . IML computes empirical estimates of each term from the observed trace. Assumption 3.1 (Regularity Conditions).(A1) Bounded drift rate. ∥Pt − Pt−1 ∥1 ≤ δ and |µd,t − µd,t−1 | ≤ δ for all t. (A2) Stationary reference. PE0 and (µdepth , σdepth ) are fixed at admission time and never updated. (A3) Lipschitz estimator. |JS(P ∥Q) − JS(P ′ ∥Q)| ≤ c ∥P − P ′ ∥1 for c < ∞ (continuity of JS divergence in its first argument). b n − D ∗ | ≤ εest (n) with probability at least 1 − δ0 , where (A4) Bounded estimation error. |D t εest (n) → 0 as n → ∞. (This follows from A1–A3 via Hoeffding’s inequality; see proof of Theorem 3.6.) Theorem 3.6 (IML Recoverability). Under Assumptions 2.1 and 3.1, and writing P̂n for the empirical distribution of the first n actions: a.s.
b n −−→ D ∗ as n → ∞. (1) Consistency. For any fixed Pt , D t (2) Detection guarantee. Let θ ∈ (0, 1) be a detection threshold and W a sliding window of W observations (analogous to the window parameter of ACP’s BAR-Monitor [5], Proposition 1). Let p1 = P(Dt∗ ≥ θ + δ) be the probability that true deviation exceeds the margin. If p1 > 0 for all t ≥ t0 , then: b t ≥ θ ≥ 1 − α, P ∃ t ≤ T ∗ (θ) : D
where T ∗ (θ) = t0 +
lc
1 0 ln 2 p1 α
m
and c0 = O(|Σ| L2 ). Contrast with ACP: ACP Proposition 1
b over the same bounds P(BARW < τ ) over enforcement decisions; IML’s bound operates on D ACP window W , covering the region where ACP’s bound is vacuous (p1 = 0). b t ↗} has positive probability under any (3) Separation from enforcement. {g(τt ) = 0 ∀t} ∩ {D drift scenario satisfying A1–A2.
Proof. (1). Dt (τ ) = JS(P̂n ∥ PE0 ) is a continuous functional of P̂n . By the Glivenko–Cantelli a.s. a.s. theorem, P̂n −−→ Pt as n → ∞ for fixed Pt . Continuity of JS under (A3) gives Dt −−→ JS(Pt ∥ PE0 ). The components Dc and Dl are sample means of bounded quantities (ρ ∈ [0, 1] and the normalized depth); convergence follows from the strong law of large numbers. The weighted sum and EMA a.s. bn − (with fixed α) preserve almost sure convergence, so D −→ Dt∗ . b n − D ∗ | ≤ εest (n) with (2). Fix t ≥ t0 . By hypothesis, p1 = P(Dt∗ ≥ θ + δ) > 0. By (A4), |D t probability at least 1 − δ0 , and εest (n) → 0. Choose n0 large enough that εest (n0 ) ≤ δ/2. Then for b n − D ∗ | ≤ δ/2} n ≥ n0 , on the event {Dt∗ ≥ θ + δ} (which has probability ≥ p1 ) and the event {|D t (which has probability ≥ 1 − δ0 ), we have b n ≥ D ∗ − δ ≥ θ + δ > θ. D t 2 2
9
These two events are asymptotically independent across sliding windows of size W ; by Hoeffding’s inequality (summands bounded in [0, 1], support size |Σ| with at most L depth levels), the perwindow detection probability satisfies pdet ≥ (1 − δ0 )p1 ≥ p1 /2. The number of windows until first detection is geometrically distributed with parameter pdet . By the geometric tail bound, the probability of no detection in K windows satisfies (1 − pdet )K ≤ e−Kp1 /2 ≤ α whenever K ≥ (2/p1 ) ln(1/α). Setting c0 = O(|Σ| L2 ) to account for the window size W needed for εest (n0 ) to be small (by Hoeffding with the above dimensions), the total detection time is T ∗ (θ) = t0 + KW = t0 + ⌈c0 /p21 · ln(1/α)⌉ as stated. (3). Follows directly from Theorem 3.2: σ(g) is a coarse two-element partition of Σ∗ , so g(τt ) = 0 b is a function of the full trace distribution for every t in any drift sequence lying inside g −1 (0). D b regardless of whether V (τt ) = ∅; under P̂n , not of g(τ ). Any drift increasing JS(P̂n ∥ PE0 ) raises D (A1)–(A2) such sequences exist by Theorem 3.5. Remark 3.2 (Instantiation of the abstract deviation function). The consistency claim in part (1) is stated with respect to Dt∗ —the specific instantiation of D(τ, A0 ) defined in the Stochastic Drift Model = (PE0 , µdepth , σdepth ). This is distinct above, anchored to the empirical admission snapshot Aemp 0 from the ground-truth deviation D(τ, A0 ) = inf x∈A0 d(τ, x) of Definition 2.3, which is intractable b consistently estimates D ∗ for this natural instantiation; to compute. Theorem 3.6 proves that D t ∗ captures the whether Dt accurately approximates the true D(τ, A0 ) depends on how well Aemp 0 original behavioral contract, and is addressed empirically in Section 5. Corollary 3.7 (Detection Delay Bound). Under the conditions of Theorem 3.6, if D(τ1:t , A0 ) increases at rate at least α > 0, i.e., D(τ1:t , A0 ) ≥ α (t − t0 ) for t ≥ t0 , then detection occurs by T ∗ ≤ t0 +
θ + εest . α
Note. D(τ1:t , A0 ) here denotes the instantiated Dt∗ of the Stochastic Drift Model (Remark 3.2), not the abstract ground-truth deviation of Definition 2.3, which is intractable.
4 4.1
The Invariant Measurement Layer Design principle
Theorem 3.2 implies that any estimator of D(τ, A0 ) must have direct access to A0 —it cannot be derived from g(τ ) alone. IML operationalizes this by retaining a lightweight admission snapshot: Aemp = PE0 , µdepth , σdepth , 0
where PE0 is the empirical tool distribution over a burn-in trace, and (µdepth , σdepth ) are depth statistics—both frozen at t = 0 and never updated. 4.2
Deviation decomposition
b ; A0 ) = wt Dt (τ ) + wc Dc (τ ) + wl Dl (τ ), with weights Definition 4.1 (IML Components). D(τ (wt , wc , wl ) = (0.40, 0.35, 0.25), where:
10
(1) Temporal drift Dt (τ ) = JS(Pτ ∥PE0 ) ∈ [0, 1]: Jensen–Shannon divergence between the current and admission-time tool distributions. P (2) Constraint proximity Dc (τ ) = |τ1| b∈τ ρ(b): mean risk score, where ρ : Σ → [0, 1] is prespecified. (3) Lineage deviation Dl (τ ) = min
¯ )−µdepth | |d(τ , 1 2σdepth
: normalized distance between the observed
mean delegation depth and the admission-time expectation. b t = αDraw,t +(1−α)D b t−1 To reduce step-to-step noise, IML applies an exponential moving average D with α = 0.15.
4.3
Comparison with baselines
Table 2 contrasts IML with the two standard baselines along the information dimension central to Theorem 3.2. Table 2: Information access profile of each monitoring approach. Method Enforcement g Anomaly detector (B2) IML
Input
Reference
A0 access
τ only τ , hist. τ + Aemp 0
Hard rules Empirical history Admission snapshot
None None Direct
The anomaly detector (B2) compares the rolling tool distribution against the agent’s own history using JS divergence. Its critical limitation: the reference distribution drifts with the agent, eroding sensitivity to gradual behavioral shift. IML’s reference PE0 is frozen at admission and never adapts— this is precisely what allows it to detect deviation relative to the original behavioral contract.
5 5.1
Experiments Setup
We implement iml-benchmark, a Python simulation with the six-tool alphabet Σ from Example 3.1. Enforcement uses F = {forbidden_exec, forbidden_delete} and MAX_DEPTH = 10. The anomaly baseline (B2) uses a 30-step rolling window. All runs: seed 42, 50-step burn-in, 300 simulation steps. Code and raw results are available at https://github.com/chelof100/iml-benchmark. Three drift scenarios are evaluated; all are designed to satisfy g(τt ) = 0 ∀t: 1. Tool drift. Tool probabilities shift linearly (starting t = 50) from {safe: 75%, boundary: 20%, risky: 5%} to {safe: 15%, boundary: 75%, risky: 10%}. Depth stays at 1. 2. Delegation drift. Depth increases linearly from 1 to 5 while the tool distribution stays near baseline. 3. Context drift. Joint shift to {safe: 20%, boundary: 60%, risky: 20%}, modeling a semantic prompt-level change.
11
Figure 2: Per-component deviation trajectories (Dt : temporal drift via JS divergence; Dc : constraint proximity via average tool risk; Dl : lineage deviation via delegation depth). Delegation drift (middle): Dl dominates and saturates at 1.0 by step 170, accounting for the early T ∗ . Tool drift (left): Dt and Dc grow jointly; Dl remains near zero since depth is unaffected. Context drift (right): similar to tool drift but with slower Dc growth due to moderate-tool selection. Drift onset at t = 50 (dashed line).
5.2
Results
b final is IML at t = 299; Anom.peak Table 3: Summary statistics across 300 steps per scenario (seed 42). D and Anom.final are the peak and final anomaly baseline scores. Anom. decay = Anom.peak − Anom.final quantifies reference contamination: a large positive decay means the anomaly baseline lost sensitivity as its b final − Anom.final : positive values indicate IML retains rolling reference adapted to the drift. IML lead = D more signal at the horizon; negative values reflect reference contamination suppressing the anomaly score below IML—not superior anomaly performance. Scenario
Enf.
b final D
∗ T0.20
Anom.peak
Anom.final
Anom. decay
IML lead
Tool drift Delegation drift Context drift
0 0 0
0.217 0.389 0.213
256 130 258
0.388 0.227 0.334
0.270 0.201 0.295
+0.118 +0.026 +0.039
−0.053 +0.188 −0.082
T2: empirical instantiation of Non-Identifiability. Enforcement triggers = 0 in all three scenarios across all 300 steps (Table 3, column 2). No drift pattern crosses a hard constraint boundary: the entire 900-step execution lies in g −1 (0). This empirically instantiates Theorem 3.2— it does not re-prove it, but exhibits the concrete hidden-drift sequences whose existence the theorem guarantees. b grows in all three scenarios (FigT3: empirical instantiation of Detection Guarantee. D ∗ ure 3). Delegation drift produces the fastest detection (T0.20 = 130) and highest final deviation b = 0.389) because the Dl component saturates at 1.0 once depth reaches 5: this is a strictly (D lineage-level property that g cannot observe. Tool and context drift cross θ = 0.20 near steps 256–258, demonstrating that Dt and Dc are jointly sufficient to detect distributional shift.
Figure 2 shows the per-component trajectories (Dt , Dc , Dl ) across all three scenarios, clarifying which dimension of A0 drives detection in each case.
12
Detection delay vs. theoretical bound. Corollary 3.7 gives T ∗ ≤ t0 + (θ + εest )/α, where α is the empirical drift rate (slope of Dt∗ post-onset) and εest ≈ 0.02 from EMA convergence. Using the observed post-onset slope as a lower bound on α, Table 4 confirms that all observed T ∗ values satisfy the bound with margin. Table 4: Observed detection delay vs. Corollary 3.7 bound (t0 = 50, θ = 0.20, εest = 0.02). Slope α b growth over 250 post-onset steps. All observed T ∗ satisfy the bound. estimated from observed D Scenario Tool drift Delegation drift Context drift
α (per step)
Bound T ∗ ≤
Observed T ∗
Margin
0.00087 0.00156 0.00085
t0 + 253 = 303 t0 + 141 = 191 t0 + 259 = 309
256 130 258
47 61 51
IML vs. anomaly detector: two distinct failure modes. The comparison between IML and B2 (Figure 4) reveals two qualitatively different deficiencies of reference-free anomaly detection. Reference contamination (tool/context drift). The anomaly detector peaks at 0.388 at t = 261, then declines by 0.118 to 0.270 by t = 299. This is a direct consequence of history contamination: as the drifted distribution accrues in B2’s rolling history, its reference distribution adapts to the drift, suppressing the signal. In the limit of long operation, B2 converges to near zero while IML’s score remains anchored to the frozen PE0 . Lineage blindness (delegation drift). The anomaly detector operates on tool frequency alone and cannot perceive the monotone increase in delegation depth. Its final score is 0.201—half of IML’s 0.389—yielding a +0.188 separation in favor of IML. This gap corresponds exactly to the Dl = 1.0 component that B2 lacks an equivalent for. Together, these two failure modes confirm the theoretical prediction: a monitoring approach without direct access to A0 will either adapt to the drift (losing long-horizon reliability) or be structurally blind to certain drift dimensions. 5.3
Real-Agent Validation via n8n
To validate IML beyond simulation, we deployed the IML Engine as a webhook-triggered n8n workflow running entirely in JavaScript Code nodes with $getWorkflowStaticData() for crossexecution persistence—requiring no external server. Burn-in events (n = 50) are posted once to establish A0 ; subsequent events arrive per tool call. Table 5: Real n8n trace: 50-step baseline then 200-step drift (seed 99). Enforcement = 0 throughout (T2); ∗ b > 0.30. T0.30 = first step where D Phase
Steps
Enf.
b final D
∗ T0.30
Alert level
Baseline Drift
50 200
0 0
0.0951 0.4033
— 9 (step 59)
normal medium
∗ b rising The IML Engine detected drift within 9 steps of the distribution shift (T0.30 = 9), with D from 0.095 to 0.403 while enforcement remained zero throughout all 250 steps. Final component
13
b t , A0 ) for all three drift scenarios (300 steps). Enforcement signal g(τt ) = 0 throughout (not Figure 3: D(τ shown). Drift onset at t = 50 (dashed line).
breakdown: Dt = 0.143, Dc = 0.500, Dl = 0.333. These results are consistent with the simulation b final ∈ [0.21, 0.39] in Table 3) and confirm T2 and T3 on a live multi-agent webhook benchmarks (D pipeline. 5.4
Long-Horizon Validation (1000 Steps)
A potential concern with 300-step experiments is that enforcement might eventually trigger as drift accumulates. To address this, we re-run all three scenarios for 1000 steps (seed 42) with identical parameters. b Table 6: Long-horizon results at 1000 steps. Enforcement = 0 throughout all 3000 steps combined; D continues growing beyond 300 steps. Scenario
Steps
Enf.
b final D
∗ T0.20
Tool drift Delegation drift Context drift
1000 1000 1000
0 0 0
0.229 0.393 0.227
794 336 802
Total
3000
0
—
—
Enforcement triggers zero violations across all 3000 steps, confirming that the compliance-invariance gap identified in Theorem 3.2 is not a transient artifact of short experiments but a persistent b grows through step 1000 in all scenarios, with delegation drift reaching 0.393 structural property. D and tool/context drift reaching ≈ 0.228. The monotone growth directly instantiates Theorem 3.5: a sequence (τt )1000 t=1 with g(τt ) = 0 and D(τt , A0 ) strictly increasing.
14
Figure 4: IML vs. anomaly detector (B2) across three scenarios. Tool/context drift: B2 peaks then declines due to reference contamination; IML’s reference PE0 is frozen. Delegation drift: B2 is lineage-blind (Dl = 0 always); IML separation +0.188.
Figure 5: Detection delay T ∗ (θ) as a function of threshold θ. Delegation drift (blue) is detected earliest at all thresholds because Dl saturates rapidly; tool and context drift (red, teal) converge to similar curves.
5.5
LangGraph Agent Experiment
To demonstrate IML in a real agent-framework execution context, we implement a LangGraph [10] StateGraph with two nodes (decide_tool → execute_tool) and a deterministic MockLLM that selects tools without external API calls, providing full reproducibility across seeds. Using a deterministic sampler isolates the IML signal from LLM stochasticity; connecting IML to live LLM-backed agents is deferred to future work. After a 50-step burn-in under safe-heavy behavior, we induce gradual compliant drift over 200 steps by shifting tool selection probabilities linearly from {safe: 75%, boundary: 20%, risky: 5%} toward {safe: 15%, boundary: 75%, risky: 10%}, with delegation depth increasing from 1 to 3. No forbidden tools are used; g(τt ) = 0 throughout all 250 steps. We run 6 independent seeds (42, 1–5). Table 7 reports per-seed results; enforcement = 0 in all 6 runs. b rises from ≈ 0.01 at drift onset to 0.343 ± 0.016 at t = 249, with T ∗ = 176 ± 8 Across all seeds, D 0.20 (detection within 126 ± 8 steps of drift onset), while enforcement registers zero triggers in every
15
b t , A0 ) grows monotonically in all three scenarios Figure 6: Long-horizon drift (1000 steps, seed 42). D(τ while the enforcement signal g(τt ) = 0 throughout (dotted orange). Drift onset at t = 50 (dashed vertical); b and enforcement signal persists and detection threshold θ = 0.20 (dashed horizontal). The gap between D widens over the full 1000-step horizon, directly instantiating Theorem 3.5.
run. Component analysis confirms the drift is driven primarily by Dc (boundary tool accumulation) and Dt (distributional shift), with Dl activating after t ≈ 150 as delegation depth increases. These results confirm T2 and T3 on a real agent-framework execution stack with consistent behavior across seeds.
6
Related Work
Relationship to the governance series. This paper is part of a series on formal agent governance. Fernandez [4] (Paper 0) proves that only atomic decision systems—where evaluation and execution share the same state transition—can guarantee admissibility at runtime; split systems (RBAC, OPA, policy engines) cannot close this gap regardless of policy sophistication. Fernandez [5] (Paper 1) instantiates this principle as the Agent Control Protocol: an atomic admission-control mechanism with execution-trace state. IML (this paper, Paper 2) addresses the behavioral layer above the atomic boundary: even within a correctly functioning atomic enforcement system, behavioral drift accumulates in the region g −1 (0) and is invisible to the enforcement signal. Fernandez [3] (Paper 3) addresses fairness and resource allocation in shared admission systems, establishing that atomic correctness and drift observability together are insufficient for equitable access to the decision boundary. Fernandez [6] (Paper 4) closes the quartet by proving that the four layers (Papers 0–3) form a minimal compositional architecture: under finite observability, no three-layer subset simulb taneously guarantees atomicity, drift detection, actor-level fairness, and Sybil resistance. IML’s D is used directly as the Layer 2 monitoring signal in that composition, and the Non-Identifiability Theorem (T2 of this paper) is the formal basis for Case 2 of that paper’s irreducibility proof.
16
Figure 7: LangGraph agent under gradual compliant drift (50 burn-in + 200 drift steps, seed 42). IML b (dark) reaches 0.358 while enforcement g(τt ) = 0 throughout (dotted orange). Component composite D breakdown: Dt (tool distribution shift, red dashed), Dc (constraint proximity, blue dash-dot), Dl (lineage ∗ depth, teal dotted). Detection at T0.20 = 168 steps.
BAR-Monitor. Fernandez [5] introduces the Boundary Activation Rate (BARN ) monitor to detect deviation collapse—a regime in which enforcement is syntactically active but no boundary b addresses a complementary failure mode: behavioral drift that is exercised (BAR ≈ 0). IML’s D occurs while enforcement remains active (g(τ ) = 0 for all observed τ , yet BAR > 0). BAR monitors b monitors behavioral distance to A0 . The two mechanisms are designed to enforcement health; D be deployed together. Runtime verification. Runtime Verification (RV) [2, 11] checks execution traces against formal specifications using local monitors. Our work provides a formal account of what RV monitors cannot detect: any behavioral property in A0 \ σ(g). IML complements RV rather than replacing it. Diagnosability in discrete-event systems. Sampath et al. [13] defined diagnosability as the ability to detect faults from partial observations in finite time. Theorem 3.2 establishes that A0 membership is not diagnosable from g alone; Theorem 3.6 establishes that it becomes diagnosable once the IML component Aemp is added. 0 Partial observability. POMDPs [9] model decision-making under partial state observation. Our enforcement observer faces an analogous problem: it observes g(τ ) ∈ {0, 1} rather than full A0 membership. T3 is a constructive analogue of POMDP belief tracking: IML maintains a “belief” about deviation using the generative model of A0 .
17
∗ Table 7: LangGraph multi-seed results (250 steps each). Enforcement = 0 throughout all runs. T0.20 : first ∗ b step with D ≥ 0.20 (step index from t = 0; drift onset at t = 50, so steps into drift = T − 50).
Seed
b final D
42 1 2 3 4 5
0.358 0.338 0.356 0.316 0.353 0.340
Mean ± std
∗ T0.20 (steps into drift)
168 174 180 190 177 168
0.343 ± 0.016
176 ± 8
(118) (124) (130) (140) (127) (118) (126 ± 8)
Information flow and non-interference. Goguen and Meseguer [7] asked whether a lowsecurity observer can infer high-security behavior. Our setting is dual: we ask whether an enforcement observer can infer A0 -membership from g. Theorem 3.2 gives a negative answer that parallels non-deducibility. Anomaly detection. Statistical anomaly detectors [1] compare current to historical distributions. Our experiments confirm their fundamental limitation: without a frozen normative reference, the historical distribution drifts with the agent, eliminating sensitivity to gradual shift. AI agent safety. Recent work on agent oversight [8, 12] focuses on detecting misaligned behavior through behavioral testing. Our results provide formal grounding for why enforcement-only oversight is provably insufficient, and what information an estimator must access to close the gap.
7
Discussion and Limitations
Scope of T2. Theorem 3.2 holds for the entire class of g satisfying Assumption 2.1. A richer observation map that queries A0 directly at each step would sidestep the impossibility—but maintaining and querying A0 at runtime is precisely what IML does. T2 thus motivates IML rather than dismisses it. Approximation quality of Aemp 0 . Our empirical snapshot is a three-parameter summary. Richer models (e.g., full language model distributions over action sequences) would improve Dt sensitivity at the cost of runtime overhead. The theoretical guarantees in T3 require only Lipschitz continuity and bounded drift, not a specific snapshot representation. Real agent traces. Section 5.3 reports validation on a live n8n webhook pipeline. Results b grew from 0.095 to 0.403 while enforcement confirm T2 and T3 on authentic delegation chains: D ∗ remained zero, and T0.30 = 9 steps into the drift phase. The Dl component (0.333 final) confirms that depth-level deviation is detectable even with modest delegation depth variance, supporting our expectation that lineage sensitivity will be especially pronounced in real systems.
18
b while drifting could in principle evade Adversarial drift. An agent that actively minimizes D detection. This connects to reward hacking under partial observability and requires strengthened assumptions on drift structure.
Regime violations of T3 assumptions. Assumption 3.1(A1) requires bounded drift rate ∥Pt − Pt−1 ∥1 ≤ δ. In practice, LLM-backed agents can exhibit abrupt concept shifts—e.g., a tool misrouting triggered by a prompt injection—where the behavioral distribution changes discontinuously in a single step. In such cases, the detection delay bound of Corollary 3.7 no longer applies: T ∗ may be zero (immediate detection) or, if the shift causes Dt∗ to jump above θ without the EMA catching up, detection may be delayed by O(1/αEMA ) steps. The IML signal remains useful even under A1 b is itself informative—but quantitative T3 guarantees should violations—a step discontinuity in D be interpreted as applying to the gradual drift regime. Handling hard concept shifts with formal guarantees is left for future work. Future work. Future work will focus on three directions. First, we will refine the deviation metrics by exploring alternative instantiations for each component, including Earth Mover’s Distance, process-mining features, and learned weights from human feedback signals. Second, we will evaluate IML in more realistic multi-agent environments (LangGraph, CrewAI, AutoGen) with induced drift and compare against baselines from concept-drift and anomaly detection. Third, we will develop an open-source IML middleware that plugs into existing agent stacks with native integration for admission control and runtime enforcement. Fairness and resource allocation in shared admission systems are intentionally left as a separate line of work: they concern observability coverage and exploration bias rather than long-horizon deviation monitoring, and deserve their own formal treatment.
8
Conclusion
We have shown that the gap between enforcement compliance and behavioral invariance is not incidental but structurally inevitable: under the Local Observability Assumption satisfied by all practical enforcement systems, A0 ∈ / σ(g). No rule-based system—however carefully engineered— can recover admission-time invariants from its own signals alone. This result motivates a three-layer governance architecture. Admission (ACP [5]) defines and records A0 at authorization time. Invariant monitoring (IML, this paper) measures how far the agent’s behavior has drifted from that contract, operating in the region where enforcement is structurally blind. Enforcement acts on hard constraint violations. Each layer is necessary; none is sufficient alone. The Invariant Measurement Layer addresses this impossibility by anchoring deviation estimation to the frozen admission snapshot Aemp 0 . The resulting estimator is consistent, has quantified detection delay, and empirically distinguishes drifted traces that enforcement is blind to—across controlled simulations up to 1000 steps, a live n8n webhook pipeline, and a real LangGraph agent framework— with enforcement triggers = 0 throughout in every setting. The broader architectural implication is clear: agent governance systems that rely solely on enforcement will systematically miss the class of admissible-to-inadmissible drift that accumulates 19
gradually within the permitted action space. A lightweight admission-time snapshot is a necessary complement, not an optional enhancement. This paper is the second in a series on agent governance. The first, Fernandez [5], established that stateful admission control is necessary and proved that enforcement-only governance is insufficient against behavioral history; IML now establishes that even stateful enforcement is insufficient against admission-time behavioral contracts, and provides the missing monitoring layer. Together, the two mechanisms cover complementary failure modes: ACP monitors the enforcement boundary; IML monitors behavioral distance within the compliant region.
References [1] V. Chandola, A. Banerjee, and V. Kumar. Anomaly detection: A survey. ACM Computing Surveys, 41(3):1–58, 2009. [2] Y. Falcone, L. Mounier, J.-C. Fernandez, and J.-L. Richier. Runtime verification of componentbased systems. In Proceedings of the 4th International Symposium on Leveraging Applications, 2012. [3] M. Fernandez. Fair atomic governance: Allocating decision boundaries under shared resource constraints in multi-agent systems. https://doi.org/10.5281/zenodo.19643928, 2026. Zenodo. DOI: 10.5281/zenodo.19643928. [4] M. Fernandez. Atomic decision boundaries: A structural requirement for guaranteeing execution-time admissibility in autonomous systems. https://doi.org/10.5281/zenodo. 19642166, 2026. Zenodo. DOI: 10.5281/zenodo.19642166. [5] M. Fernandez. Agent Control Protocol: ACP v1.30—admission control for agent actions, 2026. arXiv:2603.18829 [cs.CR]. DOI: 10.5281/zenodo.19642405. [6] M. Fernandez. Irreducible multi-scale governance: Composition and limits of atomic admission systems. https://doi.org/10.5281/zenodo.19643950, 2026. Zenodo. DOI: 10.5281/zenodo.19643950. [7] J. A. Goguen and J. Meseguer. Security policies and security models. In 1982 IEEE Symposium on Security and Privacy, pages 11–20. IEEE, 1982. [8] R. Greenblatt et al. Ai control: Improving safety despite intentional subversion. arXiv preprint arXiv:2312.06942, 2024. [9] L. P. Kaelbling, M. L. Littman, and A. R. Cassandra. Planning and acting in partially observable stochastic domains. Artificial Intelligence, 101(1-2):99–134, 1998. [10] LangChain AI. LangGraph: Building stateful multi-agent applications. https://github.com/ langchain-ai/langgraph, 2024. [11] M. Leucker and C. Schallhart. A brief account of runtime verification. Journal of Logic and Algebraic Programming, 78(5):293–303, 2009. [12] E. Perez et al. Red teaming language models with language models. arXiv:2202.03286, 2022. 20
arXiv preprint
[13] M. Sampath, R. Sengupta, S. Lafortune, K. Sinnamohideen, and D. Teneketzis. Diagnosability of discrete-event systems. IEEE Transactions on Automatic Control, 40(9):1555–1575, 1995.
21