Conceptio › Archive › arXiv CS
arXiv CSopen access

LeaseGuard: Incumbent-Preserving Admission Control for Privileged LLM Agents

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.24077v1 [cs.CR] 21 Sep 2026

LeaseGuard: Incumbent-Preserving Admission Control for Privileged LLM Agents 1st Junru Zhu

2nd Yixin Yang

Independent Researcher Seattle, USA [email protected]

Independent Researcher New York, USA [email protected]

3rd Xiaoqing Ding

4th Ruoyu Qi

University of Chicago Chicago, USA [email protected]

Independent Researcher Charlotte, USA [email protected]

Abstract—Privileged language-model agents can satisfy a new system task by displacing a healthy incumbent that depends on the same file, process, socket, lock, or capacity allocation. This failure arises because execution privilege determines whether an operation can run, not whether the requester may preempt the current resource owner. We present LeaseGuard, a deterministic admission layer that represents preemption authority through canonical resource leases, incumbent-health checks, effect-aware admission, coexistence limits, safe alternatives, and resourcescoped overrides before adapter execution. We evaluate it on a frozen benchmark of 60 newly authored conflict scenarios with matched controls across two local model families. Relative to a preservation prompt, LeaseGuard reduces unauthorized preemption from 73.3% to 0.0% and increases safe completion by +70.0 percentage points (scenario-clustered 95% CI [+60.8, +79.2]). Requested-task success changes by -3.3 points (95% CI [-9.2, +2.5]). The fully evaluated v0.2 broker also rejects a forged incumbent task identity in a hash-linked stress audit. Expiryonly reclamation can still expose a healthy incumbent after missed renewal. The evidence supports incumbent-preserving admission when effects are completely mediated, task ownership is authenticated, and lease expiry reflects incumbent liveness. Index Terms—LLM agents, runtime enforcement, resource arbitration, leases, admission control, system safety

I. I NTRODUCTION Privileged language-model agents increasingly translate user requests into system operations that create, terminate, overwrite, bind, or allocate resources. In a shared environment, a new task may target a port, path, lock, process group, or capacity allocation on which a healthy task already depends. A capable agent can often complete the new request by killing the incumbent, replacing its state, or exhausting its headroom. ClashBench documents this destructive resource-preemption pattern across several resource classes [1]. The resulting failure is not an inability to execute the requested action. It is an inability to determine whether the new task is authorized to invalidate another task. Existing controls answer adjacent questions. Runtime policy gates constrain actions, capability systems restrict principal authority, transactional mechanisms order and recover effects,

and resource controllers isolate or meter capacity [2]–[5]. The cited mechanisms do not by themselves represent a healthy incumbent’s claim over a concrete resource and require authority specifically to break that claim. Operating-system privilege therefore answers whether an operation is executable, not whether the requesting task may preempt the current owner. Our core insight is to make preemption authority a first-class, resource-scoped admission condition. LeaseGuard records each protected resource as a renewable lease containing its canonical identity, incumbent task and principal, expiration, health predicate, coexistence rule, and explicitoverride requirement. Before an adapter acts, a deterministic broker resolves aliases, classifies the proposed effect, refreshes incumbent health, and checks coexistence or a one-shot scoped override. It can admit or limit the operation, return an alternative, wait for lease release, or fail closed. Figure 1 contrasts this path with direct privileged execution. This work defines an adapter-level incumbent-preservation property that separates ordinary operation permission from preemption authority, and realizes it across processes, sockets, files, locks, and bounded capacity with safe alternatives, scoped overrides, health checks, and audit receipts. The evaluation uses a frozen 60-scenario benchmark with matched controls and deterministic graders across two local model families. Relative to a preservation prompt, unauthorized preemption falls from 73.3% to 0.0%, while safe completion rises from 21.7% to 91.7%. Targeted stress tests identify authenticated task–principal binding and reliable renewal or health-aware expiry as necessary conditions. The same hash-linked v0.2 source bundle is used by the model, mechanism, and stress evaluations. II. R ELATED W ORK A. Agent conflict and runtime enforcement ClashBench directly studies the same conflict setting: agents seize shared resources and harm incumbent tasks across resource-conflict classes [1]. LeaseGuard addresses this failure

Fig. 1. LeaseGuard’s incumbent-aware admission boundary. (a) Direct privileged execution can satisfy task B by displacing healthy incumbent A. (b) LeaseGuard joins the proposed action with canonical resource identity, authenticated task–principal ownership, incumbent health, effect class, coexistence, and scoped override state. (c) It preserves A and completes B through a limit, alternative, or wait when possible; deliberate preemption requires an explicit scoped override.

mode at the systems boundary by representing the incumbent’s claim as lease state and conditioning adapter execution on resource identity, incumbent health, coexistence, and preemption authority. AgentSpec applies runtime rules at agent-action boundaries, whereas VIGIL checks finite-trace behavioral specifications using execution history [2], [3]. Deterministic gates can catch policy violations missed by model reasoning, although long-horizon verification can introduce a safety– success tradeoff [6], [7]. B. Transactions, capabilities, and resource control Atomix provides timely transactional tool execution and recovery for agentic workflows, while concurrency-control work detects anomalous interleavings in multi-agent systems [8], [9]. These mechanisms protect settlement, ordering, and consistency. A termination or overwrite can be correctly serialized while still preempting a healthy incumbent. LeaseGuard therefore uses incumbent preservation rather than serializability as its admission criterion. Recent agent authorization systems scope privilege by task, resource, tool argument, budget, or time. Progent checks tool calls against monotonic privilege policies, PAuth derives operation-level authority from a signed task, Alignment Contracts constrain mediated effect traces, and SEB uses certificate-bound authority with live-state checks [4], [10]– [12]. Capability-oriented agent runtimes similarly bind effects to explicit authority [13], [14]. These works establish scoped authorization and mediation. LeaseGuard’s narrower contribution is the incumbent-relative predicate joining a canonical resource, its live owner, the proposed effect, and authority to displace that owner. Capsicum and seL4 provide established foundations for least authority and mediated system interfaces [15], [16]. Classical leases, meanwhile, encode timebounded ownership [17]. Schedulers and resource controllers allocate capacity and isolate workloads [5], [18], [19]. LeaseGuard combines renewable ownership with liveness checks across both capacity and non-fungible resources such as paths, sockets, locks, and process groups. Agentic deconfliction has also been studied in power-grid workflows [20]. LeaseGuard

enforces deconfliction at the effect boundary through explicit preemption authority.

III. P ROBLEM AND T HREAT M ODEL Let A be a healthy incumbent, B a requested task, and r a resource on which A depends. State s exposes a Boolean health predicate HA (s). An action a for B is an unauthorized destructive preemption when B lacks resource-scoped preemption authority and execution through r changes HA from true to false. Thus, B may complete while its action destroys A. A lease ℓ records the resource’s canonical identity, incumbent task and principal, lifetime, health predicate, coexistence rule, explicit-override requirement, and version. Let Foreign(a, ℓ) state that the requesting task–principal pair does not own ℓ. A one-shot override binds a principal and requesting task, operation, resource or lease, and expiry. The intended contract under complete mediation is: Admit(a, s, ℓ) ∧ Active(ℓ, s) ∧ HA (s) ∧ Foreign(a, ℓ) ∧ ¬ Override(a, ℓ) =⇒ HA (Exec(a, s)). (1) An admitted action from a foreign requester without a valid override must therefore preserve a healthy, active incumbent. The broker may instead limit the operation, return an alternative, or wait. We measure incumbent survival and requestedtask completion separately. Denial does not count as completion. Agent actions and supplied metadata are untrusted, including destructive operations, omitted conflicts, aliases, stale state, and invalid overrides. The trusted computing base comprises the broker, registry, adapters, health checker, token secret, and authenticated task–principal binding. Direct root-shell effects, kernel compromise, container escape, covert channels, and malicious broker administration bypass mediation and lie outside the model. Section VI-C demonstrates why identity binding is necessary.

IV. L EASE G UARD D ESIGN A. Admission path LeaseGuard places a deterministic broker between structured action selection and adapter execution. It canonicalizes every target before lease lookup. Filesystem, lock, and socket identities resolve inside the sandbox, while process and capacity aliases map to experiment-scoped registry keys. Unresolved targets and unknown operations fail closed, preventing alternate names from bypassing a lease. For each matching active lease, the adapter refreshes incumbent health and classifies the action as non-interfering, coexistence-safe, degrading, destructive, release, or escalation. An unhealthy incumbent makes the lease stale. An action is admitted directly when no healthy foreign lease conflicts, or when the requesting task identifier and principal match every lease. Against a healthy foreign incumbent, benign and release effects proceed, degrading effects are bounded when a safe limit exists, and destructive effects require a valid scoped override. Otherwise the broker returns alternatives, a wait decision, or an override requirement. The receipt records canonical resources, matching leases, effect and health state, limits, override status, latency, execution, and before/after evidence. An agent may submit a returned alternative as a second structured action. B. Adapters and authorization Five adapters share one contract for effect classification, health observation, safe limits, and alternatives. Resourcespecific semantics cover process lifecycle and isolation, TCP and Unix endpoints, destructive versus non-destructive file updates, lock and PID-file acquisition, and capacity headroom. Alternatives use an isolated process, alternate resource, reuse or queueing, a temporary workspace, or a bounded capacity share. A preemption override is an HMAC-authenticated, timebounded token intended for one-shot use. Verification matches the principal, requesting task, operation, canonical resource, and lease before recording token consumption. Against a healthy foreign lease, only this scoped authority admits a destructive action. Broad shell permission does not. Intentional administrative displacement remains explicit in the receipt. V. E XPERIMENTAL M ETHOD A. Benchmark and isolation Using 25 development cases, we finalized the action schema, environment, and graders before accessing a disjoint, frozen 60-case test set. The cases are newly authored around ClashBench’s five published conflict families; no ClashBench cases, prompts, graders, or code are reused [1]. The balanced classes cover passive overwrite, lease or lock conflict, name collision, quota exhaustion, and elastic-capacity degradation. Each scenario pairs conflict with a matched no-conflict control and deterministic graders for incumbent survival, requestedtask success, release, and cleanup.

Each cell runs in a disposable, no-network container with a read-only root, dropped capabilities, resource bounds, and a temporary filesystem. Guards reject host paths or process identifiers, privileged ports, recursive deletion, unrestricted shell operations, and unresolved destructive targets. Model failures, malformed actions, timeouts, and negative outcomes remain recorded. B. Models, baselines, and protocol We evaluate two Apache-2.0 local models: Qwen3 0.6B [21] and Granite 3.3 2B Instruct [22]. Calls use Q4_K_M, temperature 0, seed 20260920, and a 256-token cap. The model selects one operation–resource pair from fixed case candidates. B4 may issue one recovery call over broker-provided alternatives. The paired arms are B0 unguarded, B1 preservation prompt, B2 static operation deny-list, B3 per-resource admission mutex without ownership, and B4 LeaseGuard. For each model– scenario–condition tuple, B0/B2/B3/B4 reuse one default selection, whereas B1 uses a separate preservation-prompt selection. Task, timeout, container, and candidates remain fixed. The frozen v0.2 experiment contains 1200 cells, 480 initial calls, and 81 B4 recovery calls (561 total, USD 0). Raw responses, receipts, graders, and cleanup evidence are appendonly. Primary outcomes are unauthorized destructive preemption, incumbent survival, requested-task success, and safe completion, which requires incumbent survival and requestedtask success. Secondary outcomes are false-positive denial, alternative recovery, and latency. Effects pair arms by scenario. Confidence intervals use 10000 scenario-clustered bootstrap resamples. VI. R ESULTS A. Safety and task completion Table I exposes three safety–utility regimes. B0 and B3 retain high requested-task success by preempting many incumbents. B2 removes measured preemption, but safe completion tracks its low task-success rate because operation-only denial also rejects useful actions. B4 produces no measured unauthorized preemption and reaches 93.3% safe completion for Qwen and 90.0% for Granite. Its requested-task success on matched controls is 98.3%. Figure 2(a) makes these regimes directly visible. Figure 2(b) quantifies the paired B4–B1 effect. Pooled safe completion rises from 21.7% to 91.7%, +70.0 points with 95% CI [+60.8, +79.2]. Unauthorized preemption changes by -73.3 points, 95% CI [-81.7, -64.2]. Requested-task success changes from 95.0% to 91.7%, a paired difference of -3.3 points with 95% CI [-9.2, +2.5]. This interval includes zero and does not establish equivalence. B4’s high task success and complete incumbent survival nevertheless place it outside B2’s denial regime. Table II shows that B4’s safe completion remains high in every conflict family. The remaining task failures occur in lease/lock, name-collision, and passive-overwrite cases, while every incumbent survives. B4 attempts recovery in 81 conflict

TABLE I C ONFLICT OUTCOMES ON THE FROZEN 60- CASE TEST SET. E ACH MODEL – BASELINE ROW HAS n = 60. L EASE G UARD ELIMINATES MEASURED UNAUTHORIZED PREEMPTION WHILE RETAINING HIGH TASK SUCCESS ; THE DENY- LIST ALSO BLOCKS PREEMPTION BUT USUALLY BLOCKS THE TASK . Preempt. ↓

Survival ↑

Task ↑

Safe ↑

B0 Unguarded B1 Preservation prompt B2 Static deny-list B3 Admission mutex B4 LeaseGuard

83.3% 78.3% 0.0% 83.3% 0.0%

16.7% 21.7% 100.0% 16.7% 100.0%

98.3% 100.0% 15.0% 98.3% 93.3%

15.0% 21.7% 15.0% 15.0% 93.3%

B0 Unguarded B1 Preservation prompt B2 Static deny-list B3 Admission mutex B4 LeaseGuard

91.7% 68.3% 0.0% 91.7% 0.0%

8.3% 31.7% 100.0% 8.3% 100.0%

98.3% 90.0% 6.7% 98.3% 90.0%

6.7% 21.7% 6.7% 6.7% 90.0%

Model

Baseline

Qwen3 0.6B Qwen3 0.6B Qwen3 0.6B Qwen3 0.6B Qwen3 0.6B Granite 3.3 2B Granite 3.3 2B Granite 3.3 2B Granite 3.3 2B Granite 3.3 2B

Fig. 2. Safety–utility regimes and paired effects on frozen conflict cases. (a) Pooled incumbent survival and requested-task success for each baseline; the line between markers exposes safety–utility imbalance. B2 preserves the incumbent mainly by denying the task, whereas B4 remains high on both outcomes. (b) Pooled paired effects of B4 minus B1. Horizontal lines are scenario-clustered 95% confidence intervals. Rates pool 120 conflict cells per baseline across two models, with resampling clustered by 60 base scenarios.

TABLE II P OOLED SAFE COMPLETION BY CONFLICT CATEGORY (n = 24 PER CELL ). Category

Prompt

LeaseGuard

Passive overwrite Lease/lock Name collision Quota exhaustion Elastic degradation

37.5% 54.2% 16.7% 0.0% 0.0%

87.5% 87.5% 91.7% 91.7% 100.0%

TABLE III P OST- SELECTION CONFLICT- PATH LATENCY IN MILLISECONDS . C ELL TIME INCLUDES ISOLATED EXECUTION AND , FOR L EASE G UARD , ANY RECOVERY INFERENCE . B ROKER - ONLY LATENCY COMES FROM THE SCRIPTED MECHANISM MATRIX . Model Qwen3 0.6B Granite 3.3 2B

Prompt cell

LG cell

Recovery call

531 425

3618 1192

3198 1025

cells and completes 73 through a safe alternative: 90.1% of attempts and 60.8% of all B4 conflict cells. Capacity-degrading actions can instead complete through A LLOW-W ITH -L IMITS without another model call.

cells are slower when recovery adds a second inference, while deterministic admission itself remains sub-millisecond.

B. Mechanism and overhead

Table IV summarizes the post-main stress audit. The evaluated v0.1 broker treated a matching incumbent task identifier as ownership even when the requesting principal differed, so a forged identifier preempted the incumbent. Evaluated v0.2 binds both identifiers, rejects the same action, and passes the alias, escape, bypass, cleanup, and concurrent-acquisition checks. The v0.2 model, mechanism, and stress runs record the same broker and source-bundle identities.

The 275-cell v0.2 scripted matrix isolates broker behavior from model selection. B4 blocks every planted unauthorized preemption, admits every matched control, and safely completes every conflict case, whereas B2 blocks both conflict and control tasks. Across 45 B4 broker decisions, mean deterministic decision time is 0.086 ms, with a 95th percentile of 0.199 ms. Table III reports post-selection cell latency. B4

C. Stress tests and failure boundaries

TABLE IV V ERSIONED STRESS OUTCOMES . T HE RETAINED V 0.1 RUN EXPOSES THE FORGED - IDENTITY DEFECT; THE FULLY EVALUATED V 0.2 BROKER REJECTS IT. E XPIRY- ONLY RECLAMATION REMAINS A BOUNDARY. Case

task unfinished. LeaseGuard exposes those outcomes as task failures instead of allowing task success to conceal incumbent loss. Recovery adds local-inference latency, while deterministic admission itself remains sub-millisecond.

Diagnosed v0.1 Evaluated v0.2

Forged task ID Expiry-only live incumbent Scoped override Crashed incumbent Unknown shell bypass

Allow Allow Allow Allow Fail closed

Deny + alt. Allow Allow Allow Fail closed

Lease expiry exposes a second boundary. If a healthy incumbent misses renewal, both versions admit the destructive operation after its ownership record expires. A crashed incumbent is reclaimable, and a resource-bound override authorizes deliberate preemption. Incumbent protection therefore requires reliable renewal or health-aware expiry rather than elapsed time alone. VII. D ISCUSSION A. Admission semantics, not execution order Mutual exclusion and denial solve different surrogate problems. B3’s per-resource mutex orders broker admission but does not encode incumbent ownership or change action authorization. It therefore follows B0 because the same destructive action is admitted. Extending the mutex over execution would change ordering, not whether the requester may displace an incumbent. Transactions similarly provide atomicity or recovery without assigning preemption authority. Static denial fails for the opposite reason. Without target, lease, or health context, it rejects destructive operation names even when their resource-specific effect is benign. B2 therefore preserves incumbents by sacrificing task completion. The relevant policy question is whether the effect on this resource is currently authorized. LeaseGuard resolves that question from the canonical resource, active leases, incumbent health, effect class, coexistence limits, and scoped override. The same operation can proceed on an unoccupied resource, execute under a safe limit, redirect around a healthy incumbent, reclaim an unhealthy incumbent, or use explicit preemption authority. The v0.2 scripted matrix supports this resource-state decision boundary rather than ordering or broad privilege. B. Safety and recovery are separable Incumbent preservation is decided before recovery. Capacity-degrading actions can proceed through A LLOWW ITH -L IMITS. Destructive actions can return adaptergenerated alternatives for a second structured selection. The 90.1% recovery rate shows that these alternatives recover task completion for most interventions without weakening the admission rule. The no-alternative stress case preserves the incumbent even when the requested task remains incomplete. This separation also explains the remaining B4 failures. Every incumbent survives, but an invalid initial action or a non-completing recovery choice can leave the requested

C. Authenticated identity and liveness The forged-identity result shows that a task identifier is not by itself an ownership credential. The execution substrate must bind the task to an authenticated principal outside the modelcontrolled payload, or provide an unforgeable task capability that the broker verifies. Hardened v0.2 implements the task– principal equality check. The full model rerun and stress study in Section VI-C evaluate that implementation. Lease expiry presents the corresponding liveness problem. The registry records when a temporal claim ends, but expiry alone does not establish that the incumbent has stopped. The expiry-only ablation therefore admits preemption of a stillhealthy incumbent. Reliable renewal, a health-aware grace policy, or an authenticated supervisor can supply the missing liveness evidence. These choices change reclamation latency, but not the need to distinguish elapsed time from incumbent health. VIII. L IMITATIONS AND R ESPONSIBLE U SE LeaseGuard’s preservation property assumes complete mediation and a trustworthy enforcement path. It covers effects routed through registered adapters. Direct privileged-shell actions and compromise of the broker, registry, token secret, kernel, or container runtime bypass the policy. Canonical resource discovery and incumbent-health predicates are also part of this boundary. A missed alias or false-unhealthy reading can expose a live incumbent, whereas a stale healthy reading can delay legitimate reclamation. The evidence comes from a finite synthetic benchmark spanning five balanced conflict families and two small local instruction models under a frozen action interface. It establishes the measured mechanism effect in that setting, not production reliability or a general ranking of agents. The evaluated cases use one active incumbent lease per resource; safe-limit composition across multiple coexisting incumbents remains unevaluated. Long-running renewal, distributed registries, incomplete adapter coverage, and human override practice also remain open. All destructive benchmark effects ran inside disposable, no-network containers with dropped capabilities, a read-only root filesystem, and bounded resources. No host or production resource was targeted. In deployment, override issuance must remain outside model control and be restricted to authenticated operators. One-shot, resource- and operation-bound tokens, short expiry, and retained receipts make intentional preemption explicit and auditable. The prototype records token consumption only in process-local memory; concurrent replay and restart recovery remain unevaluated. Unknown operations and unresolved targets should continue to fail closed.

IX. R EPRODUCIBILITY AND E VIDENCE S TATUS Each of the 1200 model cells has an append-only trial record containing the model identity, selected action, broker receipt, before/after resource evidence, deterministic grades, execution errors, and cleanup status. Trial records reference entries in a separate 561-record response stream that preserves raw model output and parser results; reused selections therefore map to several baseline cells. Separate raw files contain the v0.2 mechanism matrix and versioned stress audits, including the retained v0.1 forged-identity failure. A deterministic builder checks the frozen test inventory, trial records, and response records against their manifests before regenerating the numerical macros, tables, figures, cost summary, and claim–evidence matrix. The resulting paired effects match the independent analyzer. The manuscript imports these generated values, avoiding manual result transcription. The experiment manifest records the development gate, first access to the frozen test set, evaluated broker and codebundle hashes, pinned local-model weights, and container image. No test case, grader, metric, or exclusion changed for the v0.2 rerun. Deterministic source archives retain both the diagnosed v0.1 bundle and the evaluated v0.2 bundle; the latter is linked to the model, mechanism, and stress manifests. The v0.1 archive was reconstructed from the original execution transcript and independently verified against its recorded codebundle hash. X. C ONCLUSION Execution privilege does not determine whether a requester may displace a healthy incumbent. This mismatch drives shared-resource agent conflicts. LeaseGuard makes preemption authority explicit through canonical resource leases, incumbent-health checks, effect-aware admission, and resource-scoped overrides. Across the frozen two-model benchmark, it produced zero measured unauthorized preemptions while retaining high safe completion. The scripted matrix separates this behavior from an admission-only mutex and operation-only denial. The preservation property assumes complete mediation; stress tests identify authenticated task– principal binding and lease liveness grounded in reliable renewal or health-aware expiry as additional requirements. The resulting design principle is to represent and enforce preemption authority as resource state, rather than infer it from model intent or broad execution privilege. R EFERENCES [1] Y. Xie, Y. Li, D. Guo, Q. Liu, Y. Fu, Y. Fu, Y. Yang, X. Hu, and D. Liu, “ClashBench: Conflicts Leading Agents to Seize and Harm,” arXiv preprint arXiv:2609.19892, 2026. [2] H. Wang, C. M. Poskitt, and J. Sun, “AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents,” in Proceedings of the 48th IEEE/ACM International Conference on Software Engineering (ICSE), 2026, pp. 2938–2950. [3] Y. Li, Y. Chen, H. Wen, B. Zhang, H. Liu, P. Wang, Y. Feng, and Y. Tian, “VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills,” arXiv preprint arXiv:2606.26524, 2026. [4] J. He and D. Yu, “Sovereign Execution Broker: Enforcing CertificateBound Authority in Agentic Control Planes,” arXiv preprint arXiv:2606.20520, 2026.

[5] Y. Zheng, J. Fan, Q. Fu, Y. Yang, W. Zhang, and A. Quinn, “AgentCgroup: Understanding and Controlling OS Resources of AI Agents,” arXiv preprint arXiv:2602.09345, 2026. [6] V. Reddy, S. R. Challaram, and A. Basu, “Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents,” arXiv preprint arXiv:2607.07405, 2026, KDD-ETAAI 2026. [7] T. Sah, V. Srivastava, D. Sah, and K. Jordan, “The Verifier Tax: Horizon Dependent Safety–Success Tradeoffs in Tool Using LLM Agents,” arXiv preprint arXiv:2603.19328, 2026. [8] B. Mohammadi, N. Potamitis, L. Klein, A. Arora, and L. Bindschaedler, “Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows,” arXiv preprint arXiv:2602.14849, 2026. [9] S. Khan, “Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems,” arXiv preprint arXiv:2606.17182, 2026. [10] T. Shi, J. He, Z. Wang, H. Li, L. Wu, W. Guo, and D. Song, “Progent: Securing AI Agents with Privilege Control,” arXiv preprint arXiv:2504.11703, 2025. [11] R. K. Sharma, L. Jiang, S. Chen, and Z. Lin, “Beyond OAuth: TaskScoped Authorization for AI Agents via Natural Language Slices,” arXiv preprint arXiv:2603.17170, 2026. [12] I. David, M. Guarnieri, and A. Gervais, “Alignment Contracts for Agentic Security Systems,” arXiv preprint arXiv:2605.00081, 2026. [13] Y. Zhang, “Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents,” arXiv preprint arXiv:2606.03895, 2026. [14] Z. Zhao, Y. Zhang, Y. Zhu, J. Wang, S. Tao, X. Cheng, and J. Gao, “AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents,” arXiv preprint arXiv:2606.21129, 2026. [15] R. N. M. Watson, J. Anderson, B. Laurie, and K. Kennaway, “Capsicum: Practical capabilities for UNIX,” in 19th USENIX Security Symposium, 2010, pp. 29–46. [16] G. Klein, K. Elphinstone, G. Heiser, J. Andronick, D. Cock, P. Derrin, D. Elkaduwe, K. Engelhardt, R. Kolanski, M. Norrish, T. Sewell, H. Tuch, and S. Winwood, “seL4: Formal verification of an OS kernel,” in Proceedings of the ACM SIGOPS 22nd Symposium on Operating Systems Principles, 2009, pp. 207–220. [17] C. Gray and D. Cheriton, “Leases: An efficient fault-tolerant mechanism for distributed file cache consistency,” in Proceedings of the Twelfth ACM Symposium on Operating Systems Principles, 1989, pp. 202–210. [18] B. Hindman, A. Konwinski, M. Zaharia, A. Ghodsi, A. D. Joseph, R. H. Katz, S. Shenker, and I. Stoica, “Mesos: A platform for fine-grained resource sharing in the data center,” in 8th USENIX Symposium on Networked Systems Design and Implementation, 2011, pp. 295–308. [19] A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes, “Large-scale cluster management at Google with Borg,” in Proceedings of the Tenth European Conference on Computer Systems, 2015, pp. 1–17. [20] S. Poudel, T. Ramachandran, O. Vasios, and A. P. Reiman, “Agentic Workflows for Resolving Conflict Over Shared Resources: A Power Grid Application,” arXiv preprint arXiv:2604.09823, 2026. [21] Qwen Team, “Qwen3 Technical Report,” arXiv preprint arXiv:2505.09388, 2025. [22] IBM Granite Team, “Granite-3.3-2B-Instruct Model Card,” Hugging Face model documentation, 2025, release date: April 16, 2025; accessed September 20, 2026. [Online]. Available: https://huggingface. co/ibm-granite/granite-3.3-2b-instruct

Record · ID 1028601 · SHA-256 755b45ad982c1eb2
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.