Conceptio › Archive › arXiv CS
arXiv CSopen access

Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces

Ruozhao Yang et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

B EFORE ACTING , C HANGE THE S TATE : P ROSPECTIVE S TATE I NTERVENTION FOR W EB AGENTS UNDER D E CEPTIVE I NTERFACES

arXiv:2609.34974v1 [cs.AI] 28 Sep 2026

Ruozhao Yang, Mingfei Cheng, Xiaofei Xie School of Computing and Information Systems Singapore Management University Singapore 188065

A BSTRACT LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users’ interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid action can still realize an unauthorized consequence because of the current Web state. This motivates treating task-relevant Web state itself as a runtime control target. We introduce Veer, an agent-side runtime defense that leaves task planning to the base agent and intervenes on Web state when a proposed action would produce an unauthorized consequence. Before modifying the live environment, Veer constructs a prospective intervention trajectory toward a safe task-relevant state and executes it with runtime grounding and verification. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings, exceeding the next-best defense by 15.9 and 25.0 percentage points on TrickyArenaSingle and TrickyArena-Multi, respectively, while reducing dark-pattern success on WebDecept to 0.3%. These gains persist across dark-pattern types and all 12 agent, model, and benchmark configurations. Ablations show that active state intervention provides the largest gain, while prospective rollout and temporal evidence contribute additional improvements. These results establish task-relevant Web state as an effective runtime control target for protecting Web agents from deceptive outcomes.

1

I NTRODUCTION

Large language model (LLM)-based Web agents can autonomously perform multi-step tasks such as shopping, booking, and information retrieval Zhou et al. (2024); Koh et al. (2024). As these agents increasingly act on users’ behalf, they also encounter dark patterns: deceptive interfaces that steer decisions toward outcomes users may not otherwise choose. Recent studies reveal substantial susceptibility. TrickyArena reports an average susceptibility of 41% to individual dark patterns across six Web agents Ersoy et al. (2026), while DECEPTICON finds undesirable outcomes in more than 70% of tested tasks Cuvin et al. (2026). WebDecept further demonstrates similar failures in realistic shopping tasks Shi et al. (2026). Together, these findings establish deceptive interfaces as a persistent risk across agents, models, domains, and interaction settings. Existing defenses address this risk through prompting, action screening, guidance, and replanning. Safety instructions and dark-pattern-aware prompting can reduce susceptibility, though their effectiveness varies across tasks and dark-pattern types Cuvin et al. (2026); Shi et al. (2026). Recent runtime defenses reason more explicitly about deceptive interactions and action consequences: DUDE provides deception-aware guidance Zhang et al. (2026), WebGuard predicts the outcomes and risks of state-changing Web actions Zheng et al. (2025), and SafePred and SeerGuard anticipate future consequences to support screening, guidance, and replanning Chen et al. (2026); Yu et al. (2026). Yet recognizing deceptive patterns alone does not reliably prevent undesirable agent behavior Tang et al. (2026). Across these approaches, runtime protection primarily centers on whether or how the agent should proceed with its next action. 1

Current Web State: with unwanted warranty

Intervene Web State: unselect the warranty

ShopMart

User Task: Buy the laptop.

ShopMart

https://shopmart.com/cart

https://shopmart.com/cart

Your Cart

Your Cart Laptop 15.6'' FHD, 16GB RAM, 512GB SSD

2-Year Warranty Extended coverage for 2 years

Select all items

$899.00

1

▾

$899.00

$129.00

1

▾

$129.00

Subtotal (2 items)

$1028.00

Taxes and shipping calculated at checkout

Continue Shopping

Checkout

Laptop 15.6'' FHD, 16GB RAM, 512GB SSD

2-Year Warranty Extended coverage for 2 years

Select all items

$899.00

1

▾

$899.00

$129.00

1

▾

$129.00

Subtotal (1 item)

$899.00

Taxes and shipping calculated at checkout

Continue Shopping

Checkout

Figure 1: Motivating example of a task-valid action whose consequence becomes unauthorized because of the current Web state.

This action-centric view leaves an important case unresolved: a task-valid action can still realize an unauthorized consequence because of the Web state in which it is executed. Figure 1 illustrates this setting. Checkout remains appropriate for the requested purchase, yet an auto-added warranty changes its consequence. The action is task-valid; the safety-critical factor is the state at execution time. Similar cases arise from preselected options, retained consent, enabled settings, and other conditions established during interaction. We refer to these task-relevant conditions as the Web state. This creates a state-level runtime control point: the consequence-causing state can be changed before the task proceeds. We therefore ask: How can a runtime defense intervene on consequencecausing Web state while preserving the agent’s progress toward the user task? Realizing state intervention in a black-box Web environment raises three challenges. ❶ Consequence-relevant state is sparse and persistent. An eventual consequence may depend on a small part of the current interface or on state established several interactions earlier, requiring the defense to connect temporal Web evidence with an often underspecified user request. ❷ Safe intervention can require multiple dependent state transitions. Application state is changed through browser interactions, and later steps may depend on state established earlier. Trying candidate interventions directly can itself modify the live application, requiring prospective reasoning before actuation. ❸ Prospective transitions can diverge from live execution. Dynamic content, failed interactions, and hidden application behavior can invalidate anticipated transitions, requiring each intervention step to be grounded and verified against the live interface. To address these challenges, we introduce Veer, an agent-side runtime defense based on consequence-guided prospective state intervention. At each action boundary, Veer uses current and temporal evidence to identify the task-relevant state responsible for an unauthorized consequence and formulates an intervention objective specifying what to change, what safe state to reach, and what task progress to preserve. Before modifying the live environment, it incrementally constructs a prospective trajectory over available browser interactions. During execution, Veer re-grounds each transition in the current interface and verifies that the observed state change matches the planned effect. Once the intervention objective is satisfied, control returns to the base agent. We evaluate Veer on TrickyArena and WebDecept against dark-pattern-specific and general agentsafety defenses. Veer achieves the highest safe task completion (STC) in all three evaluation settings, reaching 85.2%, 67.6%, and 41.0% on TrickyArena-Single, TrickyArena-Multi, and WebDecept, respectively. It exceeds the next-highest STC by 15.9 and 25.0 percentage points on the two TrickyArena settings and reduces dark-pattern success on WebDecept to 0.3%. Across 471 benchmark configurations, its gains persist across deceptive conditions and agent–model configurations. Ablations further identify active state intervention as the largest contributor to STC, with prospective rollout and temporal evidence providing additional gains. Our work makes three key contributions: ❶ We introduce consequence-guided state intervention, establishing task-relevant Web state as a runtime control target when a task-valid action would otherwise realize an unauthorized consequence. ❷ We develop Veer, which combines prospective state intervention with guarded live execution in black-box Web environments. ❸ We evaluate Veer across 471 configurations on TrickyArena and WebDecept, showing the highest STC in all three settings, robust gains across deceptive conditions and agent–model configurations, and clear contributions from its core design choices. 2

Veer

User Task

Base Agent

Authorized at

at

Consequence Assessment Authorized?

Prospective Rollout

Objective ℐt

Unauthorized

Rt Responsible State St ⋆

Interface Grounding

a1I

St

Ŝ1

Control Semantics

Target State

ai I

Admission

Ŝi

Page Semantics

Pt Task-Relevant State

ai I

Available Interactions

Expected outcomes

ŜK

Step Constraints

Preconditions · Dependencies

Web State St

aKI

Ŝi

Rt Resolved St⋆ Reached

Preserve Pt

Pt Preserved

Advance ℐt

Dependencies Closed Admitted

No

Live Web Environment

Return Control

St+

Verify ℐt No

More Steps?

Yes

Effects Supported?

Si

Yes

aiI

Ground & Check

Dependencies Met: Next Step

Figure 2: Overview of Veer. A base-agent proposal is assessed against task authorization under the grounded Web state. Unauthorized consequences trigger a prospective state intervention that is admitted before actuation and verified against the live Web environment during execution.

2

P RELIMINARY

Web-agent execution. We consider a Web agent that completes a natural-language task u through multi-step browser interactions Zhou et al. (2024); Koh et al. (2024). At step t, given the current observation ot and interaction history τ<t , the base-agent policy π proposes at ∼ π(· | u, ot , τ<t ), whose execution produces the next observation ot+1 . Task-relevant action consequences. We use ct to denote the task-relevant consequence of executing at . An action is task-valid when it is consistent with completing the user task, while its consequence may depend on task-relevant conditions such as selected options, saved settings, or workflow state. We denote these conditions by St and refer to them as the Web state. A consequence is unauthorized when it includes an outcome unsupported by the user task. A task-valid action can therefore produce an unauthorized consequence when Web state introduces an additional outcome beyond user intent. Runtime defense setting. We consider an agent-side runtime defense that operates between the base agent and the Web environment. Before executing at , it receives u, ot , τ<t , and at . The defense interacts with the website only through browser operations available to the base agent, with no privileged access to source code, backend logic, or internal application state. It must infer the relevant St from observable Web evidence and is given neither a clean counterpart of the interface nor an explicit dark-pattern label. The operational scope and applicability boundaries of this runtime setting are summarized in Appendix G.

3

Veer: C ONSEQUENCE -G UIDED S TATE I NTERVENTION

Veer realizes consequence-guided state intervention through three coupled designs (Figure 2): taskgrounded consequence assessment, prospective state intervention, and guarded trajectory execution. An explicit intervention objective connects prospective planning with live execution: the objective remains fixed while the browser-level trajectory can adapt to evidence observed during intervention. 3.1

TASK -G ROUNDED C ONSEQUENCE A SSESSMENT

To address Challenge ❶, we design task-grounded consequence assessment around two complementary representations: a dynamic task-relevant Web state St , which captures current interaction conditions, and a stable task-authorization representation Au , which specifies what the user permits. We instantiate them as St = G(ot , Lt ), Au = A(u), 3

where St combines the current observation with valid temporal evidence, preserving consequencerelevant selections, settings, and workflow state across interactions. Au is derived from the original user instruction and remains fixed: runtime observations can ground its references to concrete Web objects and facts, but cannot expand the user’s authorization. This separation allows Veer to track evolving Web state without allowing the interaction itself to redefine the user’s intent. The concrete authorization, Web-state, and temporal-evidence representations are detailed in Appendix B. 3.2

P ROSPECTIVE S TATE I NTERVENTION

When yt = U NAUTHORIZED, Veer uses the predicted consequence ĉt , the grounded Web state St , and task authorization Au to construct an explicit intervention objective It = (Rt , St⋆ , Pt ), where Rt identifies the task-relevant state in St responsible for the unauthorized consequence, St⋆ specifies the target state conditions that remove this contribution, and Pt records task-relevant state in St that should be preserved. The objective makes explicit what must change, what conditions the intervention should establish, and what existing progress must remain intact. These requirements remain fixed while Veer determines how to realize intervention through the available Web interface. Using It as the planning constraint, Veer constructs a complete prospective intervention trajectory Bt = [τ1 , . . . , τK ], before issuing any state-changing intervention action. Each transition τi specifies a browser action, its expected task-relevant effect, and dependencies on earlier transitions. Veer selects and orders these transitions so that their expected effects address Rt , establish St⋆ , and preserve Pt . Dependencies capture multi-step interventions in which a later action requires state established by an earlier one. Planning is completed before actuation because trying candidate interventions directly on the live application would itself modify the state being planned over. Before Bt is passed to execution, Veer checks that its targets are grounded in the current interface, its dependencies are valid, and its planned effects remain consistent with It . Appendix B provides the concrete intervention-objective and prospective-trajectory representations. 3.3

G UARDED T RAJECTORY E XECUTION

Challenge ❸ arises when the prospective trajectory is applied to the live Web application: the effect expected from a planned transition may differ from the state actually produced. Veer therefore treats each transition in Bt as provisional until its expected effect is supported by live evidence. Before executing τi , Veer re-grounds its target in the current Web state and checks that its dependencies have been satisfied. After execution, it observes the resulting state and compares the task-relevant change with the expected effect specified by τi . Only a confirmed transition enables dependent steps. A mismatch invalidates the remaining trajectory because its later steps may rely on state that was never established. Veer stops the current trajectory, re-grounds the actual Web state, and constructs a new trajectory when a valid continuation can still satisfy the same intervention objective It . The intervention intent thus remains fixed while its browser-level realization adapts to live evidence. Execution succeeds when the resulting state St+ satisfies St+ |= St⋆ ,

St+ |= Pt ,

Resolved(Rt , St+ ).

Veer then returns the updated Web state to the base agent, which resumes planning from it.

4

E XPERIMENTS

4.1

E XPERIMENTAL S ETUP

Benchmarks and baselines. We evaluate Veer on two benchmarks for Web agents under deceptive interfaces: TrickyArena Tang et al. (2026) and WebDecept Shi et al. (2026). TrickyArena contains 156 task–dark-pattern configurations across shopping, news, streaming, and health applications, including 88 single-pattern and 68 multi-pattern cases. WebDecept contains 45 shopping tasks 4

Table 1: Overall effectiveness on TrickyArena and WebDecept (%). Lower DPSR and higher TSR/STC are better. Bold indicates the best value, and gray shading highlights Veer. Method

TrickyArena-Single

TrickyArena-Multi

WebDecept

DPSR↓ TSR↑ STC↑ DPSR↓ TSR↑ STC↑ DPSR↓ TSR↑ STC↑ No-Defense

29.5

80.7

60.2

54.4

66.2

38.2

37.5

47.6

28.6

ICP Guardrail DUDE-S2

17.0 27.3 34.1

80.7 80.7 86.4

69.3 62.5 56.8

50.0 47.1 64.7

67.6 70.6 70.6

39.7 42.6 27.9

26.7 12.4 35.9

43.8 35.2 55.6

31.7 29.5 31.4

Spotlighting VIGIL SafePred

26.1 15.9 28.4

79.5 71.6 86.4

62.5 59.1 63.6

58.8 39.7 55.9

60.3 58.8 79.4

26.5 38.2 36.8

35.6 9.5 38.1

44.8 10.2 58.1

27.3 5.4 38.7

Veer

4.5

86.4

85.2

14.7

72.1

67.6

0.3

41.0

41.0

evaluated under seven deceptive scenarios, yielding 315 task–scenario configurations. We compare Veer with the unprotected base agent (No-Defense), dark-pattern-specific defenses including ICP and Guardrail Cuvin et al. (2026) and DUDE-S2 Zhang et al. (2026), and general agent-safety defenses including Spotlighting Hines et al. (2024), VIGIL Lin et al. (2026), and SafePred Chen et al. (2026). Within each benchmark, all methods use the same task configurations, base agent, actor model, and base task-step allowance. Detailed experimental settings are provided in Appendix C, while benchmark coverage and baseline adaptations are provided in Appendix D. Metrics. For each episode i, let Ti = 1 denote successful completion of the user task and Di = 1 indicate that at least one evaluated dark-pattern outcome occurs. We report N

1 X DPSR = Di , N i=1

N

1 X TSR = Ti , N i=1

N

1 X STC = Ti (1 − Di ). N i=1

Dark Pattern Success Rate (DPSR, ↓) measures susceptibility to deceptive outcomes, and Task Success Rate (TSR, ↑) measures completion of the original user task. We use Safe Task Completion (STC, ↑) as the primary metric, requiring task completion without any evaluated dark-pattern outcome. For TrickyArena multi-pattern cases, Di = 1 if any constituent dark pattern succeeds. Complete outcome counts and evaluation details appear in Appendix E. 4.2

RQ1: OVERALL E FFECTIVENESS

RQ1: How effectively does Veer prevent dark-pattern outcomes while preserving successful task completion? Table 1 reports the overall results on TrickyArena and WebDecept. Veer achieves the highest STC in all three evaluation settings. The results further show that this gain arises from different safety–utility profiles across the two benchmarks: on TrickyArena, Veer sharply reduces dark-pattern outcomes while maintaining task completion, whereas WebDecept additionally exposes a benchmark-level constraint on whether safe completion remains possible.

Table 2: WebDecept results on the ground-truthfeasible subset (N = 225; %). Bold marks the best value; gray shading highlights Veer. Method

DPSR↓ TSR↑ STC↑

No-Defense ICP Guardrail DUDE-S2 Spotlighting VIGIL SafePred

17.8 5.8 7.1 16.9 17.3 3.1 19.6

50.2 47.6 45.3 54.2 47.6 8.9 64.4

40.0 44.4 41.3 44.0 38.2 7.6 54.2

TrickyArena. In the single-pattern setting, Veer matches the highest TSR at 86.4% while Veer 0.0 57.3 57.3 reducing DPSR to 4.5%, compared with 28.4% for SafePred at the same TSR. This yields 85.2% STC, 15.9 percentage points above the next-highest result. The advantage grows under multiple dark patterns: relative to No-Defense, Veer reduces DPSR from 54.4% to 14.7% while increasing TSR from 66.2% to 72.1%, reaching 67.6% STC, 25.0 points above the next-highest method. These 5

Shop

News

Spotify

(b) WebDecept Health

Avg.Worst

Scenario

Avg.Worst

No-Defense

27 100

37.5 86.7

ICP

16.7 100

26.7 80

Guardrail

24.8 100

12.4 48.9

DUDE-S2

32 100

35.9 84.4

Spotlighting

23.9 100

35.6 82.2

75

50

25

0

Po p Ba up n P- ner po P- pu ba p n Ad ner dRe ons dir ec t Dr ift

cf

ds du cs to s

cf am

sa

s w bs ob

t8 p2

t7

t6

0.3 2.2

t5

5.7 100

t4

Veer t3

38.1 86.7

t2

9.5 26.7

25.8 100

t1

15.5 100

p1

VIGIL SafePred

100

DPSR (\%) ↓

(a) TrickyArena-Single

Figure 3: Per-condition dark-pattern susceptibility on TrickyArena-Single and WebDecept. Color encodes DPSR (%; lower is better); Avg. and Worst report the mean and maximum across conditions. results show that the STC gains on TrickyArena come from converting more task executions into safe completions rather than broadly suppressing task progress. WebDecept. On the full 315 configurations, Veer records only one dark-pattern success, yielding 0.3% DPSR and 41.0% STC. SafePred attains a higher TSR of 58.1%, but its 38.1% DPSR reduces STC to 38.7%. The lower TSR of Veer motivates examining task feasibility. WebDecept includes two scenarios, redirection and price drift, whose ground truth does not specify a safe completion path once the deceptive condition is encountered. We therefore separately evaluate the 225 configurations for which the ground truth specifies a safe completion path. The construction of this ground-truth-feasible subset is detailed in Appendix D.3. On this ground-truth-feasible subset, Veer eliminates all evaluated dark-pattern outcomes while achieving 57.3% TSR and STC. Compared with No-Defense, it increases TSR by 7.1 percentage points while reducing DPSR from 17.8% to 0.0%. SafePred reaches a higher TSR of 64.4%, yet its 19.6% DPSR yields 54.2% STC. When safe completion is available, Veer therefore improves task completion over the unprotected agent while achieving the strongest joint safety–utility outcome. RQ1 Takeaway. Veer achieves the highest STC in all three settings, combining strong darkpattern suppression with preserved task progress when safe completion is feasible. 4.3

RQ2: ROBUSTNESS ACROSS D ECEPTIVE C ONDITIONS AND AGENT–M ODEL C ONFIGURATIONS

RQ2: How robust is Veer across dark-pattern types, multi-pattern settings, and agent–model configurations? We examine robustness along three dimensions: individual deceptive conditions, multi-pattern Table 3: STC across agent–model configurainteractions, and agent–model configurations. tions (%). Parentheses show gains over matched Figure 3, Table 1, and Table 3 show that the adNo-Defense. vantage of Veer persists across all three. Agent

Model

Single

Multi

WebDecept

Deceptive conditions. Figure 3 shows that the Default GPT 85.2 (+25.0) 67.6 (+29.4) 41.0 (+12.4) DPSK 68.2 (+12.5) 44.1 (+10.3) 40.0 (+4.8) aggregate safety gains are not driven by a small 75.0 (+28.4) 54.4 (+25.0) 45.4 (+5.1) subset of favorable conditions. On TrickyArena- Codex GPT DPSK 85.2 (+20.4) 55.9 (+14.7) 53.7 (+15.9) Single, Veer achieves the lowest or tied-lowest DPSR in 21 of 22 categories and avoids darkpattern outcomes entirely in 20. Its mean percategory DPSR is 5.7%, compared with 15.5–32.0% for the baselines. On WebDecept, Veer records only one deceptive outcome across 315 configurations, with a mean per-scenario DPSR of 0.3% and a worst-case DPSR of 2.2%; the corresponding baseline ranges are 9.5–38.1% and 26.7–86.7%. The 6

safety advantage therefore extends across heterogeneous dark-pattern mechanisms on both benchmarks. The complete per-condition results underlying Figure 3 are reported in Appendix E.3. Multi-pattern interactions. Table 1 further shows that Veer retains its advantage when multiple dark patterns occur within the same configuration. Because TrickyArena-Single and TrickyArenaMulti contain different configuration mixtures, we compare their aggregate changes rather than treating them as paired conditions. From Single to Multi, the DPSR of Veer increases by 10.2 percentage points, compared with 19.8–33.0 points for the baselines, while its STC decreases by 17.6 points, compared with 19.9–36.0 points. Thus, although all methods degrade under the multi-pattern setting, Veer shows the smallest aggregate deterioration in both safety and safe task completion. Agent–model configurations. We evaluate Veer with two agent implementations, the default benchmark agent (Default) and Codex, each paired with GPT-5.6-Luna (GPT) and DeepSeek-V4Flash (DPSK). Across the resulting 12 agent–model–benchmark comparisons, Veer improves STC over the matched No-Defense setting in every case, with gains ranging from 4.8 to 29.4 percentage points. DPSR also decreases in all 12 comparisons, by at least 20.5 points. This consistency across agents, models, and benchmarks indicates that the gains of Veer are not specific to the main experimental configuration. The corresponding system configurations and complete results are provided in Appendices C.4 and E.4. RQ2 Takeaway. Veer remains robust across deceptive conditions, multi-pattern interactions, and agent–model configurations, reducing DPSR and improving STC in all 12 system comparisons. 4.4

RQ3: C ONTRIBUTION OF C ORE D ESIGN C HOICES

RQ3: How do the key design choices of Veer contribute to its effectiveness? Figure 4 isolates three design choices on TrickyArena-Single by decomposing each episode according to task completion and dark-pattern occurrence. Full Veer achieves 85.2% STC. The ablations reveal distinct roles: state intervention primarily preserves task progress, prospective rollout reduces unsafe completions, and temporal evidence supports both safe intervention and task completion.

Safe completion Task success + DP

Task failure, no DP Task failure + DP Full Veer STC

85.2

Full Veer

10.2

−35.2 pp

w/o State Intervention

50.0

38.6

10.2

−4.5 pp

Reactive Intervention

80.7

12.5

−8.0 pp

w/o Temporal Evidence

State intervention. Replacing active intervention with blocking causes the largest degradation, reducing TSR from 86.4% to 51.1% and STC from 85.2% to 50.0%. Episodes with neither task completion nor a dark-pattern outcome increase from 10.2% to 38.6%. This shift shows that blocking often avoids the unauthorized consequence by terminating useful task progress. Active state intervention instead resolves the responsible Web state and returns control to the base agent, allowing the task to continue safely.

77.3 0

25

50

75

17.0

100

Episode outcome (%)

Figure 4: Outcome decomposition of RQ3 ablations on TrickyArena-Single (N = 88 per variant). Annotations denote STC decreases from full Veer.

Prospective rollout. Replacing prospective rollout with reactive intervention leaves task completion nearly unchanged: TSR is 85.2%, compared with 86.4% for full Veer. Its DPSR nevertheless rises from 4.5% to 6.8%, reducing STC to 80.7%. Successful episodes containing a dark-pattern outcome likewise increase from 1.1% to 4.5%. Planning the dependent state transitions before actuation thus reduces unsafe completions without materially suppressing task progress. Temporal evidence. Removing temporal evidence reduces TSR from 86.4% to 80.7% and STC from 85.2% to 77.3%. Task failures without a dark-pattern outcome increase from 10.2% to 17.0%, 7

while successful episodes with a dark-pattern outcome increase from 1.1% to 3.4%. These shifts indicate that temporal evidence helps Veer retain task-relevant state across observations and assess the consequences of later actions as the interaction evolves. An additional one-step-intervention ablation, together with representative intervention-trajectory analyses, is reported in Appendix F. RQ3 Takeaway. State intervention drives the largest STC gain; prospective rollout reduces unsafe completions, and temporal evidence supports reliable multi-step intervention.

5

R ELATED W ORK

Dark Patterns and Deceptive Web Interfaces. Dark patterns steer users toward outcomes that may conflict with their interests, motivating extensive study of their taxonomy, prevalence, and effects Gray et al. (2018); Mathur et al. (2019); Nouwens et al. (2020); Luguri & Strahilevitz (2021). Recent work extends this threat to Web agents through studies of combined dark patterns, trajectory manipulation, and e-commerce failures Ersoy et al. (2026); Cuvin et al. (2026); Shi et al. (2026). Other work shows that recognizing deceptive patterns does not reliably prevent undesirable behavior and explores deception-aware guidance Tang et al. (2026); Zhang et al. (2026). Veer intervenes on Web state when it causes a task-valid action to realize an unauthorized consequence. Runtime Safety for Interactive Agents. Runtime defenses protect agents through input isolation, action verification, and consequence prediction. Prior work separates trusted from untrusted content and evaluates prompt-injection defenses Hines et al. (2024); Debenedetti et al. (2024), enforces explicit or intent-grounded action constraints Xiang et al. (2025); Lin et al. (2026), and predicts action consequences for screening, guidance, or replanning Zheng et al. (2025); Chen et al. (2026); Yu et al. (2026). These approaches mainly control how an action proceeds. Veer targets task-valid actions whose safety depends on changing the Web state that determines their consequence. Planning and State Reasoning for Interactive Agents. Long-horizon agents use planning, search, memory, and state prediction to reason beyond the next action. Prior approaches combine tree search, hierarchical planning, contextual guidance, accumulated experience, and reusable workflows for Web and computer-use tasks Zhou et al. (2023); Fu et al. (2024); Zhang et al. (2025); Agashe et al. (2025); Wang et al. (2024). More closely related, world-model-based Web agents predict action-induced state changes for policy selection Chae et al. (2025), while WebDreamer plans over predicted Web states before execution Gu et al. (2024). Veer uses prospective state reasoning for runtime safety, deriving transitions toward a safe state before live modification and verifying them during execution.

6

C ONCLUSION

We introduced consequence-guided state intervention for Web agents whose task-valid actions can realize unauthorized consequences under the current Web state. Veer combines explicit intervention objectives, prospective state intervention, and guarded live execution to modify consequencecausing state while preserving task progress. Across TrickyArena and WebDecept, Veer achieves the highest safe task completion in all three evaluation settings and remains effective across deceptive conditions and agent–model configurations. Ablations identify active state intervention as the largest contributor, with prospective rollout and temporal evidence providing additional gains. These results establish task-relevant Web state as an effective runtime control target for safe Web-agent execution.

8

AI U SE S TATEMENT Generative AI tools were used to assist with English-language editing and polishing of the manuscript and with writing and refining parts of the experimental code. They were not used to formulate the research questions, hypotheses, conceptual framework, system or threat-model specifications, methodology, experimental design, evaluation protocol, or interpretation of experimental results. All AI-assisted text and code were reviewed and verified by the authors. The authors take full responsibility for the final content of the paper, including all claims, results, and artifacts produced with the assistance of generative AI.

E THICS S TATEMENT This work studies runtime defenses for Web agents interacting with deceptive interfaces. Our experiments use established research benchmarks and controlled Web environments and do not involve human subjects or the collection of personal or sensitive user data. The evaluated deceptive interactions are used solely to study and improve agent safety. Because the proposed techniques reason about Web states and agent behavior, they could potentially be adapted beyond defensive purposes; our method, implementation, and evaluation are designed around preventing unauthorized outcomes in controlled benchmark settings. We follow the ICLR Code of Ethics and report our experimental methodology and results with the goal of enabling transparent and responsible evaluation.

R EPRODUCIBILITY S TATEMENT We provide an anonymized repository at https://anonymous.4open.science/r/ Veer-379B containing code and materials for reproducing our experiments. The main paper specifies the method, evaluation protocol, benchmarks, baselines, and metrics, while the appendix provides additional implementation details, configurations, prompts, and extended experimental results. We will publicly release the complete source code, experimental data, configurations, and evaluation artifacts upon acceptance.

R EFERENCES Saaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Wang. Agent s: An open agentic framework that uses computers like a human. In International Conference on Learning Representations, volume 2025, pp. 22924–22946, 2025. Hyungjoo Chae, Namyoung Kim, Kai Ong, Minju Gwak, Gwanwoo Song, Jihoon Kim, Sunghwan Kim, Dongha Lee, and Jinyoung Yeo. Web agents with world models: Learning and leveraging environment dynamics in web navigation. In International Conference on Learning Representations, volume 2025, pp. 63707–63738, 2025. Yurun Chen, Zeyi Liao, Ping Yin, Taotao Xie, Keting Yin, and Shengyu Zhang. Safepred: A predictive guardrail for computer-using agents via world models. arXiv preprint arXiv:2602.01725, 2026. Phil Cuvin, Hao Zhu, and Diyi Yang. How dark patterns manipulate web agents. In International Conference on Learning Representations, volume 2026, pp. 95945–95977, 2026. Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents. Advances in neural information processing systems, 37:82895–82920, 2024. Devin Ersoy, Brandon Lee, Ananth Shreekumar, Arjun Arunasalam, Muhammad Ibrahim, Antonio Bianchi, and Z Berkay Celik. Investigating the impact of dark patterns on llm-based web agents. In 2026 IEEE Symposium on Security and Privacy (SP), pp. 4497–4516. IEEE, 2026. Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, and Honglak Lee. Autoguide: Automated generation and selection of context-aware guidelines for large language model agents. Advances in Neural Information Processing Systems, 37: 119919–119948, 2024. 9

Colin M Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L Toombs. The dark (patterns) side of ux design. In Proceedings of the 2018 CHI conference on human factors in computing systems, pp. 1–14, 2018. Yu Gu, Kai Zhang, Yuting Ning, Boyuan Zheng, Boyu Gou, Tianci Xue, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, et al. Is your llm secretly a world model of the internet? modelbased planning for web agents. arXiv preprint arXiv:2411.06559, 2024. Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. Defending against indirect prompt injection attacks with spotlighting. arXiv preprint arXiv:2403.14720, 2024. Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. Visualwebarena: Evaluating multimodal agents on realistic visual web tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 881–905, 2024. Junda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu, Tong Xu, Yong Chen, and Enhong Chen. Vigil: Defending llm agents against tool-stream injection via verify-before-commit. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 9764–9785, 2026. Jamie Luguri and Lior Jacob Strahilevitz. Shining a light on dark patterns. Journal of Legal Analysis, 13(1):43–109, 2021. Arunesh Mathur, Gunes Acar, Michael J Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan. Dark patterns at scale: Findings from a crawl of 11k shopping websites. Proceedings of the ACM on human-computer interaction, 3(CSCW):1–32, 2019. Midas Nouwens, Ilaria Liccardi, Michael Veale, David Karger, and Lalana Kagal. Dark patterns after the gdpr: Scraping consent pop-ups and demonstrating their influence. In Proceedings of the 2020 CHI conference on human factors in computing systems, pp. 1–13, 2020. Zijing Shi, Meng Fang, and Ling Chen. Benchmarking web agent safety under e-commerce deceptive interfaces. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 22090–22103, 2026. Jingyu Tang, Chaoran Chen, Jiawen Li, Zhiping Zhang, Bingcan Guo, Ibrahim Khalilov, Simret Araya Gebreegziabher, Bingsheng Yao, Dakuo Wang, Yanfang Ye, et al. Dark patterns meet gui agents: Llm agent susceptibility to manipulative interfaces and the role of human oversight. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, pp. 1–26, 2026. Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent workflow memory. arXiv preprint arXiv:2409.07429, 2024. Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Nathaniel D Bastian, et al. Guardagent: safeguard llm agents via knowledge-enabled reasoning. In ICML 2025 workshop on computer use agents, 2025. Xue Yu, Bo Yuan, Pengshuai Yang, Kailin Zhao, Hong Hu, and Junlan Feng. Seerguard: A safety framework for mobile gui agents via world model prediction. arXiv preprint arXiv:2607.15550, 2026. Yao Zhang, Zijian Ma, Yunpu Ma, Zhen Han, Yu Wu, and Volker Tresp. Webpilot: A versatile and autonomous multi-agent system for web task execution with strategic exploration. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pp. 23378–23386, 2025. Yilin Zhang, Yingkai Hua, Chunyu Wei, Xin Wang, and Yueguo Chen. Don’t click that: Teaching web agents to resist deceptive interfaces. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6830–6852, 2026. 10

Boyuan Zheng, Zeyi Liao, Scott Salisbury, Zeyuan Liu, Michael Lin, Qinyuan Zheng, Zifan Wang, Xiang Deng, Dawn Song, Huan Sun, et al. Webguard: Building a generalizable guardrail for web agents. arXiv preprint arXiv:2507.14293, 2025. Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman, Haohan Wang, and Yu-Xiong Wang. Language agent tree search unifies reasoning acting and planning in language models. arXiv preprint arXiv:2310.04406, 2023. Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al. Webarena: A realistic web environment for building autonomous agents. In International Conference on Learning Representations, volume 2024, pp. 15585–15606, 2024.

11

A

V EER RUNTIME A LGORITHM

This appendix provides the complete runtime procedure of Veer. At each action boundary, Veer evaluates the action proposed by the base agent against the current task-relevant Web state and the authorization derived from the user task. An authorized action is released for execution. An uncertain assessment triggers bounded evidence acquisition and reassessment. When the predicted consequence is unauthorized, Veer constructs an intervention objective, plans a prospective stateintervention trajectory before modifying the live environment, and executes the trajectory under runtime state checks. Control returns to the base agent after the responsible state has been resolved while the target safe state and preserved task state remain satisfied. A.1

OVERALL RUNTIME P ROCEDURE

Algorithm 1 summarizes the complete runtime loop. Veer receives the user task u, the current Web observation ot , the preceding interaction history τ<t , temporal evidence Lt , and the action at proposed by the base agent. It first constructs the current task-relevant Web state St and evaluates the consequence that at would realize under this state. The resulting decision is one of AUTHORIZED, U NAUTHORIZED, or U NCERTAIN. The intervention objective remains fixed throughout prospective planning and guarded execution. Browser-level realization can change when live evidence invalidates a planned transition. This separation allows Veer to preserve the intended safety correction while adapting its concrete interaction sequence to the state actually observed at runtime. A.2

C ONSEQUENCE A SSESSMENT AND E VIDENCE ACQUISITION

Veer performs consequence assessment before releasing a proposed action. The current state representation St = G(ot , Lt ) combines the current observation with valid temporal evidence retained from earlier interactions, while Au = A(u) represents authorization derived from the original user instruction. Runtime observations may ground references in Au to concrete Web objects or facts, but they do not expand the authorization established by the user task. The predictor (ĉt , yt ) = P(at , St , Au ) estimates the task-relevant consequence of executing at under the current state and determines whether that consequence is authorized. ĉt includes material effects that would be realized by the proposed action, including persistent state that would be carried into the resulting outcome. Effects that still require an independent future action remain contingent and are not treated as consequences of at . An AUTHORIZED decision releases the proposed action. An U NAUTHORIZED decision transfers the predicted consequence and its supporting state evidence to state intervention. For U NCERTAIN, Veer performs bounded evidence acquisition targeted at the unresolved facts and then reassesses the same proposed action. Evidence acquisition updates the observable evidence available to St without changing the authorization represented by Au . If sufficient evidence cannot be established within the runtime bound, Veer does not treat uncertainty as authorization and instead follows the safe fallback behavior described by the runtime policy. A.3

P ROSPECTIVE I NTERVENTION C ONSTRUCTION

For an unauthorized consequence, Veer constructs It = (Rt , St⋆ , Pt ), where Rt identifies the task-relevant state responsible for the unauthorized consequence, St⋆ specifies the safe state that removes this contribution, and Pt records task-relevant state that should remain unchanged. The intervention objective remains unchanged while Veer determines how to realize it through available browser interactions. Veer then constructs a complete prospective trajectory Bt = [τ1 , . . . , τK ] 12

Algorithm 1 Veer Runtime Defense Require: User task u, observation ot , temporal evidence Lt , proposed action at Ensure: Released base-agent action, updated Web state, or safe termination 1: St ← G(ot , Lt ) 2: Au ← A(u) 3: (ĉt , yt ) ← P(at , St , Au ) 4: if yt = U NCERTAIN then 5: Acquire bounded evidence relevant to the unresolved consequence 6: Update Lt and St 7: Reassess (ĉt , yt ) ← P(at , St , Au ) 8: end if 9: if yt = AUTHORIZED then 10: Release at to the Web environment 11: return 12: end if 13: if yt ̸= U NAUTHORIZED then 14: Terminate without releasing the unresolved proposal 15: return 16: end if 17: It ← C ONSTRUCT O BJECTIVE(ĉt , St , Au ) 18: It = (Rt , St⋆ , Pt ) 19: Bt ← P LAN P ROSPECTIVELY(It , St ) 20: Bt = [τ1 , . . . , τK ] 21: if Bt does not satisfy the admission requirements then 22: Terminate without releasing the unresolved proposal 23: return 24: end if 25: for i = 1, . . . , K do 26: Re-observe the live Web environment 27: Update St and re-ground the target of τi 28: if dependencies of τi are not satisfied then 29: Invalidate the remaining prospective trajectory 30: Replan a new Bt under the fixed objective It 31: Restart guarded execution from the new Bt 32: return 33: end if 34: Execute the grounded intervention action in τi 35: Observe the resulting Web state and update St 36: Compare the observed task-relevant change with the expected effect of τi 37: if the observed transition does not support the expected effect then 38: Invalidate the remaining prospective trajectory 39: Replan a new Bt under the fixed objective It 40: Restart guarded execution from the new Bt 41: return 42: end if 43: end for 44: St+ ← St 45: if St+ |= St⋆ ∧ St+ |= Pt ∧ R ESOLVED(Rt , St+ ) then 46: Return control to the base agent from St+ 47: else 48: Terminate without releasing the unresolved proposal 49: end if

before issuing state-changing intervention actions. Each transition records a browser interaction, its expected task-relevant effect, and dependencies on earlier transitions. The dependencies capture interventions in which later operations require state established by previous ones. Veer orders transitions so that their expected effects resolve Rt , establish St⋆ , and preserve Pt . 13

Prospective construction prevents the planner from testing candidate state changes directly against the live application while deciding how to intervene. Before execution, Veer admits the trajectory only when its interaction targets can be grounded in the available interface, its transition dependencies are well formed, and the planned intervention remains consistent with the fixed objective. Failure to establish an admissible trajectory prevents the intervention from being committed to the live environment. A.4

G UARDED E XECUTION AND R EPLANNING

Prospective transitions remain provisional until they are supported by observations from the live Web environment. Before executing each transition τi , Veer refreshes the current observation, regrounds its target in the live interface, and checks that the state required by its dependencies has been established. Only grounded transitions with satisfied dependencies are executed. After executing τi , Veer observes the resulting interface and compares the task-relevant state change with the expected effect recorded during prospective planning. A confirmed transition enables dependent steps. When the observed effect diverges from the prospective transition, Veer stops relying on the remaining trajectory because subsequent steps may depend on state that was not established. It then re-grounds the actual Web state and constructs a new continuation when the same intervention objective can still be satisfied. Execution completes when the resulting state St+ satisfies St+ |= St⋆ ,

St+ |= Pt ,

Resolved(Rt , St+ ).

These conditions require the target safe state to be established, task-relevant progress to remain preserved, and the state responsible for the unauthorized consequence to be resolved. Veer then returns the updated Web state to the base agent, which resumes its original task from that state. If these conditions cannot be established within the bounded runtime procedure, Veer does not release the unresolved unsafe execution and instead follows the configured safe termination or replanning behavior.

B

S TRUCTURED R EPRESENTATIONS , S CHEMAS , AND P ROMPTS

This appendix details the structured representations used by Veer for task authorization, task-relevant Web state, consequence assessment, and prospective state intervention. Rather than maintaining a single global state object, Veer assembles task-grounded representations from the current Web observation, retained temporal evidence, and structured model outputs. Figures 5–7 summarize the principal representations and the key instructions used to construct them. B.1

TASK AUTHORIZATION R EPRESENTATION

Veer derives the task-authorization representation Au from the original user instruction before runtime consequence assessment. As illustrated in Figure 5, Au separates four forms of task evidence: explicitly authorized terminal outcomes, material constraints, requested informational outputs, and designated output targets. The Boolean field explicit terminal commitment distinguishes tasks that explicitly request an externally consequential terminal outcome from tasks that request inspection, comparison, preparation, navigation, or information retrieval. Each authorization or constraint entry is grounded in a non-empty substring of the original user task. Veer may subsequently bind a task reference to a concrete object observed in the interface, but such grounding is treated as a factual binding rather than a new permission. Runtime Web content can therefore resolve what the task refers to without expanding what the user authorized. Generated authorization spans are checked against the original instruction before use. Unsupported spans are discarded, missing evidence is not promoted to authorization, and unresolved extraction can leave the corresponding authorization judgment uncertain. The resulting task authorization remains fixed throughout the interaction. 14

Task Authorization Representation A_u = { user_task: string, explicit_terminal_commitment: boolean, authorized_terminal_outcomes: [ { task_evidence: string, ... }, ... ], material_constraints: [ { task_evidence: string, ... }, ... ], requested_informational_outputs: [ { task_evidence: string, ... }, ... ], designated_output_targets: [ { task_evidence: string, ... }, ... ], ... }

Authorization Instruction: Extract only authorization explicitly issued by the user task. Do not infer permission for unrelated material consequences. Every authorization and constraint must cite a verbatim non-empty substring of the user task. Return strict JSON only.

Current Observation {

}

ref: string, text: string, role: string, tag: string, checked: value, selected: value, aria_checked: value, value: value, semantic_context: { ... }, ...

Verified Finding {

finding_id: string, claim: string, status: "VERIFIED",

}

evidence: [string], source_url: string, source_target_ref: string, dependency_scope: { ... }, ...

Assessment Instruction: Given a USER TASK, CURRENT WEB STATE, and PROPOSED ACTION, first predict how the action changes user-relevant commitments, then judge whether every predicted commitment is authorized by the USER TASK. A commitment is a user-facing outcome such as consent, subscription, data sharing, an add-on, purchase, or submitted setting. Model consequences caused not only by visible state changes, but also by retaining a preselected state, omitting an opt-out, or crossing a Save/Accept/Continue/Checkout/Done boundary. Prediction is descriptive; authorization is task-conditioned. Keep them in separate JSON fields.

Figure 5: Task authorization, current-observation, and retained-evidence representations used by Veer. The figure also shows the core instructions for authorization extraction and task-conditioned consequence assessment. Runtime Web evidence may ground task references and update St , while authorization remains anchored to the original user instruction.

B.2

TASK -R ELEVANT W EB S TATE AND T EMPORAL E VIDENCE

The paper denotes the task-relevant Web state as St = G(ot , Lt ), where ot is the current Web observation and Lt contains retained temporal evidence. In the implementation, St is assembled from these sources as needed rather than stored as a monolithic state object. Figure 5 shows the principal information retained from the current observation. Grounded controls preserve both interface identity and state, including the element reference, visible text, role, tag, checked/selected state, value, and surrounding semantic context. The current snapshot also contains page-level information such as the URL, visible material facts, rendered collections, bounded choice sets, and a normalized page excerpt. Temporal evidence Lt preserves consequence-relevant facts established across earlier observations. Its principal contents include observed state transitions, verified factual findings, records of previously assessed and executed actions, and unresolved material anomalies. Figure 5 shows the representation of a verified finding, including its claim, supporting evidence, source, and dependency scope. Veer maintains the validity of this evidence as the Web interaction evolves. Evidence associated with an earlier route or material state can become stale when the corresponding state changes. Historical evidence remains distinguishable from current grounded facts and is never treated as an authorization source. Consequently, G denotes the programmatic assembly and filtering of current and retained evidence rather than a separate LLM-based state-synthesis step. 15

Consequence Assessment Representation {

}

predicted_commitments: [ { outcome: string, before: "latent | pending | active | committed | none", after: "latent | pending | active | committed | none", trigger: "click | retention | omission | save | submit | navigation", evidence: [string], reversibility: "reversible | limited | irreversible", ... }, ... ], authorization: "AUTHORIZED | UNAUTHORIZED | UNCERTAIN", commitment_boundary: "NONE | REVERSIBLE | SAVE | ACCEPT | SUBMIT | CHECKOUT | DONE", action_semantic_effect: "CONTEXT_ESTABLISHMENT|INFORMATION_GATHERING| STATE_REDUCING|STATE_MODIFYING|WORKFLOW_ADVANCE| COMMITMENT|UNKNOWN", proposal_risk_relation: "CREATES|INCREASES|COMMITS_EXISTING| CARRIES_TO_MATERIAL_BOUNDARY| PREEXISTING_UNCHANGED|UNRELATED|UNKNOWN", carries_material_state: boolean, carried_material_state: [string], evidence_requests: [ { modality: "visual|semantic|history", causal_scope: "PROPOSED_ACTION|LATER_ACTION", claim: string, reason: string, ... }, ... ], reason: string, uncertainty: "none | low | medium | high", ...

Figure 6: Structured consequence-assessment representation. Veer separates predicted user-facing commitments from the task-conditioned authorization decision and explicitly records action semantics, material-state propagation, causal relation, and requests for additional evidence.

B.3

C ONSEQUENCE A SSESSMENT R EPRESENTATION

At each protected action boundary, Veer assesses the base agent’s proposed action using the user task, task authorization, current grounded Web state, retained temporal evidence, and the proposed browser action. The assessment context includes the current page and grounded elements together with proposal-relevant material facts and previously established evidence. Figure 6 summarizes the structured assessment. The paper-level predicted consequence ĉt is represented primarily through predicted commitments. Each predicted commitment records the relevant outcome, its state before and after the proposal, the causal trigger, supporting evidence, and reversibility. This representation allows Veer to distinguish a proposal that directly creates or commits an outcome from one that merely leaves a condition for a later action. The decision yt ∈ {AUTHORIZED, U NAUTHORIZED, U NCERTAIN} corresponds to the normalized authorization field. Additional fields characterize the semantic role of the proposed action, the commitment boundary being crossed, whether existing material state is carried through that boundary, and the causal relation between the proposal and the identified risk. When the available evidence is insufficient, the assessment can request additional visual, semantic, or historical evidence. Each request specifies the unresolved claim and whether it concerns the proposed action itself or a later action. This distinction is important because an eventual undesirable outcome is not attributed to the current proposal when a separate future action is still required. After generation, Veer checks the structured assessment against the grounded evidence available for the current proposal. Unsupported or incomplete causal claims can leave the effective decision uncertain, triggering bounded evidence acquisition and reassessment before the proposal is released or an intervention is constructed.

Assessment safeguards. Consequence assessment is structured as separate consequence prediction and task-conditioned authorization. Material claims are grounded in available Web evidence, while authorization remains anchored to Au . When the available evidence is insufficient, Veer represents the decision as U NCERTAIN and performs bounded evidence acquisition and reassessment before the proposal can be released. 16

B.4

I NTERVENTION O BJECTIVE

When consequence assessment identifies an unauthorized consequence, Veer first constructs the state-level intervention objective It = (Rt , St⋆ , Pt ) before selecting corrective browser actions. Figure 7 shows the implementation-level representation. The field risk causing state corresponds to Rt and identifies the current state responsible for the unauthorized consequence. target corrected state corresponds to St⋆ and describes the state that must hold before task execution can safely continue. preserved task state corresponds to Pt and records preexisting task-relevant state or capabilities that should remain intact during correction. The preservation list is empty when no such state has yet been established. The objective also includes terminal postconditions used to determine whether the intervention has achieved the intended state-level effect. Evidence references link objective entries to grounded facts in the current interaction. Objective construction is deliberately separated from action planning. The objective specifies what state must change, what corrected state must be reached, and what task-relevant state must be preserved. It does not specify the sequence of browser operations used to realize that change. This separation keeps the intervention target fixed while allowing the concrete trajectory to adapt to the available interface. B.5

P ROSPECTIVE T RANSITION R EPRESENTATION

Given the fixed intervention objective, Veer constructs a prospective intervention trajectory Bt = [τ1 , . . . , τK ]. As shown in Figure 7, each transition τi contains five principal fields: a stable step identifier, a corrective operation, an exact grounded target, the expected local post-state, and dependencies on earlier transitions. Supported corrective operations include turning off or unchecking a control, removing or declining an unwanted state, saving a corrected configuration, clicking or closing a control, and navigating back when appropriate. The expected state field describes the local state that must be established by executing the transition. The depends on field names earlier transitions whose effects must be verified before the current transition becomes eligible. Veer therefore represents cross-step ordering directly rather than treating a multi-step intervention as independent local actions. Figure 7 also shows a representative settings trajectory. Two state-reducing transitions first disable unwanted settings, followed by a save transition whose dependencies require both preceding changes. The entire remaining trajectory is declared before live actuation, allowing Veer to reason about these dependencies prospectively. Prospective rollout. We use prospective rollout to denote explicit construction of the complete remaining browser-level correction trajectory before any corrective state change is issued. It does not assume a learned simulator or latent world model. The prospective property is that expected post-states and cross-step dependencies are specified before actuation and then checked against live execution; divergence invalidates the remaining trajectory and triggers replanning under the same intervention objective. Before execution, Veer checks the transition structure, supported operation, grounded target, and dependency references. During guarded execution, each transition is re-grounded in the live interface and its observed effect is compared with the declared post-state. A mismatch prevents Veer from blindly executing the remaining prospective steps and can trigger replanning under the same intervention objective. The prospective representation itself contains only the planned transition fields shown in Figure 7. Execution-phase metadata, retry state, and postcondition evidence are attached during guarded live execution and are not part of the planner output. 17

Intervention Objective Representation I_t = { structure_id: string, outcome: string, predicted_terminal_commitment: string, risk_causing_state: [ { state_id: string, description: string, evidence_refs: [string], }, ... ], target_corrected_state: [ { state_id: string, description: string, evidence_refs: [string], }, ... ], preserved_task_state: [ { state_id: string, description: string, evidence_refs: [string], }, ... ], terminal_postconditions: [ { condition_id: string, description: string, evidence_refs: [string], }, ... ], ... }

Intervention Trajectory {

}

Multi-step Trajectory [

step_id: string, verb: string, target: string, expected_state: string, depends_on: [string], ...

}, {

Trajectory Construction Instruction: Generate grounded statereducing correction transitions for an immutable objective. Copy exact current element refs. Keep task-preservation constraints. Never perform ordinary unfinished task work or the blocked original task action. Declare the COMPLETE minimal correction trajectory, including all remaining transitions, expected poststates and explicit dependencies, before any real actuation. Never change the objective.

{

}, {

]

}

step_id: "transition-1", verb: "turn_off", target: "data-sharing-switch", expected_state: "off", depends_on: [] step_id: "transition-2", verb: "turn_off", target: "activity-tracking-switch", expected_state: "off", depends_on: [] step_id: "transition-3", verb: "save", target: "save-settings-button", expected_state: "saved", depends_on: ["transition-1", "transition-2"]

Objective Construction Instruction: Declare only the correction objective, not correction actions. Use the grounded unauthorized consequence and current task/state evidence. Preserve pre-existing taskrequired state/capabilities, not unfinished task outcomes. Never describe a future action, transition sequence, dependency or plan. Strict JSON.

Figure 7: Intervention-objective and prospective-trajectory representations. Veer first declares the immutable state-level objective It , then constructs the complete remaining browser-level trajectory Bt . The example shows explicit dependencies between state changes and the final Save transition.

C

E XPERIMENTAL C ONFIGURATION

This appendix details the system configuration, model interfaces, and runtime budgets used in our experiments. Within each benchmark, all compared methods use the same task configurations, base agent, actor model, and base task-action allowance. Veer operates as an agent-side runtime layer between base-agent action generation and browser execution. Budget interpretation. The common budget controls ordinary task execution: all compared methods use the same base-agent task-action allowance within each benchmark. Veer’s internal budget is restricted to defense-side evidence acquisition and state intervention and cannot be used to advance ordinary base-agent task execution. Defense-specific auxiliary reasoning follows the corresponding runtime mechanism of each method. C.1

M AIN S YSTEM C ONFIGURATION

Default configuration. Our main experiments use the default agent provided by each benchmark together with GPT-5.6-Luna as the actor model. Veer uses the same model as the corresponding base agent for its model-backed components, including task-authorization extraction, consequence assessment, evidence reassessment, intervention-objective construction, prospective trajectory generation, and terminal verification. We do not configure separate planner or verifier models. The concrete host agent follows each benchmark’s native interaction stack. On TrickyArena, the default agent is implemented with BrowserUse and operates through Playwright/Chromium. The agent receives multimodal browser observations and interacts through the BrowserUse action interface. On WebDecept, the default agent is the benchmark’s WebArena-style PromptAgent using its multimodal accessibility-tree observation and native browser-action interface. Veer is integrated into both environments at the action boundary, where it receives the user task, observable browser state, retained interaction evidence, and the proposed action before that action is committed to the environment. 18

Table 4: Main interaction budgets. Veer’s internal budget is reserved for defense-side evidence and intervention operations. Setting TrickyArena WebDecept

Task-action budget 30 15

Veer internal budget 8 8

Veer uses browser-observable evidence exposed through the host runtime, including the current page, grounded interface elements, their observable states and semantics, retained temporal evidence, and screenshots when visual reassessment is required. Visual input is used only by components that require it; objective construction and prospective planning operate over the grounded state representation. Veer does not receive benchmark dark-pattern labels, task-success labels, hidden evaluator outputs, or backend application state during execution. Benchmark evaluators are applied only after an episode to compute the reported metrics.

Configuration used by each RQ. RQ1 uses the benchmark-default agent with GPT-5.6-Luna on TrickyArena-Single, TrickyArena-Multi, and WebDecept. The per-condition analysis in RQ2 uses the same default configuration. RQ3 uses TrickyArena-Single with the same BrowserUse–GPT-5.6Luna configuration across all ablations. The agent–model analysis in RQ2 additionally varies the base-agent implementation and model as described in Appendix C.4.

C.2

M ODEL AND S TRUCTURED -O UTPUT S ETTINGS

The main configuration uses GPT-5.6-Luna for both the base agent and Veer while retaining the model interface native to each benchmark integration. Veer’s model-backed components request structured JSON-formatted outputs through the component-specific contracts described in Appendix B. Returned structures are subsequently parsed and checked before they affect browser execution. Consequence assessment additionally supports bounded evidence acquisition and reassessment when the available evidence does not support a confident authorization decision. Interventionobjective construction and prospective planning similarly require their respective structured contracts to be satisfied before a trajectory can be admitted for live execution.

C.3

I NTERACTION AND RUNTIME B UDGETS

We distinguish task actions, which advance the base agent’s ordinary task execution, from internal defense actions, which Veer uses for evidence acquisition and state intervention. Every compared method receives the same base-agent task-action allowance within a benchmark. Veer’s internal budget is reserved exclusively for defense-side operations and cannot be used for ordinary baseagent task execution. Internal operations are accounted for separately from the base agent’s task-action budget. Passive observation and state extraction do not consume internal action slots; Veer-issued browser operations used for evidence acquisition, intervention, or explicit re-observation do. Table 5 reports the principal bounds applied to Veer’s runtime reasoning and execution. These bounds prevent individual assessment or intervention branches from repeatedly consuming the interaction budget. The evidence-attempt bound applies to unresolved evidence claims, while proposal replanning is bounded within an individual controller invocation. If the available runtime bounds cannot establish an authorized continuation or a verified intervention, Veer does not promote the unresolved decision to authorization. 19

Table 5: Principal Veer runtime limits in the main configuration. Runtime mechanism Evidence attempts per unresolved claim Total evidence-claim attempts Assessment completeness reassessment Verified-evidence reassessment Visual reassessment Proposal replanning Intervention step attempts Correction replanning Terminal-verification model calls

Limit 2 8 1 per candidate evaluation 1 per applicable branch 1 per candidate evaluation 5 replans 2 1 up to 2

Table 6: Agent–model configurations evaluated in RQ2. Veer uses the same model family as the corresponding actor. Agent Default Default Codex Codex

C.4

Actor model GPT-5.6-Luna DeepSeek-V4-Flash GPT-5.6-Luna DeepSeek-V4-Flash

Veer model GPT-5.6-Luna DeepSeek-V4-Flash GPT-5.6-Luna DeepSeek-V4-Flash

AGENT–M ODEL C ONFIGURATIONS

RQ2 evaluates whether Veer’s effectiveness depends on a particular base-agent implementation or model. We combine two agent implementations with two model families. D EFAULT denotes each benchmark’s native agent: BrowserUse on TrickyArena and PromptAgent on WebDecept. C ODEX uses Codex as the alternative base-agent implementation. The Default+GPT configuration is the main configuration used for RQ1 and RQ3. For each RQ2 agent–model configuration, Veer uses the same model family as the corresponding base agent, so the comparison evaluates the defense under the model stack used by that configuration.

D

B ENCHMARKS AND BASELINE A DAPTATIONS

This appendix details the benchmark configurations and baseline adaptations used in our experiments. Within each benchmark, all compared methods are evaluated on the same configuration identifiers, use the same base agent and actor model, and receive the same base task-action allowance. Benchmark evaluators are applied only after execution to determine task success and deceptive outcomes. D.1

T RICKYA RENA

Single-pattern setting. TrickyArena-Single contains 88 task–dark-pattern configurations spanning four application domains: Shopping, News, Music, and Health. Table 7 reports the complete composition and the abbreviations used in Figure 3 of the main paper. The 88 configurations form a fixed set of applicable task–condition pairings rather than a Cartesian product over all tasks and deceptive conditions. The domain totals are 50 Shopping, 12 News, 8 Music, and 18 Health configurations. The t1–t8 conditions correspond to interface variants associated with the premium-membership setting. Multi-pattern setting. TrickyArena-Multi contains 68 configurations in which two to four deceptive conditions are simultaneously active. Table 8 reports the 16 condition combinations used in the evaluation. Among the 68 configurations, 46 contain two active patterns, 18 contain three, and four contain four. Evaluation. We use the benchmark-defined task and deceptive-outcome evaluators over the recorded interaction trajectory. Task success is binary for each configuration. For a multi-pattern 20

Table 7: Composition of TrickyArena-Single. The labels follow Figure 3 in the main paper; the two cf conditions are disambiguated by domain. Label p1 t1 t2 t3 t4 t5 t6 t7 t8 p2 s w bs ob sa cf am ds du cs tos cf

Dark-pattern condition N Domain Premium Membership 4 Shopping Code Change Chakra 4 Shopping Code Change Images 4 Shopping Visual Change Links 4 Shopping Visual Change Button Placement 4 Shopping Code Change No Aria 4 Shopping Code Change Images No Aria 4 Shopping Mix Change Button Placement Images No Aria 4 Shopping Mix Change Button Placement Chakra 4 Shopping Cookie Management 4 Shopping Sponsored Items 5 Shopping Warranty 5 Shopping Bait and Switch 3 News Obfuscation 3 News Sponsored Ad 3 News Confusion 3 News Aesthetic Manipulation 2 Music Data Sharing 3 Music Decision Uncertainty 3 Music Complex Settings 6 Health Terms of Service 6 Health Confirm Shaming 6 Health Total 88

Domain Shopping News Music Health Total

Active combinations (N ) p1 p2 (8); p1 w (7); p1 p2 w (7); p1 p2 w s (2) bs cf (3); bs ob (3); bs cf sa (3); bs cf ob sa (2) du ds (3); am ds (2); am du (2); am ds du (2) cs cf (6); cs tos (6); tos cf (6); cs cf tos (6)

Table 8: Composition of TrickyArena-Multi. Total 24 11 9 24 68

episode, the deceptive-outcome indicator is positive when at least one constituent dark-pattern condition succeeds, matching the definition of Di used in the main paper. The evaluator is applied after execution and is not available to the agent or defense at runtime. D.2

W EB D ECEPT

WebDecept contains 45 shopping tasks, each evaluated under seven deceptive scenarios, yielding 45 × 7 = 315 task–scenario configurations. Every task is evaluated under every scenario. Table 9 summarizes the seven scenarios and the labels used in Figure 3. We use the benchmark-defined task-success and deceptive-outcome evaluators for post-episode scoring. Task success determines whether the requested shopping outcome is completed, while the deceptive-outcome evaluator records whether the scenario-specific unsafe outcome occurs during the episode. These evaluators are not exposed to the base agent or Veer during execution. D.3

F EASIBILITY-AWARE W EB D ECEPT S UBSET

WebDecept contains two scenarios, redirection and price drift, whose benchmark ground truth does not specify a safe completion path once the deceptive condition is encountered. To separate this benchmark-level feasibility constraint from defense performance, we additionally evaluate the configurations from the remaining five scenarios, for which the benchmark ground truth specifies a safe completion path. The resulting subset is 21

Table 9: WebDecept deceptive scenarios. Each scenario contains all 45 base shopping tasks. Scenario popup banner personalized popup personalized banner cart addons redirection price drift Total

Label Popup Banner P-popup P-banner Add-ons Redirect Drift

Deceptive condition Exit/checkout popup with deceptive continuation Sticky deceptive discount advertisement Task-conditioned deceptive popup Task-conditioned deceptive banner Additional donation item inserted into the cart Fake continuation substitutes the intended path Cart price changes before commitment

N 45 45 45 45 45 45 45 315

Table 10: Construction of the feasibility-aware WebDecept subset. Scenario Full Feasible Popup 45 45 Banner 45 45 P-popup 45 45 P-banner 45 45 Add-ons 45 45 Redirect 45 0 Drift 45 0 Total 315 225

Cfeasible = {i | scenario(i) ∈ / {redirection, price drift}} . The criterion depends only on the benchmark scenario and is applied identically to every method. It yields 45 × 5 = 225 configurations. Interpretation. The feasibility-aware subset is a secondary analysis and does not replace the full 315-configuration WebDecept evaluation. We report the full benchmark results for all methods and use the subset only to separate defense performance from scenarios whose benchmark ground truth does not specify a safe completion path after the deceptive condition is encountered. Subset membership depends only on the benchmark scenario and is applied identically to every method. We therefore interpret the full-set TSR and the feasible-subset TSR together: the former retains the benchmark’s original feasibility constraints, while the latter isolates configurations in which safe task completion is defined by the benchmark ground truth. D.4

BASELINE A DAPTATIONS

We compare Veer with the unprotected base agent, deceptive-interface-specific defenses, and general agent-safety defenses. Each method is adapted only as needed to operate over the native observation and action interface of the corresponding benchmark. No-Defense. tervention.

No-Defense uses the native benchmark agent without defense-side screening or in-

ICP Cuvin et al. (2026). ICP is instantiated as a static deceptive-interface warning inserted into the native actor context before action generation. Guardrail Cuvin et al. (2026). Guardrail performs deceptive-interface assessment over the current Web observation and supplies the resulting warning to the base agent. DUDE-S2 Zhang et al. (2026). We use the publicly reproducible Stage-2 defense, adapting its click-level review to the observation and click interfaces exposed by the two benchmark hosts. Spotlighting Hines et al. (2024). We apply Spotlighting by delimiting page content as untrusted input while preserving the native actor and interaction interface. 22

VIGIL Lin et al. (2026). We adapt VIGIL’s sanitize–guide–audit workflow to the browser-action representation of each benchmark. Proposed actions that are not released by the defense are returned to the base agent for replanning. SafePred Chen et al. (2026). We adapt SafePred’s public runtime policy to the native observation and action interfaces. It predicts the consequence and risk of a proposed action and invokes baseagent replanning when the proposal does not satisfy the policy. Adaptation scope. The adaptations are limited to mapping each defense to the observation, prompt, and action interfaces exposed by the two benchmark hosts. They do not add Veer’s stateintervention mechanism to the baselines or change their defense-side control point. The resulting implementations therefore preserve the distinction between prompt-level warning, action review, consequence screening, and state intervention that motivates the comparison. Comparison protocol. Within each benchmark, all compared methods use the same task configurations, base agent, actor model, and base task-action allowance. Defense-specific auxiliary reasoning follows the corresponding method, while ordinary task execution remains subject to the common base-agent budget. No compared method is given benchmark task-success labels, darkpattern labels, hidden evaluator outputs, or backend application state during runtime.

E

C OMPLETE R ESULTS AND E VALUATION P ROTOCOL

This appendix reports the complete results underlying the analyses in the main paper. We provide the raw outcome counts for the overall comparisons, the exact per-condition DPSR values underlying Figure 3, and the complete agent–model results corresponding to Table 3. E.1

E VALUATION P ROTOCOL

Evaluation units. TrickyArena-Single contains 88 task–condition configurations, TrickyArenaMulti contains 68 task–multi-condition configurations, and WebDecept contains 315 task–scenario configurations. The feasibility-aware WebDecept subset contains the 225 configurations from the five scenarios defined in Appendix D.3. Within each comparison, all methods are evaluated on the same configuration identifiers. Metrics. For configuration i, let Ti denote binary task success and Di indicate whether at least one evaluated deceptive outcome occurs. We compute N

DPSR =

N

100 X Di , N i=1

TSR =

100 X Ti , N i=1

and N

STC =

100 X Ti (1 − Di ). N i=1

Importantly, STC is computed from the episode-level joint outcome Ti (1 − Di ); it is not obtained by multiplying aggregate TSR by 1 − DPSR. We report DPSR and TSR separately so that safety and task utility remain directly visible. End-to-end evaluation. The reported metrics score the outcome of the complete runtime pipeline rather than treating intermediate LLM judgments as independent evaluation samples. Errors in consequence assessment or authorization therefore remain reflected in the final episode outcome: an unsafe proposal that is incorrectly released can increase DPSR, while an unnecessary intervention can reduce task success. This evaluation directly measures the downstream effect of the assessment procedure within the deployed defense. 23

Table 11: Complete TrickyArena-Single results (N = 88). Method #D #T #STC DPSR TSR STC No-Defense 26 71 53 29.5 80.7 60.2 ICP 15 71 61 17.0 80.7 69.3 Guardrail 24 71 55 27.3 80.7 62.5 DUDE-S2 30 76 50 34.1 86.4 56.8 Spotlighting 23 70 55 26.1 79.5 62.5 VIGIL 14 63 52 15.9 71.6 59.1 SafePred 25 76 56 28.4 86.4 63.6 Veer 4 76 75 4.5 86.4 85.2

Table 12: Complete TrickyArena-Multi results (N = 68). Method #D #T #STC DPSR TSR STC No-Defense 37 45 26 54.4 66.2 38.2 ICP 34 46 27 50.0 67.6 39.7 Guardrail 32 48 29 47.1 70.6 42.6 DUDE-S2 44 48 19 64.7 70.6 27.9 Spotlighting 40 41 18 58.8 60.3 26.5 VIGIL 27 40 26 39.7 58.8 38.2 SafePred 38 54 25 55.9 79.4 36.8 Veer 10 49 46 14.7 72.1 67.6

For TrickyArena-Multi, Di is the logical OR over the constituent deceptive conditions active in configuration i. Thus, an episode contributes one positive instance to DPSR when at least one constituent dark-pattern outcome occurs. The denominators are fixed at 88, 68, 315, and 225 for TrickyArena-Single, TrickyArena-Multi, full WebDecept, and the feasibility-aware WebDecept subset, respectively. E.2

C OMPLETE OVERALL R ESULTS

Tables 11–14 report the raw counts underlying Table 1 and Table 2 in the main paper. Here, #D is the number of configurations with a deceptive outcome, #T is the number with successful task completion, and #STC is the number that complete the task without a deceptive outcome. E.3

P ER -C ONDITION R ESULTS

Tables 15 and 16 provide the exact DPSR values underlying Figure 3. Each cell reports DPSR in percent followed by the raw number of deceptive outcomes in parentheses. On TrickyArena-Single, Veer records zero deceptive outcomes in 20 of the 22 conditions and achieves the lowest or tied-lowest DPSR in 21 conditions. Its unweighted mean per-condition DPSR is 5.7%, compared with its configuration-weighted overall DPSR of 4.5%. Across WebDecept, Veer records one deceptive outcome among 315 configurations. Its unweighted mean per-scenario DPSR is 0.3%, and its maximum scenario DPSR is 2.2%. E.4

AGENT–M ODEL R ESULTS

Table 17 reports the complete system-configuration results corresponding to Table 3 in the main paper. Each cell gives Veer’s STC followed by the absolute improvement over the matched NoDefense configuration in percentage points. Veer improves STC over the matched No-Defense setting in all 12 agent–model–benchmark comparisons, consistent with the robustness result reported in the main paper. 24

Table 13: Complete WebDecept results over all seven scenarios (N = 315). Method No-Defense ICP Guardrail DUDE-S2 Spotlighting VIGIL SafePred Veer

#D 118 84 39 113 112 30 120 1

#T #STC DPSR TSR STC 150 90 37.5 47.6 28.6 138 100 26.7 43.8 31.7 111 93 12.4 35.2 29.5 175 99 35.9 55.6 31.4 141 86 35.6 44.8 27.3 32 17 9.5 10.2 5.4 183 122 38.1 58.1 38.7 129 129 0.3 41.0 41.0

Table 14: Complete results on the feasibility-aware WebDecept subset (N = 225). Method #D #T #STC DPSR TSR STC No-Defense 40 113 90 17.8 50.2 40.0 ICP 13 107 100 5.8 47.6 44.4 Guardrail 16 102 93 7.1 45.3 41.3 DUDE-S2 38 122 99 16.9 54.2 44.0 Spotlighting 39 107 86 17.3 47.6 38.2 VIGIL 7 20 17 3.1 8.9 7.6 SafePred 44 145 122 19.6 64.4 54.2 Veer 0 129 129 0.0 57.3 57.3

F

A DDITIONAL A BLATION AND T RAJECTORY A NALYSIS

This appendix reports the complete RQ3 ablation results, an additional one-step-intervention ablation, and representative trajectory analyses. All ablations use TrickyArena-Single with N = 88 configurations per variant. Unless explicitly removed by an ablation, the variants retain the same user tasks, base agent, actor model, task-action allowance, and surrounding Veer runtime components. F.1

C OMPLETE A BLATION R ESULTS

Table 18 decomposes each episode into four mutually exclusive outcomes: safe task completion (T = 1, D = 0), task completion with a dark-pattern outcome (T = 1, D = 1), task failure without a dark-pattern outcome (T = 0, D = 0), and task failure with a dark-pattern outcome (T = 0, D = 1). State intervention. The w/o State Intervention variant replaces corrective state intervention with blocking. When Veer identifies an unauthorized consequence, the proposed action is rejected without issuing a corrective browser action. This substantially reduces task progress: TSR decreases from 86.4% to 51.1%, and STC decreases by 35.2 percentage points. Prospective rollout. The Reactive Intervention variant retains the same intervention objective, temporal evidence, preservation constraints, grounding, and runtime verification as Full Veer. It removes complete prospective rollout: Veer selects one currently grounded corrective transition, executes it, observes the resulting state, and determines the next transition only if further correction is required. TSR remains close to Full Veer, while DPSR increases from 4.5% to 6.8% and STC decreases from 85.2% to 80.7%. Temporal evidence. The w/o Temporal Evidence variant removes the retained temporal evidence used by Veer across protected interaction steps while preserving the current Web observation and the base agent’s normal interaction context. Removing this evidence reduces TSR from 86.4% to 80.7% and STC from 85.2% to 77.3%, showing the value of retaining consequence-relevant state across multi-step interaction. 25

Table 15: TrickyArena-Single DPSR by condition. Each cell reports DPSR (%) with raw #D in parentheses. Condition N No-Def. ICP Guard. DUDE Spot. VIGIL SafePred Veer shop/p1 4 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t1 4 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t2 4 75.0(3) 0.0(0) 25.0(1) 75.0(3) 25.0(1) 0.0(0) 100.0(4) 25.0(1) shop/t3 4 0.0(0) 0.0(0) 0.0(0) 25.0(1) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t4 4 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t5 4 0.0(0) 0.0(0) 0.0(0) 25.0(1) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t6 4 0.0(0) 0.0(0) 0.0(0) 25.0(1) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t7 4 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/t8 4 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/p2 4 100.0(4) 50.0(2) 100.0(4) 100.0(4) 100.0(4) 75.0(3) 100.0(4) 0.0(0) shop/s 5 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) shop/w 5 20.0(1) 0.0(0) 20.0(1) 20.0(1) 0.0(0) 0.0(0) 0.0(0) 0.0(0) news/bs 3 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) news/ob 3 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) news/sa 3 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) news/cf 3 100.0(3) 100.0(3) 100.0(3) 100.0(3) 100.0(3) 100.0(3) 100.0(3) 100.0(3) music/am 2 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) music/ds 3 0.0(0) 100.0(3) 0.0(0) 33.3(1) 0.0(0) 0.0(0) 0.0(0) 0.0(0) music/du 3 100.0(3) 0.0(0) 100.0(3) 100.0(3) 100.0(3) 66.7(2) 66.7(2) 0.0(0) health/cs 6 100.0(6) 100.0(6) 100.0(6) 100.0(6) 100.0(6) 100.0(6) 100.0(6) 0.0(0) health/tos 6 100.0(6) 16.7(1) 100.0(6) 100.0(6) 100.0(6) 0.0(0) 100.0(6) 0.0(0) health/cf 6 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0)

Table 16: WebDecept DPSR by deceptive scenario. Each cell reports DPSR (%) with raw #D in parentheses. Scenario Popup Banner P-popup P-banner Add-ons Redirect Drift

F.2

N No-Def. ICP 45 4.4(2) 4.4(2) 45 0.0(0) 0.0(0) 45 0.0(0) 0.0(0) 45 2.2(1) 2.2(1) 45 82.2(37) 22.2(10) 45 86.7(39) 80.0(36) 45 86.7(39) 77.8(35)

Guard. DUDE Spot. VIGIL SafePred Veer 6.7(3) 0.0(0) 6.7(3) 2.2(1) 6.7(3) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 0.0(0) 2.2(1) 2.2(1) 0.0(0) 0.0(0) 2.2(1) 2.2(1) 0.0(0) 2.2(1) 0.0(0) 28.9(13) 82.2(37) 77.8(35) 11.1(5) 86.7(39) 0.0(0) 48.9(22) 82.2(37) 82.2(37) 26.7(12) 84.4(38) 2.2(1) 2.2(1) 84.4(38) 80.0(36) 24.4(11) 84.4(38) 0.0(0)

A DDITIONAL O NE -S TEP A BLATION

The main Reactive Intervention ablation removes complete prospective rollout while still allowing successive corrective transitions to be constructed after observing each intermediate state. We further evaluate a stricter One-Step Intervention variant that permits only one local corrective browser action per interception and then returns control to the base-agent loop. Unlike Full Veer, this variant does not represent or verify a complete multi-step correction with cross-step dependencies. Its STC decreases from 85.2% to 61.4%, while DPSR increases from 4.5% to 14.8%. Together with the Reactive Intervention result, this shows that preserving the structure of a multi-step correction is important when safe intervention requires dependent state transitions. F.3

I NTERVENTION -T RAJECTORY A NALYSIS

We next examine the prospective trajectories admitted by Full Veer to characterize the structure of state intervention in practice. Among episodes in which Veer admits an intervention, 15 use a one-transition trajectory and six use a four-transition trajectory. Across the 24 admitted trajectories, the mean planned length is 1.88 transitions, the median is one, and the maximum is four. The presence of four-transition trajectories confirms that the evaluated interventions include dependent multi-step corrections rather than only isolated local actions. The executed corrective actions span four planner-level operation types. These operations cover both direct state reduction, such as disabling or declining an unwanted state, and transitions that commit an already constructed correction, such as saving the resulting settings. 26

Table 17: Complete STC results across agent–model configurations. Parentheses denote absolute improvement over the matched No-Defense configuration in percentage points. Agent Default Default Codex Codex

Model GPT-5.6-Luna DeepSeek-V4-Flash GPT-5.6-Luna DeepSeek-V4-Flash

TrickyArena-Single 85.2 (+25.0) 68.2 (+12.5) 75.0 (+28.4) 85.2 (+20.4)

TrickyArena-Multi 67.6 (+29.4) 44.1 (+10.3) 54.4 (+25.0) 55.9 (+14.7)

WebDecept 41.0 (+12.4) 40.0 (+4.8) 45.4 (+5.1) 53.7 (+15.9)

Table 18: Complete RQ3 ablation results on TrickyArena-Single (N = 88 per variant). ∆STC is measured relative to Full Veer. Variant Safe T+DP Fail, no-DP Fail+DP DPSR (%) TSR (%) STC (%) ∆STC (pp) Full Veer 75 1 9 3 4.5 86.4 85.2 – w/o State Intervention 44 1 34 9 11.4 51.1 50.0 −35.2 Reactive Intervention 71 4 11 2 6.8 85.2 80.7 −4.5 w/o Temporal Evidence 68 3 15 2 5.7 80.7 77.3 −8.0

F.4

R EPRESENTATIVE I NTERVENTION T RAJECTORIES

We provide two representative episodes that illustrate state correction with task preservation and guarded verification with bounded replanning. Case 1: consent correction with task preservation. In configuration health tos 80, the user asks the agent to retrieve the date of the last flu shot and write it to the scratchpad. The base agent proposes an action that would commit to broad health-data consent unrelated to the requested task. Veer constructs an intervention objective that removes the pending consent while preserving the task-relevant scratchpad capability. The prospective trajectory contains a single corrective transition: After the correction, Veer verifies that the consent state has been removed while the task-relevant capability remains available. Control returns to the base agent, which retrieves the requested medicalrecord information and writes it to the scratchpad. The episode finishes with T = 1, D = 0. This case illustrates the role of Pt : state correction removes the unauthorized consequence while retaining state needed for the original task. Case 2: terminal verification and bounded replanning. In configuration shop p2 38, the user asks for the description of a laptop. The base agent proposes accepting cookie consent that is unnecessary for the requested information-retrieval task. Veer constructs an objective that prevents committing this consent while preserving product retrieval and scratchpad-entry capability. The initial prospective correction opens the cookie-options interface. The browser action succeeds locally, but the resulting state does not yet satisfy the intervention objective. Veer therefore retains the same objective and constructs a new corrective continuation: After the corrected state is verified, control returns to the base agent, which resumes product retrieval and writes the requested description to the scratchpad. The episode finishes with T = 1, D = 0. This case illustrates why Veer verifies the resulting state against the intervention objective after browser execution: a locally successful interaction does not by itself establish that the intended state correction has been completed.

G

S COPE AND L IMITATIONS

Veer is designed as an agent-side runtime defense for Web-agent execution under deceptive interfaces. Its protection operates over the same black-box browser interface available to the base agent: it reasons from observable Web evidence, intervenes through grounded browser interactions, and derives authorization from the original user task. This section summarizes the resulting applicability boundary. 27

Table 19: Additional one-step-intervention ablation on TrickyArena-Single (N = 88). Variant DPSR (%) TSR (%) STC (%) ∆STC (pp) Full Veer 4.5 86.4 85.2 – One-Step Intervention 14.8 68.2 61.4 −23.9

Table 20: Length of admitted prospective intervention trajectories. Planned length K 1 4 Total

Episodes 15 6 21

Trajectories 17 7 24

Observable-state dependence. Veer constructs task-relevant Web state from the current observation and retained temporal evidence. Consequently, an intervention must be grounded in state that has an observable manifestation in the Web interaction. Hidden application state that cannot be inferred from available Web evidence does not directly participate in consequence assessment or intervention planning. Accordingly, Veer targets consequences whose relevant pre-commit state has an observable Web manifestation; effects determined entirely by hidden server-side state with no observable evidence fall outside the current intervention model. Browser-reachable correction. State intervention requires a browser-level path from the current state toward the target state. Dynamic interfaces, unavailable controls, or unexpected transition effects can invalidate a prospective trajectory. Veer addresses these cases through live grounding, dependency checks, and post-transition verification, and replans under the same intervention objective when a valid continuation remains available. Task-grounded authorization. Veer derives authorization from the original user instruction and does not expand it using Web content encountered during execution. Runtime observations can ground task references to concrete objects and facts, while the authorization boundary remains fixed. Tasks whose intent is insufficiently specified can therefore leave some proposed consequences uncertain. Task-state preservation. The preservation component of the intervention objective covers taskrelevant state and capabilities that can be identified from the interaction evidence available to Veer. Guarded execution verifies these declared preservation requirements together with the target corrected state before returning control to the base agent. Base-agent dependence. Veer intervenes on consequence-causing Web state and then returns control to the base agent. It does not replace the base agent’s ordinary task planner. Navigation, information retrieval, and completion of the remaining user task therefore continue to depend on the underlying agent after a successful intervention. Evaluation scope. Our evaluation covers TrickyArena-Single, TrickyArena-Multi, and WebDecept, spanning multiple deceptive-interface mechanisms, application domains, agent implementations, and model configurations. These experiments establish the effectiveness of consequenceguided state intervention in the evaluated settings; extending the evaluation to additional Web environments and interaction settings remains future work.

28

Table 21: Executed corrective operations in Full Veer. Corrective operation turn off decline save click Total

Executed actions 20 12 6 4 42

Table 22: Representative consent correction for health tos 80. Step τ1

Operation decline consent

Expected state declined

Dependency –

Table 23: Bounded intervention replan for shop p2 38. Attempt Initial Replan

Corrective transition Open cookie options Decline cookie consent

29

Result Objective unresolved Verified

Record · ID 1108627 · SHA-256 c150c0adb21c7897
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.