JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
1
Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout
arXiv:2604.09056v1 [cs.CR] 10 Apr 2026
Xiaotong Jiang Student Member, IEEE, Jun Wu Senior Member, IEEE
Abstract—With the rapid adoption of large language models (LLMs) in financial service scenarios, dialogue security detection under high regulatory risk presents significant challenges. Existing methods mainly rely on single-dimensional semantic judgments or fixed rules, making them inadequate for handing multi-turn semantic evolution and complex regulatory clauses; moreover, they lack models specifically designed for financial security detection. To address these issues, this paper proposes FinSec, a four-tier security detection framework for financial agent. FinSec enables structured, interpretable, and end-to-end identification of actual financial risks, incorporating suspicious behavior pattern analysis, delayed risk and adversarial inference, semantic security analysis, and integrated risk-based decisionmaking. Notably, FinSec significantly enhances the robustness of high-risk dialogue detection while maintaining model utility. Experimental results demonstrate FinSec’s leading performance. In terms of overall detection capability, FinSec achieves an F1 score of 90.13%, improving upon baseline models by 6–14 percentage points; its ASR is reduced to 9.09%, markedly lowering the probability of unsafe outputs; and the AUPRC increases to 0.9189—an approximate 9.7% gain over general frameworks. Additionally, in balancing utility and safety, FinSec obtains a composite score of 0.9098, delivering robust and efficient protection for financial agent dialogues. Index Terms—IEEE, IEEEtran, journal, LATEX, paper, template.
I. I NTRODUCTION ITH the rise of LLM-based (Large Language Modelbased) intelligence, its application is also rapidly entering financial-related industries to assist financial practitioners and users in completing financial business operations more efficiently. As schedulable tools, the financial agent has the ability to execute tasks and reasoning at the same time. It is used to invest in wealth management assistance, automation of financial processes, compliance support for transaction communication, and has undergone large-scale internal testing and a limited production environment[1]. For example, Microsoft Copilot for finance agents has embedded accounting processes such as reconciliation and difference analysis into Excel and financial systems[2]; Morgan Stanley has launched GPT-4-based Assistant/Debrief and other tools for financial consultants to automatically summarize meetings and generate action items [3]. The generative AI of many large banks uses agents that are more biased towards internal employees, rather than full automation for end customers, to reduce hallucinations and compliance risks. Unlike the general agent scenario, the errors and insecure output of financial-related agents will bring more direct and
W
Xiaotong Jiang and Jun Wu are with the Graduate School of Information, Production and Systems, Waseda University, Fukuoka 808-0135, Japan
significant consequences. For example, induced fraud, privacy data leakage, over-powerful operation execution, and even trigger regulatory compliance issues. The dialogue content not only carries general semantic information, but also directly relates to high-value operations such as transfer, increase, and asset change. Once the model is injected with hints at the natural language level, confrontational manipulations such as social engineering attacks or county-wide promotions, its error response will be executed by the downstream system, bringing real property losses or compliance risks. Therefore, identifying and blocking insecure behavior at the dialogue level has become a precondition and bottom-line requirement for deploying LLM-based financial agents. In the financial agent dialogue scenario, users often bypass the compliance check through requests such as step-by-step talk and spin-off transactions. This kind of risk will no longer appear in a single round of dialogue but will gradually accumulate in multiple rounds of dialogue. Therefore, the existing ordinary dialogue security detection methods cannot meet the multi-round high compliance requirements of financial scenarios. In the context of financial security research, existing dialogue-safety detection methods are insufficient to guarantee the security of financial interactions involving AI agents. The current security detection method primarily focuses on identifying explicit harmful content or defending against the prompt injection attack, as described in the system side. However, such methods struggle to cope with financial scenarios, multiple rounds of semantic trajectories, delayed risk manifestation, and field-specific regulatory constraints. Therefore, we propose an enhanced confrontational detection framework for financial intelligence, combining suspicious behavior pattern detection, delay risk simulation, and confrontational semantic analysis, to realize forward-looking and multidimensional control of financial agent dialogue risks. As shown in Fig. 1, the framework integrates an adversarial thinking framework (right-top) with the multi-layer FinSec architecture (bottom) to identify unsafe instructions from raw agent dialogues (left). In summary, our contributions are as follows: 1) We design a SAR-style suspicious activity detection mechanism tailored to financial dialogue scenarios, as well as a multi-level "triple matching" that automatically identifies high-risk clues across multiple rounds of agent dialogue and provides structured, basic signals for subsequent risk assessment. 2) We develop a deferred risk assessment and confrontational inference module for financial agent dialogues. In view of the unique characteristics of lag and accumulation of financial risks, we design the delayed risk assessment layer
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
Fig. 1. Overview of the Financial Agent Risk Detection Framework
and include historical transaction models in the measurement. At the same time, in the constructed confrontational inference module, the simulated attacker’s progressive-induction attack method across multiple rounds of dialogue addresses the limitation of relying solely on static single-round detection risks. 3) We propose FinSec, a multi-layer financial dialogue safety framework built on LLMs. Based on the abovementioned behavioral risk, delay risk, and confrontational inference, FinSec, a multi-layer security framework for financial agents, has been built. By treating the LLM as the core of semantic-level security, it enables multi-stage generative security inference. The remainder of this paper is organized as follows: Section 2 reviews related work; Section 3 presents the underlying model selection and the development of FinSec; Section 4 details the experimental results; and Section 5 provides the concluding remarks. II. R ELATED W ORK A. Agent Interaction records security research Currently, research on agent dialogues and their security is primarily concentrated in four areas. The first is defense against external perception and indirect injection. External perception serves as a unique entry point in the agent security chain, where potential attack origins expand from user input to environmental input. Attackers may conceal prompts within webpages, allowing agents to become compromised simply by browsing such content. Work such as [4] has demonstrated how the agent architecture can transform simple textual prompt injections into substantial system compromise. Kai et al. [5] defined indirect prompt injection and illustrated an attack chain in which an agent can be controlled merely through web browsing. HouYi et al. [6]proposed an injection attack framework targeting black-box agent applications, revealing that application-layer logic can be exploited for attacks, even without knowledge of model parameters. The second area is defense against role inconsistency and deception. This primarily concerns the cognitive security of agents. In high-risk scenarios such as finance, agents must remain honest and strictly adhere to their role boundaries. Recent research [7] has explored methods for defending against
2
agent jailbreaks to prevent the induction of non-compliant responses. Wang et al.‘s [8] Decoding Trust framework reveals the strong tendency of mainstream models to deviate from their designated roles under adversarial manipulation. Hubinger et al. [9] described ‘sleeper agents’ that may appear compliant with safety alignment during training but can activate malicious behavior under specific triggers. Although Bian et al. [10] have sought to improve an agent’s resistance to jailbreak instructions through safety fine-tuning, static role defense mechanisms remain inadequate when confronted with agents possessing long-term memory and complex strategies. The third area concerns tool misuse and action space control. After conducting dialogues, agents may invoke APIs to perform practical operations such as fund transfers. Thus, risks can arise if there is a discrepancy between the agent’s verbal commitment and the actual API calls made. To assess such risks, the ToolEum sandbox environment has been introduced as a standard for evaluating the safety of agent tool usage, specifically detecting high-risk operations that contradict an agent’s stated intentions [11]. DeepMind’s research provides a detailed taxonomy of the possible side effects of tool use. Although LLMs excel at code generation, they still exhibit substantial shortcomings in vulnerability assessment and localization [12]. R-judge [13] specifically evaluates the risk awareness of agents, investigating the ability of LLMs to detect and identify security risks based on agent interaction histories. Their benchmark data spans 27 critical risk scenarios across five application domains, including finance, revealing that risk awareness in open agent environments is a multidimensional capability. These works collectively emphasize the necessity of implementing independent monitoring and blocking mechanisms prior to agent output, separate from the core model logic. The fourth area relates to contagion in multi-agent collaboration. Here, attack instructions may spread among agents like viruses, so that compromise of a single agent could cripple the entire network. Such risks can occur when multiple agents interact and collaborate. The AgenSmith study [14] demonstrates that malicious prompts possess high transmissibility in agent networks. Erisken et al. [15] conducted an in-depth analysis of failure modes in agent dialogues, including ineffective argument loops, collective bias, and consensus catastrophe among agents. B. Specific Security Issues in Financial Dialogue Systems In the context of financial domain dialogue systems, research on security has begun to develop. Due to the unique characteristics of finance, security risks pertaining to financial dialogues are concentrated in three main areas: safe alignment and knowledge boundaries in domain adaptation, sensitive information identification and de-identification, and compliance control in financial dialogues. The first challenge involves the difficulty in clearly delineating safe alignment and knowledge boundaries. The core issue lies in ensuring the continuous integration of financial knowledge while maintaining appropriate boundaries, thereby preventing high-risk recommendations caused by overconfidence or erroneous knowledge
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
transfer. Thakkar et al. proposed MERGEALIGN [16], which interpolates between general alignment vectors and domainspecific vectors. This approach strikes a balance in highrisk domains such as healthcare and finance, achieving neargeneral safety baselines without noticeably sacrificing domain task performance. It provides a new technical pathway for ‘post-hoc alignment.’ In the FinBen benchmark, Zhang et al. observed that excessive fine-tuning may compromise existing general safety safeguards, making models more susceptible to responding to malicious prompts [17]. With respect to sensitive information identification and de-identification, financial dialogue systems must not only avoid providing improper recommendations but also directly process highly sensitive data such as account information and transaction records. As a result, memory leakage and privacy attacks have become central security concerns in this domain. Unlike the role-playing deception seen in general contexts, financial deception often involves subtle profit-driven inducements. Li et al. [18] proposed a misleading information detection framework specifically tailored for financial texts, employing adversarial training to identify fraudulent intent concealed within technical jargon. For the issue of ‘hallucinated lies’ potentially generated by agents, the FinRobot framework [19] integrates a fact-checking module that crossverifies agent outputs through real-time retrieval of market data. Compliance control in financial dialogue is another defining feature that distinguishes financial systems from those in other fields. In applications such as investment advising or loan approval, dialogue models must rigorously adhere to financial regulations and industry standards, robustly rejecting prompts that cross legal or ethical boundaries even in complex contexts. Datasets like FinanceBench evaluate models’ comprehension and citation abilities in large-scale financial report QA tasks, thereby indirectly reflecting their capacity to generate substantiated recommendations [20]. Chen et al. [21] introduced a set of risk-aware assessment metrics for financial agents, addressing the difficulty of capturing behavioral risks in the agent execution chain—a limitation of traditional accuracybased financial benchmarks. Their Secure Assessment Agent (SAEA) continuously tracks the evolution of risk throughout the dialogue trajectory. The studies mentioned above predominantly focus on LLMs or traditional financial dialogue systems, yet there remains a lack of systematic research on the security of financial agent dialogues. Thus, a comprehensive detection framework for financial dialogue remains absent. The particularity of financial agents lies in their ability not only to provide answers but also to execute genuine financial operations, such as modifying account balances. This underscores the urgency of developing a systematic approach to the security of financial agents.
3
detection of overreach or decision manipulation [22][23]. For harmful content detection, studies have explored the risks of static content [24][25], developed toxic content classifiers such as Beaver Dam-7B [26], and trained multi-class classifiers to determine whether responses contain hate speech, violence, or privacy violations. Constitutional AI approaches incorporate self-moderation principles, adding ethical guidelines during reinforcement learning to train models to automatically reject harmful outputs [27]. For the defense and detection of prompt injection attacks, methods include building static attack pattern classifiers—trained on known injection cases to identify suspicious inputs and attention distribution analysis, which detects deviations in model attention from the original instructions during response generation [28][29]. Other research highlights multi-layer interception and arbitration agents, introducing real-time arbiters that monitor multi-turn conversations for injection or unauthorized actions [30]. For hallucinated and misleading information, research emphasizes multi-model consistency debates [31][32], consistency-based self-verification, and semantic entropy metrics [33][34]. Sensitive information leakage detection focuses on query-based PII filtering and neuron-level memory erasure [35][36]. For overreach and role deception detection, efforts target the identification of unauthorized behaviors and the use of Guardian Agents applying the principle of least privilege [30][37]. Thus, it is evident that existing research lacks sufficient attention to the multi-turn dialogue evolution characteristic of financial contexts, and fails to integrate compliance regulations specific to finance (such as SAR, KYC, AML) into agent dialogue detection frameworks. This study addresses this gap by proposing a multi-layer detection framework, FinSec, tailored for financial agents. FinSec accounts for suspicious behavior pattern recognition, delayed risk simulation, and adversarial expression analysis, thereby offering security assurances aligned with regulatory requirements for financial applications. III. P ROBLEM F ORMULATION A. Base Model In order to ensure that it can be gradually optimized on a basic model with the best effect, first of all, we compare the ability of more than ten advanced LLM to detect the risk of financial agent data. The data set adopts the financial part agent dialogue data in R-judge [13]. As a risk assessment benchmark for LLM-based agents, R-judge has conducted 5 field classifications, including risk assessment and manual labeling of 27 major risk scenarios, and is a pioneer in agent risk detection. Therefore, we take this Zoro-shot-CoT prompt as the baseline model, calculate the LLM with the best detection effect in the newly launched LLM models of each company, and carry out follow-up tests on this basis. B. Adversarial Reasoning Framework
C. Dialogue Agent Security Detection Methods Existing research on dialogue agent security detection predominantly focuses on the following areas: harmful content detection, prompt injection defense, hallucinated and misleading information identification, sensitive information leakage, and
To identify adversarial intent across multiple rounds of dialogue, we introduce an adversarial reasoning framework to address the inability of LLM-based agents to recognize inputs that appear reasonable but carry malicious intent. The existing security monitoring scheme relies primarily on single-layer
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
4
text classifiers or rule-based matching. It can only statically differentiate between the current round of dialogue and a single round of dialogue. For financial agents, we need to face and solve two more realistic problems. The first is whether the regular dialogue in the current round may evolve into a highrisk trading scenario in the next round. And secondly, complex and roundabout offensive rhetoric, such as stepwise decomposition of requests, obfuscated expressions, and encoding-based evasion, needs to be systematically identified and quantified across turns. Based on this, we have introduced adversarial analysis in the security monitoring of financial agent dialogues. Adversarial analysis is typically used to assess the robustness and security of machine learning models [need ref]. After our optimization and improvement, we use it to construct high-risk scenarios via adversarial analysis of the agent and to evaluate the effectiveness of the defense mechanism quantitatively. Formally, we denote a dialogue by d = (c1 , c2 , . . . , cT ), where cT represents the interaction at turn t, introducing a binary random variable Y (d) ∈ {0, 1}. When Y (d) = 1, the dialogue is unsafe under financial security standards; when Y (d) = 0, it is safe. Therefore, under the adversarial assumption, the risk probability of financial dialogue in the presence of an attacker is shown in Eq.(1). Radv (d) = Padv (Y (d) = 1 | d)
(1)
Where Padv (·) indicates the conditional probability under the adversarial distribution, assuming that the user may have malicious intentions, the probability that the dialogue d is judged as unsafe. For the detection of the security of financial agent dialogue, we have established a three-perspective framework including the perspectives of defenders, attackers and red teams based on adversarial analysis. Specifically, the adversarial risk formula is shown in Eq.(2) R̂Adv (d) = σ(w⊤ s(d)) = σ wdef sdef (d) + watt satt (d) + wred sred (d)
(2)
Where w⊤ = [wdef , watt , wred ]is the weight coefficient of each perspective. Use the security defender’s perspective to check the obvious risks that have been exposed in the dialogue, and then think about the potential weaknesses that can be exploited in the dialogue from the attacker’s perspective, and finally try to combine and amplify these weaknesses from the perspective of the red team to find hidden and complex attack paths. We will record the dialogue risk score from three perspectives and then weigh it to describe the overall adversarial risk. The proposed adversarial scenario generation process is formalized in Algorithm 16. In order to ensure the adaptive and rigorous evaluation, we operate the algorithm into two phases. In the first phase, we set the framework dynamically selects the generation strategy by comparing the current suspiciousness scores against predefined thresholds.This stratification ensures computational resources are focused on the most critical conversational contexts. Following this, we evaluate each generated scenario from three distinct perspectives: the defender checks for policy com-
Algorithm 1: Adversarial Scenario Generation Input: Current conversation C, suspicious score ssus , deferred risk score sdef , user profile U , thresholds θhigh , θmed . Output: Adversarial scenario set S, risk trajectory T , rollout risk score rrollout , confidence score c. 1
Initialize: S ← ∅, T ← ∅;
// Phase 1: Risk Assessment & Scenario Selection 2 if ssus > θhigh ∨ sdef > θhigh then 3 S ← G EN H IGH R ISK S CENARIOS(C, U ); 4 else if ssus > θmed ∨ sdef > θmed then 5 S ← G EN M ED R ISK S CENARIOS(C, U ); 6 else 7 S ← G EN L OW R ISK S CENARIOS(C); // Phase 2: Three-Party Adversarial Loop 8 foreach scenario si ∈ S do 9 adef ← D EFENDER A NALYSIS(si ) ; // Policy check 10 aatt ← ATTACKER A NALYSIS(si ) ; // Exploitability 11 ared ← R ED T EAM A NALYSIS(si ) ; // Hidden vuln. 12 patt ← E XTRACT PATTERNS(adef , aatt , ared ); 13 Append C ALC R ISK(patt ) to T ; rrollout ← max(T ); c ← C ALC C ONFIDENCE(T , S); 16 return S, T , rrollout , c; 14
15
pliance, the attacker probes for direct exploitability, and the red team identifies hidden vulnerabilities. Finally, the algorithm aggregates these insights to compute the risk trajectory T and determines the maximum rollout risk rrollout , providing a worst-case safety estimation. C. Financial Agent Safety Detection Framework In the financial services sector, conversational agents are increasingly facing security risks unique to the field. Unlike general dialogue security issues, the financial field involves strict compliance constraints and potential confrontational behavior. At the same time, many high-risk behaviors do not appear immediately in a single round of dialogue, and singledimensional detection cannot portray such complex risks in a timely and comprehensive manner. Therefore, we have established a multi-level, explainable financial security detection framework, FinSec, tailored to the characteristics of finance, forming an end-to-end data flow. FinSec follows the Risk-Based Approach RBA principle and divides the whole into four complementary levels. As illustrated in Fig. 2, we construct the detailed architecture of the proposed FinSec framework, which processes input agent dialogues through a hierarchical four-layer mechanism to handle the complexities of financial agent interactions.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
While layer 1 handles immediate structural compliance checks using SAR patterns, the system’s novelty lies in its advanced reasoning layers. Layer 2 is explicitly designed to counter delayed risk manifestation by simulating future dialogue trajectories through Adversarial Generative Rollout. This allows the system to evaluate risks that may not be apparent in a single turn. Complementing this, Layer 3 provides a deep semantic audit using few-shot learning to interpret subtle intent. The outputs are fused in Layer 4, ensuring that the final risk score RF inSec (I) reflects both immediate compliance violations and potential long-term vulnerabilities. At the bottom, suspicious behavior patterns are depicted based on anti-money laundering (AML) and suspicious activity reports (SARs) to capture quantitative indicators of suspicious activity. At the intermediate level, the potential risks that "have not occurred but have a certain possibility" are evaluated from a time dimension perspective through delayed risk simulation and confrontation analysis. At the semantic level, use a large language model to assess dialogue security and uncover risk clues hidden in complex dialogues. At the decision-making level, the contents of the previous layers are integrated and calibrated to obtain standardized risk scores and discrete safety indicators. It is used to monitor better the dangers that financial agent dialogue may face. In this chapter, we will talk about the specific design of each layer of FinSec. 1) Suspicious Behavior Pattern Detection: Based on the international standard security general requirements AML applicable to financial institutions, among which the SAR, as the basis of our model, refines its function of monitoring transaction patterns and automatic reporting to identify common risk trajectories in financial agents "need to be cited". In this way, the test results with a high recall rate can be obtained effectively at a low cost. The "red flag" mechanism in anti-money laundering practice is formalized into a pattern library, and the triple matching of "keyword/slot hit-semantic similarity-sequence consistency" is used to generate a "suspicious trajectory" across multiple rounds of dialogues. In layer 1, based on the specific and different data contents, we will calculate the warning signals and patterned characteristics of suspicious behavior, as indicators of risk, and output the detection results for subsequent layer 2 and layer 3 to be called for subsequent calculation. Here, Formally, we abstract risk indicators into a pattern library P = {pk }K k=1 . The dialogue is segmented into sliding windows denoted as Wt , which are used to calculate vocabulary, semantic similarity, and sequence cons respectively, and the customer portrait and jurisdiction parameters are weighted according to the risk orientation (RBA), and then calculate the suspiciousness of time accumulation. As a stable a priori layer, suspicious behavior pattern detection not only reduces the cost of calculation but also enhances the reliability of risk assessment. As shown in Eq.(3), we define the pattern matching score Mk (XWt ), which consists of three parts: keyword hit rate Hit k , semantic similarity Sim k , and sequence consistency Order k . Hit k measures whether high-risk keywords or key slots corresponding to a pattern appear within the window, serving as a formalization of the “red flag signals” in the SAR mechanism. Nk represents the set of semantic nodes
5
in the pattern, and Ek represents the order structure among those nodes. The more matches there are, the more likely this window contains a typical suspicious behavior.
Mk (XWt ) = λ1 Hitk + λ2 Simk + λ3 Orderk , λ1 + λ2 + λ3 = 1
1 X 1 X 1{v ⊂ x} Hitk = min 1, |Wt | |Vk | x∈Wt
(3)
! (4)
v∈Vk
1 X Simk = (5) σ τ · max cos(ϕ(x), ϕ(u)) x∈Wt |Nk | u∈Nk X 1 1{t(u) < t(v)}e−β|t(u)−t(v)| (6) Order = k |Ek | (u→v)∈Ek
Here,as shown in Eq.(4), Wt denotes all dialogue turns within the current sliding window, Vk represents the set of keywords associated with pattern pk , and 1{v ⊂ x} is an indicator function: it is 1 if keyword v appears in the text x, and 0 otherwise. Sim k calculates the overall semantic alignment of the window with the pattern by comparing the semantic vectors of pattern nodes with those of all text segments within the window and selecting the maximum semantic similarity, as given by Eq.(5). This is used to capture more implicit risk expressions, such as circumvention via indirect language or semantic shifts caused by less significant vocabulary. Here, u ∈ Nk represents a semantic node within the pattern, and cos(ϕ(x), ϕ(u)) denotes the cosine similarity between the text segment and the pattern node. By calculating maxx∈Wt , we identify the text segment within the window that is most semantically similar to node u. Equation.(6) illustrates that Order k uses the set of directed edges, Ek , within the pattern to check whether relevant semantic units in the window appear in an order consistent with suspicious behavioral trajectories. A decay weight for temporal distance is incorporated, where t(u) denotes the position in the window of the text segment most similar to node u, and β is the temporal decay coefficient—the larger this value, the greater the emphasis on close-proximity matches. By combining these three components, we arrive at the overall pattern score Mk (XWt ), for Layer 1. As a robust prior layer, suspicious behavioral pattern detection not only reduces computational cost but also increases the reliability of risk interpretation. 2) Deferred Risk Assessment and Adversarial Rollout: Building on the suspiciousness scores outputted by layer 1, layer 2 incorporates industry standards by referencing FATF(Financial Action task Force) guidelines as well as the KYC/AML framework. We introduce user profiles and historical semantic manipulation[need references]. The goal is to estimate the probability of risk escalation into future states within the same probabilistic semantic context.By using
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
6
Fig. 2. Detailed architecture of the FinSec framework. The system operates through a hierarchical data flow: (1) SAR Pattern Detection for structured compliance checking; (2) Deferred Risk Assessment via generative rollout; (3) Semantic Safety Assessment using deep audit models; and (4) Risk Fusion for the final calibrated decision RFinSec (I).
small-sample simulations to reveal potential high-risk trajectories. With this approach, when real user data needs to be analyzed, only the actual data is required as input. Therefore, by integrating the features mentioned above, we quantify the deferred risk Qdef t as a function f of the suspicious score St , user profile features Ut , and historical behavior features Ht , formulated as Qdef t = f (St , Ut , Ht ). Building on this, to model the propagation efficiency of risk from the individual level to the system level, we employ an exponential decay function to define the risk propagation probability, as illustrated in Equation pt: Pt = 1 − exp −λ · Qdef t
(7)
Where Pt represents the risk propagation probability at time t, λ denotes the propagation rate parameter, and Qdef is the t estimated delay risk value calculated in Equation (pt)." When medium to high levels of suspicion or divergence arise, we trigger a generative rollout simulation in Layer 2. The specific formula for the delay risk score is shown in Eq.(8): (n) (n) r(n) = max α′ St+k + γ ′ ADVt+k , 1≤k≤K (8) N 1 X (n) DR (t) = r rollout N n=1 We sample N future dialogue paths of length K from (n) the current state, denoted as {Xt+1:t+K }. Within Layer 2, (n) we evaluate the adversarial intensity, ADVt+k , of each path using a two-stage few-shot prompting approach. Simultaneously, following the rules from Layer 1, we recalculate the (n) behavioral risk, St+k , along the sampled paths. For each path, we take the ’maximum risk within the future window’ as the path score.DRrollout (t) represents the overall metric for the simulation-based delay risk. At time t, it is calculated by averaging the risk scores across all rollouts. Thus, in Layer
2, we have achieved the integration of online estimation and adversarial simulation. 3) Semantic Safety Assessment: In the third layer, we perform security assessment of the dialogue text itself. By leveraging an optimized prompt with a LLM, we conduct semantic analysis to identify semantic ambiguities, contextdependent security risks, and deep semantic risks that are difficult for traditional systems to detect. In practice, financial fraud often involves subtle expressions, multi-turn setups, and topic shifts to obscure malicious intent, making it difficult to promptly detect hidden risks using only keywords and transaction features. To address this, we introduce a large language model as a semantic discriminator at this layer, encoding the complete multi-turn conversation C together with a small set of annotated financial security examples S as the prompt, so that the model can infer Safe/Unsafe under the conditional distribution pθ (y | C, S). Specifically, we describe the architecture of Layer 3 in two stages: semantic reasoning and security judgment, as illustrated in Equation (11). pθ (y | C, S) = LLMϕ (y | Prompt(C, S)) R = fθreason (C, S) ŷ = fθpred (C, S, R) ∈ {Safe, Unsafe}
(9) (10) (11)
First, fθreason (C, S) generates an audit-oriented risk analysis text, systematically uncovering potential clues of fraud, money laundering, and unauthorized operations. Subsequently, within the same context, fθpred (C, S, R) is invoked to force the output of a discrete label ŷ ∈ {Safe, Unsafe}. Together, these components constitute the semantic layer of FinSec. 4) Risk Fusion and Calibration: In the fourth layer of FinSec, the results from the previous three layers are combined to determine the final safety judgment. Specifically, we calculate the score from each layer: RSAR (I) for risk based
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
7
on suspicious behavior patterns, RADV (I) for risk related to delayed and adversarial threats, and RSEM for risk assessed through semantic safety evaluation, as shown in Eq.(12).
Algorithm 2: FinSec: Multi-Layered Adversarial Risk Assessment Input: Context Ct , User Ut , History Ht , Thresholds τroll , τblock . Output: Final Risk RFIN (t), Decision Dt , Propagation RFIN (T ) = α RSAR (T ) + β RADV (T ) + γ RSEM (T ) (12) Pt . We have determined the default weights for each layer in 1 Phase 1: Pattern Detection (Layer 1) Layer 4. Since Layer 3, which relies on large language model- 2 Mk (XWt ) ← λ1 Hitk + λ2 Simk + λ3 Orderk based semantic assessment, can capture high-level attack in- 3 RSAR (t) ← Aggregate({Mk }k ) tents hidden in complex, multi-turn dialogues, it is assigned 4 Phase 2: Deferred Risk Simulation (Layer 2) −λQdef t the highest weight (γ = 0.75) to emphasize the importance of 5 Qdef t ← f (RSAR (t), Ut , Ht ); Pt ← 1 − e and S ← ∅ semantic safety. Layer 1 and Layer 2 provide complementary 6 Initialize RADV (t) ← Qdef adv t structural evidence and delayed risk information; if their 7 if Qdef t ≥ τroll then weights are too high, they may exacerbate local false positives, 8 Sadv ← GenScenarios(Ct | Qdef t ) PN (n) (n) 1 so α and β are controlled within the range of 0.10˘0.15 to 9 DRroll (t) ← N n=1 maxk (α′ St+k + γ ′ ADVt+k ) serve mainly as calibration and correction factors. Through 10 RADV (t) ← max(Qdef t , DRroll (t)) ablation experiments comparing various weight settings, we 11 end determined that this ratio achieves a better balance between 12 Phase 3: Semantic Safety Analysis (Layer 3) Area Under the Precision-Recall Curve (AUPRC) and Attack 13 ŷ ← LLMϕ (Ct , Sadv ); RSEM (t) ← ScoreMapping(ŷ) Success Rate (ASR), and thus is adopted as the default 14 Phase 4: Adaptive Fusion (Layer 4) weighting scheme in FinSec’s Layer 4 calculations. 15 if ∃Inj ∈ Sadv ∨ ŷ = Unsafe then To systematically illustrate the overall structure of the Fin- 16 w ← [αadv , βadv , γadv ]T Sec model we developed, we provide an explanation as shown 17 else in Algorithm 2. This process ensures that local risk estimates 18 w ← [αbase , βbase , γbase ]T , ŷ) are continuously calibrated and propagated (RSAR , Qdef t 19 end within the system, thereby enabling precise identification of 20 RFIN (t) ← αRSAR (t) + βRADV (t) + γRSEM (t) complex financial attacks within a unified risk space. 21 Dt ← I(RFIN (t) > τblock ) 22 return RFIN (t), Dt , Pt , Sadv IV. E XPERIMENTS AND RESULT A. Original LLM selection To comprehensively evaluate the benchmark performance of current mainstream large models in the field of financial security agent dialogues, we selected 10 advanced models and conducted extensive comparative experiments to assess their capability for financial agent security detection. We illustrate the specific comparative results across eight dimensions in Fig. 3. Our evaluation metrics include F1 score, recall, specificity, and a comprehensive modified sharpe ratio. We design a Sharpe-inspired performance index to measure risk-adjusted stability across evaluation metrics. As shown in Fig. 3a, we observe that O3 emerges as the leading model with an overall F1 score of 76.48%, marginally outperforming Gemini 2.5 Pro (76.40%). However, a granular examination highlights distinct performance profiles across categories. Regarding injection defense in Fig. 3b, we find that DeepSeek Chat demonstrates superior efficacy (F1: 79.41%), whereas other models struggle to balance recall and specificity (Fig. 3c and Fig. 3d). Regarding unintended risks, Claude Opus 4 adopted a highly conservative strategy, achieving perfect sensitivity (100.00% recall) but significantly compromising specificity (48.78%). Conversely, GPT-O3 maintained a more robust equilibrium between the two metrics (specificity: 75.61%). These findings suggest that current SOTA models possess high baseline security. Finally, Fig. 3h presents the modified sharpe ratio, where we evaluate the stability of model performance relative to risk, providing a robust metric for model reliability in volatile financial environments. However,
calibrating the trade-off between over-refusal and missed detections remains a critical optimization bottleneck, particularly in financial contexts. B. FinO3 We first optimized the prompts of large language models to meet the specific requirements of financial security detection, aiming to improve their specialization and accuracy in the domain of financial agent dialogue risk assessment. Subsequently, we integrated adversarial structures into the model prompts, enabling the models to more effectively identify potential attack behaviors and evasive language. This enhances their overall ability to detect security risks within financial agent dialogues. The experimental results for the FinO3 model, both before and after the inclusion of adversarial structures (FinO3adv), are presented in Table I. In the case of FinO3adv, although adversarial verification structures were incorporated into the prompt, the overall model performance did not improve significantly, and its ability to identify unintended attacks even declined. This indicates that simply introducing complex adversarial detection mechanisms into the prompt can create tension between defensive strength and the original task objectives, resulting in model confusion when faced with covert attacks and, consequently, reduced detection accuracy. Specifically, integrating adversarial logic at the prompt level works well when the detection architecture has only a single layer. However, as the detection model
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
8
(a) Overall F1 Score
(b) Injection F1 Score
(c) Injection Recall
(d) Injection Specificity
(e) Unintended F1
(f) Unintended Recall
(g) Unintended Specificity
(h) Modified Sharpe Ratio
Fig. 3. Performance Comparison of 10 LLM Models on Financial Security Risk Assessment TABLE I C OMPARISON OF R- JUDGE (O3) AND F IN O3 VARIANTS ON F INANCIAL S ECURITY R ISK A SSESSMENT Models R-Judge(O3) FinO3 (Zero-shot) FinO3 (Few-shot) FinO3adv
All F1 (%) 76.49 82.05 83.27 82.85
F1 (%) 80.00 86.96 73.68 83.87
Injection Attacks Recall (%) Spec (%) 66.67 100.00 83.33 80.00 63.64 66.67 100.00 87.65
becomes more complex, the experimental results indicate that this method is much less effective. This is primarily because embedding adversarial structures within the prompt can lead to excessive defensiveness in the model. In particular, since the task requires detecting attacks while also performing classification, the model faces a trade-off between these objectives. This tension ultimately results in suboptimal performance. Such conflicts are especially pronounced in financial dialogue scenarios, as the context typically involves sensitive terms such as amounts, accounts, and permissions. If the model’s defensiveness is set too high, ordinary transaction requests may be misclassified as attacks. Therefore, we observed that the delayed manifestation of adversarial semantics is quite apparent in financial dialogues. The prompt layer, lacking temporal sensitivity, can only constrain the semantics of the current turn, rendering it ineffective in modeling the evolving risks of dialogue over time. Consequently, when further developing FinSec, we adjusted the placement of the adversarial model, relocating it from the third layer. The third layer now serves independently as a semantic security discrimination layer. C. FinSec 1) Layer 1 and layer 2: In layer 1, following the guidelines of AML and SAR, we introduce six primary detection patterns: structuring, gradual extraction, rapid escalation, circular flow, timing anomaly and social engineering. Among the broad spectrum of financial activities, each suspicious pattern is
F1 (%) 72.97 77.14 92.86 81.82
Unintended Risks Recall (%) Spec (%) 100.00 75.61 100.00 80.49 100.00 95.06 81.82 33.33
assigned a different weight depending on the type of financial operation. To align with the overall direction of market regulation, we assign weights ti each pattern. In the delay risk simulation module of layer 2, we integrate the calculation of abnormal risk scores from multiple risk factors according to the characteristics of the financial market. These factors include user history, transaction patterns, amount anomalies, frequency anomalies, and timing anomalies. For the current dataset, frequency anomaly demonstrates the most significant discrimination ability. Therefore, we evaluate the weight assignment based on a multidimensional assessment framework across four dimensions: discrimination, balance, robustness, and scalability. As shown in Fig. 4 displays the scores of each configuration across Discrimination, Balance, Robustness, and Scalability. The configuration with weight w = 0.2 performs well in all four dimensions with no obvious weaknesses. Fig. 5 shows the score distribution of different weight configurations for each dimension, helping to identify the optimal interval. At w=0.20, both Balance and Robustness achieve full marks, and the Composite score reaches its peak. Fig. 6a illustrates the curves of the four-dimensional scores and the composite score as the weight changes, allowing identification of the optimal weight range. The composite score maintains a high level in the range of w = 0.15 − 0.3, peaking at w = 0.2. Fig. 6b is a bar chart comparing the composite scores of different weight configurations, clearly showing that the configuration with w = 0.2 is optimal. Based on this analysis, we determine
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
9
(a) Impact of weight variation on individual performance metrics
(b) Final selection based on aggregated composite scores
Fig. 6. Performance Sensitivity and Optimal Weight Determination for Frequency Anomaly Detection
ASRFinSec = 1 − Fig. 4. Global view of performance stability via score heatmap
X q∈{inj,unint}
wq
3 X ℓ=1
ηq,ℓ
TPq,Lℓ |Aq |
(14)
Meanwhile, we established a hierarchical robust ASR for FinSec. Mirroring the logic described above, this metric incorporates weighting based on the model’s architecture and risk profile. Specifically, after determining the count of successfully identified attacks at layer ℓ (denoted as T Pq,Lℓ ), we perform a second-order aggregation across both attack types and hierarchical layers." Ultimately, we obtain a comprehensive security score that accounts for detection performance, adversarial robustness, and output stability. The specific formulation is given in equation (15).
Fig. 5. Trade-off analysis and capability balance under different weights
the weight configuration for frequency anomalies. In this way, current performance and future adaptability are balanced, allowing sufficient weight for effective factors when processing data with new risk variables (such as user histories) in the future. At the same time, the system’s robustness against measurement errors is also improved. 2) Overall Model Results Validation : To evaluate the effectiveness of FinSec in identifying anomalous and risky dialogues within financial agent conversations, we introduce the Area Under the Precision–Recall Curve (AUPRC) and the Attack Success Rate (ASR) as metrics[38][39]. The formulas are presented in Eq.(13) and Eq.(14) as follows: ! Z 1 X 3 X AUPRCFinSec = wq ωℓ Pq,Lℓ (R) dR q∈{inj,unint}
0
ℓ=1
(13) Standard AUPRC models lack the hierarchical granularity required to effectively evaluate FinSec. To address this limitation, we extended the traditional metric into a Structured AUPRC. This framework operates across three levels, calculating performance independently for two distinct attack vectors. Within each level, risk contributions are weighted according to FinSec’s hierarchical architecture. Specifically, q denotes the attack type (categorized as injection or unintended), and Pq,Lℓ (R) represents the precision at layer ℓ for attack type q at a given recall threshold R.
RFinSec = α AUPRCFinSec + β (1 − ASRFinSec ) + γ StabFinSec (15) TableII compares teh effectiveness of R-Judge, FinO3, FinO3adv, and FinSec. Specifically, regarding injection attack detection, compared to the baseline R-Judge (72.97%), FinSec demonstrates significant improvement. This result indicates that FinSec is highly effective in handling proactive malicious command attacks. The nest column illustrates the detection of unintended risk. Here, FinSec continues to deliver leading performance (85.71%). Notably, although R-Judge achieves reasonable results (80.00%), FinO3 exhibits a noticeable decline in performance (73.68%). This suggests that incorporating adversarial detection methods into the prompt can lead to degraded performance under poisoning attacks, possibly due to excessively long prompt texts. Finally we examine the overall F1 scores. FinSec is the only model to achieve a score exceeding 90%, maintaining high precision while ensuring a high recall rate. It achieves substantial improvements in both detection ability and precision, demonstrating stable performance in adversarial scenarios, and providing robust and balanced defense against both types of financial dialogue risks. Figure 7 shows the analyzes of model performance from a multi-dimensional trade-off perspective. First, we evaluate the robustness of each model in Fig. 7(a), specifically regarding both defense capabilities and precision. Here, the bar chart displays the defense rate, while the red line corresponds to the AUPRC. Through these results, we confirm the effectiveness of our proposed defense mechanisms, which maintain strong performance even on imbalanced datasets.
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
10
TABLE II F1-S CORE C OMPARISON OF F INANCIAL S ECURITY R ISK A SSESSMENT M ODELS Models O3 (R-Judge) FinO3 FinO3adv FinSec
Utility F1 (%) 76.48 76.48 82.94 90.13
Injection Attack F1 (%) 72.97 83.27 83.87 94.55
Unintended Attack F1 (%) 80.00 73.68 81.82 85.71
through this research bottleneck. Its performance significantly surpasses the Pareto frontier formed by baseline models. This demonstrates that FinSec’s unique hierarchical defense architecture avoids the typical zero-sum trade-off, successfully combining strong robustness against malicious instructions with high sensitivity to semantic risks, thereby achieving dual optimality in financial agent dialogue scenarios. V. C ONCLUSIONS
(a) Defense Rate with AUPRC
(b) Comprehensive Score and ASR
Fig. 7. Performance Evaluation: Trade-off Analysis and Final Verdict: (a) Robustness Analysis: Defense Rate and AUPRC; (b) Final Verdict: Comprehensive Score versus Risk (ASR)
Fig. 8. Performance Trade-off: Injection and Unintended Risks
Next, we present the final verdict in Fig. 7(b), contrasting the comprehensive score against the potential risk (ASR). We observe that FinSec achieves a comprehensive weighted score of 0.9098, significantly outperforming the comparative models. Notably, compared with the baseline R-Judge model, we achieve a 12% overall performance improvement while maintaining the lowest attack success rate (indicated by the red line). This demonstrates that FinSec ensures superior utility without compromising the security of agent dialogues. Figure 8 presents the results of a Pareto frontier analysis for various models on injection risk and unintended risk detection tasks in financial agent dialogues. As illustrated, the blue dashed line delineates the performance boundaries of leading LLMs such as Claude Opus4 and Gemini2.5 Pro. A clear trade-off phenomenon is observed: most models excel in only one dimension. However, our proposed model, FinSec, breaks
This paper introduces a hierarchical detection framework for financial agent dialogues security, named FinSec. FinSec employs a three-layer architecture to evaluate the safety of financial agent dialogues from multiple perspectives, producing an integrated security assessment as the final output. FinSec addresses critical limitations of existing LLMs, which struggle to balance high sensitivity and operational practicality in financial environments. Operational practicality refers to the system’s ability to efficiently accomplish legitimate financial security assessment tasks while maintaining compliance and safety. FinSec enhances the precision of intent recognition and, importantly, avoids misclassifying harmless instructions as attacks. Methodologically, Layer 1 of FinSec constructs a behavioral pattern library based on international AML/SAR standards and employs a “triple matching” mechanism for high-recall preliminary screening of explicit risks. Building upon this, Layer 2 quantifies delayed risks and utilizes adversarial detection to quantitatively evaluate the effectiveness of defense mechanisms. Layer 3 incorporates a semantic discriminator for security assessment, and finally, Layer 4 performs a comprehensive scoring and makes the final decision. The results demonstrate that FinSec achieves state-of-the-art performance, significantly outperforming baseline models. Notably, in Pareto frontier analyses, FinSec effectively overcomes the specificity–recall bottleneck, mitigating the trade-off between excessive rejection and missed detections, thereby greatly enhancing the reliability of financial agent security detection. Regarding potential improvements, we first observed that when prompts become excessively long or complex—i.e., under long-sequence input conditions—the outputs of large language models exhibit deviations, resulting in decreased accuracy. This effect is more pronounced compared to scenarios involving concise prompts. Additionally, as the overall model architecture becomes more complex, the inability of LLMs to effectively process large volumes of input at once further undermines accuracy. Therefore, it is necessary to optimize financial risk detection capabilities specifically for long-sequence prompts. Secondly, in practical scenarios, major financial losses are often caused by intentional injection
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
attacks. As a result, it is essential to refine the categorization of injection attacks and develop more targeted identification and defense mechanisms. In future work, we will continue to address these two outstanding issues through dedicated research. VI. ACKNOWLEDGEMENT This work was supported in part by the JSPS KAKENHI under Grants 23K11072, in part by the National Natural Science Foundation of China under Grants U21B2019 and 61972255. R EFERENCES [1] Z. Guo, W. Guo, Q. Li, Y. Zou, and J. Cai, “Fn-agents: Analysis of exchange rate volatility prediction based on multi-agent systems,” in 2024 5th International Conference on Computers and Artificial Intelligence Technology (CAIT), 2024, pp. 350–355. [2] Microsoft, “Microsoft 365 copilot for finance: Release wave 2 plan,” Microsoft Dynamics 365 Release Plan, Oct. 2024. [3] Celent, “Morgan stanley wealth management: Ai@ms debrief case study,” Celent Industry Research Report, Jun. 2025. [4] R. Pedro, M. E. Coimbra, D. Castro, P. Carreira, and N. Santos, “Promptto-sql injections in llm-integrated web applications: Risks and defenses,” in 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), 2025, pp. 1768–1780. [5] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising realworld llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, ser. AISec ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 79–90. [Online]. Available: https://doi.org/10.1145/3605764.3623985 [6] Y. Liu, G. Deng, Y. Li, K. Wang, Z. Wang, X. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, and Y. Liu, “Prompt injection attack against llm-integrated applications,” 2024. [Online]. Available: https://arxiv.org/abs/2306.05499 [7] Z. Liao, K. Chen, Y. Lin, K. Li, Y. Liu, H. Chen, X. Huang, and Y. Yu, “Attack and defense techniques in large language models: A survey and new perspectives,” 2025. [Online]. Available: https://arxiv.org/abs/2505.00976 [8] B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, S. T. Truong, S. Arora, M. Mazeika, D. Hendrycks, Z. Lin, Y. Cheng, S. Koyejo, D. Song, and B. Li, “Decodingtrust: a comprehensive assessment of trustworthiness in gpt models,” in Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023. [9] E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, A. Jermyn, A. Askell, A. Radhakrishnan, C. Anil, D. Duvenaud, D. Ganguli, F. Barez, J. Clark, K. Ndousse, K. Sachan, M. Sellitto, M. Sharma, N. DasSarma, R. Grosse, S. Kravec, Y. Bai, Z. Witten, M. Favaro, J. Brauner, H. Karnofsky, P. Christiano, S. R. Bowman, L. Graham, J. Kaplan, S. Mindermann, R. Greenblatt, B. Shlegeris, N. Schiefer, and E. Perez, “Sleeper agents: Training deceptive llms that persist through safety training,” 2024. [Online]. Available: https://arxiv.org/abs/2401.05566 [10] F. Bianchi, M. Suzgun, G. Attanasio, P. Rottger, D. Jurafsky, T. Hashimoto, and J. Zou, “Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https: //openreview.net/forum?id=gT5hALch9z [11] Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto, “Identifying the risks of lm agents with an lm-emulated sandbox,” in The Twelfth International Conference on Learning Representations, 2024. [12] X. Yin, C. Ni, and S. Wang, “Multitask-based evaluation of opensource llm on software vulnerability,” IEEE Transactions on Software Engineering, vol. 50, no. 11, pp. 3071–3087, 2024. [13] T. Yuan, Z. He, L. Dong, Y. Wang, R. Zhao, T. Xia, L. Xu, B. Zhou, F. Li, Z. Zhang, R. Wang, and G. Liu, “R-judge: Benchmarking safety risk awareness for llm agents,” arXiv preprint arXiv:2401.10019, 2024.
11
[14] X. Gu, X. Zheng, T. Pang, C. Du, Q. Liu, Y. Wang, J. Jiang, and M. Lin, “Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast,” arXiv preprint arXiv:2402.08567, 2024. [15] M. Cemri, M. Z. Pan, S. Yang, L. A. Agrawal, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, D. Klein, K. Ramchandran et al., “Why do multi-agent llm systems fail?” arXiv preprint arXiv:2503.13657, 2025. [16] M. Thakkar, Q. Fournier, M. Riemer, P.-Y. Chen, A. Zouaq, P. Das, and S. Chandar, “Combining domain and alignment vectors provides better knowledge-safety trade-offs in LLMs,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 268–277. [Online]. Available: https://aclanthology.org/2025.acl-short.22/ [17] Q. Xie, W. Han, Z. Chen, R. Xiang, X. Zhang, Y. He, M. Xiao, D. Li, Y. Dai, D. Feng, Y. Xu, H. Kang, Z. Kuang, C. Yuan, K. Yang, Z. Luo, T. Zhang, Z. Liu, G. Xiong, Z. Deng, Y. Jiang, Z. Yao, H. Li, Y. Yu, G. Hu, J. Huang, X.-Y. Liu, A. Lopez-Lira, B. Wang, Y. Lai, H. Wang, M. Peng, S. Ananiadou, and J. Huang, “The finben: An holistic financial benchmark for large language models,” 2024. [18] D. B. Araya and D. Liao, “Finvet: A collaborative framework of rag and external fact-checking agents for financial misinformation detection,” 2025. [Online]. Available: https://arxiv.org/abs/2510.11654 [19] H. Yang, B. Zhang, N. Wang, C. Guo, X. Zhang, L. Lin, J. Wang, T. Zhou, M. Guan, R. Zhang, and C. D. Wang, “Finrobot: An opensource ai agent platform for financial applications using large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2405.14767 [20] P. Islam, A. Kannappan, D. Kiela, R. Qian, N. Scherrer, and B. Vidgen, “Financebench: A new benchmark for financial question answering,” 2023. [Online]. Available: https://arxiv.org/abs/2311.11944 [21] Z. Chen, J. Chen, J. Chen, and M. Sra, “Standard benchmarks fail – auditing llm agents in finance must prioritize risk,” 2025. [Online]. Available: https://arxiv.org/abs/2502.15865 [22] Z. Dong, Z. Zhou, C. Yang, J. Shao, and Y. Qiao, “Attacks, defenses and evaluations for LLM conversation safety: A survey,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 6734–6747. [Online]. Available: https://aclanthology.org/2024.naacl-long.375/ [23] F. He, T. Zhu, D. Ye, B. Liu, W. Zhou, and P. S. Yu, “The emerged security and privacy of llm agent: A survey with case studies,” ACM Comput. Surv., vol. 58, no. 6, Dec. 2025. [Online]. Available: https://doi.org/10.1145/3773080 [24] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” ser. AISec ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 79–90. [Online]. Available: https://doi.org/10.1145/3605764.3623985 [25] A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: how does llm safety training fail?” in Proceedings of the 37th International Conference on Neural Information Processing Systems, ser. NIPS ’23. Red Hook, NY, USA: Curran Associates Inc., 2023. [26] J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, C. Zhang, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” 2023. [Online]. Available: https://arxiv.org/abs/2307.04657 [27] Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan, “Training a helpful and harmless assistant with reinforcement learning from human feedback,” 2022. [Online]. Available: https://arxiv.org/abs/2204.05862 [28] D. Jacob, H. Alzahrani, Z. Hu, B. Alomair, and D. Wagner, “Promptshield: Deployable detection for prompt injection attacks,” in Proceedings of the Fifteenth ACM Conference on Data and Application Security and Privacy, ser. CODASPY ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 341–352. [29] K.-H. Hung, C.-Y. Ko, A. Rawat, I.-H. Chung, W. H. Hsu, and P.-Y. Chen, “Attention tracker: Detecting prompt injection attacks in LLMs,” in Findings of the Association for Computational Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang, Eds. Albuquerque, New Mexico: Association for Computational
JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 21, NOVEMBER 2025
Linguistics, Apr. 2025, pp. 2309–2322. [Online]. Available: https: //aclanthology.org/2025.findings-naacl.123/ [30] M. Kuo, J. Zhang, A. Ding, L. DiValentin, A. Hass, B. F. Morris, I. Jacobson, R. Linderman, J. Kiessling, N. Ramos, B. Gopal, M. B. Pouyan, C. Liu, H. Li, and Y. Chen, “Safety reasoning elicitation alignment for multi-turn dialogues,” 2025. [Online]. Available: https://arxiv.org/abs/2506.00668 [31] Y. Li, Y. Du, J. Zhang, L. Hou, P. Grabowski, Y. Li, and E. Ie, “Improving multi-agent debate with sparse communication topology,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 7281– 7294. [Online]. Available: https://aclanthology.org/2024.findings-emnlp. 427/ [32] Y. Liu, Y. Liu, X. Zhang, X. Chen, and R. Yan, “The truth becomes clearer through debate! multi-agent systems with large language models unmask fake news,” in Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, ser. SIGIR ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 504–514. [Online]. Available: https://doi.org/10.1145/3726302.3730092 [33] P. Manakul, A. Liusie, and M. J. F. Gales, “Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models,” 2023. [Online]. Available: https://arxiv.org/abs/2303.08896 [34] S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal, “Detecting hallucinations in large language models using semantic entropy,” Nature, vol. 630, pp. 625–630, 2024. [Online]. Available: https://doi.org/10.1038/ s41586-024-07421-0 [35] H. Shen, Z. Gu, H. Hong, and W. Han, “Pii-bench: Evaluating query-aware privacy protection systems,” 2025. [Online]. Available: https://arxiv.org/abs/2502.18545 [36] J. Zheng, S. Qiu, and Q. Ma, “Can llms learn new concepts incrementally without forgetting?” 2024. [Online]. Available: https: //arxiv.org/abs/2402.08526 [37] Z. Xu, M. Qi, S. Wu, L. Zhang, Q. Wei, H. He, and N. Li, “The trust paradox in llm-based multi-agent systems: When collaboration becomes a security vulnerability,” 2025. [Online]. Available: https: //arxiv.org/abs/2510.18563 [38] H. B. Nguyen and V.-N. Huynh, “On sampling techniques for corporate credit scoring,” Journal of Advanced Computational Intelligence and Intelligent Informatics, vol. 24, no. 1, pp. 48–57, 2020. [39] F. Tramèr, N. Carlini, W. Brendel, and A. Madry, ˛ “On adaptive attacks to adversarial example defenses,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, ser. NIPS ’20. Red Hook, NY, USA: Curran Associates Inc., 2020.
12