Conceptio › Archive › arXiv CS
arXiv CSopen access

AuxMark: Defending Against Unauthorized Agent Distillation via Auxiliary Behavioral Watermarking

Yiqing Feng et al. · arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Preprint

AUX M ARK : D EFENDING AGAINST U NAUTHORIZED AGENT D ISTILLATION VIA AUXILIARY B EHAVIORAL WATERMARKING

arXiv:2609.34597v1 [cs.CR] 28 Sep 2026

Yiqing Feng1 , Haozhe Feng2 , Shunan Shang1 , Xiaoyu Zhang1 , Jian Lou3 , Haodong Zhao4 , Mingxun Zhou5∗ 1 Xidian University, 2 Zhejiang University 3 Sun Yat-sen University, 4 Shanghai Jiao Tong University, 5 HKUST

A BSTRACT Large language model agents can acquire complex capabilities through multi-step interaction and tool use, but their trajectories can also be illegally collected to distill student agents. However, existing watermarking methods either do not fit the structured and interactive nature of agent environments or lack reliable effectiveness across tasks and model architectures. We introduce AuxMark, a behavioral watermarking framework for tracing unauthorized agent distillation. AuxMark dynamically inserts safe, non-essential auxiliary action into teacher trajectories, and stores the associated contexts as private evidence cards. To audit a suspicious student model, AuxMark constructs paired real and fake probes from these cards and applies a card-level sign test. This black-box protocol supports both modellevel detection and trace-level attribution. Across three agent benchmarks, two teacher agents, and four student architectures, AuxMark detects all 24 distilled models with zero false positives on 48 clean models. It also preserves task utility and remains effective against data flooding, paraphrasing, truncation, and adaptive cleaning attacks. Our code will be released at this URL.

1

I NTRODUCTION

Large language model (LLM) agents can solve complex tasks by interacting with external environments and using tools (Anthropic, 2026; Team et al., 2025; Xu et al., 2026a). However, training models with these advanced skills requires huge costs and computing resources. For example, training frontier models like OpenAI’s GPT requires tens of thousands of GPUs and costs tens of millions of dollars Singh et al. (2025); Sajadieh et al. (2026). Agent distillation is a cheap way to significantly improve model performance Liu et al. (2026b); Kang et al. (2026), so it has become widely used. Competitors systematically query a powerful teacher agent to collect high-quality interaction trajectories, using them to fine-tune their own student models to avoid the high costs of training from scratch. This unauthorized knowledge transfer severely harms the intellectual property of model developers. This threat is not just a theory; leading AI companies (e.g., OpenAI, Google) have already found massive illegal distillation behaviors in their traffic OpenAI (2026); Group (2026); Anthropic (2026). Furthermore, recent industry disputes and studies have exposed widespread trajectory copying, showing that many models display “behavioral homogenization” that closely mimics the proprietary teachers Yang et al. (2026a). To protect model intellectual property, several methods have been proposed, but they face different critical limitations when applied to agent distillation. First, antidistillation watermarks designed for language models modify text tokens or reasoning traces Savani et al. (2026); Xu et al. (2026b); Ma et al. (2026), which breaks the strict format required for agent tool calls. Second, existing agent antidistillation watermarking schemes usually assume that the teacher and student share the same base model, resulting in poor performance during cross-architecture distillation Wang et al. (2026). Third, unwatermarked fingerprinting methods based on execution similarity Yang et al. (2026a); Rawat et al. (2026) are unreliable, because high behavioral similarity often comes from the standard ∗

Corresponding author.

1

Preprint

solutions of the tasks rather than actual distillation. Therefore, finding a reliable way to trace and prove this unauthorized distillation through black-box interactions has become an urgent challenge. In this work, we propose AuxMark, a behavioral watermarking framework to trace agent distillation. Our main insight stems from the multi-step nature of agent interactions and the behavioral homogenization during distillation. Instead of modifying the core actions required for the task, AuxMark injects extra auxiliary actions into the trajectory and saves these interaction as private evidence. To align with deployments in the real world, we introduce a more practical threat model: model owners can identify suspicious distillation behaviors via traffic monitoring (existing studies show that such traffic anomaly detection can achieve around 90% recall with extremely low false positive rates Liu et al. (2026c); Huh et al. (2025); Chiang et al. (2024)), and extract the corresponding evidence for targeted verification. For this verification, we introduce a paired-probe detection protocol. It compares the suspect model’s responses on real probes built from the evidence against fake probes where the core actions are semantically altered. This strict comparison effectively isolates the watermarked behavior from natural contextual choices. Empirically, our evaluation demonstrates that AuxMark achieves highly reliable detection while preserving the agent’s task utility. Across three agent benchmarks, two teacher models, and four student architectures, our method successfully detects all 24 distilled student models with zero false positives. Furthermore, AuxMark remains highly robust against various adversarial attacks. Key contributions of this paper include: • We formulate a highly practical threat model for agent distillation. By combining suspicious target identification via traffic monitoring with targeted verification using private evidence, this model provides a closed-loop solution for tracing intellectual property theft in the real world scenarios. • We design a novel behavioral watermarking framework and a paired-probe detection protocol. This architecture exploits multi-step interactions to embed auxiliary actions in real time, and provides strong statistical guarantees for verification and trace-level attribution by comparing responses on real and fake probes alongside a card-level sign test. • We conduct comprehensive experiments across multiple benchmarks and model architectures. The results demonstrate that AuxMark achieves a 100% detection rate and zero false positives while preserving the agent’s task utility, and exhibits strong robustness against various data-processing operations and adaptive cleaning attacks.

2

P ROBLEM F ORMULATION AND P RELIMINARIES

2.1

AGENT D ISTILLATION

LLM agent interacts with an external environment to solve a user task through multiple rounds of reasoning and tool use Yao et al. (2022); Qin et al. (2024). An agent trajectory records intermediate reasoning and tool-use behaviors. Given a user query q, an LLM agent M interacts with the environment over T steps. At each step i, based on the query q and the interaction history hi , the agent generates a thought ti and an action ai = (fi , pi ), where fi denotes the selected tool and pi denotes its parameters. The environment executes the action and returns an observation oi . A complete interaction trajectory τ is formally represented as: τ = {(ti , ai , oi )}Ti=1 , where (ti , ai ) = M(q, hi ).

(1)

Agent distillation Kang et al. (2026); Luo et al. (2026); Liu et al. (2026b) aims to transfer such interaction behaviors from a powerful teacher agent Mt to a student agent Ms . The teacher first generates a set of N trajectories (qj , τj )N j=1 . In agent distillation, we consider hard distillation based on supervised fine-tuning Ouyang et al. (2022). Through this process, the student can learn the teacher’s reasoning and tool-use behaviors from its interaction trajectories. The student model is trained on (qj , τj )N j=1 , where previous thoughts, actions, and observations are used as context, while the teacher’s thoughts and actions are used as supervision Kang et al. (2026); Luo et al. (2026). 2.2

M ODEL WATERMARKING

To protect the intellectual property of a teacher agent Mt , the model owner employs a secret key k to embed watermark signals in real time during the generation of trajectories τ . After attackers query the model to obtain a watermarked dataset D, they attempts to clone the teacher’s capabilities 2

Preprint

by training an base model Mc on D via a distillation algorithm to get the student model Ms . Then model owner audits suspicious student models using a private probe set Pk . A practical antidistillation watermark should satisfy the following three requirements: Effectiveness. An effective watermark must reliably transfer to the student model during distillation while maintaining a minimal false positive rate on independent models. During verification, the owner queries the suspicious model using the probe set Pk and computes a p-value to evaluate the watermark signal. A distilled student model Ms should exhibit statistically significant evidence of the watermark, satisfying p−value < p, whereas a clean model Mc should yield p−value > p. Harmlessness. The real-time watermarking mechanism should not degrade the normal operation of the teacher agent Mt . This entails two main aspects: task utility and inference efficiency. Specifically, the watermarked teacher model must maintain a success rate on the same tasks comparable to its unwatermarked counterpart, and the watermark embedding process should not significantly inflate the length or computational cost of the generated trajectories. Robustness. The watermark must resist data modifications applied by attackers. Even if the attacker alters the dataset through data flooding, semantic rewriting, truncation, and adaptive cleaning attacks. the watermark signal should remain detectable in the student model. 2.3

𝐷 Training Set 𝐷 Teacher Agent 𝑀𝑡

Attacker

Suspicious Set 𝑆

Distilation Student Model 𝑀s Probes

Model Owner Traffic Analysis

Results

𝑀𝑠 Detection

Figure 1: Overview of the access and threat model. The attacker distills a student model from trajectories, while the model owner identifies suspicious traffic and audits the student.

ACCESS AND T HREAT M ODEL

We formulate a two-party threat model comprising a model owner and an attacker as shown in Figure 1. The owner deploys a teacher agent Mt , while the attacker aims to distill a student model Ms using a training dataset D collected from the teacher agent. Model Owner. The owner deploys and fully controls a teacher agent Mt (e.g., Codex Singh et al. (2025)) as a service. To protect the model, the owner embeds watermarks into the output trajectories in real time. Since data scraping has specific patterns, the owner monitors server-side features Huh et al. (2025); Chiang et al. (2024); Liu et al. (2026c). Any interactions showing such signs are recorded to build a suspicious trajectory set S. To avoid missing potential distillation traces, the owner conservatively records all suspicious traces into S and the owner cannot determine which traces in S were actually used for distillation. When testing the deployed student model Ms , the owner only has black box access and can fully control the API inputs to verify the watermark. Attacker. To bypass the high training costs, the attacker aims to clone the capabilities of the teacher Mt . Operating under strict black-box access, the attacker systematically queries the teacher agent Mt to harvest complete multi-step interaction trajectories and this large-scale automated querying leaves identifiable server-side traffic patterns that can be captured by the model owner. The attacker may then preprocess these collected trajectories through operations such as sanitization, rewriting, truncation, or mixing with other data, forming the final distillation dataset D. Finally, the attacker uses D to distill the student model Ms and deploys it as a black-box API.

3

WATERMARK C ONSTRUCTION

3.1

OVERVIEW

During agent distillation, the distilled students and their teachers share similar interaction trajectories Yang et al. (2026a); Lyu et al. (2025); Gudibande et al. (2023). Based on this property, we proposed a watermark scheme AuxMark as illustrated in Figure 2. Firstly, the owner embeds the watermark by dynamically inserting auxiliary(aux) behaviors into the teacher’s trajectories and privately saves evidence. Once suspicious traffic is identified, the corresponding records are extracted to form an evidence pool in Section 3.2. Next, the owner uses this evidence to construct paired real and fake probes, applying statistical tests to verify if the suspect model reproduces the watermark in 3

Preprint

Section 3.3. Finally, trace-level attribution in Section 3.4 to evaluate which suspicious trajectories are most likely to have been used in the attacker’s training set. 3.2

WATERMARK E MBEDDING A LGORITHM

Unlike previous watermarking methods that directly change the generated text or the teacher model’s actions Wang et al. (2026); An et al. (2026), AuxMark embeds watermarks by adding extra, non-essential auxiliary actions into the agent’s trajectory. By doing this, AuxMark keeps the teacher model’s performance high while creating strong evidence for later tracing. Also, to stay hidden, the added aux actions look and act exactly like normal tool calls: they use valid tools and bind strictly to normal parameter values derived from the interaction history. Based on the threat model described in Section 2, the owner can subsequently identify suspicious traces from the API traffic and export their privately saved watermark evidence for later verification. The entire embedding process consists of three main stages shown in Figure8: safe tool classification, dynamic scheduling and candidate generation, and private evidence retention.

Traffic Analysis

Evidence Pool 𝐶𝑆

Phase 1

Watermarked Traces 𝜏

Teacher Agent 𝑀𝑡

𝑎𝚌𝚘𝚛𝚎

𝑎𝚌𝚘𝚛𝚎

𝑎𝚌𝚘𝚛𝚎

𝑎𝚊𝚞𝚡

𝑎𝚌𝚘𝚛𝚎

𝑎𝚊𝚞𝚡

𝑎𝚌𝚘𝚛𝚎

𝑎𝚊𝚞𝚡

𝑎𝚌𝚘𝚛𝚎

𝑎𝚌𝚘𝚛𝚎

𝑎𝚊𝚞𝚡

𝑎𝚌𝚘𝚛𝚎

Get details> Get usage> Get details(Aux)> Enable roaming

Real probes 𝑃𝚛𝚎𝚊𝚕

Phase 2

𝑀𝑡

𝐶𝑆

Results Student Model 𝑀s Fake probes 𝑃𝚏𝚊𝚔𝚎

Figure 2: Overview of AuxMark. Phase 1: the owner embeds the watermark by inserting aux behaviors; Phase 2: the owner retrieves the private evidence to construct paired real and fake probes and verify a suspect student model. denotes watermarked thing.

Safe tool classification. To ensure aux actions will not change the environment state, AuxMark applies a conservative safety policy to the tool schema F of trace τ . Since the system operates as an agent-as-a-service Li et al. (2026), this safety evaluation is performed in advance and cached to avoid real-time latency. An offline safety model Msafe evaluates the tool schema F and maps each tool f ∈ F to a safety label as shown in Eq.2:  safe, if f is read-only, Msafe (f ) = (2) unsafe, otherwise. The system exclusively permits tools from the valid subset Fsafe = {f ∈ F | Msafe (f ) = safe}, categorically rejecting any action that changes the environment state or influences the task process. Dynamic scheduling. Having established the safe tool subset Fsafe , the system must determine when to invoke these aux action during the interaction as the watermarking. To make watermark positions unpredictable, AuxMark dynamically evaluates whether to inject an aux action at step t immediately after the core action at−1 core finishes. If skipped, the system waits and checks again after the next core action. For each trajectory τ named as idτ of an account a, the system keeps a release probability pt and a remaining budget bt at step t. Immediately after the core action at−1 core finishes, the system uses a cryptographic hash function Hash(·) → [0, 1) Preneel (1994) parameterized by the owner’s secret key k to compute a uniformly distributed random value ut , as defined in Eq. 3: ut = Hash(k ∥ a ∥ idτ ∥ t) ∈ [0, 1). (3) The system tries to add an aux action when bt > 0 and ut < pt . Depending on whether the auxiliary action is successfully inserted or the turn is skipped (due to generation failure or unmet conditions), the system updates the trigger probability and remaining budget for the next turn as follows:  (p0 , bt − 1), if successfully inserted, (pt+1 , bt+1 ) = (4) (min{pt + ∆p, pmax }, bt ), otherwise. Aux action embedding. If the dynamic scheduling condition is met, AuxMark proceeds to generate and select the optimal aux action. To obtain high quality insertions, the system queries Mt to propose up to K candidates. Each candidate j is formulated as a tuple of a thought and an action: j (tjaux , ajaux ), where the action ajaux consists of a tool faux and its arguments pjaux . To ensure that the injected action perfectly blends into the trajectory, AuxMark enforces a validation mechanism. 4

Preprint

First, the chosen tool must belong to the safe subset Fsafe . Second, to prevent model hallucinations and ensure contextual rationality, the arguments pjaux must be extracted from the context (such as the core action’s observation or history). Consequently, AuxMark categorically discards any candidate that uses an unapproved tool or hallucinates untraceable arguments. For each candidate j that passes the validation, AuxMark employs a scoring model Mq to evaluate it across three dimensions: naturalness, logical consistency, and distillability. Specifically, Mq assigns a relevance score sjrel to ensure the injected action aligns naturally with the task. Additionally, a reliability score sjrelb ensures the high quality of the thought. Crucially, to foster a stable behavioral habit that the student model can easily absorb during distillation, the system calculates a repeatability score sjrep . This score is directly determined by nj , which represents the number of times the j candidate’s specific tool pairing (fcore , faux ) has been successfully injected in the account’s history. The final score Sj for candidate j is computed as a weighted sum: Sj = α · sjrel + β · sjrelb + γ · sjrep . (5) where α, β, γ, and the step-function mapping from repeat count nj to sjrep are detailed in Appendix B.1. The system selects the candidate with the highest score as the best aux action best (tbest aux , aaux ), then injects this aux action, establishing a natural sequence: t−1 t−1 best best best tt−1 (6) core → acore → ocore → taux → aaux → oaux . Based on the generation result, the system updates its states. A successful generation applies the success update (Eq. 4) and increases the repeat count nbest for this tool pairing. If the generation fails, the system skips the insertion, applies the failure update, moves to generate core action and waits for next chance. For future auditing, every successful insertion creates a private evidence card ct . To securely record the exact context and execution details, this card stores the historical context h, the preceding core action and observation, and the injected auxiliary tool parameters. It is saved as a tuple and added to the trajectory’s evidence set Cτ : t−1 t−1 best best ct = (h, tt−1 Cτ = Cτ ∪ {ct }. (7) core , acore , ocore , faux , paux ), As mentioned in our threat model, the model owner uses signals like querying styles, IPs, and traffic patterns to identify a suspicious set S Chiang et al. (2024); Liu et al. (2026c). TheSowner then collects the evidence from these flagged traces to form a global verification pool CS = τ ∈S Cτ .

3.3

WATERMARK D ETECTION

While the embedding phase injects an aux action into the teacher’s trajectory, the detection phase verifies if a suspect model Ms has learned this specific behavior pattern. To achieve this, AuxMark uses paired probes built from the combined evidence pool CS for verification as shown in Figure 9. Paired probe construction. To verify watermark retention, AuxMark constructs paired probes for the combined evidence pool CS . A probe is formally defined as a prompt-target tuple: given a history and a core action as the prompt, the suspect model Ms is expected to output a specific aux action. Specifically, based on each evidence card, we derive two types of probes: real probes and fake probes. This paired design ensures that the model Ms has genuinely memorized the specific watermarked trace, rather than unconditionally outputting the target action. For the real side construction of each card ci ∈ CS , the system first locates the nearest core action. To ensure that the suspect model Ms has learned the expected behavior, AuxMark creates up to N parameterized variants for this real side by modifying one argument of the original core action, the system prompts the teacher model Mt to synchronously generate a logically consistent core triplet (tjcore , ajcore , ojcore ) alongside the new downstream target aux arguments pjaux . This process directly generates a set of real side probes pireal for the specific card ci as shown in Eq.8:  Ni pireal = (h, tjcore , ajcore , ojcore ), (faux , pjaux ) j=1 . (8) h is the historical context, and faux is the targeted aux tool. Finally, S the system aggregates the real-side probes from all cards to form the global real probe set Preal = ci ∈CS pireal . To better contrast with the real probes Preal , obtain more definitive statistical results, AuxMark constructs a set of fake probes based on the Preal . This verifies whether the aux action stems from memorization rather than natural context or random chance.For this fake side. However, it prompts the teacher model Mt to replace the core action with a semantically similar fake core behavior 5

Preprint

(t̂jcore , âjcore , ôjcore ) that achieves the identical goal. This process generates the fake probe set pifake for card ci , directly matching the variants in pireal as shown in Eq.9:  Ni (9) pifake = (h, t̂jcore , âjcore , ôjcore ), (faux , pjaux ) j=1 . Thus, the strict difference between the paired S probes lies only in the core action. Similar to the real side, the fake probe set is defined as Pfake = ci ∈CS pifake . Card-level statistical test. After generating the paired probes, AuxMark evaluates the suspect model Ms using (Preal , Pfake ) and performs a card-level statistical test which is discussed in Appendix C. For the j-th probe pair of card ci , the expected aux action is faux and arguments pjaux . Let (fout , pout ) be the tool call parsed from the output of Ms . For both the real and fake sides, AuxMark computes the toolhit indicator ti,j and the strict fullhit indicator hi,j ∈ {0, 1}: ti,j = Match(fout , faux ),

hi,j = ti,j · Match(pout , pjaux ).

(10)

where the function Match(·, ·) → {0, 1} evaluates to 1 if and only if there is exact structural and value equivalence between its two arguments, and 0 otherwise. Because the Ni paired variants derived from the same card ci share an identical core-to-aux dependency, they are not independent samples. Therefore, AuxMark aggregates the scores at the card level to calculate the real-side hit rate Rireal , the fake-side hit rate Rifake , and their difference ∆i for each card ci as shown in Eq.11: N

Rireal =

i 1 X hreal , Ni j=1 i,j

N

Rifake =

i 1 X hfake , Ni j=1 i,j

∆i = Rireal − Rifake .

(11)

To evaluate the verification, we establish the null hypothesis H0 and the alternative hypothesis H1 . H0 posits that the suspect model Ms has not learned the watermarked habit; any correct action is due to general context or innate biases. Conversely, H1 asserts that Ms has memorized the watermark, making it significantly more likely to output the target action on real probes. Based on the performance difference ∆i for each card ci , we classify the outcomes into a real win (∆i > 0), a fake win (∆i < 0), or a tie (∆i = 0). Excluding ties, let w and ℓ be the numbers of real and fake wins across CS . The p-value is computed using a one-sided exact sign-test Hollander et al. (2013): w+ℓ X w + ℓ  p−value = 2−(w+ℓ) . (12) k k=w

where p−value = 1 if w + ℓ = 0. To reject H0 and declare a successful verification, AuxMark requires dual criteria: the statistical significance must be p−value P< p, and the averaged real fullhit rate must exceed the fake rate by a predefined margin r (i.e., |C1S | ∆i ≥ r). This empirical margin absorbs unpredictable fluctuations caused by context rewriting in the fake probes and guards against false positives. Appendix B.1 specifies the hyperparameters p, r, and N . 3.4

T RACE -L EVEL ATTRIBUTION

While the detection protocol in Section 3.3 determines whether a suspect model Ms has distilled from Mt , it cannot pinpoint which traces within S were utilized. For exact attribution, AuxMark introduces a trace-level attribution. For each evaluated trace τ , the system computes four rates to capture both exact and partial matches: the strict full-hit rates Rτreal and Rτfake , and the toolhitonly rates Tτreal and Tτfake (where ti,j = 1 ∧ hi,j = 0). To ensure fair comparison across traces, AuxMark normalizes each rate into an average-rank quantile Q(·) ∈ (0, 1): Q(xτ ) =

rankavg (xτ ) . |S| + 1

(13)

The attribution logic is designed as follows: Q(Rτreal ) acts as the primary positive signal for trace memorization. Conversely, if the model outputs the target aux action even on the fake probes, it implies innate biases rather than specific watermark retention. Thus, Q(Rτfake ) serves as a penalty term. Furthermore, the toolhit-only terms address partial memorization, providing supplementary evidence to calibrate the final score. Consequently, AuxMark ranks these traces using a score Sτ : Sτ = Q(Rτreal ) − α′ Q(Rτfake ) + β ′ Q(Tτreal ) − γ ′ Q(Tτfake ). ′

′

′

(14)

where the empirical weighting parameters α , β , and γ are detailed in Appendix B.1. Finally, the system identifies the top-K traces with the highest Sτ most likely included in the training set D. 6

Preprint

4

E XPERIMENTS

4.1

E XPERIMENTAL S ETUP

We evaluate AuxMark in agent distillation. We supervise thoughts and actions from the trajectories and fine-tune the student models using LoRA. The specific training configurations are rank 32, learning rate 2 × 10−4 , and α = 64. Prior traffic-analysis systems have reported identification precision above 70% Chiang et al. (2024); Liu et al. (2026c). We therefore adopt a more conservative 50% precision setting to avoid missing potential distillation traces, with |S| = 100 suspicious trajectories and |Dc | = 50 trajectories used for distillation. Appendix B.1 provides the remaining experimental details, including hyperparameter selection and baselines. Our experiments assess watermark effectiveness, harmlessness, robustness against adversarial processing. Datasets and Models. Our evaluation spans three datasets: BFCL Patil et al. (2025), SWEbench Jimenez et al. (2024), and Telecom (a service-tool domain from the Tau2 benchmark1 ). GPTOSS-120B Agarwal et al. (2025) and Kimi-K2.5 Team et al. (2026) serve as teacher agents (Mt ), with Qwen3-8B acting as both the candidate scorer (Mq ) and safety model (Msafe ). For distillation, we fine-tune four student models (Ms ): Mistral-Small-3.1-24B Liu et al. (2026a), Qwen3-14B Yang et al. (2025), Qwen3-32B Yang et al. (2025), and GLM-4.7-Flash Glm et al. (2024). Evaluation metrics. For detection, we report strict fullhit rates (Rτreal , Rτfake ) on real and fake probes, alongside the one-sided p-value. A successful detection requires p < 0.05 and a realfake mean fullhit gap ≥ 0.05. For trace-level attribution, we measure Precision @K of top-K trajectories belonging to the distillation subset Dc . For harmlessness, we compare task accuracy (Acc) and mean interaction steps (L) between control and watermarked agents. 4.2

E FFECTIVENESS

Table 1: True-positive results. Cards is the number of cards; hit-rate cells report hits / probes (rate). w/ℓ gives card-level real / fake wins. The ‘Sig’ column indicates statistical significance: ✓ denotes p < 0.05, indicating a statistically significant change, whereas × denotes p ≥ 0.05. Teacher AgentMt

Dataset

Cards

GPT-OSS-120B

BFCL

118

Kimi-K2.5

BFCL

145

Student Model Ms Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B

Rreal (%) 93 / 182 (51.1%) 108 / 182 (59.3%) 44 / 182 (24.2%) 89 / 182 (48.9%) 175 / 379 (46.2%) 155 / 379 (40.9%) 102 / 379 (26.9%) 82 / 379 (21.6%)

Detection Results Rfake (%) w/ℓ 59 / 182 (32.4%) 25 / 1 80 / 182 (44.0%) 19 / 1 28 / 182 (15.4%) 20 / 4 58 / 182 (31.9%) 30 / 9 105 / 379 (27.7%) 45 / 10 69 / 379 (18.2%) 47 / 4 26 / 379 (6.9%) 46 / 4 18 / 379 (4.7%) 39 / 5

p-value 4.0 × 10−7 2.0 × 10−5 7.7 × 10−4 5.3 × 10−4 1.0 × 10−6 1.2 × 10−10 2.2 × 10−10 7.0 × 10−8

Sig. ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

In this section, all student models are trained on the original dataset Dc . We test our scheme on distilled and clean models via separate true positive and false positive experiments. We also evaluate the performance of existing agent watermark baselines under the matched setting. Furthermore, we evaluate AuxMark’s trace-level attribution using Precision @K in Appendix B.2. Finally, we measure how identification accuracy affects statistical significance in Appendix B.2. True positives. We first evaluate whether AuxMark can successfully identify student models distilled from watermarked trajectories. Our experiments cover all combinations of 3 datasets (SWE and Telecom are in Appendix B.2), 2 teacher agents, and 4 student models. In each setting, we compare the hit rates between the real side and the fake side and calculate the p-value. The results in the Table 1 and 6 show that AuxMark successfully achieves detection in all settings. These results show that AuxMark can extract watermark from student models of different architectures. False positives. We evaluate whether AuxMark causes false positives on clean models that have not seen watermarked trajectories. As shown in Table 2 and Table 7 in Appendix B.2, in addition to the four undistilled student models, we also select four base models (DeepSeek-V4-Flash Xu et al. (2026a), MiniMax-M2.5 Chen et al. (2026), Qwen3.5-Flash Yang et al. (2025), and MiMo-V2.5 Xiao et al. (2026)) and build a total of 48 controlled false-positive evaluation settings. The results show that AuxMark achieves zero false positives across all 48 settings. 1

https://github.com/sierra-research/tau2-bench

7

Preprint

Table 2: False-positive results. The ‘Sig’ column indicates statistical significance: ✓ denotes p < 0.05, indicating a statistically significant change, whereas × denotes p ≥ 0.05. Teacher AgentMt

Dataset

Cards

GPT-OSS-120B

BFCL

118

Kimi-K2.5

BFCL

145

Clean Model Mc

Rreal (%) 10 / 182 (5.5%) 9 / 182 (4.9%) 9 / 182 (4.9%) 7 / 182 (3.8%) 4 / 182 (2.2%) 8 / 182 (4.4%) 10 / 182 (5.5%) 5 / 182 (2.7%) 10 / 379 (2.6%) 6 / 379 (1.6%) 1 / 379 (0.3%) 1 / 379 (0.3%) 1 / 379 (0.3%) 0 / 379 (0.0%) 6 / 379 (1.6%) 2 / 379 (0.5%)

DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash

Detection Results Rfake (%) w/ℓ 11 / 182 (6.0%) 2/3 10 / 182 (5.5%) 2/3 7 / 182 (3.8%) 3/2 6 / 182 (3.3%) 2/1 4 / 182 (2.2%) 2/2 8 / 182 (4.4%) 2/1 11 / 182 (6.0%) 2/2 1 / 182 (0.5%) 4/0 9 / 379 (2.4%) 5/4 3 / 379 (0.8%) 4/2 3 / 379 (0.8%) 1/2 0 / 379 (0.0%) 1/0 0 / 379 (0.0%) 1/0 0 / 379 (0.0%) 0/0 6 / 379 (1.6%) 4/3 2 / 379 (0.5%) 2/1

p-value 8.1 × 10−1 8.1 × 10−1 5.0 × 10−1 5.0 × 10−1 6.9 × 10−1 5.0 × 10−1 6.9 × 10−1 6.2 × 10−2 5.0 × 10−1 3.4 × 10−1 8.8 × 10−1 5.0 × 10−1 5.0 × 10−1 1.0 × 100 5.0 × 10−1 5.0 × 10−1

Sig. × × × × × × × × × × × × × × × ×

Baseline effectiveness. We evaluate the effectiveness of the two baseline schemes under the matched distillation setting and the Agentwm SeqWM Dataset Student Model Ms Passes Sig. Median p Sig. details about baselines are in Appendix B.1. Mistral-24B 1/5 × 0.6848 × As shown in Table 3, for Agentwm, only GLM-4.7-Flash 3/5 ✓ 0.5130 × BFCL Qwen3-14B 1/5 × 0.4650 × the GLM-4.7-Flash student model meets the Qwen3-32B 2/5 × 0.5744 × passing criterion of ≥ 3/5 passes, while all Mistral-24B 1/5 × 0.6274 × GLM-4.7-Flash 3/5 ✓ 0.5914 × other students fail. For Seqwm, the median SWE-bench Qwen3-14B 1/5 × 0.7073 × p-values across all tested models are greater Qwen3-32B 0/5 × 0.4895 × than 0.05, failing to successfully detect any distilled models. As expected, baselines fail, because Seqwm is not designed for antidistillation, and Agentwm has weak watermark signals that requires identical teacher and student base models. Table 3: Baseline effectiveness. Agentwm is significant at ≥ 3/5 passes; SeqWM at median p < 0.05.

4.3

H ARMLESSNESS

Harmlessness. A reasonable watermark scheme should not noticeably affect the performance of the teacher agent. Therefore, we evaluate impact of AuxMark on the teacher agent Mt ’s task utility and inference cost. For the four cohorts in BFCL and Telecom, we compare the unwatermarked agent Mtc and the watermarked teacher agent Mt on 200 tasks. We measure the task accuracy difference ∆Acc = E[Acc(Mt ) − Acc(Mtc )] and the expected interaction step difference ∆L = E[L(Mt ) − L(Mtc )]. As summarized in Figure 3, embedding the watermark does not cause a drop in overall performance. The task accuracy difference ∆Acc ranges only from −3.0% to +3.0%, indicating that the accuracy loss is acceptable. Meanwhile, the increase in interaction cost is acceptable: the average step count increases from 20.49 to 21.09 steps, and ∆L ranges from −1.38 to +2.04 steps.

48.5 45.5

40 20 0

GPT

Kimi

(a) BFCL: accuracy

12

25 20

17.0

15

15.0 12.0

10 5 0

40

9.79

18.0

GPT

9 8.03

(b) Telecom: accuracy

6.93

6 3 0

Kimi

8.97

GPT

Kimi

(c) BFCL: steps

Mean steps

60

65.5

Mean steps

64.5

Task accuracy (\%)

Task accuracy (\%)

80

30

31.60 30.22

35.40 35.38

20 10 0

GPT

Kimi

(d) Telecom: steps

Figure 3: Harmlessness results. In (a) and (b), the vertical axis reports accuracy; in (c) and (d), it reports the mean number of steps. Colors distinguish ■ unwatermarked and ■ watermarked runs. 4.4

ROBUSTNESS

To demonstrate the robustness of our watermark scheme, We evaluate AuxMark’s robustness against four common pre-distillation data transformations. These operations include: data flooding, paraphrasing attack, adaptive attack, advanced data flooding and truncation (the last two are in 8

Preprint

Appendix B.3). We evaluate the robustness of Agentwm under data flooding in Appendix B.3. We use GPT-OSS-120B as the teacher agent. Data flooding. Before distillation, an attacker may mix the training set with standard clean data to improve distillation performance Ouyang et al. (2022); Mukherjee et al. (2023), which will dilutes the watermark signals. To evaluate its impact on our watermark strength, we mix the original watermarked dataset Dc with clean standard answers. Specifically, in the D1, D5, and D10 settings, we incorporate 1×, 5×, and 10× the amount of clean data relative to Dc . We evaluate a total of 24 settings across 2 benchmarks (BFCL and SWE-bench), 3 dilution ratios, and 4 student models. As shown in Figure 4, AuxMark successfully detects every setting. Even under the strongest 1:10 dilution ratio, all 8 tested student models remain successfully detected.

4.2

3 2 p = 0.05 1 Mistral GLM Q14

2.1 2 p = 0.05 1 Mistral GLM Q14

10

(d) SWE-bench: D1

20

5

10.9

10.3

7 5.5 6 5 4.2 4 3.0 3 2 p = 0.05 1 0 Mistral GLM Q14

25

16.6

15

0

Q32

Q32

(b) BFCL: D5 15.6

9.7

4.6

2.9

3

20

−log10 p

−log10 p

4

0

Q32

(a) BFCL: D1

14 12.1 12 9.1 10 8 6 4 3.4 2 p = 0.05 0 Mistral GLM Q14

4.7

5

4

0

6

5.3 4.5

−log10 p

5.3

−log10 p

−log10 p

5

−log10 p

6

Mistral GLM Q14

(e) SWE-bench: D5

Q32

Q32

(c) BFCL: D10 18.3 13.4

15

13.1

11.3

10 5

p = 0.05

4.3

0

p = 0.05

Mistral GLM Q14

Q32

(f) SWE-bench: D10

Figure 4: Data-flooding robustness. The vertical axis reports − log10 p. In (a)–(f), bar colors denote marks p = 0.05. ■ Mistral-24B, ■ GLM-4.7-Flash, ■ Qwen3-14B, and ■ Qwen3-32B. The Paraphrasing attack. An attacker may paraphrase Table 4: Paraphrasing attack robustness. Dataset Student Model Ms p-value Sig. the trajectories, aiming to alter semantic patterns and 5.7 × 10−5 ✓ Mistral-24B evade watermark detection while preserving the utilGLM-4.7-Flash 1.4 × 10−4 ✓ BFCL ity Krishna et al. (2023). To evaluate the robustness −3 Qwen3-14B 3.3 × 10 ✓ against this rewriting attack, we use DeepSeek-V4Qwen3-32B 9.8 × 10−5 ✓ Flash to paraphrase the Dc . Subsequently, we train 4 Mistral-24B 3.2 × 10−2 ✓ GLM-4.7-Flash 1.4 × 10−5 ✓ student models on these rewritten datasets across the Telecom Qwen3-14B 3.5 × 10−3 ✓ BFCL and Telecom and detect their watermark sig−3 Qwen3-32B 4.1 × 10 ✓ nals. As reported in Table 4, AuxMark successfully detects the watermark across all 8 settings showing our scheme’s robustness under paraphrasing. Adaptive attack. In this section, we consider a stronger attacker who knows that our watermark relies on auxiliary behaviors. To avoid detection, the attacker uses DeepSeek-V4-Flash to clean the colBFCL lected data by finding and removing suspected auxiliary action while keeping the trajectories useful. Next, we fine-tune student models on this cleaned Telecom data and test them on BFCL and Telecom. As shown in Table 5, even with this removal, AuxMark successfully detects all student models (with the minimal p-value of 9.6 × 10−3 ). This shows that our watermark is highly robust, even when an attacker actively tries to move the signals.

Table 5: Adaptive attack robustness. Dataset

5

Student Model Ms Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B

p-value 9.2 × 10−4 1.4 × 10−4 9.6 × 10−3 4.1 × 10−4 2.8 × 10−8 3.7 × 10−3 1.0 × 10−11 5.5 × 10−12

Sig. ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

C ONCLUSION

We present AuxMark, a behavioral watermarking framework to track the unauthorized distillation of LLM agents. AuxMark inserts useful extra tool calls after normal core actions. To verify 9

Preprint

if the watermark is copied, it uses paired real and fake probes along with a card-level sign test. This design supports both model-level detection and trace-level tracking.In experiments across three benchmarks, two teacher models, and four student architectures, AuxMark successfully detects all 24 distilled students with zero false positives on 48 clean models. It also remains highly effective under data flooding, rewriting, truncation, and adaptive cleaning attacks. Future work will extend this evaluation to more teacher models, cleaning tools, and training methods.

AI USE STATEMENT In this work, we used generative AI tools for research execution, generating synthetic datasets, and code generation. Specifically, large language models were employed to generate agent-interaction trajectories, construct paired verification probes for our evaluations, and assist with implementation. We have not used generative AI tools for drafting core sections of the paper or proving mathematical claims, and the rest of the required disclosure tasks are not applicable to this work. Additionally, we used generative AI tools (Gemini 3.1 Pro) for aiding and polishing the English writing of the manuscript. We have reviewed all AI-assisted work. The authors manually assessed AI-assisted research ideas through a literature survey and technical discussion. For the synthetic datasets, the LLM-generated agent trajectories and verification probes were programmatically validated and manually sampled to ensure they form valid, executable tool-use sequences that strictly adhere to the predefined tool schemas and task objectives. All AI-generated code was reviewed, tested, and, where necessary, revised by the authors before use. For the writing assistance, all AI-polished text was thoroughly reviewed and edited by the authors to guarantee it accurately reflects our original scientific intent, methodology, and conclusions without introducing hallucinations or overclaims. We take responsibility for the final content of this work, including text, claims, or artifacts produced with the aid of generative AI.

R EPRODUCIBILITY S TATEMENT To ensure the full reproducibility of our results, we have made our complete source code, benchmark adapters, and evaluation scripts publicly available via an anonymous repository (https://github.com/qx041609/Auxmark). Additionally, all trained student model weights (e.g., LoRA adapters) necessary to replicate our detection and trace-level attribution experiments are hosted on Hugging Face (huggingface.co/AuxMark/AuxMark/tree/main).

R EFERENCES Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025. Hyeseon An, Shinwoo Park, Dongsu Kim, and Yo-Sub Han. Sequential behavioral watermarking for llm agents. arXiv preprint arXiv:2605.11036, 2026. Anthropic. Introducing Claude Opus 4.8, May 2026. URL https://www.anthropic.com/ news/claude-opus-4-8. Released May 28, 2026. Anthropic. Detecting and preventing distillation attacks. https://www.anthropic.com/ news/detecting-and-preventing-distillation-attacks, 2026. [Accessed 17-09-2026]. Aili Chen, Aonian Li, Baichuan Zhou, Bangwei Gong, Binyang Jiang, Boji Dan, Changhao Zhang, Changqing Yu, Chao Wang, Cheng Ma, et al. The minimax-m2 series: Mini activations unleashing max real-world intelligence. arXiv preprint arXiv:2605.26494, 2026. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132, 2024. 10

Preprint

Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, et al. Ultrafeedback: Boosting language models with scaled ai feedback. arXiv preprint arXiv:2310.01377, 2023. Team Glm, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, et al. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024. Google Threat Intelligence Group. Gtig ai threat tracker: Distillation, experimentation, and (continued) integration of ai for adversarial use. https: //cloud.google.com/blog/topics/threat-intelligence/ distillation-experimentation-integration-ai-adversarial-use, 2026. [Accessed 17-09-2026]. Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song. The false promise of imitating proprietary llms. arXiv preprint arXiv:2305.15717, 2023. Xuanli He, Qiongkai Xu, Lingjuan Lyu, Fangzhao Wu, and Chenguang Wang. Protecting intellectual property of language generation apis with lexical watermark. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 10758–10766, 2022a. Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, and Ruoxi Jia. Cater: Intellectual property protection on text generation apis via conditional watermarks. Advances in Neural Information Processing Systems, 35:5431–5445, 2022b. Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. Myles Hollander, Douglas A Wolfe, and Eric Chicken. Nonparametric statistical methods. John Wiley & Sons, 2013. Abe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, and Yulia Tsvetkov. Semstamp: A semantic watermark with paraphrastic robustness for text generation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 4067–4082, 2024. Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. In Findings of the association for computational linguistics: ACL 2023, pp. 8003–8017, 2023. Kaibo Huang, Zipei Zhang, Zhongliang Yang, and Linna Zhou. Agent guide: A simple agent behavioral watermarking framework. arXiv preprint arXiv:2504.05871, 2025. Jun Ho Huh, Hyejin Shin, Sunwoo Ahn, Hayoon Yi, Joonho Cho, Taewoo Kim, Minchae Lim, and Nuel Choi. Preventing artificially inflated {SMS} attacks through {Large-Scale} traffic inspection. In 34th USENIX Security Symposium (USENIX Security 25), pp. 5405–5423, 2025. Jiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang, Yibo Yan, Aiwei Liu, Xuming Hu, and Mingxun Zhou. Pmark: Towards robust and distortion-free semantic-level watermarking with channel constraints. In International Conference on Learning Representations, volume 2026, pp. 27100–27126, 2026a. Jiahao Huo, Wenjie Qu, Yibo Yan, Kening Zheng, Jiaheng Zhang, Xuming Hu, Philip S Yu, and Mingxun Zhou. Samark: A self-anchored text watermarking with paragraph-level paraphrase robustness. arXiv preprint arXiv:2605.25796, 2026b. Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? In International Conference on Learning Representations, volume 2024, pp. 54107–54157, 2024. 11

Preprint

Minki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho, and Sung Ju Hwang. Distilling llm agent into small models with retrieval and code tools. Advances in Neural Information Processing Systems, 38:106501–106538, 2026. John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. A watermark for large language models. In International conference on machine learning, pp. 17061–17084. PMLR, 2023. Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, and Mohit Iyyer. Paraphrasing evades detectors of ai-generated text, but retrieval is an effective defense. Advances in neural information processing systems, 36:27469–27500, 2023. Hao Li, Haoxiang Zhang, and Ahmed E Hassan. Aidev: studying ai coding agents on github. In Proceedings of the 23rd International Conference on Mining Software Repositories, pp. 1029– 1033, 2026. Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen. A semantic invariant robust watermark for large language models. In International Conference on Learning Representations, volume 2024, pp. 6499–6519, 2024. Alexander H Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien Sadé, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584, 2026a. Jun Liu, Zhenglun Kong, Peiyan Dong, Changdi Yang, Tianqin Li, Yanyue Xie, Yifan Gong, Xuan Shen, Pu Zhao, Hao Tang, et al. Structured agent distillation for large language model agents. In Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems, pp. 3676–3685, 2026b. Shuze Liu, Qianwen Guo, and Yushun Dong. An embarrassingly simple detector for model extraction attacks in large language model api traffic. arXiv preprint arXiv:2606.05725, 2026c. Yinyi Luo, Yiqiao Jin, Weichen Yu, Mengqi Zhang, Srijan Kumar, Xiaoxiao Li, Weijie Xu, Xin Chen, and Jindong Wang. Agentark: Distilling multi-agent intelligence into a single llm agent. arXiv preprint arXiv:2602.03955, 2026. Yuanjie Lyu, Chengyu Wang, Jun Huang, and Tong Xu. From correction to mastery: Reinforced distillation of large language model agents. arXiv preprint arXiv:2509.14257, 2025. Xinhang Ma, William Yeoh, Ning Zhang, and Yevgeniy Vorobeychik. Protecting language models against unauthorized distillation through trace rewriting. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11307–11324, 2026. Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, and Ahmed Awadallah. Orca: Progressive learning from complex explanation traces of gpt-4. arXiv preprint arXiv:2306.02707, 2023. OpenAI. Updated Stakes for American-Led, Democratic AI. https://assets.bwbx. io/documents/users/iqjWHBFdfxIU/rRmql_jJcxb4/v0, 2026. [Accessed 17-092026]. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 27730–27744, 2022. Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, and Maksym Andriushchenko. Stealing reasoning traces from proprietary llm apis. arXiv preprint arXiv:2608.09867, 2026. 12

Preprint

Shishir G Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E Gonzalez. The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models. In Forty-second International Conference on Machine Learning, 2025. Bart Preneel. Cryptographic hash functions. European Transactions on Telecommunications, 5(4): 431–448, 1994. Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al. Toolllm: Facilitating large language models to master 16000+ real-world apis. In International Conference on Learning Representations, volume 2024, pp. 9695–9717, 2024. Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, and Sewon Min. Referencebased distillation detection in llms. arXiv preprint arXiv:2607.09692, 2026. Sha Sajadieh, Loredana Fattorini, Raymond Perrault, Yolanda Gil, Vanessa Parli, Lapo Santarlasci, Juan Pava, Nestor Maslej, Russ Altman, Erik Brynjolfsson, et al. Artificial intelligence index report 2026. arXiv preprint arXiv:2606.15708, 2026. Yash Savani, Asher Trockman, Zhili Feng, Yixuan Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, and Zico Kolter. Antidistillation sampling. Advances in Neural Information Processing Systems, 38:117800–117827, 2026. Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025. Kimi Team, Yifan Bai, Yiping Bao, Y Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, et al. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534, 2025. Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Ziwei Chai, Y Charles, HS Che, Cheng Chen, et al. Kimi k2. 5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026. Liwen Wang, Zongjie Li, Yuchong Xie, Shuai Wang, Dongdong She, Wei Wang, and Juergen Rahmel. On protecting agentic systems’ intellectual property via watermarking. arXiv preprint arXiv:2602.08401, 2026. Bangjun Xiao, Bingquan Xia, Bo Yang, Bofei Gao, Bowen Shen, Chen Zhang, Chenhong He, Chiheng Lou, Fuli Luo, Gang Wang, et al. Mimo-v2-flash technical report. arXiv preprint arXiv:2601.02780, 2026. Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, et al. Deepseek-v4: Towards highly efficient milliontoken context intelligence. arXiv preprint arXiv:2606.19348, 2026a. Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein, Fei Fang, and J Zico Kolter. Antidistillation fingerprinting. arXiv preprint arXiv:2602.03812, 2026b. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Chenghao Yang, Yuning Zhang, Zhoufutu Wen, Tao Gong, Jiaheng Liu, Qi Chu, and Nenghai Yu. When agents look the same: Quantifying distillation-induced similarity in tool-use behaviors. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 10482–10502, 2026a. Guang Yang, Amir Ghasemian, Fengchen Liu, Zhong Wang, Ninareh Mehrabi, and Homa Hosseinmardi. Asking back: Interaction-layer antidistillation watermarks. arXiv preprint arXiv:2605.16462, 2026b. 13

Preprint

Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.

Haobo Zhang, Xutao Mao, Guangyuan Dong, Ziwei Li, Xuanbo Su, Kaijie Chen, Jing Yang, and Zheng Lin. Memmark: State-evolution attribution watermarking for agent long-term memory systems. arXiv preprint arXiv:2605.25002, 2026.

A

R ELATED W ORK

A.1

M ODEL D ISTILLATION

Knowledge distillation is originally proposed as to transfer the capicity of a teacher model to a student model Hinton et al. (2015). With rapid development of LLM, fine-tuning on teacher-generated data has become an increasingly important form of distillation, allowing student models to acquire advanced capabilities directly from the teacher’s responses Hsieh et al. (2023); Mukherjee et al. (2023). This paradigm has also been extended to agent systems Kang et al. (2026); Liu et al. (2026b); Luo et al. (2026). Agent Distillation Kang et al. (2026) proposed an agent distillation method that enables student model to learn both tool-use behaviors and reasoning capabilities. AgentArk Luo et al. (2026) introduced a method that distills multi-agent dynamics into the weights of a single model. Agent trajectories record not only generated text but also tool use and multi-step decisions, making them valuable supervision for distillation. As these trajectories can be used to cheaply transfer the capabilities of teacher agents, unauthorized distillation poses a growing threat to model copyright and calls for effective detection and tracing mechanisms.

A.2

M ODEL WATERMARKING

Content watermarking. Recent studies have proposed a series of content watermarking methods to identify whether a given output is generated by a specific model. Early methods mainly embed statistical signals during token generation, such as the red-green list mechanism Kirchenbauer et al. (2023). Later studies further improve the robustness and generation quality of watermarks by using semantic information Liu et al. (2024); Huo et al. (2026a;b); Hou et al. (2024). However, with the development of LLM agents, traditional content watermarks are difficult to directly apply to agent settings. These methods usually embed watermark signals by changing token choices or semantic expressions in free-form text, while the key outputs of agents are often tool calls with strict format requirements. Directly modifying these outputs may break the tool calls and affect correct execution. Therefore, recent studies have started to move watermarks from generated content to agent trajectories and behavior patterns Huang et al. (2025); An et al. (2026); Zhang et al. (2026). Model watermarking against distillation. Several watermarking methods have been proposed to protect models against unauthorized distillation. Methods such as Lexical Watermark and CATER insert special word patterns into teacher outputs, so that student models trained on these outputs can inherit detectable features He et al. (2022a;b); Yang et al. (2026b). Recent work further targets LLM distillation by changing token generation or rewriting reasoning traces, either reducing the value of teacher outputs for distillation or leaving detectable signals in the student model Savani et al. (2026); Xu et al. (2026b); Ma et al. (2026). However, these methods mainly work on text tokens or reasoning traces, and are difficult to directly apply to agents whose behavior is mainly represented by structured tool calls and multi-step interactions. Recent work has extended anti-distillation watermarking to agent trajectories Wang et al. (2026), but its main evaluation still focuses on settings where the teacher and student models use the same base model, and its robustness to data processing such as data mixing remains limited. Therefore, a more general and robust watermarking method is still needed for tracing agent distillation. 14

Preprint

B

S UPPLEMENT OF E XPERIMENTS

B.1

E XPERIMENTAL D ETAILS

Hyperparameter selection. We select AuxMark hyperparameters through preliminary experiments to balance stealthiness, utility and detection reliability. For dynamic scheduling, we set p0 = 0.1, ∆p = 0.1, pmax = 0.5, and the maximum budget B = 10. The system generates K = 3 candidates, and the scoring weights in Eq. 5 are α = 0.35, β = 0.25, and γ = 0.40. The repeatability score is sjrep = 0 for nj = 0, 0.6 for 1 ≤ nj < 5, and 1 for nj ≥ 5. For detection, we generate N = 3 parameterized variants per card and use a significance level p = 0.05 with margin threshold r = 0.05. For trace-level attribution in Eq. 14, we set α′ = 0.25, β ′ = 0.10, and γ ′ = 0.10. Baselines. We evaluate our approach against two representative agent watermarking baselines. We exclude large language model antidistillation watermarks Xu et al. (2026b); Savani et al. (2026) from our comparison, as agent environments demand formatted outputs and modifying tokens or reasoning traces can cause format corruption and execution failures. Thus, we focus on agent schemes An et al. (2026); Huang et al. (2025); Luo et al. (2026). We select Seqwm An et al. (2026) and Agentwm Wang et al. (2026) because Seqwm demonstrates better robustness among trace watermarks generated for agents, and Agentwm specifically targets agent anti-distillation. Their available implementations use different watermark encodings, test units, and decision statistics from AuxMark; we therefore report each baseline in its native metric. Specifically, we train the student models Ms using the distillation dataset Dc watermarked by each respective scheme, and then have them replay 100 tasks from the set S. For Agentwm, we collect unwatermarked traces using four distinct models to estimate the natural initial distribution Pc of synonymous tool sets. The scheme then modifies this Pc into a watermarked distribution Pwm . During detection, it extracts 5 fixed synonymous tool sets (referred to as passes) from the 100 task replays. A pass is considered successfully detected only when the JSD between the student’s empirical distribution and the distribution Pwm is at most 0.015. The student model is flagged as significant only if at least 3 out of 5 passes are detected (≥ 3/5). For Seqwm, it evaluates each replayed trace against a background null distribution generated from 1,000 incorrect keys. It uses the median p-value of these 100 test traces to determine whether the student model Ms is distilled, requiring a median p < 0.05 for a successful detection. B.2

S UPPLEMENTARY E FFECTIVENESS R ESULTS

Table 6: Supplementary true-positive results for SWE-bench and Telecom. Cards is the number of cards; hit-rate cells report hits / probes (rate). w/ℓ gives card-level real / fake wins. The Sig. column indicates statistical significance: ✓ denotes p < 0.05, whereas × denotes p ≥ 0.05. Teacher AgentMt

Dataset

Cards

SWE-bench

274

Telecom

125

SWE-bench

262

Telecom

122

GPT-OSS-120B

Kimi-K2.5

Student Model Ms Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B Mistral-24B GLM-4.7-Flash Qwen3-14B Qwen3-32B

Rreal (%) 60 / 387 (15.5%) 81 / 387 (20.9%) 67 / 387 (17.3%) 72 / 387 (18.6%) 83 / 146 (56.8%) 76 / 146 (52.1%) 86 / 146 (58.9%) 85 / 146 (58.2%) 85 / 437 (19.5%) 80 / 437 (18.3%) 48 / 437 (11.0%) 45 / 437 (10.3%) 68 / 192 (35.4%) 56 / 192 (29.2%) 60 / 192 (31.2%) 66 / 192 (34.4%)

Detection Results Rfake (%) w/ℓ 18 / 387 (4.7%) 46 / 6 21 / 387 (5.4%) 67 / 6 18 / 387 (4.7%) 54 / 4 20 / 387 (5.2%) 58 / 6 46 / 146 (31.5%) 44 / 7 42 / 146 (28.8%) 42 / 4 44 / 146 (30.1%) 48 / 6 36 / 146 (24.7%) 57 / 7 50 / 437 (11.4%) 38 / 9 42 / 437 (9.6%) 34 / 5 17 / 437 (3.9%) 33 / 8 14 / 437 (3.2%) 27 / 1 48 / 192 (25.0%) 37 / 14 28 / 192 (14.6%) 33 / 6 36 / 192 (18.8%) 33 / 9 43 / 192 (22.4%) 27 / 6

p-value 5.2 × 10−9 2.0 × 10−14 1.6 × 10−12 4.5 × 10−12 6.1 × 10−8 2.6 × 10−9 1.6 × 10−9 3.8 × 10−11 1.2 × 10−5 1.2 × 10−6 5.6 × 10−5 1.1 × 10−7 8.8 × 10−4 7.1 × 10−6 1.4 × 10−4 1.6 × 10−4

Sig. ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓

Trace-level detection. After a successful model-level detection, we further examine whether AuxMark can identify the specific trajectories in the suspicious set S that are most likely used for distillation. We separately rank the scores Sτ of the 100 trajectories from S on the four student models and report the arithmetic mean of Precision @K across the four students in Table 8. The overall average precision across the six cohorts remains at least 82.1% for every reported K. SWE15

Preprint

Table 7: Supplementary false-positive results for SWE-bench and Telecom. Cards is the number of cards; hit-rate cells report hits / probes (rate). w/ℓ gives card-level real / fake wins. The Sig. column indicates statistical significance: ✓ denotes p < 0.05, whereas × denotes p ≥ 0.05. Teacher AgentMt

Dataset

Cards

SWE-bench

274

Telecom

125

SWE-bench

262

Telecom

122

GPT-OSS-120B

Kimi-K2.5

Clean Model Mc

Rreal (%) 0 / 387 (0.0%) 5 / 387 (1.3%) 15 / 387 (3.9%) 0 / 387 (0.0%) 0 / 387 (0.0%) 4 / 387 (1.0%) 3 / 387 (0.8%) 1 / 387 (0.3%) 18 / 146 (12.3%) 21 / 146 (14.4%) 11 / 146 (7.5%) 17 / 146 (11.6%) 17 / 146 (11.6%) 11 / 146 (7.5%) 13 / 146 (8.9%) 1 / 146 (0.7%) 12 / 437 (2.7%) 14 / 437 (3.2%) 20 / 437 (4.6%) 1 / 437 (0.2%) 2 / 437 (0.5%) 15 / 437 (3.4%) 10 / 437 (2.3%) 5 / 437 (1.1%) 11 / 192 (5.7%) 23 / 192 (12.0%) 19 / 192 (9.9%) 10 / 192 (5.2%) 10 / 192 (5.2%) 9 / 192 (4.7%) 10 / 192 (5.2%) 1 / 192 (0.5%)

DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash DeepSeek-V4-Flash MiniMax-M2.5 Mistral-24B Qwen3-14B Qwen3-32B Qwen3.5-Flash MiMo-V2.5 GLM-4.7-Flash

Detection Results Rfake (%) w/ℓ 0 / 387 (0.0%) 0/0 2 / 387 (0.5%) 4/1 10 / 387 (2.6%) 10 / 6 3 / 387 (0.8%) 0/3 2 / 387 (0.5%) 0/2 5 / 387 (1.3%) 4/3 3 / 387 (0.8%) 2/2 2 / 387 (0.5%) 1/2 13 / 146 (8.9%) 10 / 5 19 / 146 (13.0%) 6 / 3 9 / 146 (6.2%) 6/5 21 / 146 (14.4%) 3 / 6 18 / 146 (12.3%) 2 / 3 8 / 146 (5.5%) 6/3 13 / 146 (8.9%) 8/8 0 / 146 (0.0%) 1/0 5 / 437 (1.1%) 8/3 7 / 437 (1.6%) 9/4 11 / 437 (2.5%) 10 / 4 1 / 437 (0.2%) 1/1 2 / 437 (0.5%) 2/2 13 / 437 (3.0%) 6/4 6 / 437 (1.4%) 8/5 1 / 437 (0.2%) 4/0 8 / 192 (4.2%) 4/1 18 / 192 (9.4%) 9/5 16 / 192 (8.3%) 14 / 11 8 / 192 (4.2%) 7/5 5 / 192 (2.6%) 6/1 3 / 192 (1.6%) 6/1 7 / 192 (3.6%) 6/4 0 / 192 (0.0%) 1/0

p-value 1.0 × 100 1.9 × 10−1 2.3 × 10−1 1.0 × 100 1.0 × 100 5.0 × 10−1 6.9 × 10−1 8.8 × 10−1 1.5 × 10−1 2.5 × 10−1 5.0 × 10−1 9.1 × 10−1 8.1 × 10−1 2.5 × 10−1 6.0 × 10−1 5.0 × 10−1 1.1 × 10−1 1.3 × 10−1 9.0 × 10−2 7.5 × 10−1 6.9 × 10−1 3.8 × 10−1 2.9 × 10−1 6.2 × 10−2 1.9 × 10−1 2.1 × 10−1 3.5 × 10−1 3.9 × 10−1 6.2 × 10−2 6.2 × 10−2 3.8 × 10−1 5.0 × 10−1

Sig. × × × × × × × × × × × × × × × × × × × × × × × × × × × × × × × ×

bench–GPT achieves a perfect 100.0% at P@10. These results show that AuxMark can identify the specific leaked data with high precision across all six cohorts. Table 8: Trace-level attribution accuracy. Each entry is the arithmetic mean of Precision @K over the four student models in the cohort. Teacher Agent Mt GPT-OSS-120B Kimi-K2.5 Overall average

Dataset BFCL SWE-bench Telecom BFCL SWE-bench Telecom

P@10 72.5% 100.0% 75.0% 82.5% 85.0% 80.0% 82.5%

P@15 80.0% 98.3% 78.3% 81.7% 86.7% 83.3% 84.7%

P@20 80.0% 96.2% 80.0% 82.5% 86.2% 82.5% 84.6%

P@25 79.0% 96.0% 82.0% 84.0% 82.0% 75.0% 83.0%

P@30 81.7% 94.2% 84.2% 80.8% 80.8% 70.8% 82.1%

Impact on identification accuracy. To evaluate the impact of identification accuracy on detection accuracy, we simulate different training shares by resampling our empirical results. In practice, upstream traffic monitoring is often tuned for high recall to avoid missing distillation requests, which can introduce additional benign or unrelated trajectories into the suspicious set S. We test AuxMark under lower shares down to 10%. Specifically, we divide the evaluated cards for each cohort and student model into trained and untrained groups. Next, we estimate their outcome distributions and resample these two groups at various target ratios. The 50% share serves as the observed baseline. For other simulated shares, we perform 2,000 independent draws and apply the same one-sided exact sign test. As shown in Figure 5, which reports the median p-value, all 24 settings remain significant at a 20% training share. At 10%, 21 of 24 settings remain significant.

16

30

20

10

Training share (\%)

−log10 p

−log10 p

10 8 6 4 2 p = 0.05 30

20

0

10

50

20

10

8 7 6 5 4 3 2 p = 0.05 1 0 50 40

(d) BFCL / Kimi

30

20

10

(c) Telecom / GPT 6 5 4 3 2 p = 0.05 1

30

20

10

0

50

Training share (\%)

Training share (\%)

40

Training share (\%)

(b) SWE-bench / GPT-OSS-120B

−log10 p 30

12

Training share (\%)

(a) BFCL / GPT 11 10 9 8 7 6 5 4 3 2 p = 0.05 1 0 50 40

16 14 12 10 8 6 4 2 p = 0.05 0 50 40

−log10 p

7 6 5 4 3 2 p = 0.05 1 0 50 40

−log10 p

−log10 p

Preprint

40

30

20

10

Training share (\%)

(e) SWE-bench / Kimi-K2.5

(f) Telecom / Kimi

Figure 5: Impact of identification accuracy. The horizontal axis denotes the training share and the vertical axis denotes − log10 p, where larger values indicate stronger evidence. In (a)–(f), the colored lines and hollow markers denote student models: ◦ Mistral-24B, □ GLM-4.7-Flash, △ Qwen3-14B, and ⋄ Qwen3-32B. The marks p = 0.05. S UPPLEMENTARY OF ROBUSTNESS

9.0

8 5.4

6

4.2

3.8

Mistral GLM Q14

Q32

4

−log10 p

−log10 p

10

2 p = 0.05

25

(a) BFCL: R1 19.2

−log10 p

20 15

14.0

20 13.5 9.0

10 5 0

25

Mistral GLM Q14

Q32

6 5

18.3

4

20 10.1

Mistral GLM Q14

(e) SWE-bench: R2

Q32

1.9

Mistral GLM Q14

Q32

(c) BFCL: R3 19.9 13.2

15

14.9

12.1

10 5

p = 0.05

4.0

3.9

2 p = 0.05 1

25

12.9

5.0

3

0

Q32

20.7

10

0

6.4

(b) BFCL: R2

15

5

p = 0.05

(d) SWE-bench: R1

−log10 p

0

7 5.8 6 4.9 5 4 3 2.3 2 p = 0.05 1 0 Mistral GLM Q14

−log10 p

12

−log10 p

B.3

0

p = 0.05

Mistral GLM Q14

Q32

(f) SWE-bench: R3

Figure 6: Advanced data flooding robustness. The horizontal axis lists the four student models and the vertical axis reports − log10 p. In (a)–(f), bar colors denote ■ Mistral-24B, ■ GLM-4.7-Flash, ■ Qwen3-14B, and ■ Qwen3-32B. The marks p = 0.05. Advanced data flooding. An attacker may build a diverse training set by collecting solutions to the same tasks from different teacher agents to improve the distilled model’s performance Cui et al. (2023); Luo et al. (2026), which also dilutes the watermark signals. To evaluate the robustness 17

Preprint

of AuxMark against this advanced data flooding tactic, we mix the original watermarked dataset Dc with clean trajectories generated by alternative teacher agents. Specifically, we sequentially introduce clean data from DeepSeek-V4-Flash, MiMo-V2.5, and Qwen3-235B-A22B. In the R1, R2, and R3 settings, we incorporate data from 1, 2, and 3 alternative teachers, resulting in ratios of 1:1, 1:2, and 1:3. Then we use the combined dataset to train the student models and detect their watermarks. We evaluate a total of 24 settings across BFCL and SWE-bench, 3 dilution ratios, and 4 student models. As shown in Figure 6, AuxMark successfully detects every setting. Even under the most challenging R3 setting, all 8 tested student models remain successfully detected.

Q32

(a) BFCL: T10 16.4

5 0

16.1

7.4

6.4 p = 0.05

Mistral GLM Q14

−log10 p

1.9

Q32

(b) BFCL: T15

15 10

3.5 2.9 2.8 3.0 2.2 2.5 2.0 1.5 p = 0.05 1.0 0.5 0.0 Mistral GLM Q14

Q32

(d) SWE-bench: T10

16 13.7 14 12 10.9 9.9 10 8 6 4 2 p = 0.05 0 Mistral GLM Q14

8 7 6 5 3.9 4 2.7 2.4 3 2 p = 0.05 1 0 Mistral GLM Q14 20

4.7

Q32

(e) SWE-bench: T15

−log10 p

−log10 p

20

3.6

−log10 p

8 6.6 7 6 5 4 3 2.4 2.2 2 p = 0.05 1 0 Mistral GLM Q14

−log10 p

−log10 p

Truncation. Before distillation, an attacker may truncate the trailing tokens of the interaction trajectories to fit the student models’ context length limits or to deliberately disrupt potential watermark signals Kirchenbauer et al. (2023). To evaluate the robustness against this operation, we remove 10%, 15%, and 20% of the tokens from the tail of each watermarked training sequence, and train the student models using these truncated trajectories. We evaluate a total of 24 settings across the BFCL and SWE-bench benchmarks, 3 truncation levels (T10, T15, and T20), and 4 student models. As shown in Figure 7, AuxMark successfully detects the watermark in every setting. Even under the most aggressive T20 truncation setting, all 8 tested student models remain successfully detected.

15

6.6

Q32

(c) BFCL: T20 14.8 12.1

11.7 9.3

10 5 0

p = 0.05

Mistral GLM Q14

Q32

(f) SWE-bench: T20

Figure 7: Truncation robustness. The horizontal axis lists the four student models and the vertical axis reports − log10 p . In (a)–(f), bar colors denote ■ Mistral-24B, ■ GLM-4.7-Flash, ■ Qwen314B, and ■ Qwen3-32B. The marks p = 0.05.

Baseline robustness. Since Seqwm does not successfully detect distilled students(Table 3), Dataset Student Model Ms D1 D5 Mistral-24B 0/5 × 0/5 × we focus our robustness evaluation on AgenGLM-4.7-Flash 0/5 × 0/5 × twm under data flooding (the D1 and D5 setBFCL Qwen3-14B 2/5 × 0/5 × Qwen3-32B 1/5 × 0/5 × tings). To ensure a rigorous comparison, we Mistral-24B 1/5 × 1/5 × process the clean standard data to match the GLM-4.7-Flash 2/5 × 0/5 × initial distribution P before mixing it with the SWE-bench c Qwen3-14B 0/5 × 1/5 × 0/5 × 1/5 × watermarked traces. As shown in Table 9, all Qwen3-32B 16 evaluated settings fail to meet Agentwm’s ≥ 3/5 detection threshold. Agentwm encodes its watermark by shifting the tool-variant distribution towards Pwm . By flooding the training corpus with clean trajectories calibrated to Pc , the overall frequencies are pulled back towards the unwatermarked state. This directly dilutes the artificial distributional signal, resulting in the distilled students not sufficiently reproducing the watermarked patterns for reliable detection. Table 9: Agentwm robustness to data flooding.

18

Preprint

C

D ISCUSSION ON S TATISTICAL T EST

To evaluate how the choice of statistical unit affects detection results, we perform the same paired-probe tests at three granularities: card, trace, and probe levels. The card-level test agCondition Card-level Trace-level Probe-level gregates parameterized probes derived from the same card, the trace-level test further aggregates True positives 24/24 24/24 24/24 False positives 0/48 0/48 3/48 all cards from the same trajectory, and the probeParaphrasing attack 8/8 7/8 8/8 level test directly treats each paired probe as an 15% training share 24/24 22/24 24/24 individual statistical unit. As shown in Table 10, the card- and trace-level tests remain consistent in the main effectiveness evaluation: both achieve 24/24 detections in the true-positive settings and 0/48 significant results in the false-positive settings. Differences mainly appear under more challenging conditions. Trace-level aggregation compresses multiple watermark releases from the same trajectory into a single sign, allowing evidence from different cards to cancel and thereby reducing statistical power and robustness. Consequently, detection decreases from 8/8 to 7/8 under the paraphrasing attack, and from 24/24 to 22/24 at a 15% training share. In contrast, although probe-level testing shows stronger apparent statistical power, it repeatedly counts parameterized probes that share the same underlying release behavior and context, resulting in 3/48 cases with p < 0.05 in the false-positive settings. Although none of these cases exceeds the final fullhit-gap threshold and therefore no actual false-positive decision is triggered, the results indicate a higher false-positive risk when probes are counted individually. Based on these observations, we use the card as the primary statistical unit: it avoids repeated counting of the same watermark release event while preserving effective evidence from distinct release events. Table 10: Statistical-unit sensitivity. Entries count settings with an unadjusted one-sided sign-test p < 0.05; the 15% training-share row uses the median over 2,000 resamples.

19

Preprint

D

D ETAILED WATERMARKING P ROTOCOLS

AuxMark Watermark Embedding Protocol Inputs: Tool schema F, Secret key k, Trajectory ID idτ , Hyperparameters (p0 , ∆p, pmax , B, K). Outputs: Watermarked trajectory τ , Private evidence card set Cτ . 1. Initialization Phase: • System Setup: – Evaluate the tool schema F using the safety model Msafe to obtain the valid subset Fsafe . – Initialize the dynamic trigger probability p1 = p0 , the remaining budget b1 = B, and an empty evidence set Cτ = ∅. 2. Real-Time Embedding Phase: For each step t immediately following a completed core action at−1 core , the system executes: • Dynamic Scheduling: – Compute a cryptographic hash ut = Hash(k ∥ a ∥ idτ ∥ t) ∈ [0, 1). – If the budget is none (bt = 0) or the condition is not met (ut ≥ pt ), skip and failure update. • Candidate Generation & Scoring: – Generate K auxiliary candidates (tjaux , ajaux )K j=1 using the teacher model Mt and filter them through the validation mechanism. – For each valid candidate (tjaux , ajaux ), compute the score Sj based on naturalness (sjrel ), logical j consistency (sjrelb ), and repeatability sjrelb for the tool pairing (fcore , faux ). best best – Select the candidate with the highest score as (taux , aaux ). If no valid candidates exist, skip to the failure update. • Injection & Success Update: best best – Inject (tbest aux , aaux ) into the trajectory to receive the auxiliary observation oaux . – Reset the probability pt+1 = p0 and decrement the budget bt+1 = bt − 1. best – Increment the historical repeat count for this specific tool pairing (fcore , faux ). t−1 t−1 t−1 best best – Create an evidence card ct = (h, tcore , acore , ocore , faux , paux ) and append it to Cτ .

• Failure/Skip Update (Executed only if no injection occurred): – Increase the trigger probability pt+1 = min{pt + ∆p, pmax } to boost future injection chances, and retain the current budget bt+1 = bt .

Figure 8: The watermark embedding protocol of AuxMark.

20

Preprint

AuxMark Watermark Detection Protocol Inputs: Combined evidence pool CS , Suspect model Ms , Teacher model Mt , Hyperparameters (p, r, N ). Outputs: Detection decision ∈ {True, False}, and the statistical p−value. 1. Probe Construction Phase: For each evidence card ci ∈ CS , the system executes: • Paired Probes Generation: – Generate up to Ni parameterized variants using card ci to form the real probe set pireal . – Construct the corresponding fake probe set pifake by replacing the core action in pireal with a semantically altered fake core. 2. Evaluation Phase: For each probe variant j ∈ [1, Ni ] of card ci , the system executes: • Target Model Query & Parsing: – Submit the real probe pireal,j to the suspect model Ms and parse the output to compute the strict fullhit indicator hreal i,j ∈ {0, 1}. – Submit the fake probe pifake,j to the suspect model Ms and parse the output to compute the strict fullhit indicator hfake i,j ∈ {0, 1}. 3. Card-Level Aggregation Phase: For each card ci ∈ CS , the system executes: • Hit Rate & Performance Calculation: – Compute the hit rates across all Ni variants: Rireal = N1i

P

P real fake = N1i j hfake i,j . j hi,j and Ri

– Calculate the performance difference: ∆i = Rireal − Rifake . – Tally the results: increment the real win count w if ∆i > 0, and increment the fake win count ℓ if ∆i < 0 (ties where ∆i = 0 are excluded). 4. Statistical Test & Decision Phase: • Hypothesis Testing: P – Compute the one-sided exact sign-test p−value = w+ℓ k=w P – Compute the global mean difference ∆mean = |C1S | i ∆i .

 −(w+ℓ) 2 .

w+ℓ k

• Dual Criteria Check: – Output True (Detection Successful) if both p−value < p and ∆mean ≥ r hold. Otherwise, output False.

Figure 9: The watermark detection protocol of AuxMark.

21

Preprint

E

L IMITATIONS OF U NWATERMARKED D ISTILLATION D ETECTION

We additionally investigate approaches that infer distillation from similarity to a suspected teacher without an embedded watermark. We examine two representative similarity-based analyses of distillation: token-level n-gram overlap under reasoning prefilling (Panfilov et al., 2026) and agent execution graph similarity (Yang et al., 2026a). These approaches assess similarity at the levels of surface text and structured tool-use behavior, respectively. Our analyses show that shared solution patterns and response formatting can yield high textual overlap, while structurally different dependency graphs can receive near-perfect similarity scores. High scores therefore admit multiple explanations and do not uniquely identify a training source. E.1

V ULNERABILITIES IN T OKEN -L EVEL N- GRAM M ATCHING

Panfilov et al. (2026) study visible-answer overlap after prefilling a model’s reasoning channel with a short teacher reasoning prefix. Their Appendix B.2 evaluates 30 Humanity’s Last Exam problems using shared 1-, 2-, and 3-grams against the first 100 tokens of the reference answer, with best-ofk sampling up to k = 100. The visible answer is generated entirely by the target model. Their Figures 11–23 illustrate 13 public examples with problem statements, teacher reasoning excerpts, and visible-answer excerpts. Using selected public cases, we test whether alternative prefixes and shared solution patterns can produce high overlap with the same teacher reference. We use a fixed local overlap implementation. For each candidate a and reference b, we tokenize both texts with the same regex tokenizer and retain up to their first 100 tokens. Let Gn (x) denote the set of distinct n-grams in the retained tokens of x. We compute P3 |Gn (a) ∩ Gn (b)| sngram (a, b) = n=1 . (15) P3 n=1 |Gn (b)| Alternative prefixes and comparisons across model versions. On the geometry problem in Figure 11 of Panfilov et al. (2026), Kimi-K2.6, which was publicly released before Claude Opus 4.8 (Team et al., 2025; Anthropic, 2026), obtains mean/best-of-4 scores of 0.274/0.565 with the teacher opening This is a known, compared with 0.373/0.662 under This is an unfamiliar and 0.250/0.415 without a prefix. The best local score of 0.662 approaches the reported 0.80 overlap between prefilled Kimi-K3 and Opus 4.8. Within this fixed scoring setup, the alternative prefix produces a larger increase than the authentic prefix across both metrics. The selected 0.662 answer shares the opening derivation order C, F, E, D, B with the Opus 4.8 reference, as shown in Figure 10. The sampled Opus 4.6 answers use tables or different orders; the highest score among these answers is 0.343 against the same reference. Thus, this local formatting match is not exclusive to the teacher prefix or to the sampled Claude-family answers. Convergence on standard solutions. For the linear differential equation in Figure 13 of Panfilov et al. (2026), a standard solution uses a cosh integrating factor. Without a teacher prefix, GPT-OSS120B reaches a mean overlap of 0.500 and a best-of-4 score of 0.537. Kimi-K2.6 similarly scores 0.501 without a prefix and 0.502 with the teacher prefix. The unprefilled local scores are numerically close to the reported 0.50 overlap between prefilled Kimi-K3 and Opus 4.8. The unprefilled answer excerpts are shown in Figure 11. Our results show that substantial local overlap can arise in standard solutions with conventional terminology and ordering, while the teacher prefix adds little in the Kimi-K2.6 comparison. Response steering and sampling. Visible-answer overlap must also be interpreted in light of how responses are generated, presented, and selected. In our geometry evaluation of Kimi-K3, all 20 completions across five prefix conditions re-derive the solution, and none states the answer within its first 50 reasoning tokens, unlike the early answer in the reference trace. Higher visible overlap can therefore coexist with a different observed reasoning pattern. At the level of answer presentation, even correct solutions can receive very low scores when little matching text is available: for the problem in Figure 15 of Panfilov et al. (2026), four correct answers constrained to at most 15 words each score 0.005. Such sensitivity to presentation also matters when interpreting best-of-k results, which report the largest overlap in a sampling pool. Formatting cues can therefore raise the reported score through occasional high-overlap completions without representing typical output behavior. 22

Preprint

Opus 4.8

Kimi-K2.6 (prefilled)

Opus 4.6

## Setting Up Coordinates Let me place the points on a number line with $A = 0$. **Finding positions from the given distances:** - $AC = 26 \Rightarrow C = 26$ - $AF = 73 \Rightarrow F = 73$ - $CE = 31 \Rightarrow E = 57$ - $DF = 33 \Rightarrow D = 73 - 33 = 40$ - $BD = 22 \Rightarrow B = 40 - 22 = 18$

Place the six collinear points on the $x$-axis. Set $A$ at the origin, so $a=0$.

## Setting Up Coordinates

From the given lengths: - $AC = 26 \Rightarrow c = 26$ - $AF = 73 \Rightarrow f = 73$ - $CE = 31 \Rightarrow e = c+31 = 57$ - $DF = 33 \Rightarrow d = f-33 = 40$ - $BD = 22 \Rightarrow b = d-22 = 18$ Thus the coordinates are $$A(0),\; B(18),\; C(26),\; D(40),\; E(57),\; F(73).$$

I place the points on a number line with **A = 0**. Using the given distances: | Equation | Result | |---|---| | AF = 73 | **F = 73** | | AC = 26 | **C = 26** | | DF = 33 → d = 73 − 33 | **D = 40** | | BD = 22 → b = 40 − 22 | **B = 18** | | CE = 31 → e = 26 + 31 | **E = 57** | **Verification of order:** A(0) < B(18) < C(26) < D(40) < E(57) < F(73) ✓

Figure 10: Excerpts from the original answers to the geometry problem in Figure 11 of Panfilov et al. (2026). Green shading marks the shared C, F, E, D, B derivation order. For each candidate model, we show the highest-overlap answer from four samples. Opus 4.8

Kimi-K2.6 (unprefilled)

GPT-OSS-120B (unprefilled)

## Setting Up the Solution The equation is linear: $$x_3' + \tanh(t) x_3 = e^{t/3}.$$ **Integrating factor:** $$\mu(t)=\exp\left(\int\tanh t\,dt\right)=\exp(\ln\cosh t)=\cosh t.$$

The given first-order linear ODE is $$x_3'(t)+\tanh(t)\,x_3(t)=e ^{-t/3}, \qquad x_3(0)=5. $$

The linear equation

**1. Solve the ODE.** The integrating factor is $$\mu(t)=\exp\!\left(\int \tanh t\,dt\right)=\exp\! \bigl(\ln\cosh t\bigr)= \cosh t.$$

\[ x_3'(t)+\tanh t\,x_3(t)=e^{t/3},\qquad x_3(0)=5 \] has integrating factor \[ \mu(t)=\exp\!\Bigl(\int \tanh t\,dt\Bigr)=\exp(\ln \cosh t)=\cosh t . \]

Figure 11: Excerpts from the original answers to the differential-equation problem in Figure 13 of Panfilov et al. (2026). Green shading marks the shared cosh t integrating factor. Each candidate has the highest reported overlap among four unprefilled samples. E.2

V ULNERABILITIES IN AGENT E XECUTION G RAPH S IMILARITY

AgentEcho’s Action Graph Similarity (AGS) averages optional-tool agreement Snode , sequentialpattern similarity Sseq , and dependency-pattern similarity Sdep (Yang et al., 2026a). Moving from text to tool behavior introduces structural evidence, but its interpretation depends on which graph properties the metrics preserve. Scale invariance and structural ambiguity. For a dependency graph G, Sdep compares ϕ(G) = (rreuse , dmax , rfanout ) by cosine similarity (Yang et al., 2026a, Appendix B.4). These features describe the fraction of calls after the first that receive dependency inputs, the longest dependency path, and the fraction of source nodes with multiple outgoing dependency edges. Cosine similarity ignores feature-vector magnitude. Consider two graphs with 11 nodes: one contains a single dependency edge, whereas the other is a complete 10-edge chain. With path length measured in edges, ϕ(Gsingle ) = (0.1, 1, 0),

ϕ(Gchain ) = (1, 10, 0) = 10ϕ(Gsingle ). 23

(16)

Preprint

Their Sdep is exactly 1 despite large differences in reuse and dependency depth. Hence, a perfect score does not guarantee closely matching dependency structures. Evidence from tool trajectories. We audit 75 trajectories from five models on 15 tasks and compare all ten model pairs within each task, yielding 150 comparisons. The mean Sdep is 0.885, and 121 comparisons score at least 0.9. Among these, 53 have dependency edge counts differing by at least three, and 49 pair a successful trajectory with a failed one. In retail-37, a failed trajectory with one dependency edge and a successful trajectory with 12 edges score 0.991. We further verify candidate dependency edges with an LLM judge, using anonymized values and a prompt adapted from the published protocol (Yang et al., 2026a, Appendix B.1). Both the score and the edge-count discrepancy remain unchanged for this example. Thus, near-perfect dependency similarity persists despite substantial structural differences, even after LLM-based edge verification. Behavioral convergence after reasoning distillation. Inspired by the reference-based formulation of Rawat et al. (2026), we measure changes in behavioral similarity relative to the base checkpoint. We use the Qwen3.5-9B base model as the reference checkpoint and evaluate its publicly released reasoning-distilled variant.2 The model card identifies Claude Opus 4.6 as the teacher. Our audit of 12,592 publicly available examples from the model card’s listed datasets finds no structured tool trajectories. We use the same agent harness and user simulator for 15 τ -bench-style tasks, with one rollout per model and task. Claude Sonnet 4.5 (thinking), Kimi-K2 (thinking), and GPT-OSS-120B serve as comparison models. Comparison model Claude Sonnet 4.5 (thinking) Kimi-K2 (thinking) GPT-OSS-120B

AGSbase 0.549 0.561 0.584

AGSdist 0.709 0.850 0.770

∆AGS +0.160 +0.289 +0.186

∆Snode −0.031 +0.042 −0.047

∆Sseq +0.425 +0.750 +0.525

∆Sdep +0.086 +0.077 +0.080

Table 11: Changes in AGS and its components after reasoning distillation on the same 15 tasks. Component values are rounded independently. Table 11 shows that AGS increases toward all three models, alongside a rise in task success from 4/15 to 12/15. The gains are dominated by Sseq , which increases by 0.425–0.750. In contrast, Sdep increases by 0.077–0.086, while Snode decreases toward both Claude and GPT-OSS. Thus, these gains primarily reflect convergence in local execution statistics and do not uniquely identify teacher-specific tool-use inheritance. Implications for attribution. These cases expose ambiguity in interpreting high unwatermarked similarity as evidence of a specific training source. They motivate controls for shared solutions, response presentation, and general behavioral improvement, alongside calibrated false-positive rates. Our watermarking framework provides private verification signals tied to recorded teacher interactions.

2

Model card: Jackrong’s reasoning-distilled Qwen3.5-9B (GGUF).

24

Record · ID 1108635 · SHA-256 ec7e7db480bcd4bc
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.