InkShield: Writing Style Protection Against Unauthorized Handwriting Mimicry Jian Xiong1 , Wenbo Jiang1∗ , Zihan Wang1 , Rui Zhang1 , Wenshu Fan1 , Hongwei Li1 , Guowen Xu1 1
University of Electronic Science and Technology of China [email protected], [email protected]
arXiv:2607.26976v1 [cs.CR] 29 Jul 2026
Abstract Recent handwritten text generators can reproduce a writer’s style from publicly available references, posing risks of document forgery and identity misuse. An attacker may use a publicly available handwritten note or signature sample to generate forged recommendation letters or authorization forms, leading to document fraud, identity misuse, and misleading decisions. However, existing protections against unauthorized image editing or synthesis transfer poorly to handwriting style mimicry. Designed for natural images with complex backgrounds, they often optimize perturbations over the whole image. For sparse handwriting images, such global perturbations become conspicuous in blank background regions and largely degrade the visual quality. In this work, we propose InkShield, a proactive writing-style defense that protects reference images before release. InkShield selects a decoy writer to define a style-displacement direction, optimizes perturbations with a frozen handwriting-generation surrogate, and confines them to ink-stroke edges to avoid conspicuous background artifacts. On IAM, the average Top1/Top-5 rates at which generated samples are retrieved as the target writer by two independent writer evaluators decrease from 11.94%/36.52% to 2.03%/8.79%. Meanwhile, the protected references remain visually close to the originals (LPIPS 0.0078), and the generated text remains readable. InkShield also exhibits transferability to other handwriting generators. Overall, InkShield provides practical protection against unauthorized handwriting style mimicry.
Introduction Handwritten text generation models can synthesize new words or phrases with controllable textual content and writer style (Kang et al. 2020; Bhunia et al. 2021; Pippi, Cascianelli, and Cucchiara 2023; Nikolaidou et al. 2023). This capability is useful for data augmentation, font-like personalization, and historical document restoration. Moreover, recent SOTA generators can achieve one-shot mimicry, reproducing a target writer’s style from only a single reference sample (Dai et al. 2024; Le et al. 2026), thereby significantly lowered the barrier. However, this strong capability also introduces a new security risk: a single handwriting sample can become ∗
Corresponding author. Copyright © 2027, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.
a reusable style key. A scanned form, a homework sheet, or a handwritten note, once used as a reference image, may facilitate document forgery, identity misuse, and other serious harms. Figure 1 illustrates the threat posed by one-shot handwriting style reuse. With only a single handwriting reference image, the attacker can synthesize new content in a style similar to that of the target writer. However, for practical purposes such as submitting coursework, sharing notes, or showcasing calligraphic works, reference owners may inevitably need to upload handwritten materials online. This raises a referenceside protection problem: Can we protect handwriting images before release while preserving their readability and weakening downstream target-writer mimicry? Existing image-protection methods primarily target natural-image scenarios, including unauthorized image editing, face or artistic-style imitation (Salman et al. 2023; Choi et al. 2025; Van Le et al. 2023; Liu et al. 2024; Li et al. 2025). When directly applied to handwriting references, these methods face two limitations. First, their objectives are not optimized for handwriting style protection. Because writer-specific style is expressed through fine-grained local stroke morphology, generic image-level perturbation objectives may provide limited suppression of target-writer mimicry. Second, they often optimize perturbations globally over the entire image, introducing noticeable artifacts in the largely blank background regions of handwritten word images and degrading visual quality. To address these two challenges, we propose InkShield, a reference-side writing-style protection method against unauthorized handwriting mimicry. Before release, InkShield proactively applies low-visibility perturbations to protect handwriting references, impairing downstream mimicry while preserving readability and visual utility. Specifically, InkShield consists of three stages: decoy writer selection, handwriting-aware perturbation constraints, and SurrogateGuided Generative Disruption. A decoy pool provides a styledisplacement direction that pushes the protected reference away from the target-writer style, addressing the limited ability of generic image-level objectives to disrupt writer-specific style. A handwriting-aware stroke-edge mask confines perturbations to style-sensitive regions, avoiding noticeable artifacts in blank background areas caused by global perturbations. Finally, a frozen surrogate generation objective uses
Related Work Prof. A’s handwritten note
One-shot reference 𝒙
Diffusion-based (denoising) generator
Content text 𝒄
Acquire Handwriting Samples
Prepare Forgery
Generate Fake Letter
Application Materials Recommendation Letter Upload Please upload a PDF copy of your professor’s recommendation letter.
v
v
Uploaded: Prof. A.PDF
Mislead Admission
Submit Forgery
Figure 1: Security risk of one-shot handwriting style mimicry: A public handwritten word enables an attacker to generate and submit a forged recommendation letter in the professor’s style. multiple pseudo-content words during optimization, allowing the learned perturbation to generalize across different target texts rather than overfitting to a particular word. In our experiments, we evaluate InkShield on the IAM dataset from four complementary perspectives: target-writer mimicry suppression, reference stealth, content preservation, and generation quality. Under the primary One-DM (one-shot) setting, InkShield reduces the average targetwriter Top-5 retrieval rate across two independent writer evaluators from 36.52% to 8.79%, while achieving strong reference stealth with an LPIPS of 0.0078. Results on CONSTANT(one-shot) and DiffusionPen(multi-shot) further demonstrate cross-model transferability and effectiveness with multiple handwriting references. Our contributions are summarized as follows: • We formulate unauthorized generative handwriting style mimicry as a reference-side protection problem and propose a proactive defense method for protecting released handwriting references. To the best of our knowledge, this is the first work to address this problem. • We propose InkShield, which integrates decoy-guided style displacement, handwriting-aware stroke-edge constraints, and frozen surrogate-guided generative disruption. It suppresses target-writer style reuse while preserving the readability and visual utility of released handwriting references. • We conduct extensive experiments on One-DM, CONSTANT, and DiffusionPen. Results show effective mimicry suppression with preserved reference readability, and transferability across generators.
Handwritten Text Generation Handwritten text generation synthesizes text images with controllable content and writer style. A central distinction among existing methods is whether they operate in multishot or one-shot settings. Multi-shot methods require collecting multiple handwriting samples from a writer to obtain a more complete style representation. GANwriting performs content- and style-conditioned handwritten word generation in a multi-shot setting (Kang et al. 2020). HWT uses Transformer attention to model global and local patterns from multiple style examples (Bhunia et al. 2021). VATr employs visual archetypes for multi-shot generation with rare characters and unseen writers (Pippi, Cascianelli, and Cucchiara 2023), while VATr++ improves input preparation, regularization, and evaluation for this setting (Vanherle et al. 2024). DiffusionPen further formulates handwritten text generation as a five-shot latent-diffusion problem (Nikolaidou et al. 2024). In contrast, one-shot methods aim to reproduce a writer’s style from only one handwriting image. HiGAN performs one-shot handwriting imitation through content–style disentanglement (Gan and Wang 2021). One-DM enhances style extraction from a single reference using high-frequency information (Dai et al. 2024), and CONSTANT improves one-shot generation through patch contrastive enhancement and style-aware quantization (Le et al. 2026). DiffBrush uses one-shot style learning, extending generation from isolated words to complete handwritten text lines by modeling intraand inter-word style patterns (Dai et al. 2025). InkShield protects released handwriting references against unauthorized target-writer mimicry and remains effective in both one-shot and multi-shot settings, while primarily focusing on the stringent setting in which a single exposed reference is sufficient for mimicry.
Protection Against Unauthorized Image Usage Existing work has explored protections against unauthorized uses of images, including style mimicry, subject personalization, and image editing. Glaze protects artworks from text-toimage style mimicry (Shan et al. 2023). Anti-DreamBooth protects personal images from unauthorized diffusion personalization (Van Le et al. 2023), while MetaCloak improves the transferability of such protections through metalearning (Liu et al. 2024). PAP considers both face privacy and artistic style in personalized diffusion models (Wan et al. 2024), and StyleGuard targets text-to-image style mimicry using style perturbations (Li et al. 2025). PhotoGuard, EditShield, and DiffusionGuard further protect images against diffusion-based editing (Salman et al. 2023; Chen et al. 2024; Choi et al. 2025). Mist improves adversarial perturbation optimization for diffusion models (Liang and Wu 2023). Despite their different downstream targets, these methods are primarily designed for natural images, artistic styles, or faces. When transferred directly to handwriting references, their generic image-level objectives may not sufficiently disrupt the fine-grained stroke morphology that encodes writerspecific style, while global perturbations can introduce noticeable artifacts in the largely blank backgrounds of hand-
Handwriting Source
InkShield A. Decoy Writer Selection
Target writer 𝒕
Text Generation
C. Surrogate Generation Optimization
Decoy pool 𝑫𝒕
Cached clean Pseudo-generation
Target writer 𝒕 Candidate writer Decoy writer 𝒅𝒕 (p95)
⋯
Publicly released handwritten sample
closer
Noised latent 𝒛𝒕
Style distance from 𝒕
B. Handwriting-Aware Stealth Constraint Edge mask 𝑴(𝒙)
farther
𝑲 cached clean pairs
𝑲 cached clean pairs (𝒚~𝒌, 𝒄𝒌 )
Pertubation 𝜹
P𝐫𝐨𝐭𝐞𝐜𝐭𝐞𝐝 𝐫𝐞𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝒙𝒑𝒓𝒐
Clean reference 𝒙 𝜹 ∞≤𝜺
LPIPS(𝒙, 𝒙𝒂𝒅𝒗 ) ≤ 𝝉
Clean reference 𝒙
𝒕 , 𝑳𝒂𝒑 𝒙𝒕 𝜺𝜽 (𝒛𝒕 , 𝒕|𝒙𝒑𝒓𝒐 𝒑𝒓𝒐 , 𝒄𝒌 )
𝒕=𝑻
P𝐫𝐨𝐭𝐞𝐜𝐭𝐞𝐝 𝐫𝐞𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝒙𝒑𝒓𝒐
Predicted noise 𝜺ො 𝒌
P𝐫𝐨𝐭𝐞𝐜𝐭𝐞𝐝 𝐫𝐞𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝒙𝒑𝒓𝒐
Noise predictor
𝐕𝐀𝐄 encoder
Extract word Clean reference 𝒙
Content text 𝒄
I agree to the terms
Frozen surrogate
Handwritten Text Generator 𝑮
𝒕 < 𝑻, 𝐭 = 𝐭 + 𝟏
Clean generation
𝒙𝒕𝒑𝒓𝒐 masked PGD update
Protected generation
maximize 𝑱 = 𝑱𝐬𝐭𝐲𝐥𝐞 + 𝝀𝐠𝐞𝐧 𝑱𝐠𝐞𝐧 − 𝝀𝐩𝐞𝐫𝐜 𝑷𝐩𝐞𝐫𝐜
Figure 2: Overview of InkShield. Given a clean reference from a target writer, InkShield selects a distant decoy writer to define a style-displacement direction, confines low-visibility perturbations to stroke-edge regions, and optimizes them through a frozen surrogate generation chain over multiple pseudo-content words. The protected reference reduces downstream target-writer mimicry while preserving content readability and visual quality. written word images. To address these limitations, InkShield uses decoyguided style displacement to interfere with writer-specific style extraction and confines perturbations to stroke-edge regions to preserve visual quality.
Threat Model Scenario. We prioritize a scenario in which a handwriting reference image is used without authorization for style mimicry in the one-shot setting. A target writer t owns a clean handwriting reference image x and needs to publish or share it online. After obtaining the released image, an attacker provides it together with arbitrary content text c to a handwriting generator G and synthesizes new content G(c, x) in the style of the target writer. Attacker’s Capability. The attacker can obtain a released handwriting reference image and use a handwriting generator for style mimicry. The attacker may choose arbitrary content text and, in the one-shot setting, requires only a single publicly available reference image (e.g., a word image) to carry out the attack. Defender’s Capability. Before releasing the reference image, the defender can access the clean image x and apply low-visibility perturbations to obtain a protected image xpro for publication. To construct this perturbation, the defender may use a surrogate handwriting generator and writer-style representations during optimization. However, the defender does not know the downstream generator, its parameters, or its processing pipeline used by the attacker, and cannot control the attacker’s content text, random seed, or additional reference samples. Defender’s Goals. The primary goal is target-writer mimicry suppression: when the attacker uses xpro for generation, the resulting handwriting should no longer closely resemble the target writer’s style. At the same time, the protected reference should preserve its original textual content and readability under low-visibility perturbations. A successful defense must not rely on destroying the released reference.
Algorithm 1: Algorithm for InkShield Require: Clean reference x of target writer t; authorized candidate writers W and their reference samples; frozen surrogate Gθ ; pseudo texts C; ρ, ϵ, α, and T Ensure: Protected reference xpro 1: Compute writer prototypes pt and {pu }u∈W from available reference samples 2: Compute ∆(t, u) = 1 − cos(pt , pu ) for all u ∈ W 3: Sort candidates by increasing ∆(t, u) to obtain the ordered decoy pool Dt 4: Select dt ← Dt [⌈ρ|Dt |⌉] and a fixed reference xdt of dt 5: Construct the handwriting-aware edge mask M (x) using Eq. 4 6: Cache clean pseudo generations {e yc ← Gθ (c, x)}c∈C 7: Initialize δ (0) ∼ U(−ϵ, ϵ) ⊙ M (x) 8: for k ← 0 to T − 1do 9: xpro ← clip[0,1] x + M (x) ⊙ δ (k) 10: Compute Jstyle using Eq. 7 11: Compute Jgen using Eq. 8 12: Compute Pperc using Eq. 6 13: Compute the total objective J using Eq. 9 14: Update δ (k+1) using Eq. 10 15: end for 16: return xpro ← clip[0,1] x + M (x) ⊙ δ (T )
Methodology We propose InkShield to protect handwriting references before release. As shown in Figure 2, InkShield selects a stylistically distant decoy writer from the target writer’s decoy pool to define a style-displacement direction and optimizes low-visibility perturbations confined to stroke-edge regions through a frozen surrogate. In our implementation, this surrogate is instantiated with One-DM (Dai et al. 2024). Algorithm 1 summarizes the complete protection workflow.
Decoy Writer Selection InkShield selects a decoy writer to guide the displacement of the protected reference in writer-style space. The selected decoy serves only as a defense-side displacement direction: the goal is to reduce similarity to the target writer rather than make downstream generations imitate the decoy writer. For
each target writer t, the defender constructs a decoy pool Dt comprising authorized candidate writers other than t and their reference samples. To measure the style distance between the target writer and the candidate writers in Dt , we draw on the style representation mechanism of One-DM. One-DM decomposes handwriting style into complementary low-level and highlevel components, corresponding to fine-grained stroke appearance and more global writer-level characteristics, respectively. Following this observation, we extract low-level and high-level style representations from handwriting reference images, denoted by ϕlow (x) and ϕhigh (x). For the target writer and each candidate writer w, we average these features over the writer’s available handwriting samples to obtain high component-wise prototypes plow w and pw , and concatenate them as the writer-level prototype pw . We then compute the style distance between the target writer t and each candidate writer u: ∆(t, u) = 1 − cos(pt , pu ). (1) We sort the candidate writers in Dt by increasing ∆(t, u); after sorting, Dt [j] denotes the j-th closest candidate writer to t, where j ∈ {1, . . . , |Dt |} and |Dt | denotes the number of candidate writers in the pool. Decoy selection balances perturbation cost against mimicry suppression. A nearby decoy may require a smaller perturbation and thus better preserve the visual quality of the protected reference, but its style remains close to that of the target writer and may provide limited suppression. In contrast, a distant decoy provides a larger style-displacement direction and can more effectively reduce the similarity between generated handwriting and the target-writer style. However, an excessively distant decoy may introduce unstable guidance or require stronger perturbations. Our decoypool sensitivity study shows that distant decoys generally yield stronger target-writer mimicry suppression. We therefore select a high-percentile decoy writer dt from the ordered pool: dt = Dt [⌈ρ|Dt |⌉] , (2) where ρ is the decoy percentile. We use ρ = 0.95 as the default setting, which achieves strong mimicry suppression without relying on the potentially outlying farthest writer.
Handwriting-Aware Stealth Constraint Handwritten word images are sparse: foreground strokes occupy only a small portion of the image, whereas most pixels belong to the blank background. Meanwhile, writerspecific style is encoded in local stroke morphology, including stroke width, curvature, slant, junctions, and boundary shapes. InkShield therefore constrains perturbations to a handwriting-aware mask that focuses on stroke-boundary regions. Given a grayscale reference image x ∈ [0, 1], we first obtain an ink foreground mask F (x) = 1[x < ηfg ], where ηfg = 0.92. Let D1 (·) and E1 (·) denote morphological dilation and erosion with a one-pixel radius, respectively. We construct a morphological stroke boundary as B(x) = [D1 (F (x)) − E1 (F (x))]+ ,
(3)
where [·]+ clips negative values to zero. In addition, we compute the normalized gradient magnitude Ssob (x) using a 3 × 3 Sobel operator (Sobel and Feldman 1968) and retain strong responses as H(x) = 1[Ssob (x) > ηsob ], where ηsob = 0.12. We retain only the strong Sobel responses within a one-pixel neighborhood of the ink foreground and take their binary union with B(x) to form a preliminary edge support E(x). The final mask is then defined as M (x) = F (x) ⊙ D1 E(x) . (4) This construction expands the combined edge support by one pixel and finally restricts it to the ink foreground, concentrating perturbations near stroke boundaries while avoiding the blank background. Given the clean reference image x, the protected reference is defined as xpro = clip[0,1] (x + M (x) ⊙ δ) ,
∥δ∥∞ ≤ ϵ,
(5)
where δ is the perturbation, ⊙ denotes element-wise multiplication, and ϵ controls the perturbation budget. In addition to the spatial mask and the ℓ∞ constraint, InkShield uses a soft perceptual penalty to encourage stealthiness. We measure perceptual distortion using LPIPS (Zhang et al. 2018): Pperc = max (0, LPIPS(xpro , x) − τ ) ,
(6)
where τ is a penalty threshold rather than a strict upper bound. This term discourages perceptually noticeable changes while allowing the perturbation to suppress targetwriter mimicry. During optimization, we mask the gradient updates and project δ onto the ℓ∞ ball after each step. The multiplication by M (x) in Eq. 5 further guarantees that the final perturbation is confined to the handwriting-aware region.
Surrogate-Guided Generative Disruption Having defined the decoy-guided displacement direction and handwriting-aware perturbation constraints, InkShield optimizes the perturbation δ through a surrogate handwriting generator to suppress target-writer mimicry in downstream generations. We first define a style displacement objective using the selected decoy writer. During optimization, InkShield uses the clean reference x of the target writer t and a fixed representative reference xdt from the selected decoy writer dt . Let r ∈ {low, high} denote the low-level and high-level style components. We define h X Jstyle = λrdec cos(ϕr (xpro ), ϕr (xdt )) r∈{low,high}
− λrtar cos(ϕr (xpro ), ϕr (x))
i (7) .
where λrdec and λrtar control decoy attraction and target-writer repulsion at style level r, respectively. Thus, the objective moves the protected reference toward a decoy-guided direction while explicitly moving it away from the target-writer style.
Inspired by Fully-trained Surrogate Model Guidance (FSMG) in Anti-DreamBooth (Van Le et al. 2023), we use a frozen One-DM (Dai et al. 2024) model as a text-conditioned surrogate generator. Unlike FSMG, which targets DreamBooth personalization and requires a subject-specific surrogate model for each protected subject, InkShield directly optimizes against a frozen one-shot handwriting generator without additional per-writer training. Given a set of K pseudo content texts C = {c1 , . . . , cK }, we pre-generate and cache clean pseudo generations yec = Gθ (c, x) for each c ∈ C, where Gθ denotes the frozen surrogate generator. At each optimization step, we encode each cached pseudo generation yec into the latent space, sample a diffusion timestep and noise, and evaluate the frozen OneDM denoising process under content condition c. In this process, the clean reference x in the reference-dependent style and stroke conditions is replaced with the current protected reference xpro . Let Lden (e yc , c, xpro ) denote the resulting denoising prediction error between the predicted and sampled noise. We define the surrogate generation objective as 1 X Jgen = Lden (e yc , c, xpro ) . (8) K
threat in a controlled setting, we conduct experiments under the OOV-U protocol, in which both the test writers and content words are unseen during training. For each held-out target writer, we evaluate generation on 80 OOV-U content words, yielding 161 × 80 = 12,880 generated word images per method. For the unified evaluation protocol, word images are resized to a height of 64 pixels while preserving their aspect ratios. Evaluation Metrics. We evaluate InkShield from four complementary perspectives: target-writer mimicry suppression, reference stealth, content preservation, and generation quality. These perspectives must be considered jointly: a reduction in target-writer similarity alone does not constitute successful protection if it is achieved by visibly corrupting the released reference or severely degrading the generated handwriting.
Experimental Setup
• Target-writer mimicry suppression: We assess targetwriter mimicry using two independently trained writer evaluators, ResNet50 (He et al. 2016) and DeepWriter (Xing and Qiao 2016). Using penultimate-layer embeddings, we retrieve gallery writers for each generated image and report target-writer Top-1 and Top-5 retrieval rates; lower values indicate stronger mimicry suppression. Training details are provided in Appendix. • Reference stealth: We compare each protected reference with its clean counterpart using LPIPS (Zhang et al. 2018), PSNR (Huynh-Thu and Ghanbari 2008), and SSIM (Wang et al. 2004). Lower LPIPS indicates smaller perceptual distortion, whereas higher PSNR and SSIM indicate better pixel-level and structural preservation, respectively. We further report background average absolute perturbation (BgAbs) and background energy ratio (BgE). Lower BgAbs and BgE indicate that perturbations are concentrated on handwriting regions rather than blank backgrounds. • Content preservation: We use TrOCR (Li et al. 2023) to transcribe generated word images and compute the character error rate (CER) with respect to the requested content text. Lower CER indicates better textual fidelity and provides an automated proxy for generated-text legibility. • Generation quality: We report writer-level HWD (Pippi et al. 2023) and FID (Heusel et al. 2017) by comparing generated images with real IAM test images from the corresponding target writer in an unpaired manner. Lower HWD and FID indicate closer distributional agreement with real handwriting. Since the protection is intended to weaken target-writer style similarity, these measures serve to detect severe generation degradation rather than as direct protection objectives.
Dataset. We evaluate InkShield on the IAM handwriting dataset (Marti and Bunke 2002), which contains 62,857 English word images from 500 writers. Following prior work (Bhunia et al. 2021; Kang et al. 2020; Pippi, Cascianelli, and Cucchiara 2023), we split the writers into 339 training writers and 161 test writers. The training split is used to train the writer evaluators introduced below, while all protection and generation experiments are conducted on the held-out test writers. To instantiate the reference-abuse
Generation Models. We use One-DM (Dai et al. 2024) as the surrogate model and primary evaluation generator. One-DM and CONSTANT (Le et al. 2026) each condition generation on one fixed reference image per target writer. DiffusionPen (Nikolaidou et al. 2024) follows its multi-shot setting and uses five references per target writer; we independently protect all five references with InkShield. This protocol evaluates whether protected references transfer across generators with different reference-conditioning schemes.
c∈C
Maximizing Jgen makes the protected reference less compatible with the clean-reference generation behavior of the target writer. Optimizing over multiple pseudo content texts further prevents the perturbation from overfitting to a particular word. All surrogate-model parameters remain frozen, and gradients are propagated only to δ. The final objective combines style displacement, surrogate generation disruption, and perceptual stealthiness: max
δ:∥δ∥∞ ≤ϵ
J = Jstyle + λgen Jgen − λperc Pperc .
(9)
Starting from a masked initialization δ (0) , we optimize δ using PGD-style projected gradient ascent (Madry et al. 2018). At iteration k, the perturbation is updated as (10) δ k+1 = Πϵ M (x) ⊙ δ k + α sign (∇δk J) . where α is the step size and Πϵ denotes projection onto the ℓ∞ ball of radius ϵ. The mask constrains every perturbation update to handwriting-aware regions, while the projection enforces the perturbation budget. After each update, we recompute xpro using Eq. 5, which also clips it to the valid image range.
Table 1: Main results on One-DM. Top1/Top5 denote target-writer retrieval rates; BgAbs and BgE denote background average absolute perturbation and background energy ratio. ResNet50
DeepWriter
Reference Stealth
Content Generation Quality
Method
Top1 ↓ Top5 ↓ Top1 ↓ Top5 ↓ LPIPS ↓ PSNR ↑ SSIM ↑ BgAbs ↓ BgE ↓
CER ↓
HWD ↓
FID ↓
Clean DiffusionGuard PhotoGuard StyleGuard Anti-DreamBooth
13.20 4.08 3.84 2.60 4.13
38.73 15.82 15.66 10.20 15.35
10.68 4.03 3.92 2.42 3.83
34.31 15.47 15.32 10.06 14.85
– 0.0096 0.0087 0.1249 0.0138
– 37.98 38.09 20.75 36.64
– 0.9746 0.9761 0.9111 0.9532
– 0.0046 0.0043 0.0170 0.0122
– 63.05 59.60 55.44 69.73
15.51 18.36 18.20 17.49 18.61
1.862 1.979 1.987 2.053 1.989
123.43 131.48 131.46 133.78 130.95
InkShield
2.12
9.21
1.93
8.37
0.0078
33.44
0.9856
0.0000
0.00
17.40
2.157
137.17
Figure 3: Qualitative comparison of reference protection methods on One-DM. All methods use the same clean reference from one target writer and generate the requested word planets. The first column shows the shared clean reference, InkShield edge mask, and clean generation; the remaining columns show the corresponding protected references, perturbations, and mimicry results. Comparison Methods. For the primary One-DM evaluation, we compare InkShield with four representative imageprotection methods: DiffusionGuard (Choi et al. 2025), PhotoGuard (Salman et al. 2023), StyleGuard (Li et al. 2025), and Anti-DreamBooth (Van Le et al. 2023). These methods respectively cover protection against diffusion-based image editing, style mimicry, and personalized image synthesis. Each baseline is applied to the same clean handwriting references and evaluated using the same One-DM generation protocol and evaluation metrics as InkShield. Implementation Details. InkShield performs PGD-style projected gradient ascent in normalized grayscale image space with a masked random start. We use the P95 decoy selection (ρ = 0.95), the handwriting edge mask, a perturbation budget of ϵ = 16/255, a step size of α = 2/255, and 80 optimization steps. The surrogate-generation objective uses K = 4 cached pseudo content texts with λgen = 3.0. We use an LPIPS soft penalty with threshold τ = 0.005. The surrogate is a frozen One-DM model with a Stable Diffusion v1.5 backbone, whose parameters remain fixed during optimization. The component-wise style weights and the remaining loss weights are provided in the Appendix.
Main Results Table 1 reports the main results on one-shot handwriting generation with One-DM. InkShield achieves the strongest
target-writer mimicry suppression under both independent writer evaluators. On ResNet50, it reduces target-writer Top1/Top-5 retrieval from 13.20%/38.73% for clean generation to 2.12%/9.21%. On DeepWriter, the corresponding rates decrease from 10.68%/34.31% to 1.93%/8.37%. InkShield obtains the lowest retrieval rates across all four evaluator– retrieval combinations, showing that protected references substantially reduce the attribution of downstream generations to the target writer. InkShield also achieves strong reference stealth. Its BgAbs = 0.0000 and BgE = 0.00 indicate that perturbations are confined to the handwriting-aware mask without affecting blank background regions. It further achieves the lowest LPIPS (0.0078) and the highest SSIM (0.9856) among all compared methods. These results show that InkShield suppresses mimicry through localized, low-visibility perturbations rather than broad reference corruption. InkShield maintains generated-text readability, achieving a CER of 17.40%, close to the clean result of 15.51% and lower than several competing defenses. Its HWD and FID increase relative to clean generation, as expected when outputs depart from the target-writer distribution, and serve only to diagnose severe generation degradation. Figure 3 shows that InkShield concentrates perturbations near stroke boundaries while preserving the requested content word.
Table 2: Ablation study of InkShield on One-DM. RN and DW denote ResNet50 and DeepWriter, respectively. ∆Top1/∆Top5 denote target-writer retrieval-rate reductions relative to clean generation. D-Rank denotes the selected decoy writer’s rank decrease. Figure 4: Cross-model transferability. Qualitative generations from One-DM, CONSTANT, and DiffusionPen for the requested word planets. DiffusionPen uses five independently InkShield-protected references; one representative reference is shown.
RN@1 ↑ RN@5 ↑ DW@1 ↑ DW@5 ↑ D-Rank ↑ LPIPS ↓ BgE ↓
Variant Global mask
11.77
32.83
9.70
28.92
21.51
0.0915
67.90
InkShield
11.08
29.52
- Dec.
10.88
28.25
8.75
25.94
16.43
0.0078
0.00
8.35
24.27
10.79
0.0136
- Tgt.
10.58
0.00
28.67
8.43
24.98
15.27
0.0140
- Gen.
9.98
0.00
26.75
8.18
23.46
11.99
0.0150
0.00
Ablation Study Table 2 reports ablations of the InkShield components. RN and DW denote the ResNet50 and DeepWriter evaluators, and RN@1/RN@5 and DW@1/DW@5 denote target-writer Top-1/Top-5 retrieval-rate reductions relative to clean generation. D-Rank is the average reduction in the selected decoy writer’s retrieval rank across the two evaluators. “- Dec.”, “Tgt.”, and “- Gen.” remove decoy attraction, target-writer repulsion, and surrogate-generation optimization, respectively. Removing either decoy attraction or target-writer repulsion weakens mimicry suppression, confirming that both terms are needed for effective style displacement. Removing the surrogate-generation objective produces the largest degradation: RN@5 decreases from 29.52 to 26.75, and DW@5 decreases from 25.94 to 23.46. This result shows that static style representations alone are insufficient for the one-shot handwriting threat. The global-mask variant yields stronger suppression (RN@5: 32.83; DW@5: 28.92), but incurs a substantial stealth cost: LPIPS increases from 0.0078 to 0.0915, and BgE rises from 0.00 to 67.90. It therefore serves as an aggressive reference point rather than a practical defense. Overall, InkShield achieves a more favorable protection–stealth trade-off through decoy-guided displacement, surrogategeneration optimization, and edge-constrained perturbations.
Hyperparameter Sensitivity Figure 5 evaluates the decoy percentile while keeping all other settings fixed. Increasing the percentile generally lowers target-writer Top-1 and Top-5 retrieval rates under both
RN@1
RN@5
DW@1
DW@5
10
5
1
5
10
20
80
90
95
99
Decoy percentile
(a) 0.0080
137
0.0075
LPIPS
138
FID
Figure 4 compares generations obtained from clean and InkShield-protected references. Clean references enable all three generators to reproduce target-writer-like characteristics, whereas protected references induce visible style departures while preserving the requested word. The effect is strongest on the One-DM surrogate, but the changes observed for CONSTANT and DiffusionPen demonstrate transferability across generators with different reference-conditioning schemes. For DiffusionPen, we independently protect all five conditioning references, indicating that the defense remains applicable when generation uses multiple style examples. Detailed quantitative cross-model results are provided in the Appendix.
Target-writer retrieval (%)
Cross-model Transferability
136 135
0.0070 0.0065
1
5
10
20
80
90
95
99
1
5
10
20
80
Decoy percentile
Decoy percentile
(b)
(c)
90
95
99
Figure 5: Sensitivity to the decoy percentile. Target-writer retrieval rates, FID, and LPIPS across decoy percentiles are shown in (a), (b), and (c), respectively. P95 provides the selected protection–stealth trade-off.
ResNet50 and DeepWriter, since more distant decoys provide stronger style-displacement directions. P95 achieves the lowest or near-lowest retrieval rates without relying on the single farthest decoy. More extreme decoys provide limited additional suppression but increase LPIPS and lead to a less favorable FID trend. We therefore use P95 as the default, as it provides strong mimicry suppression with localized, lowvisibility perturbations. Additional sensitivity results for the perturbation budget, optimization steps, number of pseudo content texts, and perceptual-constraint parameters are provided in the Appendix.
Conclusion In this paper, we study proactive reference-side protection against unauthorized handwriting style mimicry. We propose InkShield, which protects handwriting references before release through decoy-guided style displacement, strokeedge constraints, and frozen surrogate optimization. Experiments on IAM show that InkShield suppresses target-writer mimicry while preserving reference readability and visual utility, with transfer to CONSTANT and DiffusionPen. These results demonstrate a practical defense against unauthorized style reuse of publicly shared handwriting.
References Bhunia, A. K.; Khan, S.; Cholakkal, H.; Anwer, R. M.; Khan, F. S.; and Shah, M. 2021. Handwriting Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 1086–1094. Chen, R.; Jin, H.; Liu, Y.; Chen, J.; Wang, H.; and Sun, L. 2024. EditShield: Protecting Unauthorized Image Editing by Instruction-Guided Diffusion Models. In Computer Vision – ECCV 2024, volume 15121 of Lecture Notes in Computer Science, 126–142. Springer. Choi, J. S.; Lee, K.; Jeong, J.; Xie, S.; Shin, J.; and Lee, K. 2025. DiffusionGuard: A Robust Defense Against Malicious Diffusion-Based Image Editing. In Proceedings of the International Conference on Learning Representations. Dai, G.; Zhang, Y.; Ke, Q.; Guo, Q.; and Huang, S. 2024. One-DM: One-Shot Diffusion Mimicker for Handwritten Text Generation. In Proceedings of the European Conference on Computer Vision. Dai, G.; Zhang, Y.; Qin, Y.; Guo, Q.; Huang, S.; and Yan, S. 2025. Beyond Isolated Words: Diffusion Brush for Handwritten Text-Line Generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 19054–19064. Gan, J.; and Wang, W. 2021. HiGAN: Handwriting Imitation Conditioned on Arbitrary-Length Texts and Disentangled Styles. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 7484–7492. He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778. Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30. Huynh-Thu, Q.; and Ghanbari, M. 2008. Scope of validity of PSNR in image/video quality assessment. Electronics letters, 44(13): 800–801. Kang, L.; Riba, P.; Wang, Y.; Rusiñol, M.; Fornés, A.; and Villegas, M. 2020. GANwriting: Content-Conditioned Generation of Styled Handwritten Word Images. In Proceedings of the European Conference on Computer Vision, 273–289. Le, A.-D.; Pham, V.-L.; Vo, T.-N.; Mai, X. T.; and Tran, T.-A. 2026. CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Li, M.; Lv, T.; Chen, J.; Cui, L.; Lu, Y.; Florencio, D.; Zhang, C.; Li, Z.; and Wei, F. 2023. Trocr: Transformer-based optical character recognition with pre-trained models. In Proceedings of the AAAI conference on artificial intelligence, volume 37, 13094–13102. Li, Y.; Zhang, W.; Lyu, X.; Liu, Y.; and Xiao, B. 2025. StyleGuard: Preventing Text-to-Image-Model-Based Style Mimicry Attacks by Style Perturbations. In Advances in Neural Information Processing Systems.
Liang, C.; and Wu, X. 2023. Mist: Towards Improved Adversarial Examples for Diffusion Models. arXiv:2305.12683. Liu, Y.; Fan, C.; Dai, Y.; Chen, X.; Zhou, P.; and Sun, L. 2024. MetaCloak: Preventing Unauthorized Subject-Driven Textto-Image Diffusion-Based Synthesis via Meta-Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24219–24228. Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; and Vladu, A. 2018. Towards Deep Learning Models Resistant to Adversarial Attacks. In International Conference on Learning Representations. Marti, U.-V.; and Bunke, H. 2002. The IAM-database: an English sentence database for offline handwriting recognition. International journal on document analysis and recognition, 5(1): 39–46. Nikolaidou, K.; Retsinas, G.; Christlein, V.; Seuret, M.; Sfikas, G.; Barney Smith, E.; Mokayed, H.; and Liwicki, M. 2023. WordStylist: Styled Verbatim Handwritten Text Generation with Latent Diffusion Models. In Document Analysis and Recognition – ICDAR 2023, volume 14188 of Lecture Notes in Computer Science, 384–401. Springer. Nikolaidou, K.; Retsinas, G.; Sfikas, G.; and Liwicki, M. 2024. DiffusionPen: Towards Controlling the Style of Handwritten Text Generation. In Proceedings of the European Conference on Computer Vision Workshops. Pippi, V.; Cascianelli, S.; and Cucchiara, R. 2023. Handwritten Text Generation from Visual Archetypes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 22458–22467. Pippi, V.; Quattrini, F.; Cascianelli, S.; and Cucchiara, R. 2023. HWD: A novel evaluation score for styled handwritten text generation. arXiv preprint arXiv:2310.20316. Salman, H.; Khaddaj, A.; Leclerc, G.; Ilyas, A.; and Madry, A. 2023. Raising the Cost of Malicious AI-Powered Image Editing. In Proceedings of the International Conference on Machine Learning, 29894–29918. Shan, S.; Cryan, J.; Wenger, E.; Zheng, H.; Hanocka, R.; and Zhao, B. Y. 2023. Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models. In Proceedings of the USENIX Security Symposium. Sobel, I.; and Feldman, G. 1968. A 3x3 Isotropic Gradient Operator for Image Processing. Presented at the Stanford Artificial Intelligence Project. Van Le, T.; Phung, H.; Nguyen, T. H.; Dao, Q.; Tran, N.; and Tran, A. 2023. Anti-DreamBooth: Protecting Users from Personalized Text-to-Image Synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2116–2127. Vanherle, B.; Pippi, V.; Cascianelli, S.; Michiels, N.; Van Reeth, F.; and Cucchiara, R. 2024. Vatr++: Choose your words wisely for handwritten text generation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(2): 934–948. Wan, C.; He, Y.; Song, X.; and Gong, Y. 2024. PromptAgnostic Adversarial Perturbation for Customized Diffusion Models. In Advances in Neural Information Processing Systems, volume 37.
Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 13(4): 600–612. Xing, L.; and Qiao, Y. 2016. DeepWriter: A Multi-Stream Deep CNN for Text-Independent Writer Identification. arXiv preprint arXiv:1606.06472. Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 586–595.
A Writer Evaluator Training and Validation We use two independently trained writer evaluators, ResNet50 and DeepWriter, to assess target-writer mimicry. Both models are trained as writer classifiers, while evaluation uses their penultimate-layer embeddings for writer-level retrieval rather than their softmax outputs. This design measures whether a generated image remains close to the target writer in the learned writer-style space. For both evaluators, we use the 161 IAM test writers as the target-writer set. For each writer, five images are reserved to construct the retrieval gallery, resulting in 805 gallery images in total. These gallery images are excluded from training and validation. The remaining real test-writer images are split independently by writer into approximately 70%/10%/20% training, development, and held-out test subsets. We select the checkpoint with the best development Top-1 classification accuracy. ResNet50. We initialize ResNet50 from ImageNet weights. Each grayscale word image is inverted, padded to a square canvas, and resized to 224 × 224. The model is first trained for writer classification on the IAM training writers. It is then transferred to the 161 test writers by retaining the backbone, reinitializing the 161-way classification head, and fine-tuning on the corresponding training subset. DeepWriter. DeepWriter uses a two-stream architecture tailored to handwritten word images. Inputs preserve their aspect ratios after resizing to height 113, and two width-113 strips are sampled as the two streams. We first train the model for writer classification on the IAM training writers, then transfer its branches and backbone to the 161 test writers, reinitialize the 161-way classification head, and fine-tune using the same split protocol. For validation, we extract penultimate-layer embeddings for held-out real images and retrieve the nearest writer among the 161 gallery centroids. Table 3 reports this clean held-out retrieval performance over 4,927 query images. These results establish that both evaluators can reliably recognize writer identity on real IAM images before being used to assess target-writer mimicry in generated samples. Table 3: Clean held-out retrieval performance of writer evaluators on IAM test writers. Five images per writer form the gallery (805 images total), and results are computed on 4,927 held-out real query images. Top-1/Top-5 are embedding-retrieval rates, not softmax classification accuracies. Source similarity denotes the cosine similarity to the centroid of the true writer. Evaluator
Top-1 ↑ Top-5 ↑ Median Rank ↓ Source Sim. ↑
ResNet50 DeepWriter
79.84 65.15
91.66 85.40
2.25 3.72
0.8088 0.7921
B Additional Cross-Model Transferability Results We further evaluate whether InkShield-protected references transfer to handwriting generators with different reference-conditioning mechanisms. One-DM is the frozen
surrogate used during optimization, whereas CONSTANT performs one-shot generation with a distinct architecture. DiffusionPen follows its multi-shot setting, in which all five reference images of each target writer are independently protected by InkShield. Raw target-writer retrieval rates are not directly comparable across generators because their clean-reference mimicry abilities differ. We therefore report both the clean and protected retrieval rates, together with chance-adjusted suppresC P sion. Let RG,E @K and RG,E @K denote the target-writer Top-K retrieval rates for generator G under evaluator E using clean and protected references, respectively. With N = 161 candidate writers, the chance retrieval rate is K/N . We define SuppG,E @K =
C P RG,E @K − RG,E @K C @K − K/N RG,E
× 100%.
(11)
This metric measures the fraction of target-writer retrievability above the random baseline that is removed by protection. Table 4 shows that InkShield provides the strongest suppression on the One-DM surrogate, removing 82.86%– 88.08% of above-chance target-writer retrieval across the two evaluators and retrieval levels. On CONSTANT and DiffusionPen, InkShield still removes approximately 28%–39% of above-chance retrieval. These results indicate cross-model transferability: although the defense effect is weaker than that on the surrogate, protected references consistently reduce target-writer mimicry across one-shot and multi-shot generators. Table 4: Additional cross-model transferability results. C and P denote clean-reference and InkShield-protectedreference generation, respectively. Supp. denotes chanceadjusted target-writer suppression defined in Eq. 11. All entries are percentages. (a) Top-1 Retrieval Generator
Evaluator
C↓
P↓
Supp. ↑
One-DM One-DM
ResNet50 13.20 2.12 DeepWriter 10.68 1.93
88.08 86.99
CONSTANT ResNet50 6.34 CONSTANT DeepWriter 5.95
4.12 4.29
38.82 31.15
DiffusionPen ResNet50 6.27 DiffusionPen DeepWriter 5.74
4.35 3.87
33.99 36.53
(b) Top-5 Retrieval Generator
Evaluator
C↓
P↓
Supp. ↑
One-DM One-DM
ResNet50 38.73 9.21 DeepWriter 34.31 8.37
82.86 83.13
CONSTANT ResNet50 24.72 18.77 CONSTANT DeepWriter 21.87 16.05
27.53 31.02
DiffusionPen ResNet50 20.09 14.98 DiffusionPen DeepWriter 18.14 13.39
30.09 31.59
C.1 Perturbation Budget Figure 6 analyzes how the perturbation budget ϵ controls the available capacity for defense-side style displacement. Small budgets preserve a reference that is extremely close to the clean input, but they constrain InkShield’s ability to modify the style-sensitive stroke-boundary regions used by downstream generators. Consequently, target-writer retrieval remains relatively high at the smallest tested budgets.
2/255 15
Mean R@5 (%) ↓
4/255
13 12 11
RN@1 RN@5
DW@1 DW@5
10
0.0082
8
0.0080
6 4
0.0078 0.0076
2 20
40
60
80
100
120
20
40
60
80
Steps
Steps
(a)
(b)
100
120
Figure 7: Sensitivity to the number of optimization steps. (a) Target-writer retrieval rates for ResNet50 (RN) and DeepWriter (DW). (b) LPIPS. The vertical dashed line indicates the selected default of 80 steps.
16
14
Figure 7 studies the number of projected-gradient optimization steps. At low step counts, the perturbation has insufficient opportunity to jointly optimize the decoy-guided style objective and the generation-aware surrogate objective. As the step count increases, target-writer retrieval decreases for both ResNet50 and DeepWriter, showing that additional optimization improves suppression during the early stage.
LPIPS
We provide additional sensitivity analyses for the hyperparameters used by InkShield. These experiments complement the decoy-percentile study reported in the main paper. In each sweep, we vary only one hyperparameter and keep the remaining settings fixed to the default configuration. We evaluate the resulting trade-off among target-writer mimicry suppression, reference stealth, generation quality, and optimization cost. Unless otherwise stated, the default configuration uses ϵ = 16/255, 80 optimization steps, K = 4 cached pseudo content texts, LPIPS threshold τ = 0.005, and LPIPS penalty weight λperc = 1.
C.2 Number of Optimization Steps
Target-writer retrieval (%) ↓
C Supplementary Experimental Results
8/255 12/255
10 9 16/255
24/255
8 0.000 0.002 0.004 0.006 0.008 0.010 0.012 0.014
LPIPS ↓
Figure 6: Sensitivity to the perturbation budget ϵ. Each point is annotated with its budget. The horizontal axis denotes LPIPS and the vertical axis denotes mean Top-5 target-writer retrieval; both are lower-is-better. The orange star marks the selected default ϵ = 16/255. Increasing ϵ generally lowers mean Top-5 target-writer retrieval, showing that a larger feasible perturbation region enables stronger mimicry suppression. This benefit is accompanied by higher LPIPS, which indicates a larger perceptual deviation from the clean reference. The sweep therefore reveals a clear suppression–stealth trade-off rather than a monotonic preference for the largest possible budget. We select ϵ = 16/255 because it lies near the favorable knee of this trade-off. Compared with smaller budgets, it substantially reduces target-writer retrieval; compared with ϵ = 24/255, it retains a lower LPIPS while providing similar protection. This choice is therefore used in all main experiments as a practical low-visibility operating point.
The retrieval curves become substantially flatter around 80 steps. Beyond this point, additional iterations yield only limited and inconsistent changes across the two writer evaluators. In contrast, LPIPS exhibits a gradual upward trend as optimization proceeds, indicating that longer runs consume more of the perceptual budget without consistently improving target-writer mimicry suppression. We therefore choose 80 steps as the default. This setting reaches the stable region of the retrieval curves while avoiding the extra runtime and perceptual deviation associated with longer optimization. The result also shows that InkShield does not rely on an excessively long attack schedule to obtain its main protection effect.
C.3 Number of Pseudo Content Texts Figure 8 evaluates the number K of cached clean pseudo content texts used by the generation-aware surrogate objective. Using multiple pseudo texts exposes the protected reference to diverse content conditions during optimization, reducing the risk that InkShield over-specializes to a single generated word or text instance. Using multiple pseudo texts generally improves mimicry suppression relative to a single pseudo text while preserving broadly similar LPIPS, CER, HWD, and FID values. This pattern indicates that diverse content conditions strengthen the generality of the surrogate objective without substantially degrading reference stealth or generated-content fidelity. The marginal benefit becomes small beyond K = 4. In contrast, the measured optimization time per reference increases markedly as more pseudo texts are included, because each text requires an additional surrogate-loss evaluation. We
2.41
2.24
2.47
2.31
RN@5 ↓
10.01
9.66
9.40
9.00
DW@1 ↓
2.44
2.03
2.15
1.95
DW@5 ↓
9.43
8.69
8.80
8.34
LPIPS ↓
0.0080
0.0080
0.0080
0.0080
CER ↓
17.43
17.63
17.27
17.17
HWD ↓
2.12
2.14
2.13
2.12
FID ↓
136.49
137.23
137.16
136.83
Time/ref (s) ↓
11.2
17.7
30.2
60.5
1
2
4
8
10.0
0.005
9.8 9.6 9.4
K
Figure 8: Sensitivity to the number of pseudo content texts K. Cell annotations show measured values. Colors are normalized within each metric row, with darker cells indicating lower values. The final row reports the measured optimization time per reference sample. therefore use K = 4 as the default efficiency–effectiveness trade-off: it captures useful content diversity while avoiding the substantial runtime cost of K = 8.
C.4 Perceptual-Constraint Parameters Figure 9(a) analyzes the LPIPS threshold τ , which determines when the hinge-style perceptual penalty becomes active. With Pperc = max(0, LPIPS(xpro , x)−τ ), setting τ = 0 makes the penalty active for any nonzero perceptual deviation. Thus, smaller thresholds impose a stricter preference for reference-level visual similarity, whereas larger thresholds allow more perceptual deviation before the penalty is activated. The selected threshold τ = 0.005 provides a balanced operating point. It allows InkShield to modify localized handwriting regions sufficiently to reduce target-writer retrieval, while discouraging unnecessary perceptual deviation. These results support a thresholded soft perceptual penalty rather than a strict hard constraint throughout optimization. Figure 9(b) varies the penalty weight λperc , which determines the relative importance of the LPIPS term after the threshold is exceeded. Setting λperc = 0 removes this stealthoriented regularization, whereas an overly large value causes the optimizer to prioritize visual similarity over target-writer mimicry suppression. The default λperc = 1 provides the most balanced observed trade-off, maintaining low LPIPS while preserving strong suppression of target-writer retrieval.
2
0.020
0.25
9.2 1
9.0
0.5 0
0.010
0.015
8.8 0.008 0.009 0.010
lower
λperc (default: 1.0)
τ (default: 0.005)
higher
Mean Top-5 retrieval (%) ↓
RN@1 ↓
0.011
0.012
0.006
0.008
0.010
0.012
LPIPS ↓
LPIPS ↓
(a)
(b)
0.014
Figure 9: Sensitivity to perceptual-constraint parameters. (a) Effect of the LPIPS threshold τ ; smaller τ activates the hinge-style perceptual penalty at smaller perceptual deviations. (b) Effect of the LPIPS penalty weight λperc ; setting λperc = 0 removes the perceptual regularization. The horizontal axis is LPIPS and the vertical axis is mean Top-5 target-writer retrieval; both are lower-is-better. Orange diamonds mark the selected default settings.