Conceptio › Archive › arXiv CS
arXiv CSopen access

CFE-PPAR: Compression-friendly encryption for privacy-preserving action recognition leveraging video transformers

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2605.05692v1 [cs.CV] 7 May 2026

CFE-PPAR: COMPRESSION-FRIENDLY ENCRYPTION FOR PRIVACY-PRESERVING ACTION RECOGNITION LEVERAGING VIDEO TRANSFORMERS 1

Haiwei Lin, 2 Shoko Imaizumi

1,2

Graduate School of Informatics Chiba University Chiba, Japan

ABSTRACT Privacy-preserving action recognition (PPAR) enables machines to understand human activities in videos without revealing sensitive visual content. Among the various strategies for PPAR, encryption-based methods achieve strong privacy protection while maintaining high recognition performance. However, these methods lead to a catastrophic decrease in recognition performance and visual quality when the encrypted videos are compressed. That is, the previous methods are not compression-friendly. To address these issues, in this paper, we propose the first compression-friendly encryption method for PPAR, called CFE-PPAR. In CFE-PPAR, videos encrypted with secret keys can be directly recognized by a video transformer, which uses parameters transformed by the same keys as those used for video encryption. In experiments, it is verified that CFE-PPAR outperforms previous methods on the UCF101 and HMDB51 datasets under Motion-JPEG and H.264 compression. Index Terms— Privacy-preserving, action recognition, video transformers, video encryption, video compression 1. INTRODUCTION Action recognition is a fundamental task in computer vision that aims to identify and classify human activities from videos or image sequences. With the rapid development of deep neural networks (DNNs), the applications and usage environments of DNNs are becoming more widespread. In practice, action recognition systems are often deployed in cloud and edge environments. However, data transmission and storage raise serious concerns such as data privacy leakage and increasing data volume. Accordingly, privacy-preserving action recognition (PPAR) has become an urgent challenge. PPAR can be broadly divided into two types: quality reduction-based and encryption-based. For quality reductionbased methods [1, 2, 3, 4, 5], sensitive information in frames is obfuscated by converting it into low-fidelity representations while preserving task-relevant features; however, these This work was supported in part by JSPS KAKENHI under Grants JP25K07733 and JP25K07750.

3∗

Hitoshi Kiya

3

Faculty of System Design Tokyo Metropolitan University Tokyo, Japan No encryption Plain accuracy Accuracy drops slightly Plain videos

Model

If uncompressed inputs

If compressed inputs

(a) Plain accuracy Accuracy drops significantly Model Lossless Existing encryption

Distorted Decryption Plain accuracy Accuracy drops slightly

(b)

Model Lossless CFE-PPAR (Ours)

Clear Decryption

Fig. 1: Utility comparison of encryption-based methods under uncompressed and compressed conditions. (a) Previous method. (b) Our CFE-PPAR.

methods commonly suffer from degraded recognition performance and incomplete privacy protection. In comparison, encryption-based methods [6, 7] leave scarcely any perceptible visual cues. In addition, protected content can be restored (i.e., decrypted) by authorized clients who need to conduct forensic investigations in surveillance applications. However, existing encryption-based methods do not consider applying video compression, so videos have to be transmitted in uncompressed formats. Accordingly, we propose a compression-friendly encryption method for PPAR for the first time, called CFE-PPAR. As shown in Fig. 1 (a), when videos encrypted with existing methods are compressed [6, 7, 8], recognition accuracy significantly drops, and the decryption process fails to recover faithful content. In contrast, the proposed method, CFEPPAR, aims to address the above limitations, as shown in

Fig. 1 (b). CFE-PPAR is carried out with a video transformer (VT) [9, 10] and two training-free components: compressionfriendly encryption (CFE) and key-dependent domain adaptation (KDDA). As a result, the VT can perform accurate inference on the encrypted videos without requiring any specific fine-tuning or architectural modifications. In experiments, the effectiveness of CFE-PPAR is verified on the UCF101 [11] and HMDB51 [12] datasets under Motion-JPEG and H.264 compression. The experimental results show that CFE-PPAR can maintain the same accuracy as recognition on plain videos without compression and effectively mitigate catastrophic performance degradation under lossy compression. 2. PRELIMINARY

Given a video clip X, the cube embedding layer divides X into patches with a spatial resolution of Hp × Wp pixels. For every T consecutive frames, the patches with identical spatial coordinates are stacked to form a cube with dimensions of T × Hp × Wp . Subsequently, all cubes are linearly projected into embedding vectors through 3D convolution with a kernel E. To retain positional information, positional embeddings Epos are added to the embedding vectors. Finally, the embedding vectors are passed through the transformer encoder to produce logits for action recognition. 3. PROPOSED METHOD The proposed method, CFE-PPAR, is described below.

2.1. Privacy-Preserving Action Recognition

3.1. Overview and Threat Model

As mentioned, methods for PPAR can be broadly categorized into quality reduction-based methods and encryption-based methods. Quality reduction-based methods, such as downsampling [1, 2], quantization [3], and perturbation [4, 5], leverage various forms of visual distortion to suppress sensitive attributes. However, these methods remain unable to sufficiently decouple sensitive attributes from task-relevant features [13]. This inevitably compromises recognition performance and often leaves visual cues that are exploitable by reconstruction attacks. As a promising alternative, encryption-based methods excel in recognition performance and security. Such methods have primarily explored two approaches: homomorphic encryption (HE) [6] and perceptual encryption (PE) [7]. HE [6] enables inference to be directly performed on encrypted data that is compatible with arithmetic computations. However, HE-based methods are computationally intensive and rely on specialized DNNs. In contrast, [7] is a PE-based method, which uses keydependent pixel shuffling to encrypt videos. While [7] does not match the provable security of HE, it can fully preserve recognition performance on encrypted videos without incurring additional latency or requiring architectural modifications to the model. Nevertheless, the method does not focus on compressing video data. Specifically, inference accuracy degrades significantly when encrypted data undergo compression. Moreover, once subjected to compression, the encrypted data cannot be faithfully recovered through decryption.

Fig. 2 presents an overview of the proposed method. Our scenario comprises three types of entities: clients, a model developer, and a public server, where the public server is untrusted. As shown in Fig. 2, clients generate encrypted videos and send them to the public server after compressing the encrypted data. A block-wise encryption, which will be explained in Section 3.2 in detail, is carried out using secret keys KST and KM S . The keys are pre-shared between the authorized clients and the model developer through secure channels. The model developer prepares a domain-adaptation VT model in which a portion of the model parameters in a trained VT are transformed with keys KST and KM S as in [15]. The public server decompresses the received data and inputs the encrypted data into the domain-adaptation VT to recognize the videos uploaded by the clients. Note that the public server has access to neither the secret keys nor the visual information of the videos. A threat model includes a set of assumptions, such as an attacker’s goals, knowledge, and capabilities. The aim of an adversary is to restore visual information from encrypted data. We assume that the attacker can access the encrypted data and the encryption algorithm but does not have the keys. Accordingly, the attacker can only perform ciphertext-only attacks (COAs). The proposed method is discussed by taking these environments into consideration in this paper, similar to still images.

2.2. Video Transformer

3.2. Compression-Friendly Encryption

Our methodology is built upon VTs [9, 10], which have recently emerged as the dominant architecture for video understanding. For simplicity, we consider a standard VT consisting of a cube embedding layer and a vanilla transformer encoder [14]. This model handles video inputs according to the pipeline described below.

As described in Section 3.1, clients encrypt video inputs, and the encrypted videos are compression-friendly, but those encrypted with existing encryption methods are not. Accordingly, we propose an encryption method for PPAR that achieves high recognition accuracy in addition to excellent compression characteristics.

Clients

Model developer Spatially subdivide

Frame sampling and resize

3D Conv kernel 𝑬 (𝑇×𝐻! ×𝑊" ) Plain video

1 3 1 3

1 2 3 4 Divide

Frame

Main-blocks (𝐻! ×𝑊" )

Sub-block-level transformations 𝑲𝑺𝑻

Subdivide

4 2 4 2

1 3 1 3

4 2 4 2

1 3 1 3

2 4 2 4

1 3 1 3

𝑲𝑴𝑺

4 2 4 2

1 3 1 3

4 1 4 1 2 3 2 3

Transformed ' kernel 𝑬

𝑲𝑺𝑻

2 4 1 3 Rearranged ' 𝒑𝒐𝒔 embeddings 𝑬

𝑲𝑴𝑺

(b) Key-dependent domain adaptation

Sub-blocks (𝐻# ×𝑊# )

Main-block scrambling

4 2 4 2

Rearrangement

Positional embeddings 𝑬𝒑𝒐𝒔

2 4 2 4

Sub-block-level transformations

Sub-blocks (𝐻# ×𝑊# )

Clip

1 3 1 3

2 4 2 4

Public server (Untrusted) 1 3 1 3

4 2 4 2

1 3 1 3

Encrypted frame

Domain-adaptation Video Transformer

Upload Compress and decompress

(a) Compression-friendly encryption

Logits 3D Conv Encrypted video

Add

Cube embedding layer

Transformer encoder

(c) Inference on encrypted videos

Fig. 2: Overview of CFE-PPAR. As illustrated in Fig. 2 (a), the proposed method starts with preprocessing such as frame sampling and frame size conversion of original videos to align the dimensions of query videos with those of the VT on the server. After that, clients carry out the following steps to generate compression-friendly encrypted videos. A. 1 Divide each resized frame into main-blocks (MBs) with a size of Hp × Wp pixels, which is equivalent to the patch size used for cubes in VTs. A. 2 Subdivide each MB into sub-blocks (SBs) with a resolution Hs × Ws . A. 3 Randomly apply five sub-block-level transformations to each SB using key KST : rotation, flipping, negativepositive inversion of pixel values, RGB-channel shuffling, and sub-block scrambling (permutation). A. 4 Carry out main-block scrambling (block permutation) within each frame where the scrambling order of MBs is determined by key KM S . As shown in Fig. 3, frames encrypted with the above steps can protect the visual information of original videos. In addition, the blocks in encrypted videos still maintain a high correlation, so the encrypted data is compressible. 3.3. Key-Dependent Domain Adaptation To prevent the encryption from influencing the recognition accuracy, we apply a key-dependent domain adaptation technique to models trained with plain videos as proposed in [15]. As shown in Fig. 2 (b), the parameters of the cube embedding layer in VT are transformed by modifying the 3D convolution

kernel E and the positional embeddings Epos . The procedure is given below. B. 1 Subdivide the spatial components of E into sub-blocks with a size of Hs × Ws , which is the same size as that used in Section 3.2. B. 2 To avoid the influence of A.3, apply the five sub-blocklevel transformations used in A.3 to 3D convolution kernel E with key KST . B. 3 To avoid the influence of A.4, modify the positional embeddings Epos with key KM S . The modified models, called domain-adaptation VTs, are used for action recognition on the public server. 4. EXPERIMENT 4.1. Experimental Setup Experiments were conducted on two benchmark datasets for action recognition, UCF101 and HMDB51, where UCF101 contains 3,783 test videos spanning 101 action classes, and HMDB51 is composed of 1,530 test videos from 51 action classes. We used vivit-b-16x2 [9] as a base model in our experiments. This model is composed of a vanilla transformer encoder with a convolutional cube embedding layer using a cube size of 2 × 16 × 16 (i.e., T = 2, Hp = 16, and Wp = 16). It operates on videos with a duration of 32 frames and a spatial resolution of 224 × 224. To meet this input specification, we adopted uniform frame sampling and bicubic interpolation to sample and resize the frames, respectively. We initial-

Table 1: Comparison of PPAR methods in terms of recognition accuracy and visual quality under compression. Bold indicates the best score for each condition, excluding Plain.

Method

MJPEG

Uncomp

H.264 without inter-frame prediction

High bitrate

Mid bitrate

Low bitrate

High bitrate

Mid bitrate

High bitrate

Mid bitrate

Low bitrate

Acc

Acc

PSNR

Acc

PSNR

Acc

PSNR

Acc

Acc

PSNR

Acc

PSNR

Acc

PSNR

Acc

PSNR

Acc

PSNR

92.92 82.52 92.92 92.92 92.92

91.99 79.32 12.18 91.07 89.88

35.75 N/A 18.09 32.24 31.31

90.88 77.00 6.61 87.73 85.59

33.25 N/A 17.89 29.88 29.17

81.05 69.07 2.34 71.42 66.40

29.84 N/A 17.39 26.95 26.66

89.29 79.11 19.56 90.88 90.06

88.43 77.42 9.44 89.29 86.33

33.61 N/A 17.91 30.64 29.64

87.02 73.09 6.15 77.35 70.45

30.31 N/A 17.77 27.56 26.42

92.36 80.41 20.06 92.52 92.47

38.01 N/A 18.20 36.77 36.61

92.33 78.35 12.79 92.04 91.83

36.55 N/A 18.01 35.12 34.96

92.18 75.60 4.50 91.09 90.67

34.44 N/A 17.63 32.77 32.61

69.15 66.99 4.44 62.06 58.95

35.45 N/A 19.05 30.21 30.02

68.63 65.10 2.94 49.08 47.19

33.36 N/A 18.77 27.25 27.00

70.00 67.06 14.77 68.76 68.75

39.65 N/A 19.26 36.62 36.61

69.54 67.12 6.86 68.43 67.65

38.92 N/A 18.89 35.02 35.02

69.47 66.14 4.58 66.73 65.62

37.75 N/A 18.72 32.75 32.74

PSNR

Low bitrate

H.264 with inter-frame prediction

UCF101 Plain BDQ [3] LCVE [7] CFE-PPAR (V1) CFE-PPAR (V2)

35.78 N/A 18.06 32.50 36.42

HMDB51 Plain BDQ LCVE CFE-PPAR (V1) CFE-PPAR (V2)

Plain

70.26 68.23 70.26 70.26 70.26

BDQ [3]

69.08 67.52 7.84 67.12 64.11

35.53 N/A 19.17 32.18 31.31

LCVE [7]

68.16 66.67 4.57 64.90 60.20

33.57 N/A 18.95 30.36 29.85

63.85 64.71 2.48 51.63 46.01

CFE-PPAR (V1)

31.02 N/A 18.74 30.21 26.88

69.48 67.32 11.24 66.54 65.03

CFE-PPAR (V2)

Fig. 3: Frame samples of test videos produced by different PPAR methods in experiments.

ized the model with parameters pretrained on Kinetics-400 and fine-tuned it on the target datasets. For simplicity and reproducibility, we adopted a relatively naive fine-tuning configuration without task-specific tricks such as robust training. Our method was compared with three methods: Plain, LCVE [7], and BDQ [3], where Plain serves as a standard reference that directly takes the original videos as input and involves no modifications to the model. LCVE and BDQ are two representative current state-of-the-art methods for PPAR, where LCVE is a perceptual encryption-based method and BDQ follows a quality reduction-based strategy. BDQ trains a privacy-preserving encoder, namely BDQ encoder, to obfuscate video inputs while preserving motion features for downstream recognition. In CFE-PPAR, we set 16 × 16 as the MB size and 8×8 (i.e., Hs = Ws = 8) as the SB size. We considered two variants of CFE-PPAR, namely CFE-PPAR (V1) and CFE-PPAR (V2), according to their key assignment schemes for the MBs in a frame. Specifically, CFE-PPAR (V1) transforms all MBs using the same keys, while CFE-PPAR (V2) transforms each MB using different keys. Fig. 3 shows frame samples for the methods mentioned above. As the experimental protocol, we first generated test videos processed by BDQ, LCVE, and CFE-PPAR, and then compressed the processed test videos and plain videos. Here, two widely used video coding methods, Motion-JPEG (MJPEG) and H.264, were used for video compression. For

36.94 N/A 19.19 32.52 32.41

H.264, two prediction modes, H.264 without inter-frame prediction and H.264 with inter-frame prediction, were used for evaluation. The latter uses both intra-frame and inter-frame predictions. 4.2. Recognition Performance without Compression The experimental results are summarized in Tab. 1. The accuracy of action recognition without compression (Uncomp) is given as Acc (%) in the table. The results show that both LCVE and our method, CFE-PPAR, had the same accuracy as that of Plain, but BDQ had lower accuracy due to the use of low-quality videos. The results verified that CFE-PPAR does not cause any performance degradation when the encrypted videos are uncompressed. 4.3. Recognition Performance with Compression When lossy compression is applied to videos, the action recognition accuracy degrades in general due to compressioninduced noise. To evaluate the recognition performance for each method under lossy compression, we compressed videos at several bitrates as shown in Tab. 1, in which High, Mid, and Low bitrates corresponded to average bitrates of 0.80, 0.60, and 0.40 bpp (bit per pixel), respectively. LCVE showed significant accuracy degradation across all bitrates on both datasets compared with Plain. In contrast, at High and Mid bitrates, CFE-PPAR showed substantially smaller degradation than LCVE across all video coding methods. Notably, under H.264 with inter-frame prediction, CFEPPAR remained nearly on par with Plain across all bitrates. Although noticeable accuracy degradation was observed at Low bitrates, CFE-PPAR achieved a substantial improvement over LCVE. Both variants of CFE-PPAR exhibited similar trends in accuracy, where V1 consistently outperformed V2 by a slight margin under all conditions. While BDQ showed

Compressed and then decrypted Encrypted

(a)

(b)

(c)

LCVE

Bitrate: 0.41 bpp PSNR: 16.27 dB

Bitrate: 0.40 bpp PSNR: 10.66 dB

Bitrate: 0.39 bpp PSNR: 10.64 dB

CFE-PPAR (V1)

Bitrate: 0.39 bpp PSNR: 31.01 dB

Bitrate: 0.39 bpp PSNR: 26.96 dB

Bitrate: 0.40 bpp PSNR: 29.53 dB

Plain

Bitrate: 3.32 bpp

(a) Plain

(b) Encrypted

CFE-PPAR (V1)

CFE-PPAR (V2)

(c) Directly restored by EJPS

CFE-PPAR (V1)

CFE-PPAR (V2)

(d) Compressed and then restored by EJPS

CFE-PPAR (V1)

CFE-PPAR (V2)

Fig. 5: Encrypted videos of CFE-PPAR restored by EJPS under different compression conditions. CFE-PPAR (V2)

Bitrate: 0.39 bpp PSNR: 30.07 dB

Bitrate: 0.40 bpp PSNR: 26.70 dB

Bitrate: 0.40 bpp PSNR: 29.49 dB

Fig. 4: Decrypted frame samples from encrypted videos compressed at Low bitrate (≈ 0.40 bpp). (a) MJPEG. (b) H.264 without inter-frame prediction. (c) H.264 with inter-frame prediction.

competitive performance on HMDB51, CFE-PPAR achieved consistently better results on UCF101. This trend indicates that CFE-PPAR scales more effectively with increasing numbers of action classes. 4.4. Reconstruction of Original Frames PPAR methods are often required to reconstruct the visual information of original data for authorized clients. To assess the reconfigurability, we decrypted videos encrypted under different compression conditions and then calculated PSNR (peak signal-to-noise ratio) values compared with the uncompressed videos. Tab. 1 shows the average PSNR values of test videos for each condition, where PSNR values for BDQ were excluded since videos protected by BDQ are irreversible. From Tab. 1, the videos produced by LCVE had PSNR values below 20 dB. In contrast, the videos produced by CFE-PPAR effectively avoided such catastrophic degradation. Fig. 4 illustrates the effects of different video coding methods on the quality of decrypted frames. Accordingly, our method outperformed the previous PPAR methods in terms of reconfigurability. 4.5. Attack Resistance As an encryption-based method, CFE-PPAR has to be robust against various attacks, including known-plaintext attacks (KPAs), chosen-plaintext attacks (CPAs), and ciphertext-only attacks (COAs). To mitigate the risks of KPAs and CPAs, CFE-PPAR adopts a one-time key policy, where each video is encrypted using a set of unique keys. In this paper, we mainly consider COAs in the threat model, which allows attackers to

access the encrypted videos and the encryption algorithm, but not the secret keys. Many studies have been conducted on the COA resistance of perceptual image encryption [16, 17]. However, since CFE-PPAR is a block-wise encryption method that maintains a high correlation within a block, it is potentially vulnerable to attacks based on the jigsaw puzzle solver (JPS) [18, 19]. Accordingly, we adopted the extended JPS (EJPS) [19], a state-of-the-art attack method that is particularly effective against encryption schemes using small block sizes. We considered two video conditions to conduct the attack: one where the encrypted videos are not compressed, and another where the encrypted videos undergo lossy compression. Under the first condition, the videos encrypted with CFE-PPAR (V1) were almost restored by EJPS, as illustrated in Fig. 5 (c). In contrast, no meaningful visual information was restored for CFE-PPAR (V2). Accordingly, CFE-PPAR (V2) was demonstrated to be more robust than CFE-PPAR (V1). For the second condition, we compressed the encrypted videos to Low bitrate (≈ 0.40 bpp) using MJPEG and then applied EJPS to them. In this case, reconstructing visual information became much more difficult as illustrated in Fig. 5 (d), regardless of whether the videos were encrypted with CFEPPAR (V1) or CFE-PPAR (V2). 5. CONCLUSION In this paper, we proposed a novel method for privacypreserving action recognition, called CFE-PPAR. CFE-PPAR protects the visual information of videos while enabling the encrypted videos to be effectively compressed using standard compression methods such as MJPEG and H.264. Even under severe lossy compression, CFE-PPAR substantially mitigates accuracy degradation and allows authorized clients to reconstruct the original videos with reliable quality. In experiments, the effectiveness of CFE-PPAR was verified on the UCF101 and HMDB51 datasets in terms of recognition accuracy, reconstruction quality of original frames, and attack resistance.

6. REFERENCES [1] Michael Ryoo, Brandon Rothrock, Charles Fleming, and Hyun Jong Yang, “Privacy-preserving human activity recognition from extreme low resolution,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 31, no. 1, Feb. 2017.

[10] Zhan Tong, Yibing Song, Jue Wang, and Limin Wang, “Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training,” in Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds. 2022, vol. 35, pp. 10078– 10093, Curran Associates, Inc.

[2] Michael Ryoo, Kiyoon Kim, and Hyun Yang, “Extreme low resolution activity recognition with multi-siamese embedding learning,” in Proceedings of the AAAI conference on artificial intelligence, 2018, vol. 32.

[11] Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah, “UCF101: A dataset of 101 human actions classes from videos in the wild,” CoRR, vol. abs/1212.0402, 2012.

[3] Sudhakar Kumawat and Hajime Nagahara, “Privacypreserving action recognition via motion difference quantization,” in European Conference on Computer Vision. Springer, 2022, pp. 518–534. [4] Filip Ilic, He Zhao, Thomas Pock, and Richard P. Wildes, “Selective interpretable and motion consistent privacy attribute obfuscation for action recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2024, pp. 18730–18739. [5] David Schneider, Sina Sajadmanesh, Vikash Sehwag, Saquib Sarfraz, Rainer Stiefelhagen, Lingjuan Lyu, and Vivek Sharma, “Activity recognition on avataranonymized datasets with masked differential privacy,” arXiv preprint arXiv:2410.17098, 2024. [6] Miran Kim, Xiaoqian Jiang, Kristin Lauter, Elkhan Ismayilzada, and Shayan Shams, “Secure human action recognition by encrypted neural network inference,” Nature communications, vol. 13, no. 1, pp. 4799, 2022. [7] Yuchi Ishikawa, Masayoshi Kondo, and Hirokatsu Kataoka, “Learnable cube-based video encryption for privacy-preserving action recognition,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), January 2024, pp. 7003– 7013. [8] Koki Madono, Masayuki Tanaka, and Masaki Onishi, “Scramblemix: A privacy-preserving image processing for edge-cloud machine learning,” in Image and Video Technology, Wei Qi Yan, Minh Nguyen, Parma Nand, and Xuejun Li, Eds., Singapore, 2024, pp. 326–340, Springer Nature Singapore. [9] Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid, “Vivit: A video vision transformer,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021, pp. 6836–6846.

[12] H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre, “Hmdb: A large video database for human motion recognition,” in 2011 International Conference on Computer Vision, 2011, pp. 2556–2563. [13] Zhenyu Wu, Zhangyang Wang, Zhaowen Wang, and Hailin Jin, “Towards privacy-preserving visual recognition via adversarial training: A pilot study,” in Proceedings of the European Conference on Computer Vision (ECCV), September 2018. [14] Alexey Dosovitskiy, “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [15] Teru Nagamori, Sayaka Shiota, and Hitoshi Kiya, “Efficient fine-tuning with domain adaptation for privacypreserving vision transformer,” APSIPA Transactions on Signal and Information Processing, vol. 13, no. 1, pp. 1–15, 12 2024. [16] Banafsheh Saber Latibari, Najmeh Nazari, Muhtasim Alam Chowdhury, Kevin Immanuel Gubbi, Chongzhou Fang, Sujan Ghimire, Elahe Hosseini, Hossein Sayadi, Houman Homayoun, Soheil Salehi, and Avesta Sasan, “Transformers: A security perspective,” IEEE Access, vol. 12, pp. 181071–181105, 2024. [17] Chengqing Li, Xianhui Shen, and Sheng Liu, “Solving the block-wise puzzles in an encryption-thencompression system,” IEEE Transactions on Multimedia, pp. 1–15, 2026. [18] Dror Sholomon, Omid David, and Nathan S. Netanyahu, “A genetic algorithm-based solver for very large jigsaw puzzles,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2013. [19] Tatsuya Chuman and Hitoshi Kiya, “A jigsaw puzzle solver-based attack on image encryption using vision transformer for privacy-preserving dnns,” Information, vol. 14, no. 6, 2023.

Record · ID 160743 · SHA-256 c2e264c90cc2a8f1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.