Conceptio › Archive › arXiv CS
arXiv CSopen access

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Highlights Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in ResourceConstrained Agricultural Field Conditions Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith, Gangireddy Rahul Jogi, Sudheesh Manalil, Arnab Raha, Amitava Mukherjee, Parthasarathy Seethapathy, G. Gopakumar • H-BAC enables curvature-aware adaptive block pruning of Vision Transformers. • H-BAC cuts ViT-B/16 FLOPs by 49% while retaining 95.53% OOD accuracy. • The full pipeline compresses ViT-B/16 by 54.5× to a 6.01 MB model.

arXiv:2609.05334v1 [cs.CV] 4 Sep 2026

• Cross-village and cross-device testing evaluates real-world generalization.

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions⋆ Mahadev Sunil Kumara , Bhavika Gondia , Desaisetty Venkata Satya Sai Swapnitha , Gangireddy Rahul Jogia , Sudheesh Manalilb , Arnab Rahac , Amitava Mukherjeed , Parthasarathy Seethapathye and G. Gopakumara,∗ a Department of Computer Science and Engineering, Amrita School of Computing, Amrita Vishwa Vidyapeetham, Amritapuri, Kerala, India b Amrita School of Agricultural Sciences, Amrita Vishwa Vidyapeetham, Coimbatore, 642109, Tamil Nadu, India c Intel Corporation, Santa Clara, California, USA d Department of Computer Science, Birla Institute of Technology and Science, Pilani Dubai Campus, Dubai, United Arab Emirates e Agrotechnology Division, CSIR–Institute of Himalayan Bioresource Technology, Palampur, 176061, India

ARTICLE INFO

ABSTRACT

Keywords: Vision Transformer Model Compression Edge AI Plant Disease Detection Knowledge Distillation

Chilli (Capsicum annuum) is one of India’s most economically significant crops, yet its productivity is persistently threatened by diseases that are difficult to identify without expert intervention. While Vision Transformers (ViTs) have achieved high classification accuracy, their large computational footprint makes deployment on resource constrained devices challenging. Existing compression approaches typically address pruning, quantization, and knowledge distillation in isolation, leaving the potential benefits and interactions of their combined application insufficiently explored. We propose a unified Vision Transformer compression framework that combines Hessian-Balanced Adaptive Block Pruning (H-BAC), guided by second-order sensitivity estimation, with quantization and attentionbased knowledge distillation. To systematically identify the most effective configuration within each compression family, each technique is first evaluated independently through controlled ablation studies, after which the best-performing components are integrated into a sequential deployment pipeline tailored to real-world agricultural constraints. On a chilli 3-class village-split dataset with a genuine cross-village, cross-device out-of-distribution test split, the resulting compressed models match or exceed the 95.13% FP32 baseline’s accuracy, alongside 74–98% model size reduction, and the fully integrated compression pipeline achieves a 54.5× size reduction (327.42 MB to 6.01 MB) at 95.13 ± 2.32% accuracy across four tested configurations. A direct comparison further reveals that, on this dataset, a directly-trained student of the same final size, without pruning or distillation, reaches comparable accuracy of 94.87%, at the same 6.01 MB INT8 size, indicating where H-BAC and knowledge distillation are, and are not yet shown to be, worth their computational cost.

1. Introduction The major portion of India’s economy is reliant on agriculture that is the main source of livelihood for nearly 55% of the population. Among the crop varieties produced ⋆

This work was supported by the ANRF, New Delhi, India, under Grant No. SUR/2022/004268, “Light Weight Deep Learning Based Mobile Application for the Early Detection, Identification, and Spatiotemporal Monitoring of Plant Diseases.” ∗ Corresponding author [email protected] ( Mahadev Sunil Kumar); [email protected] ( Bhavika Gondi); [email protected] ( Desaisetty Venkata Satya Sai Swapnith); [email protected] ( Gangireddy Rahul Jogi); [email protected] ( Sudheesh Manalil); [email protected] ( Arnab Raha); [email protected] ( Amitava Mukherjee); [email protected] ( Parthasarathy Seethapathy); [email protected] ( G. Gopakumar) https://www.mahadevsunilkumar.com ( Mahadev Sunil Kumar); https://sites.google.com/site/arnaverse ( Arnab Raha); https://www.bits-pilani.ac.in/dubai/dr-amitava-mukherjee/ ( Amitava Mukherjee); https://www.amrita.edu/faculty/gopakumarg/ ( G. Gopakumar) ORCID (s): 0009-0005-4368-6537 ( Mahadev Sunil Kumar); 0000-0002-8848-1069 ( Arnab Raha); 0000-0003-1694-3911 ( Amitava Mukherjee); 0000-0003-0660-9311 ( Parthasarathy Seethapathy); 0000-0003-1559-5150 ( G. Gopakumar)

M.S Kumar et al.: Preprint submitted to Elsevier

throughout the country, chilli (Capsicum annuum) is of major importance because the country is the largest producer of chilli in the world and Andhra Pradesh on its own produces 49% of the total amount at the national level [34]. However, this crucial crop is constantly at the risk of diseases. Leaf Curl Virus, Anthracnose, and Bacterial Wilt are some of the diseases that result in a loss of 20% to total crop failure depending on the severity of infection [34], and pest pressure compounds this risk further [45]. What makes this even more challenging is that many of these diseases look the same visually in the very early stages. Traditional diagnosing by manual expert checking or laboratory-based serological testing is very slow, far too costly, and highly dependent on having specialists available, which makes it not feasible at the scale of rural Indian farming [12]. To address this problem, it is necessary to have an automated system that can not only detect plant diseases but also identify them at the earliest stage in a direct manner in the field. The last ten years have witnessed tremendous strides in this area. At first, the use of traditional machine learning methods based on hand-crafted features like color histograms, texture descriptors and support vector machine classifiers were setting the stage but unfortunately, they did Page 1 of 16

Lightweight ViT Compression for Plant Disease Detection

not generalize well to actual field conditions [41, 1]. It was the advent of Convolutional Neural Networks (CNNs) that transformed the scene: [8] reached 95% classification performance over 25 plant species, deeply rooting image-based deep learning as the prime approach. Further crop-specific designs even improved on this; for example, [34] reached 99.12% on a chilli dataset. Vision Transformers (ViTs) [7], which model relationships across all image patches simultaneously through self-attention, have since pushed performance beyond 99% on standard benchmarks [24, 9], and now represent the state of the art for plant disease classification. However, this state-of-the-art representational power comes at a steep computational cost. A standard ViT-B/16 occupies 327 MB and takes nearly 28 ms per CPU inference even on modern consumer hardware (Section 4), making it incompatible with the low-end mobile devices that represent the primary computing resource available to most rural farmers [24, 9]. The gap between laboratory performance and field deployability is not a minor inconvenience; it is the central unsolved problem in applied agricultural AI. Model compression techniques (pruning, quantization, and knowledge distillation) offer a good guide forward, but previous works see them as separated interventions [27, 21]. No one has, so far, conducted a joint evaluation of all three for ViT compression in agriculture where, besides variable lighting and class imbalance, device inference limitations add up to changes in what the general benchmark can depict [19]. Although pruning, quantization, and knowledge distillation have been individually explored, existing works largely treat them as independent techniques and do not account for the varying sensitivity of transformer blocks. Moreover, conventional distillation approaches often suffer from feature mismatch when applied to heterogeneous architectures, and prior evaluations are typically limited to generalpurpose benchmarks without considering real-world deployment constraints. The main contributions of this work are as follows: • We propose H-BAC, a curvature-aware adaptive block pruning method for Vision Transformers that combines Hutchinson-estimated block-level Hessian curvature with first-order Taylor-based component pruning within each block, and is, to our knowledge, the first such method evaluated as part of a joint compression pipeline for agricultural disease detection. • We introduce an attention-based knowledge distillation scheme that transfers relational reasoning through attention maps, bypassing the feature-dimension mismatch that causes conventional feature-based distillation to fail when teacher and student differ substantially in capacity. • We present the first joint evaluation of pruning, quantization aware training, and knowledge distillation for ViT compression under real-world agricultural constraints, providing a systematic comparison across all three techniques. M.S Kumar et al.: Preprint submitted to Elsevier

2. Related work and research gaps 2.1. From traditional ML to vision transformers The detection of crop diseases using automated methods has evolved considerably over the past decade. Early approaches relied on traditional machine learning with hand-crafted features such as color histograms, texture descriptors, and SVM classifiers [4, 35, 13]. Earlier plant disease studies combined classical machine learning with CNN-based feature learning, demonstrating that hybrid approaches can improve detection performance over handcrafted feature methods [39]. These methods, although feasible, were unable to generalize under real field conditions. Deep learning solved this problem by allowing an end-to-end feature learning that directly extracts features from images without the need for intermediate representation [41, 1]. Image-based deep learning frameworks have also been proposed for plant disease prediction, showing that endto-end models can effectively learn discriminative disease features directly from leaf images [22]. [8] found that CNNs could achieve 95% accuracy in classifying 25 different plant species in the PlantVillage dataset [32], establishing the convolutional approach as the reference standard in the field of agricultural image classification. Following that, domainspecific modifications: [34] through the integration of a Squeeze-and-Excitation CNN into a chilli dataset, reached an accuracy of 99.12%, while [10] developed a particular network architecture for chilli disease detection. Recent work has shown that CNN-based transfer learning can significantly improve plant disease classification accuracy [31]. More recent crop-specific studies show that transformerbased and modern deep learning models are increasingly being adapted for agricultural diagnosis [30]. Nevertheless, CNN-based approaches are still limited by their local receptive fields, which is one of the reasons for changing the focus on attention-based architectures. Vision Transformers (ViTs), introduced by [7], overcome this limitation by considering an image as a set of patches where selfattention is applied to model local as well as global contexts. This characteristic is very advantageous for the task of plant disease detection, which requires relating the infection pattern to the overall leaf structure. [42] claimed that their GAN-augmented ViT could reach an accuracy of 99.92% on PlantVillage, whereas [2] demonstrated tomato disease detection through a ViT-SmartAgri app on a smartphone at 90.99% accuracy. [12] supplemented attention mechanisms to architectures used for multi-crop classification, beating previous best results on several benchmarks. [14] proposed a multitask ViT that jointly localizes and classifies plant disease regions, demonstrating that attention-based architectures can support diagnostic tasks beyond simple classification. However, standard ViTs require very high computational resources [43]. Crop-specific mobile disease diagnosis has also been explored for eggplant little leaf disease, reinforcing the need for deployment-aware model selection [40]. [9] proposed TrIncNet, replacing Multi-Layer Perceptron (MLP) blocks with Inception modules to reduce Page 2 of 16

Lightweight ViT Compression for Plant Disease Detection

model complexity. Newer framework-level studies continue to confirm the promise of deep learning for early detection of plant diseases, but most still do not directly address efficient deployment under strict edge-device constraints [44].

2.2. Lightweight architectures and model compression Plant-based Mobile ViT (PMVT) [24] fused the efficiency of MobileViT with Convolutional Block Attention Modules (CBAM), demonstrating acceptable accuracyefficiency trade-offs. SLViT [26] aimed at sugarcane leaf disease diagnosis, introduced shuffle operations to reduce computational complexity. A recent Edge-AI for agriculture study [19] showed that full-scale ViTs (323 MB) can reach 96% accuracy but are unsuitable for resource limited settings, demonstrating that architectural efficiency alone is insufficient and post-training compression methods are required. Model compression covers pruning, quantization, and knowledge distillation [27]. The work by [11] on a combination of magnitude-based pruning with quantization and Huffman coding [16] laid the groundwork, resulting in significant storage reductions with almost no accuracy loss. [33] further enhanced this with importance estimation based on Taylor expansion, including first-order gradient information for more principled pruning decisions. Closest to our pruning approach, [48] (NViT) derives a Hessianbased structural saliency criterion for global ViT pruning with latency-aware regularization, redistributing parameters across blocks and within-block structures. H-BAC differs in three respects: it estimates block-level curvature via the Hutchinson trace estimator specifically to keep second-order estimation tractable at ViT-B/16 scale without a differentiable latency-regularization term, combines this with a separate first-order Taylor criterion for within-block component (attention head and MLP neuron) selection rather than a single unified saliency score, and is evaluated as one stage of a joint pruning-distillation-quantization deployment pipeline on a real-world agricultural edge-deployment task rather than as a standalone pruning method on generalpurpose benchmarks. Knowledge distillation [15] allows small student models to learn deeper inter-class relations from the teacher’s soft probability outputs. [37] revealed that the student can get closer to the teacher model’s performance if temperature scaling and intermediate layer alignment are combined. Quantization-Aware Training (QAT) imitates low numerical precision during fine-tuning, so the model can update its weights before INT8 conversion [3]. Each of these methods has been analyzed in isolation, meaning that whether their combination yields cumulative compression benefits remains an open research question. Recent work demonstrates that combining compression techniques yields results beyond what any single method achieves. [21] proposed PQK, a sequential pruning, quantization, and distillation pipeline. [29, 28] showed that joint knowledge distillation and channel pruning consistently outperforms either applied individually. [46] extended this to M.S Kumar et al.: Preprint submitted to Elsevier

unsupervised domain adaptation through iterative KD and pruning under label scarcity. [25] demonstrated that OpenVINO’s [18] Joint Pruning, Quantization, and Distillation toolkit achieves a 5.24× compression ratio with a 4.19× performance gain on BERT-base at under 1% accuracy loss. In the context of agricultural disease detection using Vision Transformers, a joint evaluation of these combined compression methods remains unexplored. Despite these promising results, each approach treats pruning, quantization, and distillation as largely separable steps, optimized without regard for how one technique influences the behavior of the others. These interdependencies mean that sequentially applying well-performing individual techniques does not guarantee a well-performing joint solution, and can lead to compounding accuracy losses. Beyond this methodological limitation, none of the above works are evaluated under conditions representative of real agricultural deployment.

2.3. Research gaps and our contributions Despite these advancements, there is still a major research gap in the systematic co-design of pruning, quantization, and distillation for ViTs in agricultural contexts. Most existing compression works are aimed at generalpurpose benchmarks and overlook domain-specific challenges such as changing light conditions, uneven class distributions, and the need for explanation in farmer-facing tools [27, 19]. Furthermore, conventional ViT pruning techniques score first-order importance identically for all transformer blocks, thereby neglecting significant differences in block-level sensitivity. Conventional feature-based distillation similarly breaks down when teacher and student differ substantially in hidden dimension, as linear adapters cannot adequately bridge the feature-space mismatch. This work addresses these gaps through a compression pipeline co-designed around ViT geometry rather than a sequential application of independent methods. The blocklevel curvature gap is addressed by H-BAC, which employs second-order Hessian curvature to assign non-uniform pruning rates across transformer blocks. The attention-based KD gap is targeted by an Attention-Based Knowledge Distillation scheme that transfers relational reasoning through multi-head attention maps, bypassing feature-dimension incompatibility entirely. Finally, the joint evaluation gap is filled by evaluating H-BAC, Attention-Based KD, and PTQDynamic INT8 quantization together as a single deployment pipeline rather than as independent techniques.

3. Methodology 3.1. Dataset and preprocessing The dataset [20] consists exclusively of locally captured images of diseased chilli leaves, taken on-site across smallholder farms in Coimbatore, Tamil Nadu, India, spanning three classes: Healthy control, Initial Symptoms of chilli leaf curl virus (ChiLCV), and Severe Symptoms of ChiLCV. These images reflect real-world acquisition conditions that

Page 3 of 16

Lightweight ViT Compression for Plant Disease Detection

include natural lighting variation, partial occlusion, complex backgrounds, and mixed disease symptom presentation: characteristics largely absent from controlled laboratory datasets. A selection of representative images from this collection is shown in Fig. 1.

Figure 1: Indigenously collected chilli leaf images captured in field conditions.

Unlike a pooled random split, every split in this dataset is drawn from a physically distinct village, and the Out of Distribution (OOD) split used for every reported result in this paper additionally spans different smartphone camera devices than training. The training split (17,655 images) was collected in Arasampalayam; the validation split (2,207 images) in Vadasithur; an in-distribution test split (2,207 images) in Myleripalayam; and a fourth, an out-ofdistribution split (760 images) across two further villages, Kuladupalayam and Andipalayam, captured on three different phones (Realme C2, Redmi Note 12, Samsung Galaxy S23) not used to collect any other split. All four splits were collected between June and November 2024. Fig. 2 visualizes the per-class composition of each split. As every split is a physically distinct village, rather than a random stratified sample of a pooled collection, this dataset directly tests cross-domain generalization. Therefore, a model can no longer succeed by memorizing villagespecific lighting, soil background, or camera-sensor characteristics that happen to recur across a random train and test portions. We exploit this directly by reporting all held-out accuracy, precision, recall, and F1 figures throughout this paper on the OOD split rather than the in-distribution Test split, since OOD additionally varies the acquisition device. This provides stricter and more deployment-realistic generalization test than an in-distribution village would provide on its own. All images were resized to 224×224 pixels using bicubic interpolation to meet the patch embedding input requirements of the ViT-B/16 architecture [7]. During training, random resized cropping with scale range (0.8, 1.0) and random horizontal flipping were applied to improve robustness to scale variation and viewpoint changes. Color jitter augmentation with brightness and contrast factors of 0.2 was applied to simulate natural lighting variation encountered in field settings. All images were normalized using ImageNet M.S Kumar et al.: Preprint submitted to Elsevier

Figure 2: Per-class image counts across the four villagepartitioned splits (log scale).

statistics [5] (mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225]).

3.2. Baseline model ViT-B/16 was chosen as the high-accuracy baseline model pre-trained on ImageNet-1K [5] due to its wellestablished performance on image classification tasks [7]. To adapt the model for 3-class chilli disease classification, the final classification head was replaced with a fully connected layer with 3 output units, while all other pretrained weights were retained and fine-tuned end-to-end on the Train split (Arasampalayam). The fine-tuned baseline, evaluated on the OOD split, achieved an accuracy of 95.13%, a (weighted) precision of 95.55%, recall of 95.13%, and F1score of 95.10%. All latency figures in this paper use batched throughput timing: each timed call processes a batch of 128 images, and wall-clock time is divided by 128 to obtain a per-sample figure measurement. Reported values are the mean over independent runs (3 trials on CPU, 10 on GPU). Model training/fine-tuning and CPU inference latency measurement used an Apple M4 Pro (12-core CPU) MacBook Pro; GPU inference latency was measured separately on a NVIDIA RTX 4060 (8 GB) instance. All CPU latency figures are measured through Apple’s Core ML runtime (models converted with coremltools, executed on the CPU compute unit), which on this hardware is substantially faster than PyTorch’s eager-mode CPU execution for every architecture tested here; GPU figures remain PyTorch/CUDA measurements, since Core ML targets only Apple’s own compute units. The baseline performance and deployment characteristics are summarized in Table 1. The baseline fine-tuning recipe (AdamW, LR= 10−4 , weight decay 10−4 , label smoothing 0.1, cosine schedule with a 3-epoch warmup, early stopping on validation loss with patience 5) was used. We compared it, on the same Page 4 of 16

Lightweight ViT Compression for Plant Disease Detection Table 1 Baseline ViT-B/16 Performance and Deployment Characteristics Metric Value OOD Accuracy 95.13% Precision (weighted) 95.55% Recall (weighted) 95.13% F1-Score (weighted) 95.10% Model Size 327.42 MB Parameters 85.80M CPU Latency (Apple M4 Pro, Core ML) 7.18 ± 0.01 ms GPU Latency (NVIDIA RTX 4060) 5.56 ± 0.004 ms

held-out run, against four alternatives: a 3× and 10× lower learning rate (3 × 10−5 and 10−5 ), the lower learning rate combined with a dropout layer (0.1) before the classification head, and the lower learning rate combined with a 5× higher weight decay (5 × 10−4 ). None of the four alternatives improved on the existing recipe (96.05% OOD accuracy in this comparison). Although the baseline achieves strong classification accuracy on genuinely unseen villages and devices, its large parameter count (85.80M) makes it a poor fit for the low-computation-power smartphones that most smallholder farmers in India rely upon. This motivates the compression pipeline described in the following subsections. To handle the large model, we have applied three complementary model compression methods: Hessian-Balanced Adaptive Block Pruning (H-BAC), PTQ-Dynamic INT8 quantization, and Attention-Based Knowledge Distillation (KD). Fig. 3 illustrates the general principle of a size/accuracybudget-driven deployment workflow: compress, check against the target size and minimum-accuracy constraints, and escalate to a more aggressive configuration if the constraints are not yet met. The paper reports two concrete uses of this principle. The primary one, evaluated throughout Section 4, is the fixed Integrated Compression Pipeline described below (H-BAC at a single pre-selected pruning ratio, then KD, then PTQ-Dynamic quantization), which does not itself iterate. The second is a constrained search mode where we start from a size budget and an accuracy-drop budget and want the least aggressive configuration that meets both automatically; Section 4 reports this mode’s concrete staged algorithm and grid-search results separately from the fixed pipeline’s own numbers.

3.3. H-BAC pruning We present Hessian-Balanced Adaptive Block Pruning (H-BAC), the curvature-aware pruning method that reduces Vision Transformer FLOPs while preserving classification accuracy for edge deployment. Unlike prior work that applies uniform pruning across all transformer blocks, H-BAC incorporates second-order Hessian curvature information to allocate non-uniform, curvature-weighted pruning rates across blocks, and calculates importance using first-order Taylor series for component-level pruning within each block.

M.S Kumar et al.: Preprint submitted to Elsevier

Figure 3: Workflow of a size/accuracy-budget-driven compression workflow: compress, check against target constraints, escalate if not met. Section 4 details the concrete 4-stage cascade used by the constrained search mode and its results.

The method proceeds through four phases: obtaining blockwise curvature via Hessian trace estimation, assigning pruning resources based on curvature, executing structural pruning on attention heads and MLP neurons, and performing fine-tuning for accuracy recovery. Unlike first-order pruning methods that rely solely on gradient magnitudes, Hessian-based sensitivity measure captured through the Hessian matrix provides a more principled measure of parameter sensitivity. However, exact computation of the full Hessian is computationally intractable for large models such as ViT-B/16, as memory and compute requirements grow quadratically with the number of parameters [48]. To make second-order sensitivity estimation tractable, H-BAC employs the Hutchinson trace Page 5 of 16

Lightweight ViT Compression for Plant Disease Detection

estimator [17], which approximates the Hessian trace using random Rademacher vectors without explicitly forming the full Hessian matrix. This reduces the per-block curvature estimation cost from (𝑝2 ) to (𝐾𝑝), where 𝑝 is the number of parameters in a block and 𝐾 is the number of Hutchinson samples. In practice, 𝐾 = 10 samples per block were found sufficient for stable curvature estimates. We used the Hutchinson estimator [17] to approximate the curvature of each transformer block by calculating the trace of the Hessian: 1 ∑ 𝑇 𝑣 𝐻𝑣𝑘 Tr(𝐻) ≈ 𝐾 𝑘=1 𝑘 𝐾

(1)

where 𝑣𝑘 are random Rademacher vectors. For each transformer block, we compute the curvature separately for attention parameters and MLP parameters. The total block curvature is the weighted sum: 𝑏 𝑏 + 𝜆 ⋅ 𝐻mlp 𝐶𝑏 = 𝐻attn

(2)

where 𝜆 balances the relative importance of the Multi-Layer Perceptron (MLP) component. This coefficient is essential because the significant differences in parameter counts and gradient scales between the attention and MLP layers would otherwise cause the MLP block to dominate the overall sensitivity estimate. To make rough second-order estimates more reliable, we take the absolute value of the curvature and average it over a small calibration dataset. Unlike conventional pruning methods that apply uniform or magnitude-based criteria across all blocks, H-BAC uses block-level curvature as a sensitivity metric to determine how much each block should be pruned. Having curvatures of blocks 𝐶1 , 𝐶2 , … , 𝐶 𝐵 , we (√ ) normalise the pruning weights √ C . Given a global pruning rate 𝑟, as 𝑊𝑏 = 𝐶𝑏 ∕mean the pruning ratio 𝑅𝑏 assigned to block 𝑏 is: 𝑅𝑏 = 𝑟 ⋅ 𝑊𝑏 Block ratios are clipped to [0.05, 0.95], then rescaled so the mean ratio across blocks equals the global pruning target 𝑟, and finally re-clipped to the same bounds (clipping alone does not preserve the mean when curvature values are highly skewed across blocks, so an explicit renormalization step after clipping is required, and the rescaling can in turn push individual ratios back outside the bounds, so a second clip is applied; see Algorithm 1). Under this direct (noninverted) weighting, higher-curvature blocks receive larger pruning ratios and lower-curvature blocks are pruned less. This was verified empirically. At a 50% global ratio, prefinetune accuracy under this direct weighting (64.87%) exceeds a uniform-pruning baseline at the same ratio (41.32%) by 23.55 percentage points. This direct-weighting formulation runs counter to the inverse-curvature intuition (protect high-curvature/high-sensitivity blocks) that motivates most Hessian-based pruning heuristics in the literature. A theoretically-motivated inverse-weighting variant, tested under identical conditions, performed worse on both ends (34.21% pre-finetune, 94.21% post-finetune) – even underperforming uniform pruning pre-finetune – confirming that M.S Kumar et al.: Preprint submitted to Elsevier

direct weighting empirically identifies blocks that tolerate more pruning, not less. At the component level, each transformer layer contains multiple attention heads that are rated using a first-order Taylor importance criterion [23, 33]. For a particular attention head ℎ, the importance score is: ⟨ ⟩ 𝐵 𝜕 1 ∑ (𝑏) ,𝑓 (3) 𝐼𝑗 = 𝐵 𝑏=1 𝜕𝑓 (𝑏) 𝑗 𝑗

where 𝑓𝑗(𝑏) denotes the feature map produced by the 𝑗-th filter for the 𝑏-th sample in a mini-batch of size 𝐵,  is the classification loss, and ⟨⋅, ⋅⟩ denotes the Frobenius inner product. For block 𝑏, the number of heads pruned 𝑘𝑏 is: 𝑘𝑏 = ⌈𝑅𝑏 ⋅ 𝐻⌉

(4)

where 𝐻 is the total number of heads in the block. The least important heads are masked by zeroing the corresponding Query, Key and Value slices. MLP neuron pruning is performed at the neuron level in the first fully connected layer of each transformer block. Neuron importance is estimated from a first-order Taylor approximation: | 𝜕 | 𝐼𝑛 = || ⊙ 𝑤𝑛 || | 𝜕𝑤𝑛 |

(5)

where 𝑤𝑛 denotes the weights of neuron 𝑛. Neurons are ordered according to the L2 norm of their output weights, and those with the least importance are eliminated. H-BAC uses mask-based structured pruning, where deactivated attention heads and MLP neurons are found by zeroing their weights while keeping tensor dimensions intact. During fine-tuning, the masks act as gradient masks to control the effect of pruning on training. An iterative shortretraining phase is conducted after the pruning steps to help the network recover from potential performance degradation. Fine-tuning is performed using a reduced learning rate with all remaining parameters unfrozen, and a single epoch was used in our testing. The hybrid design of H-BAC is crucial because it incorporates second-order curvature to reveal global sensitivity differences among the blocks, and uses first-order Taylor pruning as a lightweight mechanism for within-block pruning, producing higher pruning granularity and a better accuracy-efficiency trade-off. H-BAC exploits the curvature of the loss function to decide the pruning ratio of each transformer block, giving more pruning to high-curvature blocks and less to low-curvature, non-critical ones (see Section 3.3 for the empirical validation of this direct weighting).

3.4. Quantization strategies To further reduce memory usage and accelerate inference, two Post-Training Quantization (PTQ) variants were investigated: dynamic and static quantization, both applied directly to an already-trained FP32 checkpoint with no additional fine-tuning. Page 6 of 16

Lightweight ViT Compression for Plant Disease Detection

Algorithm 1: H-BAC: Hessian-Balanced Adaptive Block Pruning Input: Pretrained ViT model  with 𝐵 transformer blocks; Calibration dataset 𝑐𝑎𝑙𝑖𝑏 ; Global pruning ratio 𝑟; Hutchinson samples 𝐾; MLP balance coefficient 𝜆; Per-block ratio bounds 𝜌min = 0.05, 𝜌max = 0.95 Output: Pruned model ′ 𝐵 1 Initialize block curvature vector 𝐂 ← 𝟎 ∈ ℝ ; 2 for 𝑏 = 1 to 𝐵 do 3 𝑐𝑏 ← 0; 4 foreach (𝑥, 𝑦) ∈ 𝑐𝑎𝑙𝑖𝑏 do 5 𝓁 ← ((𝑥), 𝑦); ( ) 2 6 𝐻attn ← Tr ∇ 𝑏 𝓁 estimated using 𝐾 𝜃attn

7

8

Hutchinson samples; ( ) 𝐻mlp ← Tr ∇2𝑏 𝓁 estimated using 𝐾 𝜃mlp

Hutchinson samples; 𝑐𝑏 ← 𝑐𝑏 + 𝐻attn + 𝜆 ⋅ 𝐻mlp ;

𝐶𝑏 ← max(𝑐𝑏 , 0)∕|𝑐𝑎𝑙𝑖𝑏 |; √ √ 10 𝐖 ← 𝐂∕mean( 𝐂); 11 𝐑 ← clip(𝑟 ⋅ 𝐖, 𝜌min , 𝜌max ); 12 𝐑 ← 𝑟 ⋅ 𝐑∕mean(𝐑); 13 𝐑 ← clip(𝐑, 𝜌min , 𝜌max ); 14 for 𝑏 = 1 to 𝐵 do 15 Compute first-order Taylor importance | 𝜕𝓁 || 𝐼𝑖 = ||𝑤𝑖 ⋅ 𝜕𝑤 | for attention heads in block 𝑏; 𝑖| | 16 Prune ⌈𝑅𝑏 ⋅ 𝐻⌉ least-important heads by zeroing corresponding QKV projection slices and output projection channels; 17 Compute first-order Taylor importance | 𝜕𝓁 || 𝐼𝑖 = ||𝑤𝑖 ⋅ 𝜕𝑤 | for MLP neurons in block 𝑏; 𝑖| | 18 Prune ⌈𝑅𝑏 ⋅ 𝑁⌉ least-important neurons by zeroing corresponding rows in fc1 and columns in fc2; 9

19

return ′ ;

PTQ-Dynamic converts trained models to low precision without the need for retraining. Dynamic quantization was applied to the trained ViT-B/16 checkpoint, converting selected layer weights from FP32 to INT8 at inference time [36]. All Linear layers, including attention projections, MLP layers, and the classifier head were targeted, with weights quantized to QINT8 to reduce memory footprint and speed up linear operations. Activations remain in FP32, with scale factors computed dynamically at inference, requiring no calibration dataset.

PTQ-Static converts both weights and activations to INT8 before inference using a calibration dataset to compute fixed scale factors, offering greater speedup than dynamic quantization [36]. However, static quantization remains challenging for complex Vision Transformer architectures, primarily due to missing quantization operators for attention mechanisms and unsupported layer normalization operations. To address this, we propose a calibrated hybrid solution in which all linear layers are quantized to INT8, and a calibration dataset of 1,000 training images sampled using Weighted Random Sampling ensures class-balanced calibration across all 3 disease classes. We quantize every checkpoint in this study with PTQ-Dynamic applied directly to its final trained weights, with no extra fine-tuning stage, and no accuracy is given up once the checkpoint being quantized is already well-trained. Full performance results across both quantization approaches are reported in Table 6 (Section 4). Between the two quantization candidates, PTQ-Dynamic (94.34% accuracy, 84.42 MB) and PTQ-Static (94.61% accuracy, 85.78 MB) both sit close to the FP32 baseline (95.13%). PTQ-Static’s 0.27-point edge is within normal run-to-run noise, and PTQ-Dynamic is what the rest of this pipeline (KD ablation, capacity sweep, constrained search) already uses consistently, so PTQ-Dynamic is adopted throughout for its simplicity. On CPU latency, the INT8-quantized model measures 7.21 ms (Apple M4 Pro via Core ML), showing no significant improvement over the FP32 baseline’s 7.18 ms.

3.5. Knowledge distillation Knowledge distillation (KD) transfers the knowledge of a large and powerful teacher model to a small student model. We evaluated three distillation approaches: responsebased, feature-based, and attention-based, for compressing the ViT-B/16 teacher (85.80M parameters, 327.42 MB) to a TinyViT student (5.52M parameters, 21.15 MB), resulting in a 15.48× compression ratio. The student model vit-tiny-patch16-224 [7, 47] was initialized from ImageNet21k [6] pre-training with 5.52M parameters and 21.15 MB. Response-Based KD: Response-based KD uses the final output probability distributions (logits) for knowledge transfer without requiring internal model representations. For each training batch 𝑡, teacher logits 𝑧𝑡 ∈ ℝ3 are obtained by passing the input images through the frozen teacher model, and the same images are passed through the student model to obtain student logits 𝑧𝑠 ∈ ℝ3 . Both sets of logits are divided by temperature 𝑇 = 4.0 before softmax for softer probability distributions: exp(𝑧𝑠𝑖 ∕𝑇 ) exp(𝑧𝑡𝑖 ∕𝑇 ) , 𝑝𝑠𝑖 = ∑ (6) 𝑝𝑡𝑖 = ∑ 3 3 𝑡 ∕𝑇 ) 𝑠 ∕𝑇 ) exp(𝑧 exp(𝑧 𝑗=1 𝑗=1 𝑗 𝑗 Here, 𝑧𝑡𝑖 and 𝑧𝑠𝑖 denote the teacher and student logits, respectively, and 𝑝𝑡𝑖 and 𝑝𝑠𝑖 denote the corresponding softened probabilities. The combined loss balances learning from the ground truth and soft targets: RKD = 𝛼hard + (1 − 𝛼)soft

M.S Kumar et al.: Preprint submitted to Elsevier

(7) Page 7 of 16

Lightweight ViT Compression for Plant Disease Detection

∑ The hard loss hard = − 3𝑖=1 𝑦𝑖 log(𝑝𝑠𝑖 ) employs the true labels to guarantee correct prediction; the soft loss soft = ∑ 𝑝𝑡 𝑇 2 ⋅ 3𝑖=1 𝑝𝑡𝑖 log 𝑝𝑠𝑖 (Kullback-Leibler divergence) expresses 𝑖

to what extent the student imitates the teacher’s distribution; and 𝛼 = 0.5 weights both targets equally. 𝑇 2 scaling compensates for the gradient magnitude decrease that results from the temperature division. Feature-Based KD: Feature-based KD targets the student model’s learning at both the output level and the internal representation levels of the teacher model. Considering the hierarchical feature learning of Vision Transformers [38, 7], we chose 4 intermediate transformer layers spaced at Layers 3, 6, 9, and 11 (out of 12). Since teacher features have 768 dimensions and the student has only 192 dimensions, adaptation layers are needed to bridge this gap, following practical guidelines established for ViT feature-based distillation [49]. For each matched layer, we introduce a learnable adapter network consisting of a two-layer MLP with ReLU activation: Linear(192 to 768) followed by ReLU non-linearity, and Linear(768 to 768) for final projection. The combined feature KD loss FKD is: FKD = 𝛼FKD hard + 𝛽feature + 𝛾soft

(8)

∑ where the feature loss feature = 41 𝑙∈{3,6,9,11} MSE(𝑓̂𝑙𝑠 , 𝑓𝑙𝑡 ) averages Mean Squared Error across 4 layers, and 𝛼 = 0.5, 𝛽 = 0.3, 𝛾 = 0.2, 𝑇 = 4.0. Attention-Based KD. Unlike prior feature-based approaches, attention-based KD works by transferring knowledge via the distribution of attention weights rather than raw feature values. The attention weights reveal which patches interact strongly (e.g., diseased region attending to healthy regions for contrast), enabling the student to learn the teacher’s visual reasoning process. For each transformer layer, self-attention computes: ( ) 𝑄𝐾 𝑇 Attention(𝑄, 𝐾, 𝑉 ) = softmax √ 𝑉 (9) 𝑑𝑘 √ The attention matrix 𝐴 = softmax(𝑄𝐾 𝑇 ∕ 𝑑𝑘 ) ∈ ℝ197×197 represents pairwise token relationships. Since the teacher has 12 attention heads per layer and the student has only 3, the attention weights are averaged across all heads within each layer: 12 1 ∑ 𝑡 𝐴̄ 𝑡 = 𝐴 , 12 ℎ=1 ℎ

3 1∑ 𝑠 𝐴̄ 𝑠 = 𝐴 3 ℎ=1 ℎ

(10)

For each matched layer 𝑙 ∈ {3, 6, 9, 11}, teacher and student multi-head attention maps are collected and averaged over heads to produce 𝐴̄ 𝑡𝑙 and 𝐴̄ 𝑠𝑙 . The combined attention KD loss AKD is: AKD = 𝛼AKD hard + 𝛽attention + 𝛾soft

(11)

∑ where the attention loss attention = 41 𝑙∈{3,6,9,11} and MSE(𝐴̄ 𝑠𝑙 , 𝐴̄ 𝑡𝑙 ) averages Mean Squared Error between attention maps, and 𝛼 = 0.5, 𝛽 = 0.3, 𝛾 = 0.2, 𝑇 = 4.0. M.S Kumar et al.: Preprint submitted to Elsevier

Full performance results across all knowledge distillation approaches are reported in Table 7 (Section 4). Table 2 H-BAC 50% Pruning – Classification Performance (postfinetune) Model ViT-B/16 Baseline ViT-B/16 + H-BAC (50%)

Accuracy (%) 95.13 95.53

FLOPs 35.13 GFLOPs 17.86 GFLOPs

3.6. Integrated Compression Pipeline The proposed compression framework follows a twostage optimization strategy designed to ensure principled component selection before pipeline integration. Initially, the three compression families – pruning, quantization, and knowledge distillation, are evaluated independently through controlled ablation experiments. This step determines the maximum performance achievable and side-effects related to the deployment for each individual technique, which allows for a fair comparison of compression families without the influence of confounding interactions. PTQ-Dynamic is used for the pipeline’s final quantization stage. An extra fine-tuning pass before quantization only matters on checkpoints that have not yet had adequate training, and every checkpoint reaching this stage of the pipeline already has KD’s 15 distillation epochs. PTQ-Dynamic is applied directly with no added training step. Attention-Based KD is selected among distillation strategies for its capacity to transfer relational reasoning. In the second stage, the selected components are composed sequentially to construct the final deployment-ready model. The full-precision ViTB/16 teacher undergoes H-BAC structural pruning to reduce its computational footprint while preserving important transformer blocks. Attention-Based Knowledge Distillation is subsequently applied, imparting the reasoning patterns of the frozen pruned teacher to the student model. Finally, PTQ-Dynamic quantization is applied directly to the distilled student to obtain an INT8-precision model suitable for deployment on resource-constrained hardware. This sequential composition ensures that after each stage, the working representation is further compressed and that the handling of compression technique interactions is done in a way that is controlled and reproducible rather than the joint optimization disregarding their interdependencies. Table 4 reports the stage-by-stage outcome of this pipeline at the 50% H-BAC operating point, measured on the OOD split throughout. Starting from the pipeline’s own 95.13% / 327.42 MB FP32 entry point, H-BAC pruning followed by one epoch of recovery fine-tuning yields 94.34% accuracy at unchanged mask-based size (43.70M of 85.80M parameters active). Attention-Based KD distills this into a 21.15 MB TinyViT student at 93.68% accuracy, and the final PTQDynamic quantization step, applied directly with no further fine-tuning, produces a 6.01 MB deployed model at 91.97% accuracy, corresponding to a 54.5× size reduction relative to the FP32 checkpoint size. Page 8 of 16

Lightweight ViT Compression for Plant Disease Detection Table 3 H-BAC 50% Pruning – Deployment Characteristics, uncompacted mask-based model (CPU: Apple M4 Pro via Core ML; GPU: NVIDIA RTX 4060 via PyTorch) Model ViT-B/16 Baseline ViT-B/16 + H-BAC (50%)

Active Params 85.80M 43.70M

Size (MB) 327.42 327.42

CPU (ms) 7.18 7.18

GPU (ms) 5.560 5.560

Latency here is measured on the mask-based model before structural compaction (pruned rows/columns still physically present, zeroed), matching the pipeline stages of Table 4. The two rows are, by construction, the same architecture (mask-based pruning zeroes weights without changing tensor shapes), so they get an identical measured latency here – this is deliberately not the deployment-facing number, which Table 5 reports after structural compaction.

Table 4 Integrated Compression Pipeline: stage-by-stage results (H-BAC 50%, OOD split; CPU: Apple M4 Pro via Core ML; GPU: NVIDIA RTX 4060 via PyTorch) Stage FP32 Baseline† H-BAC Pruned (pre-finetune) H-BAC Pruned (finetuned) + Attention-Based KD + PTQ-Dynamic (deployed) Direct-training alternative ∗

Accuracy (%) 95.13 64.87 95.53 93.68 91.97 94.87

Size (MB) 327.42 327.42 327.42 21.15 6.01 6.01

Active Params (M) 85.80 43.70 43.70 5.52 5.52 5.52

CPU (ms) 7.18 7.18 7.18 1.02 1.02 1.02

GPU (ms) 5.560 5.560 5.560 0.660 – —

∗ Skips H-BAC pruning and KD entirely: TinyViT trained directly on ground-truth labels, then quantized with PTQ-Dynamic, applied directly with no extra fine-tuning. † Measured through the pipeline’s own data loader, which applies the same background-suppression preprocessing used to originally train this checkpoint – see the baseline-model note earlier in this section. The pre-finetune row is measured directly from the saved intermediate checkpoint. All latency figures use batched throughput timing (batch of 128, wall time divided by 128) to amortize fixed per-call dispatch overhead; CPU and GPU latency are pending re-measurement on the target hardware (see baseline-model note).

Repeating the full pipeline across three additional configurations yields accuracies of 91.97% (50% pruning, Table 4), 97.24% and 96.45% (two further 50% runs), and 94.87% (70% pruning). Across the four tested configurations, the final-stage accuracy is therefore 95.13 ± 2.32%. The traced 50%-pruning configuration gives the lowest accuracy at 91.97%, suggesting some variability in final-stage performance across runs. We additionally tested the simplest alternative to this pipeline, skipping H-BAC pruning and KD entirely, by training a TinyViT student directly on labels and subsequently quantizing it. As shown in the final row of Table 4, this direct-training alternative reaches 94.87% accuracy at the same 6.01 MB model size. This is comparable to the 95.13 ± 2.32% accuracy obtained across the four pipeline configurations, while avoiding the additional pruning and distillation stages. Section 4 discusses this comparison and its implications in detail.

Model training/fine-tuning and CPU inference latency measurement used an Apple M4 Pro (12-core CPU) MacBook Pro; GPU inference latency was measured separately on a NVIDIA RTX 4060 (8 GB). These measurements characterize each compression technique’s relative efficiency gain (size reduction, FLOPs reduction, and cross-method latency ranking) under controlled, reproducible conditions, not the absolute latency a farmer’s device would see. Validating the deployed model on actual ARM smartphone hardware is a scope boundary of the present study, discussed further in Section 4.

4. Results 4.1. Experimental setup All experiments are carried out with the ViT-B/16 backbone fine-tuned on the chilli 3-class village-split dataset described in Fig. 2 (17,655 training images; all reported accuracy/precision/recall/F1 figures below are measured on the 760-image OOD split). H-BAC serves as a posttraining, mask-based structured pruning method. Global pruning ratios from 10% to 90% with a step of 10% were evaluated; ratios beyond 90% were not swept in this study. M.S Kumar et al.: Preprint submitted to Elsevier

Figure 4: OOD accuracy (%) of H-BAC pruned ViT-B/16 at each global pruning ratio, before and after one epoch of recovery fine-tuning. Both curves are measured.

Page 9 of 16

Lightweight ViT Compression for Plant Disease Detection Table 5 Ablation of H-BAC pruning ratio. FLOPs and active parameters (mask-based model) decrease with pruning ratio; CPU/GPU latency measured on a structurally compacted copy (CPU: Apple M4 Pro via Core ML, 3 trials; GPU: NVIDIA RTX 4060 via PyTorch, 10 trials). Pruning ratio (%) Active params (M) FLOPs (G) Acc. pre-FT (%) Acc. post-FT (%) CPU (ms) GPU (ms) 0 (baseline) 85.80 35.13 95.13 93.95 7.13 5.560 10 77.14 31.57 94.87 96.45 6.76 5.034 20 68.56 28.04 89.34 96.32 6.13 4.554 30 60.34 24.68 73.03 94.61 5.50 4.027 40 51.72 21.13 68.68 97.37 4.88 3.494 50 43.70 17.86 64.87 95.53 4.27 3.003 60 36.30 14.81 53.42 94.08 3.62 2.489 70 29.37 11.98 34.21 91.97 3.27 2.138 80 25.75 10.50 33.03 93.68 2.96 1.888 90 22.88 9.32 34.21 91.84 2.62 1.649 Sparsity (%) is omitted as it is directly derivable from active params: 1 − active∕85.80. All latency figures use batched throughput timing (batch of 128, wall time divided by 128).

Fig. 4 shows the pre-/post-finetune accuracy trends discussed below. Because H-BAC’s mask-based pruning zeroes weights without shrinking the underlying tensors, latency was measured on a structurally compacted copy of each pruned model, consistent with the deployment-facing sizes reported below. Table 5 reports both the pre-finetune and post-finetune accuracy at each global pruning ratio, since HBAC’s recovery fine-tuning stage substantially changes the accuracy-ratio trade-off. All accuracy figures are measured on the OOD split. Pre-finetune accuracy declines steeply and steadily with pruning ratio, from 95.13% at 0% to a near-random floor of 33–34% by 80–90% pruning. This is the expected signature of removing function before any recovery step is applied: heavier pruning strips away more of the network’s learned capacity, and pre-finetune accuracy tracks that loss almost monotonically across the sweep. Post-finetune, the picture changes entirely: accuracy is remarkably stable across the full ratio sweep, ranging only 91.84–97.37% with no visible degradation trend even at 90% pruning (91.84%, within 2 points of the least-aggressive configurations tested). One epoch of recovery fine-tuning is sufficient to erase nearly all of the pre-finetune ratiodependence on this dataset. In practice, this gives a practitioner considerable headroom: once recovery fine-tuning is part of the pipeline (as it always is in the integrated pipeline of Table 4), the pruning ratio’s dominant effect is on FLOPs and active-parameter count rather than on final accuracy, so the ratio can be pushed well past this study’s chosen 50% operating point – trading further FLOPs reduction for essentially no measured accuracy cost. H-BAC’s curvature-aware allocation was compared against a uniform-pruning baseline at the same 50% global ratio (identical first-order Taylor ranking within each block, but every block assigned the same pruning rate rather than HBAC’s curvature-weighted allocation). H-BAC outperforms uniform pruning at both stages of the pipeline: pre-finetune, 64.87% vs. 41.32% (a 23.55-point margin) at near-identical M.S Kumar et al.: Preprint submitted to Elsevier

active parameter counts (43.70M vs. 43.30M); post-finetune, 95.53% vs. 93.16% (a 2.37-point margin). Curvature-guided budget allocation therefore protects accuracy more effectively than a flat allocation both before and after recovery training, confirming that the block-level sensitivity information H-BAC estimates via the Hutchinson trace is worth the extra computation it costs.

4.2. Quantization results Two quantization strategies were evaluated against the FP32 baseline: dynamic post-training quantization (PTQDynamic) and calibrated static post-training quantization (PTQ-Static). Both target INT8 precision on all Linear layers and are applied directly to the trained FP32 checkpoint, with no additional fine-tuning step. Table 6 summarizes the results. Both methods produce a ∼74% model size reduction, as the change in weight precision is the primary driver of compression, not structural modifications (PTQ-Static’s slightly larger footprint reflects the small overhead of its embedded quantization observers/scale metadata). Each costs a small amount of accuracy relative to the FP32 baseline (PTQ-Dynamic −0.79 points, PTQ-Static −0.52 points) – the expected signature of converting an already-fixed set of FP32 weights to INT8 with no further adaptation. PTQStatic’s 0.27-point edge over PTQ-Dynamic is within normal run-to-run noise; we adopt PTQ-Dynamic throughout this study (Table 4 onward) for its simplicity, since it requires no calibration dataset. On latency, INT8 quantization shows no significant improvement over FP32 inference on this hardware (1.00×; 7.21 ms vs. 7.18 ms). Because these figures use batched throughput timing (batch of 128, wall time divided by 128), fixed per-call dispatch overhead is already amortized out of the measurement, so this reflects the compute cost of Core ML’s weight-only INT8 path, which dequantizes weights back to floating point for the matrix multiply rather than executing native INT8 arithmetic, leaving per-operation cost Page 10 of 16

Lightweight ViT Compression for Plant Disease Detection Table 6 Ablation of quantization methods (CPU: Apple M4 Pro via Core ML, 3 trials). Both methods achieve a ∼74% model size reduction; the difference is in accuracy retention. Method

Size (MB) Acc. (%) ΔAcc. (%) Latency (ms) Speedup

FP32 Baseline PTQ-Dynamic PTQ-Static

327.42 84.42 85.78

95.13 94.34 94.61

— −0.79 −0.52

7.18 7.21𝑎 7.21𝑎

1.00× 1.00× 1.00×

𝑎 Core ML applies a single weight-only INT8 scheme, so both PTQ variants convert to the same INT8 weight representation and share one

measured latency; the accuracy figures are from the respective PyTorch PTQ implementations, which do differ.

Table 7 KD Ablation Study: Model Size, Performance, and Speedup Comparison (CPU: Apple M4 Pro via Core ML, 3 trials), mean ± std over 3 runs Method

Size (MB)

Compression Ratio

Acc. (%)

CPU Speedup𝑏

Teacher (ViT-B/16) TinyViT (no KD)𝑎 Response KD Feature KD Attention KD

327.42 21.15 21.15 21.15 21.15

1.00× 15.48× 15.48× 15.48× 15.48×

95.13 96.58 ± 0.26 96.27 ± 0.65 96.18 ± 0.60 96.71 ± 1.03

1.00× 7.07× 7.07× 7.07× 7.07×

𝑎 Trained using ground truth labels only. 𝑏 Speedup computed from batched-throughput CPU inference latency (batch of 128, wall time divided by 128); the teacher’s CPU latency in this benchmark run was 7.18 ms, TinyViT’s 1.02 ms. All four student configurations share the identical TinyViT architecture, so CPU speedup is identical across rows by construction; only the learned weights (and therefore accuracy) differ.

essentially unchanged. This is an important consideration for practitioners targeting ARM-based edge hardware: on this platform, INT8 quantization’s benefit is almost entirely in model size (74% reduction, critical for storage- and bandwidth-constrained devices), not inference speed.

4.3. Knowledge distillation results Three distillation strategies, plus a no-distillation control, were evaluated for compressing the ViT-B/16 teacher (327.42 MB, 85.80 M parameters, 95.13% OOD accuracy) to a TinyViT student (21.15 MB, 5.52 M parameters): responsebased KD, feature-based KD, and attention-based KD. All configurations achieve an identical 15.48× compression ratio. Table 7 reports mean ± standard deviation over 3 independent runs per method. All four training recipes cluster tightly, within a 0.53point band of each other (96.18–96.71%), with AttentionBased KD achieving the highest mean accuracy of the four (96.71±1.03%), narrowly ahead of direct training on groundtruth labels alone (96.58 ± 0.26%, no KD), response KD (96.27±0.65%), and feature KD (96.18±0.60%). AttentionBased KD’s advantage over the no-KD control is modest (0.13 points) but consistent with attention transfer being the more effective of the three distillation signals tested here: it is also the only KD variant with a higher mean than the no-KD control, while response KD and feature KD both trail it. Attention-Based KD does carry the widest run-to-run spread of the four (std 1.03 vs. 0.26 for no-KD), meaning its advantage should be read as a consistent edge across configurations rather than a guaranteed per-run win.

M.S Kumar et al.: Preprint submitted to Elsevier

Beyond the standalone comparison above, the question that matters for deployment is whether Attention-Based KD earns its place inside the full integrated pipeline, where it distills from an H-BAC-pruned-and-finetuned teacher rather than the full unpruned one, and is followed directly by PTQDynamic quantization. Training a fresh TinyViT student directly on ground-truth labels and quantizing it with the same PTQ-Dynamic step used by the pipeline’s final stage yields a 6.01 MB model at 94.87% accuracy (Table 4’s directtraining alternative row). Compared against the matched single pipeline run (91.97%), this direct-training alternative is 2.89 points ahead; compared against the mean of four independent pipeline configurations (95.13 ± 2.32%, three at a 50% pruning ratio and one at 70%), it is instead 0.26 points behind. The full pipeline’s run-to-run variance is therefore the more important effect here than any fixed accuracy cost or benefit from chaining H-BAC and Attention-Based KD ahead of quantization: averaged over multiple runs, the integrated pipeline is modestly ahead of simply training the target-sized architecture directly and quantizing it, even though any single pipeline run can land above or below that simpler alternative. To test whether Attention-Based KD’s advantage over direct training grows as student capacity shrinks below TinyViT’s 5.52 M parameters, we swept four smaller student configurations (1.31 M to 3.16 M parameters), each trained from scratch (no ImageNet-pretrained checkpoint exists for these custom architectures, so both the no-KD and Attention-KD runs at a given size start from identical random initialization, a fair like-for-like comparison) and quantized directly with PTQ-Dynamic, across 3 runs Page 11 of 16

Lightweight ViT Compression for Plant Disease Detection

tried in both FP32 and INT8 form – twelve candidate stages in total. Within families 1 and 2, where a continuous pruning ratio must be chosen, we exploit the fact that both size and accuracy are monotonically non-increasing in pruning ratio (Table 5): the ratio that just barely satisfies the size budget is therefore also the ratio with the best accuracy among every ratio that fits, so a bisection search over the ratio finds Params No-KD (%) Attn-KD (%) Advantage (pp) the optimal configuration within a family in 𝑂(log(1∕𝜖)) 1.31 M 77.24 ± 13.75 70.26 ± 19.15 −6.97 ± 32.59 evaluations rather than a linear sweep. Each capacity-tier 1.89 M 70.57 ± 19.35 76.80 ± 15.86 +6.23 ± 34.25 stage trains its candidate once per target with a freshly drawn 2.63 M 77.15 ± 6.01 81.32 ± 8.92 +4.17 ± 5.41 random initialization rather than the fixed initializations 3.16 M 90.35 ± 0.88 82.11 ± 8.07 −8.25 ± 8.06 used elsewhere in this study, since this mode answers "can a feasible model be found" as a live search rather than reproducing a specific prior run. per size. Table 8 reports the resulting means. Run-to-run We evaluated 5 size targets (10, 15, 20, 25, 30 MB) variance dominates this sweep: at three of the four sizes × 2 accuracy-drop budgets (2% and 3%), 10 combinations (1.31 M, 1.89 M, 2.63 M) A fairly consistent effect emerges in total. All 10 were satisfied, but the tightest and loosest is in the largest tested, 3.16 M parameters, where Attentionbudgets in this grid are met by different mechanisms. At the KD trails direct training by a margin close to its own noise two loosest size budgets (25 and 30 MB), the cascade stops (−8.25±8.06 points across runs). Training these small, nonat stage 3: the plain FP32 KD-distilled student (21.15 MB, pretrained architectures from scratch is evidently a much 93.68% accuracy, a 1.45-point drop from baseline) already less stable process than fine-tuning an ImageNet-pretrained fits, so no quantization is needed. At the three tightest backbone. Deploying at these sizes should budget for mulbudgets (10, 15, and 20 MB), stage 4 – the same KD student, tiple training runs rather than trusting a single run, and quantized to INT8 – is not enough: at 6.01 MB it comfortably should not expect a clean, monotonic capacity-dependent fits the size budget, but its 3.16-point accuracy drop exceeds trend from this recipe on this dataset. even the looser 3% target, so the search escalates further. Attention-Based KD from the pruned-and-finetuned teacher Stages 5 through 8 (the two largest capacity-tier students, reaches 93.68%, 1.84 percentage points below AttentionFP32 and INT8) are each tried and rejected on accuracy, Based KD run from the unpruned teacher under otherwise not size: all four reach their size targets easily but sit 8.7– identical settings (95.53%). This suggests that, at least on 11.8 points below baseline. The search succeeds at stage this dataset, whatever generalization capacity is lost through 9, the 1.89 M-parameter capacity-tier student (FP32, noH-BAC pruning and its single-epoch recovery fine-tune KD strategy): 7.27 MB at 94.08% accuracy, only a 1.05is not fully passed on for the student to recover during point drop – comfortably inside every budget from 10 MB distillation. Quantizing the KD student directly with PTQupward. Table 9 summarizes the grid; Table 10 traces stages Dynamic reaches 91.97%, matches or slightly exceeds a fine3 through 9 of the escalation for the tightest target (10 MB, tune-then-quantize variant we tested on the same checkpoint 2% drop). No target in this grid was satisfied at stage 1 or 2 (91.45%). This is consistent with the pattern established in (H-BAC pruning alone, with or without quantization): H3. The extra fine-tuning pass is only useful on checkpoints BAC’s mask-based pruning does not reduce stored model that have not yet had adequate training, and the KD student, size at all, so a pruned-but-unquantized ViT-B/16 stays at which already distilled for 15 epochs, does not need. 327.42 MB regardless of ratio, far above every budget tested here. 4.4. Constrained search results Both search modes are provided for different use cases: Beyond the fixed pipeline evaluated above, we implethe fixed pipeline (Table 4) is the right tool when the demented a constrained search mode for the scenario where ployment target is already the smallest, most aggressive a a maximum deployment size and a maximum acceptable configuration and the goal is to characterize its accuracy cost accuracy drop from the FP32 baseline is specified. The least precisely; the constrained search is the right tool when the aggressive configuration that satisfies both automatically, deployment target is a size or accuracy constraint from which rather than committing to the full H-BAC+KD+quantization the appropriate compression level – and, as this grid shows, pipeline by default is needed. The search cascades through potentially an entirely different model family – should be candidate families of increasing aggressiveness, stopping derived automatically. The escalation from the KD-distilled at the first one that satisfies both constraints: (1) H-BAC TinyViT family to the capacity-tier student family at the structural pruning alone (FP32), (2) H-BAC pruning plus tightest budgets in this grid is a direct demonstration of that dynamic INT8 quantization, (3) the KD-distilled TinyViT value: a search restricted to the fixed pipeline’s own architecstudent (FP32), (4) the KD-distilled student plus INT8 ture (stages 1–4 only) would have reported no feasible model quantization, and, should none of these four satisfy the under 21 MB at all, despite one existing. target, four additional from-scratch student architectures Table 8 Capacity sweep: Attention-Based KD vs. no-KD direct training (both taken through PTQ-Dynamic quantization, applied directly with no extra fine-tuning), four from-scratch (nonpretrained) student sizes below TinyViT’s 5.52 M parameters, mean ± std over 3 runs

below TinyViT’s capacity (the 3.16 M, 2.63 M, 1.89 M, and 1.31 M-parameter students of Table 8, largest first), each M.S Kumar et al.: Preprint submitted to Elsevier

Page 12 of 16

Lightweight ViT Compression for Plant Disease Detection Table 9 Constrained search: full 10-target grid (5 size budgets × 2 accuracy-drop budgets). All 10 combinations were satisfied. Size budget

Stage reached

Family

Acc. (%)

Size (MB)

10–20 MB 25–30 MB

9 3

1.89M-param student (FP32, no-KD) KD student (FP32)

94.08 93.68

7.27 21.15

Table 13 consolidates the best configuration from each compression family against the FP32 baseline (95.13% OOD accuracy, Table 1), providing a unified view of the accuracyefficiency space navigated by each method. Three of the four families in this table meet or exceed the FP32 baseline’s 95.13%: H-BAC 50% pruning by 0.40 points, Attention KD by 1.58 points, and the inteStage Family Acc. (%) Size (MB) Result grated pipeline, averaged over its four tested configurations, 3 KD student (FP32) 93.68 21.15 size fails 4 KD student + INT8 91.97 6.01 drop fails matches it exactly. PTQ-Dynamic quantization alone is the 5 3.16M student (FP32) 86.32 12.11 drop fails one family that costs accuracy against the baseline (−0.79 6 3.16M student + INT8 86.45 3.64 drop fails points), the expected small penalty of converting an already7 2.63M student (FP32) 83.42 10.08 drop fails fixed set of FP32 weights to INT8 with no further adaptation. 8 2.63M student + INT8 83.29 3.08 drop fails H-BAC achieves its computational reduction through struc9 1.89M student (FP32) 94.08 7.27 succeeds tured FLOP elimination (49.1% active-parameter reduction, 49.2% FLOPs reduction) while preserving important transformer blocks; PTQ-Dynamic cuts the model size by 74.2% Table 11 at a small, expected accuracy cost; and Attention-Based KD Comparison of our optimized model to existing models (CPU: Apple M4 Pro via Core ML, 3 trials; GPU: NVIDIA RTX 4060 gives the largest single-technique reduction in size (93.5%, via PyTorch, 10 trials) 327.42 MB to 21.15 MB) with no measured accuracy cost. H-BAC, Attention-Based KD, and PTQ-Dynamic quanModel Mean Latency (ms) Std. Dev. (ms) tization were combined sequentially into the integrated comCPU GPU CPU GPU pression pipeline described in Section 4 (Table 4). The ResNet50 2.320 1.712 0.010 0.0001 EfficientFormer 1.386 1.031 0.010 0.0002 combined pipeline outcome is shown in the last row of Our Model 1.015 0.660 0.002 0.0003 Table 13: a 98.2% size reduction (327.42 MB to 6.01 MB, CPU std. dev. is across 3 trials, GPU across 10; CPU trial count a 54.5× compression ratio) at 95.13 ± 2.32% accuracy, averwas reduced from GPU’s for wall-clock practicality (see aged over four independent runs, exactly matching the FP32 Methodology baseline-model note). baseline’s 95.13% despite compounding all three techniques, and notably higher than PTQ-Dynamic’s own standalone −0.92-point cost, since KD’s accuracy headroom more than 4.5. Latency comparison offsets quantization’s small penalty in the full pipeline. The We compared our model after optimization – the Attentionfinal model size (6.01 MB) is set almost entirely by KD’s Based KD-distilled TinyViT student, FP32, prior to the architectural change to TinyViT followed by quantization, final INT8 quantization step – against other lightweight since H-BAC contributes FLOPs reduction rather than fileimage classification architectures trained and evaluated on size reduction. the same chilli 3-class village-split dataset: ResNet50 and As detailed in the Knowledge Distillation Results secEfficientFormer. The results are tabulated in Table 11. CPU tion, this same 6.01 MB size is also reachable by skipping benchmarks were run on an Apple M4 Pro (12-core CPU) H-BAC and KD entirely, training a TinyViT of the same MacBook Pro; GPU benchmarks were run on a rented target size directly on labels and quantizing it, reaching NVIDIA RTX 4060 (8 GB) instance; both use batched 95.13% accuracy on this dataset. Averaged over its four throughput timing (batch of 128, wall time divided by configurations, the full pipeline (95.39%) is modestly ahead 128). On CPU, our model is faster than ResNet50 (2.3×) of this simpler alternative, though any single pipeline run and EfficientFormer (1.4×). The same ordering holds on can land above or below it. The full pipeline’s additional GPU, where our model again leads ResNet50 (2.6×) and value beyond matching this simpler alternative’s accuracy EfficientFormer (1.6×). lies in what the direct-training alternative cannot provide on its own: an H-BAC-pruned intermediate checkpoint that 4.6. Comparative analysis and discussion keeps the full ViT-B/16 backbone at near-full accuracy while Table 12 situates this study’s baseline accuracy against cutting its FLOPs roughly in half (Table 2), useful whenever two prior plant-disease-detection works. This study evalua deployment target needs the larger backbone’s represenates exclusively on locally collected, village-partitioned field tational capacity but not its full compute cost, alongside images (Fig. 2). Table 10 Escalation trace for the tightest target in Table 9 (10 MB, 2% accuracy-drop budget), stages 3 through 9. Stages 1–2 (HBAC pruning, with or without quantization) already fail every budget in this grid on size alone and are omitted; stages 5–8 fit the size budget but fail the accuracy budget.

M.S Kumar et al.: Preprint submitted to Elsevier

Page 13 of 16

Lightweight ViT Compression for Plant Disease Detection Table 12 Comparison of Existing Research on Plant Disease Detection with the Current Study Research Work Current Study (Ours)

Dataset Acc. (%) Architecture Chilli 3-Class, Village-Split Field Im95.13 ViT-B/16 ages (OOD split) Deep Learning-Based Disease Detection PlantVillage 99.69 GoogLeNet Model in Plants Multiclass Plant Disease Detection via Dense PlantVillage 99.25 DenseNet201 CNNs

Table 13 Summary of the best configuration from each compression family, plus the fully integrated pipeline (last row). H-BAC reduces effective computation (FLOPs) without reducing stored model size, while quantization and KD reduce both model size and inference latency. CPU: Apple M4 Pro via Core ML; GPU: NVIDIA RTX 4060 via PyTorch. Configuration ViT-B/16 (FP32) H-BAC (50% pruning) PTQ-Dynamic Attention KD Integrated Pipeline

Size (MB) 327.42 327.42 84.42 21.15 6.01

Reduction (%) — 49.1∗ 74.2 93.5 98.2

Accuracy (%) 95.13 95.53 94.34 96.71 ± 1.03 95.13 ± 2.32‡

CPU Latency (ms) 7.18 4.27† 7.21 1.02 1.02§

GPU Latency (ms) 5.560 3.003† — 0.660 —

∗ Reduction in active parameters (43.70M of 85.80M); H-BAC’s mask-based pruning does not shrink the stored checkpoint, so its FLOPs reduction (49.2%) is the operative efficiency gain, not file size. † Measured on a structurally compacted copy (pruned dimensions physically removed), consistent with Table 5. ‡ Mean ± std over 4 independent configurations (three at pruning ratio 0.5, plus one at ratio 0.7); the single ratio-0.5 configuration traced stage-by-stage in Table 4 reaches 91.97%. § Identical to the Attention KD row above: this row is the same distilled student with PTQ-Dynamic quantization additionally applied, which on this Apple M4 Pro CPU shows no significant latency change either way (Table 6) – the size reduction (21.15 MB → 6.01 MB) is real, but it does not come with a latency improvement on this hardware. Both PTQ-Dynamic’s and the Integrated Pipeline’s GPU latency are omitted because the INT8 models were benchmarked through Core ML’s CPU compute unit, and no CUDA-executable INT8 build was produced for the RTX 4060 used for this study’s GPU measurements.

the fully compressed 6.01 MB endpoint for targets that need both. It is important to note that the ablation above evaluates each compression technique independently, while the pipeline described in Section 4 introduces interactions between techniques that cannot be predicted from individual results alone. For instance, the final PTQ-Dynamic quantization step is applied to a pruned-then-distilled student’s weight distribution rather than the original dense teacher’s, so its accuracy cost is not necessarily the same as quantizing the unpruned baseline directly (Table 6). Such interdependencies are why we followed a sequential composition strategy, wherein each compression stage is applied to the output of the preceding stage in a controlled, reproducible order, and why we report the integrated pipeline’s measured outcome directly (Table 4) rather than inferring it from the individual ablations.

5. Conclusion This paper introduces H-BAC, a compact and efficient pipeline for plant disease detection using the Vision Transformer, tailored for real-time application on resource-limited edge devices. H-BAC combines attention-driven knowledge distillation, pruning of structured parts, and low-precision quantization into a unified system that shrinks the model size and computational footprint while maintaining accurate diagnostic capability. Our experimental results show M.S Kumar et al.: Preprint submitted to Elsevier

that each compression method contributes differently: HBAC pruning at the 50% operating point reduces FLOPs by 49.2% (35.13 GFLOPs to 17.86 GFLOPs), reaching 95.53% accuracy after recovery fine-tuning (0.40 points above the 95.13% baseline); PTQ-Dynamic quantization achieves a 74.2% size reduction (327.42 MB to 84.42 MB) at 94.34% accuracy, a small expected cost of converting an alreadytrained FP32 checkpoint to INT8 with no further adaptation; and Attention-Based KD achieves 96.71±1.03% accuracy at a 93.5% size reduction (327.42 MB to 21.15 MB) and lower CPU latency than ResNet50 and EfficientFormer. Chaining all three techniques into a single integrated pipeline reduces the model size from 327.42 MB to 6.01 MB, corresponding to a 54.5 × compression ratio (98.16% reduction) and at 91.97% accuracy for the traced configuration. Also this result is broadly stable across three additional pruning-ratio configurations (95.13±2.32%, 𝑛 = 4) – exactly matching the 95.13% baseline despite compounding all three compression techniques, with KD’s accuracy headroom offsetting quantization’s small standalone cost. Several considerations qualify these results. First, on the Apple M4 Pro CPU used for all CPU latency measurements in this study, INT8 quantization shows no significant improvement over FP32 inference (Table 6), unlike the speedup typically reported on server-class x86 hardware; INT8’s benefit here is in model size, not latency, and practitioners targeting other ARM/edge platforms should remeasure latency on their actual target device rather than Page 14 of 16

Lightweight ViT Compression for Plant Disease Detection

assume quantization speedups transfer across hardware. This holds under Core ML’s own weight-only INT8 quantization, which dequantizes weights to floating point for the matrix multiply rather than executing native INT8 arithmetic, so per-operation cost is essentially unchanged; absent dedicated INT8 compute kernels, weight-only quantization buys size, not speed, on this platform. No GPUnative (ONNX/TensorRT) INT8 comparison was built for this study, so no claim is made either way about GPU quantization behavior. Second, the compressed model is not the fastest option in this comparison on either device class: it leads ResNet50 and EfficientFormer on both CPU (2.3× and 1.4×) and GPU (2.6× and 1.6×). Third, our KD ablation across 3 independent runs shows all four training recipes (no distillation, response-, feature-, and attention-based KD) clustered tightly within a 0.53-point band, with AttentionBased KD narrowly the highest (96.71±1.03%) and no clear instability in any single strategy (Table 7). This underscores that distillation’s accuracy benefit over direct training is modest at best on this task, and should not be assumed a priori for a new deployment target. Fourth, we directly tested whether the full H-BAC + KD + quantization pipeline is actually necessary: training a TinyViT student directly on labels (skipping pruning and distillation entirely) and quantizing it with PTQ-Dynamic yields a 6.01 MB model at 94.87% accuracy, 2.89 points above the matched pipeline configuration, but 0.26 points below the pipeline’s own 4configuration mean (95.13 ± 2.32%). The full pipeline’s runto-run variance is the larger effect here; averaged over multiple configurations it is modestly ahead of the simplest directtraining-plus-quantization alternative, though any single run can land on either side of it. Distilling from the H-BACpruned teacher rather than the full teacher cost 1.84 points here, a genuine trade-off of the sequential H-BAC → KD design against the FLOPs savings pruning provides. Fifth, we tested whether KD’s advantage over direct training grows as student capacity shrinks below TinyViT’s 5.52 M parameters, with a capacity sweep across four smaller, from-scratch student sizes (1.31 M to 3.16 M, Table 8) replicated across 3 runs per size; that table and Section 4 discuss the resulting capacity-dependent pattern in detail. Sixth, and most directly relevant to this paper’s stated deployment motivation, every latency figure reported here was measured on an Apple M4 Pro laptop CPU and a rented NVIDIA RTX 4060 GPU instance, neither of which is representative of the low-end Android smartphones that motivate this work (Section 1). This is a genuine scope boundary of the present study, not an oversight discovered late: these devices were chosen because they give controlled, reproducible measurements suitable for comparing compression techniques against each other, and the resulting size and FLOPs reductions (which are device-independent) are the primary claims this paper rests its deployment argument on. The absolute latency numbers, however, should not be read as predictions of on-device farmer-facing performance. Validating the final deployed model on actual ARM smartphone-class hardware (e.g. Snapdragon or MediaTek SoCs) via TensorFlow Lite, M.S Kumar et al.: Preprint submitted to Elsevier

ONNX Runtime Mobile, or CoreML conversion is necessary before any latency-based deployment claim can be made with confidence, and is left to future work.

Acknowledgment The authors would like to acknowledge Mr. Kannan, Project Associate, ANRF Grant No. SUR/2022/004268, for his valuable assistance in organizing and making the dataset publicly available through Figshare (https://doi.org/10.6084/m9.figshare.32820110), which was used in this study. The authors are grateful to the ANRF, New Delhi, India, for funding the research on “Light Weight Deep Learning Based Mobile Application for the Early Detection, Identification, and Spatiotemporal Monitoring of Plant Diseases” (Grant No. SUR/2022/004268).

References [1] Ahmad, A., Saraswat, D., El Gamal, A., 2023. A survey on using deep learning techniques for plant disease diagnosis and recommendations for development of appropriate tools. Smart Agricultural Technology 3, 100083. [2] Barman, U., Sarma, P., Rahman, M., Deka, V., Lahkar, S., Sharma, V., Saikia, M.J., 2024. Vit-smartagri: vision transformer and smartphone-based plant disease detection for smart agriculture. Agronomy 14, 327. [3] Cheng, Y., Wang, D., Zhou, P., Zhang, T., 2018. Model compression and acceleration for deep neural networks: The principles, progress, and challenges. IEEE Signal Processing Magazine 35, 126–136. [4] Cortes, C., Vapnik, V., 1995. Support-vector networks. Machine Learning 20, 273–297. doi:10.1007/BF00994018. [5] Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009a. Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee. pp. 248– 255. [6] Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L., 2009b. Imagenet: A large-scale hierarchical image database, in: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248– 255. doi:10.1109/CVPR.2009.5206848. [7] Dosovitskiy, A., 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 . [8] Ferentinos, K.P., 2018. Deep learning models for plant disease detection and diagnosis. Computers and electronics in agriculture 145, 311–318. [9] Gole, P., Bedi, P., Marwaha, S., Haque, M.A., Deb, C.K., 2023. Trincnet: a lightweight vision transformer network for identification of plant diseases. Frontiers in Plant Science 14, 1221557. [10] Hamim, S.A., Jony, A.I., 2024. Enhanced deep learning model architecture for plant disease detection in chilli plants. Journal of Edge Computing 3, 136–146. [11] Han, S., Mao, H., Dally, W.J., 2015. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv preprint arXiv:1510.00149 . [12] Haque, M.A., Deb, C.K., Gole, P., Karmakar, S., Dheeraj, A., Shah, M.U.D., Dutta, S., Kumar, M.P., Marwaha, S., 2025. An enhanced vision transformer network for efficient and accurate crop disease detection. Expert Systems with Applications , 127743. [13] Haralick, R.M., Shanmugam, K., Dinstein, I., 1973. Textural features for image classification. IEEE Transactions on Systems, Man, and Cybernetics SMC-3, 610–621. doi:10.1109/TSMC.1973.4309314. [14] Hemalatha, S., Jayachandran, J.J.B., 2024. A multitask learningbased vision transformer for plant disease localization and classification. International Journal of Computational Intelligence Systems 17, 188.

Page 15 of 16

Lightweight ViT Compression for Plant Disease Detection [15] Hinton, G., Vinyals, O., Dean, J., 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 . [16] Huffman, D., 2024. A method for the construction of minimumredundancy codes. 1952. Proceedings of the IRE 40, 1098. [17] Hutchinson, M., 1989. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines. Communication in Statistics- Simulation and Computation 18, 1059–1076. doi:10.1080/ 03610919008812866. [18] Intel Corporation, 2023. Intel® distribution of OpenVINO™ toolkit. https://github.com/openvinotoolkit/openvino. Open-source toolkit for optimizing and deploying deep learning models. [19] Joshi, H., 2024. Edge-ai for agriculture: lightweight vision models for disease detection in resource-limited settings. arXiv preprint arXiv:2412.18635 . [20] Kannan, M., Seethapathy, P., 2026. Chilli leaf curl virus dataset final ver. URL: https://doi.org/10.6084/m9.figshare.32820110, doi:10. 6084/m9.figshare.32820110. [21] Kim, J., Chang, S., Kwak, N., 2021. Pqk: model compression via pruning, quantization, and knowledge distillation. arXiv preprint arXiv:2106.14681 . [22] Kirola, M., Joshi, K., Chaudhary, S., Singh, N., Anandaram, H., Gupta, A., 2022. Plants diseases prediction framework: a imagebased system using deep learning, in: 2022 IEEE World Conference on Applied Intelligence and Computing (AIC), IEEE. pp. 307–313. [23] Lecun, Y., Denker, J., Solla, S., 1989. Optimal brain damage, in: Advances in Neural Information Processing Systems, pp. 598–605. [24] Li, G., Wang, Y., Zhao, Q., Yuan, P., Chang, B., 2023a. Pmvt: a lightweight vision transformer for plant disease identification on mobile devices. Frontiers in Plant Science 14, 1256773. [25] Li, J., Zhang, L.L., Xu, J., Wang, Y., Yan, S., Xia, Y., Yang, Y., Cao, T., Sun, H., Deng, W., et al., 2023b. Constraint-aware and ranking-distilled token pruning for efficient transformer inference, in: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 1280–1290. [26] Li, X., Li, X., Zhang, S., Zhang, G., Zhang, M., Shang, H., 2023c. Slvit: Shuffle-convolution-based lightweight vision transformer for effective diagnosis of sugarcane leaf diseases. Journal of King Saud University-Computer and Information Sciences 35, 101401. [27] Li, Z., Li, H., Meng, L., 2023d. Model compression for deep neural networks: A survey. Computers 12, 60. [28] Malihi, L., Heidemann, G., 2023. Efficient and controllable model compression through sequential knowledge distillation and pruning. Big Data and Cognitive Computing 7, 154. [29] Malihi, L., Heidemann, G., 2024. Matching the ideal pruning method with knowledge distillation for optimal compression. Applied System Innovation 7, 56. [30] Megalingam, R.K., Menon, G.G., Binoj, S., Sai, D.A., Kunnambath, A.R., Manoharan, S.K., 2024. Cowpea leaf disease identification using deep learning. Smart Agricultural Technology 9, 100662. [31] Menon, V., Ashwin, V., Deepa, R.K., 2021. Plant disease detection using cnn and transfer learning, in: 2021 international conference on communication, control and information sciences (ICCISc), IEEE. pp. 1–6. [32] Mohanty, S.P., Hughes, D.P., Salathé, M., 2016. Using deep learning for image-based plant disease detection. Frontiers in plant science 7, 1419. [33] Molchanov, P., Mallya, A., Tyree, S., Frosio, I., Kautz, J., 2019. Importance estimation for neural network pruning, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11264–11272. [34] Naik, B.N., Malmathanraj, R., Palanisamy, P., 2022. Detection and classification of chilli leaf disease using a squeeze-and-excitationbased cnn model. Ecological Informatics 69, 101663. [35] Ojala, T., Pietikainen, M., Maenpaa, T., 2002. Multiresolution grayscale and rotation invariant texture classification with local binary patterns. IEEE Transactions on Pattern Analysis and Machine Intelligence 24, 971–987. doi:10.1109/TPAMI.2002.1017623.

M.S Kumar et al.: Preprint submitted to Elsevier

[36] Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S., 2019. Pytorch: An imperative style, high-performance deep learning library, in: Advances in Neural Information Processing Systems, pp. 8026–8037. [37] Paula, E., Soni, J., Upadhyay, H., Lagos, L., 2025. Comparative analysis of model compression techniques for achieving carbon efficient ai. Scientific Reports 15, 23461. [38] Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., Dosovitskiy, A., 2021. Do vision transformers see like convolutional neural networks?, in: Advances in Neural Information Processing Systems, pp. 12116– 12128. [39] Sajitha, N., Nema, S., Bhavya, K., Seethapathy, P., Pant, K., 2022. The plant disease detection using cnn and deep learning techniques merged with the concepts of machine learning, in: 2022 2nd International conference on advance computing and innovative technologies in engineering (ICACITE), IEEE. pp. 1547–1551. [40] Seethapathy, P., Kannan, M., Gopakumar, G., Manivasagam, V., Manalil, S., 2025. Comparative analysis of pretrained cnn architectures for mobile-based early detection of brinjal little leaf disease, in: AI and Sustainable Transformations. CRC Press, pp. 108–114. [41] Shoaib, M., Shah, B., Ei-Sappagh, S., Ali, A., Ullah, A., Alenezi, F., Gechev, T., Hussain, T., Ali, F., 2023. An advanced deep learning models-based plant disease detection: A review of recent research. Frontiers in Plant Science 14, 1158933. [42] Singh, A.K., Rao, A., Chattopadhyay, P., Maurya, R., Singh, L., 2024. Effective plant disease diagnosis using vision transformer trained with leafy-generative adversarial network-generated images. Expert Systems with Applications 254, 124387. [43] Thai, H.T., Tran-Van, N.Y., Le, K.H., 2021. Artificial cognition for early leaf disease detection using vision transformers, in: 2021 International conference on advanced technologies for communications (ATC), IEEE. pp. 33–38. [44] Vinay, K., Surya, V., Thushar, S., Singh, T., Sahay, A., 2025. A deep learning framework for early detection and diagnosis of plant diseases. Procedia Computer Science 258, 1435–1445. [45] Wang, X., Liu, J., Chen, Q., 2025a. An advanced deep learning method for pepper diseases and pests detection. Plant Methods 21, 70. [46] Wang, Z., Shi, L., Mei, Z., Zhao, X., Wang, Z., Li, J., 2025b. Iterative knowledge distillation and pruning for model compression in unsupervised domain adaptation. Pattern Recognition 164, 111512. [47] WinKawaks, 2022. vit-tiny-patch16-224. https://huggingface.co/ WinKawaks/vit-tiny-patch16-224. [48] Yang, H., Yin, H., Shen, M., Molchanov, P., Li, H., Kautz, J., 2023. Global vision transformer pruning with hessian-aware saliency, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 18547–18557. [49] Yang, Z., Li, Z., Zeng, A., Li, Z., Yuan, C., Li, Y., 2022. Vitkd: Practical guidelines for vit feature knowledge distillation. arXiv preprint arXiv:2209.02432 .

Page 16 of 16

Record · ID 660796 · SHA-256 f642802f567f9d64
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.