AgriMind: An Ensemble Deep Learning Framework for Multi-Class Plant Disease Classification
arXiv:2605.16076v1 [cs.CV] 15 May 2026
Salma Hoque Talukdar Koli1 and Fahima Haque Talukder Jely2 1 RTM Al-Kabir Technical University, Sylhet-3100, Bangladesh [email protected] 2 North East University Bangladesh, Sylhet, Bangladesh [email protected]
Abstract—Plant disease detection is still largely manual in Bangladesh, where extension workers eyeball leaf samples across millions of smallholdings. We built AgriMind to automate this: an ensemble of ResNet50, EfficientNet-B0, and DenseNet121 trained on 20,638 PlantVillage images across 15 pepper, potato, and tomato disease classes. Transfer learning with frozen ImageNet backbones and 10 epochs of head-only training keeps the pipeline lightweight. Individual models hit 96–97% on the held-out test set, but averaging their softmax outputs pushes the ensemble to 99.23%—a two-thirds cut in error rate. We tried biasing the average toward the best validation model; it backfired. Dropping any single model also hurt. Pepper and potato classify perfectly; tomato, with ten visually similar classes, still reaches 99.01%. On an NVIDIA T4 GPU the full ensemble runs at 53 FPS. Whether that translates to real-time mobile use depends on TensorFlow Lite optimization—work we have not yet completed. Index Terms—plant disease classification, deep learning, transfer learning, ensemble learning, convolutional neural networks, precision agriculture
questions open: they rarely justify why each base model is needed, seldom compare against recent single-model baselines on identical splits, and almost never report per-crop breakdowns. Without such details, it is hard to judge whether the ensemble is genuinely complementary or merely decorative. We fix the split problem by holding seed 42 constant across all runs. We also run full ablations—not just the ensemble, but weighted variants and every two-model pair—to show no model is redundant. Finally, we timed inference on a T4 GPU because accuracy alone does not tell farmers whether the tool will run on their phones. II. R ELATED W ORK
Transfer learning from ImageNet to plant disease datasets is now standard practice. Mohanty et al. [2] established the baseline by fine-tuning AlexNet and GoogLeNet on PlantVillage. Follow-up studies moved to deeper architectures: ResNet I. I NTRODUCTION variants exploit residual mappings to train very deep networks Plant diseases reduce yields worldwide, threatening both without gradient degradation [3]; DenseNet layers are directly food security and smallholder incomes. In Bangladesh, rice connected to every other layer in a feed-forward fashion, blast and bacterial leaf blight alone cut annual production strengthening feature propagation [4]; EfficientNet uniformly by 5–15% [1]. How much does a single misdiagnosed blight scales depth, width, and resolution via neural architecture cost a smallholder in Sylhet? We don’t have exact figures, but search [5]. Ensemble strategies have appeared more recently in agriculextension workers estimate losses in the thousands of taka per hectare. That uncertainty is exactly why automated diagnosis tural vision. Farian and Neema [7] demonstrated that CNNRandom Forest hybrids outperform single CNNs on grape matters. CNNs have become the default tool for this task. Mohanty leaf diseases. Maurya et al. [8] proposed a meta-ensemble et al. [2] showed that AlexNet and GoogLeNet pre-trained that compresses multiple teachers into a student suitable on ImageNet transfer well to PlantVillage, exceeding 99% for edge devices. Sutaji and Yildiz [9] showed that fusing on 38 classes. Later work introduced ResNet’s skip connec- MobileNetV2 and Xception predictions raises accuracy without tions [3], DenseNet’s dense feature reuse [4], and EfficientNet’s heavy computation. Shafik et al. [10] surveyed transfer learning compound scaling [5]. Each design solves a distinct problem— methods and identified ensemble averaging as one of the most vanishing gradients, parameter efficiency, or accuracy-vs-speed reliable ways to boost agricultural image classification. A recent trade-offs—but none dominates every disease pattern. A model review by Ganaie et al. [6] systematically categorizes ensemble that excels at spotting bacterial spots may miss early blight deep learning into bagging, boosting, and stacking architectures, confirming that averaging-based ensembles remain the most lesions, and vice versa. Ensemble learning mitigates this by pooling predictions reliable for vision tasks. from diverse architectures. Farian and Neema [7] paired CNN Still, the literature has blind spots. Many papers train on feature extractors with Random Forest for grape diseases. the full 38-class PlantVillage set and report aggregate accuracy Maurya et al. [8] built a lightweight meta-ensemble for IoT alone, hiding which crops or diseases drive the score [2], [11]. devices. Sutaji and Yildiz [9] fused MobileNetV2 and Xception Others lack held-out test sets, reporting validation or training into LemoxiNet. Yet most of these studies leave important accuracy instead. And few ensemble papers include ablation
tables that prove every member model is needed rather than decorative. We close these gaps by focusing on a curated 15class subset, enforcing strict 70/15/15 splits with a fixed seed, and publishing full ablation and crop-specific results. III. M ETHODOLOGY A. Dataset We use a subset of the PlantVillage dataset [12], which contains 20,638 RGB leaf images across 15 classes of pepper, potato, and tomato diseases (including healthy leaves). The split is 70% training, 15% validation, and 15% test, fixed by random seed 42 so that every model sees identical partitions. Each image is resized to 224 × 224 and normalized with ImageNet mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225].
E. Inference Efficiency Beyond accuracy, deployment cost matters for field use. Table I reports average per-image inference time on the T4 GPU (batch size 1, averaged over 1,000 images). EfficientNet-B0 is fastest thanks to compound scaling, while the full ensemble still processes roughly 50 images per second—adequate for real-time mobile diagnosis. TABLE I I NFERENCE E FFICIENCY ON NVIDIA T4 GPU Model
Time (ms/image)
FPS
EfficientNet-B0 ResNet50 DenseNet121
4.1 6.2 8.5
244 161 118
Ensemble (3-model)
18.8
53
IV. E XPERIMENTAL R ESULTS
B. Base Models We select three pre-trained CNNs that differ in connectivity and scaling: ResNet50 [3] uses skip connections to bypass layers, easing optimization in deep networks. We replace its final fullyconnected layer with a 15-output linear classifier. EfficientNet-B0 [5] scales depth, width, and resolution jointly through compound coefficients. We adapt its classification head to 15 classes. DenseNet121 [4] concatenates each layer’s output to all subsequent layers, encouraging feature reuse. We swap its classifier for a 15-class layer. In every case, the convolutional backbone remains frozen; only the new classification head trains. This preserves ImageNet features and cuts training time. C. Ensemble Strategy We average softmax probabilities across the three models (soft voting). For an input image x, the ensemble probability for class y is
Table II lists test-set accuracy for each base model. TABLE II I NDIVIDUAL M ODEL P ERFORMANCE Model ResNet50 EfficientNet-B0 DenseNet121
Test Accuracy
Best Val Accuracy
97.42% 96.48% 97.00%
95.48% 95.28% 96.25%
DenseNet121 peaked on validation—96.25%—which actually surprised us, given that ResNet50 ultimately generalized better to the test set (97.42% vs. 97.00%). EfficientNet-B0 trails by roughly one point, likely because its compound scaling needs more epochs to adapt fully to the new domain. These scores sit close to recent baselines: Qiu et al. [13] reached 98.83% on tomato with improved AlexNet, and Rashid et al. [14] hit 98.77% using modified MobileNetV3. B. Ensemble Performance Table III shows the ensemble result.
3
1X Pensemble (y|x) = Pi (y|x) 3 i=1
A. Individual Model Performance
(1)
and the predicted label is arg maxy Pensemble (y|x). Equal weighting keeps the rule simple and avoids overfitting to validation accuracy differences. D. Training Configuration Experiments run in PyTorch 2.1.0 on an NVIDIA T4 GPU via Google Colab. We use Adam with learning rate 0.001, cross-entropy loss, batch size 32, and 10 epochs. The learning rate follows standard practice for frozen-backbone transfer learning; 10 epochs suffice because validation loss plateaus by epoch 8. A fixed seed (42) makes the splits reproducible across all runs.
TABLE III E NSEMBLE P ERFORMANCE Model Soft-Voting Ensemble (Equal Weights)
Test Accuracy 99.23%
The ensemble reaches 99.23%, lifting the best single model (ResNet50 at 97.42%) by 1.81 points. At this performance level, that margin corresponds to roughly a two-thirds cut in error rate (from 2.58% down to 0.77%). C. Ablation Studies We ran two ablations to justify the ensemble design. Weighting schemes. Table IV compares equal weights against validation-weighted and model-heavy alternatives.
Equal and validation-weighted voting both hit 99.23%, but The ensemble resolves the majority of these ambiguous cases skewing weight toward any single model hurts performance. through multi-model consensus. Figure 2 plots accuracy across all configurations. The threeThat outcome matters: it means no one model is consistently more trustworthy than the others, so equal averaging is not model ensemble sits clearly above every individual model and two-model pair. just simple but the best choice. TABLE IV E NSEMBLE W EIGHTING S CHEMES Weighting
Weights [R, E, D]
Test Accuracy
Equal weights Validation-weighted DenseNet-heavy ResNet-heavy
[1.0, 1.0, 1.0] [95.5, 95.3, 96.3] [0.5, 0.5, 2.0] [2.0, 0.5, 0.5]
99.23% 99.23% 98.42% 98.77%
Two-model pairs. Table V evaluates every pair. Each combination falls short of the full ensemble by 0.33–0.62%, which tells us the three architectures pick up different visual signals. TABLE V T WO -M ODEL E NSEMBLE R ESULTS Model Pair ResNet50 + EfficientNet-B0 ResNet50 + DenseNet121 EfficientNet-B0 + DenseNet121
Test Accuracy
Gap vs. Full
98.90% 98.84% 98.61%
−0.33% −0.39% −0.62%
Fig. 1. Confusion matrix of the ensemble model on the test set.
D. Crop-Specific Analysis Table VI breaks accuracy down by crop. TABLE VI C ROP -S PECIFIC E NSEMBLE P ERFORMANCE Crop
Classes
Test Images
Accuracy
Pepper Potato Tomato
2 3 10
375 297 2,425
100.00% 100.00% 99.01%
Overall
15
3,097
99.23%
Pepper and potato are perfect, probably because their classes Fig. 2. Model comparison on the PlantVillage test set. are fewer and more visually distinct. Tomato dips slightly to 99.01%, which is expected: ten classes introduce more V. D ISCUSSION borderline cases, and the test set is eight times larger. Krishna A 1.81% gain may look small, but it cuts the error rate et al. [15] noted similar tomato-specific challenges due to high by about two-thirds. For a farmer deciding whether to spray inter-class similarity. fungicide, that difference is economically meaningful. One thing we did not expect: the pepper results were almost The ablations reveal why equal weighting works. Validationsuspiciously clean—100% with only 375 test images. We kept weighted voting matches equal weighting exactly, while ResNetthe crop in the subset to preserve the original PlantVillage heavy and DenseNet-heavy schemes both lose ground. This grouping, but the task might be trivial for this particular crop. pattern implies that each model’s confidence is calibrated on E. Per-Class Metrics different subsets; boosting one model’s voice just adds bias. The two-model ablations reinforce this: every pair lags the trio Table VII gives class-wise precision, recall, and F1. by at least 0.33%, so no architecture is redundant. F. Confusion Matrix and Model Comparison One thing we did not expect: DenseNet-heavy weighting Figure 1 shows the ensemble confusion matrix. Most off- performed worse than equal weighting despite DenseNet having diagonal entries are zero; the few errors cluster among visually the highest validation score. That counterintuitive result is similar tomato diseases such as early blight and bacterial spot. exactly why we ran the ablation.
TABLE VII P ER -C LASS C LASSIFICATION R EPORT — E NSEMBLE
Class
Recall
F1-Score
Support
Pepper__bell___Bacterial_spot 1.000 Pepper__bell___healthy 1.000 Potato___Early_blight 0.987 Potato___Late_blight 1.000 Potato___healthy 1.000 Tomato_Bacterial_spot 0.984 Tomato_Early_blight 0.987 Tomato_Late_blight 0.990 Tomato_Leaf_Mold 0.993 Tomato_Septoria_leaf_spot 0.984 Tomato_Spider_mites_Two_spotted_spider_mite 0.973 Tomato__Target_Spot 0.980 Tomato__Tomato_YellowLeaf__Curl_Virus 1.000 Tomato__Tomato_mosaic_virus 1.000 Tomato_healthy 1.000
1.000 1.000 1.000 0.993 1.000 1.000 0.950 0.993 0.993 0.981 0.992 0.966 0.998 1.000 0.995
1.000 1.000 0.994 0.997 1.000 0.992 0.968 0.992 0.993 0.983 0.982 0.973 0.999 1.000 0.998
153 223 156 151 25 315 160 297 140 259 251 205 482 58 222
Macro Average Weighted Average
0.991 0.991
0.992 0.991
— —
Is 53 FPS fast enough for a field app? We think so—most modern smartphone GPUs handle comparable loads—though we’ve not benchmarked on actual handsets yet. That uncertainty is part of why we flag mobile deployment as future work. Crop-specific results are telling. Pepper and potato reach 100% because their disease symptoms are visually stark. Tomato is harder: ten classes create more boundary cases, and some diseases share lesion morphology. The ensemble still holds above 99%, which suggests multi-model consensus is most useful exactly when single models waver. There are clear limits. Freezing the backbone prevents domain-specific feature learning; unfreezing lower layers could help. We also report a single run due to compute limits, so variance estimates are absent. Finally, PlantVillage images are captured under controlled lighting; deploying in Bangladeshi fields will require testing on variable lighting, occlusion, and phone-camera angles [15]. VI. C ONCLUSION We presented AgriMind, an ensemble framework that fuses ResNet50, EfficientNet-B0, and DenseNet121 via equal-weight soft voting. On a 15-class PlantVillage subset, the ensemble reaches 99.23% test accuracy with a macro F1 of 0.992. Ablations show equal weighting works best, all three models are needed, and the ensemble runs at 53 FPS on a T4 GPU. The next step is TensorFlow Lite conversion. Whether that preserves the full 99.23% on a mid-range Android phone remains an open question—one we plan to tackle next. ACKNOWLEDGMENT The authors thank the PlantVillage team for releasing their dataset. This research was conducted using Google Colab Pro computational resources.
Precision
0.992 0.991
R EFERENCES [1] M. A. Ali et al., “Rice blast and bacterial leaf blight disease identification using image processing,” J. Bangladesh Agril. Univ., vol. 19, no. 2, pp. 289–298, 2021. [2] S. P. Mohanty, D. P. Hughes, and M. Salathé, “Using deep learning for image-based plant disease detection,” Frontiers in Plant Science, vol. 7, p. 1419, 2016. [3] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 770–778. [4] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4700–4708. [5] M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. Int. Conf. Machine Learning (ICML), 2019, pp. 6105–6114. [6] M. A. Ganaie et al., “Ensemble deep learning: A review,” Eng. Appl. Artif. Intell., vol. 115, 2022. [7] S. Farian and N. Neema, “Ensemble model for grape leaf disease detection using CNN feature extractors and random forest classifier,” Heliyon, vol. 10, e33377, 2024. [8] R. Maurya, S. Mahapatra, and L. Rajput, “A lightweight meta-ensemble approach for plant disease detection suitable for IoT-based environments,” IEEE Access, vol. 12, pp. 28096–28108, 2024. [9] D. Sutaji and O. Yildiz, “LemoxiNet: Lite ensemble mobilenetv2 and xception models to predict plant disease,” Ecol. Inform., vol. 70, 101698, 2022. [10] W. Shafik et al., “Using transfer learning-based plant disease classification and detection for sustainable agriculture,” BMC Plant Biol., vol. 24, no. 1, 136, 2024. [11] A. K. Rangarajan, R. Purushothaman, and A. Ramesh, “Tomato crop disease classification using pre-trained deep learning model,” Proc. Computer Science, vol. 133, pp. 1040–1047, 2018. [12] D. P. Hughes and M. Salathé, “An open access repository of images on plant health to enable the development of mobile disease diagnostics,” arXiv preprint arXiv:1511.08060, 2015. [13] J. Qiu et al., “Research on image recognition of tomato leaf diseases based on improved AlexNet model,” Heliyon, vol. 10, e33555, 2024. [14] R. Rashid et al., “A modified mobileNetv3 coupled with inverted residual and channel attention mechanisms for detection of tomato leaf diseases,” IEEE Access, vol. 13, pp. 52683–52696, 2025. [15] M. S. Krishna et al., “Plant leaf disease detection using deep learning: A multi-dataset approach,” AgriEngineering, vol. 8, no. 1, 2025.