Configurable Multi-Stage Vision Pipeline for Crop Disease and Pest Diagnosis Naga Ganesh*
Chandrashekar M S* Lakshmi Pedapudi† Vineet Singh
Aakash Singh
Digital Green
* Equal contribution.
† Corresponding author.
arXiv:2609.21651v1 [cs.CV] 18 Sep 2026
[email protected], [email protected], [email protected], [email protected], [email protected]
Abstract FarmerChat is Digital Green’s farm advisory service for smallholder farmers. When something looks wrong with a crop, farmers usually send a photograph as their entire query: no symptoms described, no crop named, and often no text at all. The system must determine whether the image is usable, identify the crop, and diagnose the disease or pest from photographs taken on cheap phones in real field conditions. The current production system offers little control over these decisions: quality thresholds cannot be adjusted, new crops and problems cannot be added, and there is no configurable confidence threshold or fallback when the diagnosis is uncertain. We study about 1.16 million photographs sent to FarmerChat from Ethiopia, India, Kenya, and Nigeria. The production quality gate rejected 46.8% of the images it judged, more than a quarter of images reaching diagnosis received no crop label, and 35.8% of labelled problems classified as “disease” were pests. We therefore split diagnosis into three independently evaluated stages: image quality (M0), crop detection (M1), and disease or pest detection (M2). We evaluate two routes: Route A uses a single fine-tuned vision-language model (Qwen3-VL-4B), while Route B uses smaller specialist models (DaViT and YOLO26). Each stage can be replaced independently and its thresholds can be configured. We replace the production GPT-4o quality gate with a MobileNetV3 gate that reaches 86.9% F1 with 12 ms latency. On a common test set, hierarchical DaViT-Base identifies crops with 95.41% accuracy, compared with 91.46% for the production baseline. The same backbone also leads on disease and pest identification and does not decline to answer, while the language models leave a substantial share of rows without a diagnosis. The specialist route also has lower hosting cost at the measured query volume. The fine-tuned VLM provides two capabilities that the specialist route does not: it handles all three stages in a single call and can request a more informative photograph when the available image is insufficient for diagnosis.
1
Introduction
Table 1: Research questions.
Farmers using FarmerChat [1] mostly report crop health problems by sending a photograph. These photographs look nothing like the tidy images in research datasets. The light is poor, the camera moves, and the subject changes from one photo to the next: a single leaf fills one frame, and the next holds a whole field, a hand, or a farm animal. Giving a useful answer therefore means making five decisions, not one: 1. Is the image usable? 2. What crop is it? 3. What disease or pest, if any, is present? 4. How confident is that call? 5. What structured output does the downstream advisory system need? The system in production lets us change almost nothing. It offers none of the following: • adjustable thresholds for photograph rejection; • addition of new crops, diseases or pests; • control over how the model behaves, or a confidence cut-off; • a choice of what happens when the answer is weak: reject it, try again, or send it to a person. A system we cannot adjust throws away photographs
ID
Question
RQ1
Can a configurable local quality stage match the paid gate’s decisions at lower cost and latency? Does separating crop detection from disease and pest detection improve diagnosis? How does a fine-tuned VLM compare with orchestrated specialist models? What do per-country submission patterns imply for crop and problem coverage? What is the accuracy, coverage, cost and latency trade-off?
RQ2 RQ3 RQ4 RQ5
a tunable one would keep, and it cannot be pointed at the crops and problems that matter in one country. Section 2 measures each of these on real farmer queries. Table 1 lists the five questions this paper addresses. This paper contributes four things: • A failure analysis of about 1.16 million farmer photographs from four countries. • A MobileNetV3 quality gate that matches the paid gate’s decisions at low memory and latency. • A case for splitting pest detection from disease detection, and a four-head model that does it, released as a trained checkpoint.1 • A benchmark of seven systems, specialist computer1
Table 2: Existing production path: stage facts from five production traces recorded on one day in August 2026. Property
Value
Quality gate Diagnosis Caller controls
GPT-4o, one call per image, ten checks; $0.0106 to $0.0132 per image across the five traces Plantix, one call per image; crop and disease returned in one answer; 81-crop vocabulary None. The gate returns pass or fail per check and Plantix returns a likelihood word (unlikely, likely, very likely); no threshold or numeric score can be set or read Quality gate 6.2 to 8.2 s, Plantix diagnosis 2.8 to 4.4 s, advisory text 1.5 to 2.9 s, total 10.5 to 13.7 s per image across the five traces; the quality gate is the slowest stage in every trace
Server-side latency
vision models and vision-language models, on one test set. The four-head benchmark’s labels and label space are released with it.2
2
2.1.1
Table 3 follows the photographs through the pipeline. The gate judged seven in ten of them and rejected nearly half of those it judged. Most rejected photographs were never sent to Plantix, though some were sent anyway. Table 4 gives the fail rate of each check. Figure 2 shows what a rejected photograph actually looks like: six photographs across four of the ten checks. • The highest-volume check, Dominant Content, is semantic (“what is in frame”), as is Plant Detected. Blur and contrast statistics cannot assess either check, which is why §3.2 uses a learned model rather than calibrated thresholds alone. • Separating usable images that were wrongly rejected from the rejection rate as a whole needs its own measurement, listed in §6.
Background, Production Failures and Limitations
Figure 1 shows the production pipeline and Table 2 summarises its stages. A farmer’s photograph first goes through a GPT-4o quality gate, which performs ten checks in a single call. Images that pass are then sent to Plantix, which returns the crop and disease or pest in one response. Neither stage exposes settings that the caller can change. • Neither stage has a documented or adjustable threshold. • The paid quality-gate call is the slowest stage in every measured trace (Table 2). Plant disease recognition from images has been studied extensively [2], using datasets such as PlantVillage [3] and field datasets such as PlantDoc [4]. These datasets also show the gap between controlled images and photographs taken in real field conditions. Our setting differs in three ways. First, our images are photographs submitted by farmers to a live service and were not curated for this study. Second, our labels come from a panel of models rather than experts labelling images at scale. Third, we study how the diagnosis pipeline can be configured and improved, rather than the accuracy of a single classifier. We use established architectures for the proposed pipeline: MobileNetV3 [5] for the quality gate, DaViT [6] and YOLO [7] for the specialist route, and Qwen3VL [8] and Gemma 3 [9] as fine-tuning candidates. The FarmerChat platform is described in Singh et al. [1]. Our evaluation approach also builds on our earlier work on conversational-AI evaluation [10] and agricultural ASR benchmarking [11]. Unless otherwise stated, the measurements in this paper cover all photographs sent to FarmerChat up to August 2026. Each table specifies the rows included in its analysis.
2.1
Image Quality Failures
2.1.2
Crop Coverage
Of the photographs sent for diagnosis, more than a quarter came back with no crop named (Table 3). Among those that did get a crop, Table 5 lists the ten sent most often. Figure 3 shows how much of the farmer query volume the top crops cover. 2.1.3
Disease and Pest Coverage
We sorted every label in the problem vocabulary into a type using keyword rules (Table 6). More than a third of the labelled problems are pests, not pathogens. Rare classes are also poorly represented: 116 of the 250 disease classes in the evaluation label set have fewer than 10 test images. And the problems named most often mix insects and pathogens together without saying which is which (Table 8).
2.2
Limitations
The measurements above point at six limitations of the system in production: • Quality thresholds are neither visible nor adjustable. • Nearly half of the images the gate judged were rejected, and more than a quarter of all submissions never reached diagnosis (Table 3). • An 81-crop vocabulary: images showing crops outside this list receive no diagnosis. • No per-stage failure visibility: a wrong answer cannot be traced to quality, crop or disease. • No country-specific tuning. • No structured path for a human correction to reach model training.
Production Failures
The three subsections below measure what the existing pipeline does with the photographs it receives: which images it throws away, which crops it can name, and which problems it can name.
1 Checkpoint: https://huggingface.co/DigiGreen/crop-disease-pest-detection-dg
2 Benchmark labels: https://huggingface.co/datasets/DigiGreen/Crop-Disease-Image-Eval-Synthetic
2
Quality gate
Farmer image
Crop + disease
GPT-4o, 10 checks
Advisory text
Plantix, one call, one answer
Figure 1: Existing production path: two paid calls, no intermediate stage exposed to configuration. Table 3: Submission funnel over all photographs submitted to FarmerChat up to August 2026. Each share names its denominator. Stage
Images
Submitted No quality decision recorded Judged by the gate Rejected Withheld from the diagnosis call Sent for diagnosis anyway Sent for diagnosis Crop named No crop named
Plant Detected A model animal, not a plant
1,163,658 341,871 821,787 384,271 328,520 55,751 833,681 608,742 224,939
Plant Detected Artificial flowers indoors
Share 29.4% of submitted 70.6% of submitted 46.8% of judged 85.5% of rejected; 28.2% of submitted 14.5% of rejected 71.6% of submitted 73.0% of sent 27.0% of sent
Dominant Content A wall fills the frame
Dominant Content
Lighting
Focus
A whole field, no subject
Half the frame in shadow
Too soft to read a leaf
Figure 2: Photographs the quality gate rejected, across four of the ten checks. Held-out quality-gate test split of Table 15. Each image failed exactly one of the ten checks, so the label names that check.
2.2.1
therefore fits some countries better than others. • Two of the three most-named problems in India and Nigeria are missing nutrients, not diseases, and Fall Armyworm, an insect, is in the top three in three of the four countries. All of this is what one “disease” head is currently asked to cover.
Country-Level Funnel
Table 7 and Figure 4 follow the photographs through the pipeline for the four countries that send almost all of them. Table 8 shows what each country sends. • India’s rejection rate is more than double Kenya’s (Table 7). One flat quality threshold does not fit both. • India’s healthy share among crop-named images is under half of Ethiopia’s (Table 7). What a “reasonable” disease rate looks like differs by country. • India’s five most-submitted crops cover under half of its crop-named images, against more than fourfifths in Ethiopia (Table 8). A single crop vocabulary
2.2.2
Human Review Labelling
Human review is critical for building models that are reliable and grounded in real-world conditions. However, producing ground-truth labels from scratch is timeconsuming and limits how much data can be reviewed. A more scalable approach is to have reviewers assess 3
top 20 100%
89.2% of crop-named images 75%
50%
wheat + maize, 41.8%
25%
1
10
20
30
40
57
Crops, ranked by submission volume
Figure 3: Cumulative coverage of crop-named images by the top-N crops, ranked by volume. Accepted, diagnosis returned
Ethiopia
Rejected by quality gate
65.0%
Unresolved, no crop or health unclear
20.2%
14.7%
452,995
India
27.5%
45.0%
27.4%
389,597
Kenya
57.6%
18.3%
24.0%
237,241
Nigeria
64.9%
18.0%
17.1%
71,035
Share of each country's submissions. Verdict recorded but never sent (0.1% per country) omitted.
Figure 4: Outcome shares by country. Table 4: How often each quality check fails. Check
Type
Dominant Content Resolution Focus Obstruction Motion Blur Plant Detected Lighting Orientation Color Balance Noise
semantic photometric photometric semantic photometric semantic photometric semantic photometric photometric
Table 6: Problem vocabulary by type.
Fail rate
Problem type
35.7% 32.8% 32.2% 27.7% 27.6% 26.0% 25.3% 25.2% 21.6% 20.5%
Crop
Images
Cumulative share
Wheat Maize Cabbage Rice Potato Tomato Pepper Coffee Bean Banana
157,526 97,176 36,933 35,280 34,566 27,728 23,768 15,945 14,839 12,715
25.9% 41.8% 47.9% 53.7% 59.4% 63.9% 67.8% 70.5% 72.9% 75.0%
Disease (pathogen) Pest (insect) Nutrient deficiency Virus Unclassified Weed
37.1% 35.8% 10.3% 9.1% 7.5% <0.1%
1,548 1,907 153 408 1,632 39
All types
100%
5,689
propose an annotation that the reviewer verifies or corrects. This shifts human effort from generating labels to validating them, allowing substantially more samples to be reviewed within the same time and creating a larger pool of reliable labels for model evaluation and training. Section 6 describes how this human-in-the-loop approach can be incorporated into the pipeline.
Table 5: Top crops by volume with cumulative share of the 608,742 crop-named images.
1 2 3 4 5 6 7 8 9 10
Distinct labels
Shares are of the 401,594 labelled occurrences; the last column counts distinct canonical labels. Two labels marked unspecified are counted in the total only, and shares are rounded.
On the quality-gate training split of Table 15, half gate-rejected images and half gate-accepted. The labels are the GPT-4o gate’s own.
Rank
Share of occurrences
3
Pipeline Design
Each principle in Table 9 answers a finding in §2.
3.1
AI-annotated samples rather than generate every label independently. Instead of asking a reviewer to identify the crop and problem from a blank form, the model can
Three-Stage Architecture
Figure 5 shows the proposed pipeline. M0 determines whether the photograph is usable and sends recoverable images for enhancement. M1 identifies the crop. M2.1 4
Table 7: Submission outcomes for the four countries that send 98.9% of all photographs. Country
Total
Rejected
Rejected share
Crop named
Healthy
Healthy share
Disease
Livestock signal
Ethiopia India Kenya Nigeria
452,995 389,597 237,241 71,035
91,288 175,125 43,374 12,772
20.2% 45.0% 18.3% 18.0%
300,117 116,521 141,369 47,059
206,201 38,057 81,880 32,923
68.7% 32.7% 57.9% 70.0%
88,328 69,239 54,796 13,163
1,815 4,740 4,192 344
Rejected share is of the country’s own total, not of judged images as in Table 3. Healthy share is of its crop-named images. Table 8: Top five crops and top three named problems by country. Country
Top five crops
Ethiopia
wheat 51.4%, maize 17.2%, cabbage 7.1%, potato 6.6%, tomato 2.9% rice 11.6%, pepper 10.6%, cucumber 7.9%, tomato 7.7%, eggplant 7.2% maize 25.2%, cabbage 12.0%, potato 8.1%, tomato 7.9%, coffee 7.1% rice 38.0%, maize 14.4%, millet 5.6%, pepper 5.6%, sorghum 5.2%
India Kenya Nigeria
Top-5 share 85.2% 45.0% 60.3% 68.7%
Top three named problems Fall Armyworm 14.5%, Fusarium Head Blight 13.7%, Northern Leaf Blight 12.1% Nitrogen Deficiency 13.1%, Tobacco Caterpillar 6.6%, Potassium Deficiency 6.6% Cercospora Leaf Spot of Beet 15.3%, Fall Armyworm 13.7%, Tobacco Caterpillar 6.2% Nitrogen Deficiency 19.5%, Fall Armyworm 10.9%, Red Cotton Bug 8.4%
Crop shares are of that country’s crop-named images; problem shares are of its problem-named images.
M2.1
Disease detect
Farmer image
M0
Quality detect / enhance
M1
Crop detection
M2.2
Confidence / decision layer thresholds, routing
Structured diagnosis
Pest detect reject / re-ask farmer
no crop / unknown
Figure 5: Proposed three-stage pipeline with named reject branches. Table 9: Six design principles. ID
Principle
P1
Separate image quality from diagnosis. A diagnosis model is never asked to compensate for an unusable image. Tune the quality gate to keep usable images: reject what cannot be read, and rarely reject what can. Break the problem into separate decisions: quality, crop, disease, pest. Make every threshold a setting of ours rather than a provider default we inherit. Let production data decide coverage: the crops, countries, diseases, pests and photo conditions the pipeline actually sees. Measure every stage independently, not only the final answer.
P2 P3 P4 P5 P6
3.2
Module M0: Image Quality Detection and Enhancement
Let as many useful images through to M1 and M2 as possible, while staying fast and cheap. 3.2.1
Quality Gate A: VLM Reference
The production gate runs GPT-4o once per image with ten checks (Motion Blur, Lighting, Focus, Obstruction, Color Balance, Dominant Content, Resolution, Noise, Orientation, Plant Detected) at the per-image price in Table 2. We treat its answers as the labels Gate B learns from, not as the truth: called twice on the same image it does not always say the same thing. 3.2.2
identifies the disease, while M2.2 identifies the pest. A final decision layer applies the configured thresholds and determines how each image is routed.
Quality Gate B: Lightweight CV Model
Gate B is the replacement we propose: a small model trained to copy Gate A’s decisions, so that the threshold is ours to move and no paid call is made per image. We built five candidates, ranging from simple threshold rules to a convolutional network: calibrated threshold rules over blur, exposure and contrast statistics; gradientboosted trees over the same statistics; a hybrid of those trees with a tiny convolutional network; a tiny custom CNN sized to fit on a phone; and MobileNetV3-small, a standard mobile backbone.
If an image is rejected by M0 or M1, the pipeline records the reason so the farmer can be asked for a better photograph instead of receiving no diagnosis. Both routes (§3.5, §3.6) use this same architecture; they differ only in whether the stages are implemented as separate model calls or as a single model call. 5
Each candidate was scored against Gate A’s answers on the quality-gate split of Table 15, whose train, validation and test images never share an identifier, both overall and one check at a time. The metric is F1, which balances precision and recall for rejection: how often a rejection is correct and how many images that should be rejected are identified. The target, fixed before the runs, was 88% F1. Two things are asked of the winner besides agreement with Gate A: a median latency low enough to sit in front of every request, and a file small enough that the same gate could later run on the phone rather than on a server. Section 5.1 reports all three, per candidate and per check. 3.2.3
• Pest detection (M2.2). Input: image plus the crop label from M1, taken as one more input rather than as a filter. Output: pest name, confidence. 3.4.2
The system in production sends every image through the Plantix service after a quality gate. Plantix detects the crop and the disease or pest together, in one answer. Our design separates the two because crop-based narrowing helps disease detection but hurts pest detection. For pathogens narrowing matches the biology: the same disease looks different on different plants, so knowing the plant helps, and the disease list is narrowed to what occurs on that crop. For insects the plant helps less and narrowing hurts. A caterpillar looks like a caterpillar whatever plant it sits on, and a crop found on the wrong plant would rule the right insect out. So the crop label goes to both heads, and only the disease head is allowed to narrow its list by it. The pest head keeps all 92 choices open, which is what stops a wrong M1 call from losing the answer. Figure 6 and Table 10 give the split we propose. How much of the query volume each branch carries is in Table 6: pathogens and viruses together are the larger share, insects the next.
Image Enhancement and Routing
M0’s answer sends an image one of three ways. A good image goes straight to M1. A fixable one is cleaned up first and then goes to M1. An unusable one is rejected, and the farmer is asked for a better photograph. The cleaning step is a proposal only: we have not tested any method for it yet (§6).
3.3
Module M1: Crop Detection
• Input: a quality-passed image, plus the farmer’s registered crop profile where one exists. • Output: crop name, confidence, and an “unknown or unsupported” flag when confidence falls below the configured threshold. Three models can fill M1. They differ on one thing: whether the crop is answered on its own or in the same call as the diagnosis. • Qwen3-VL-4B, fine-tuned. The Route A model (§3.5). Answers the crop in the same call as the diagnosis, in its own words rather than from a list. Scored in Table 21. • DaViT-Base, fine-tuned. The Route B reference model (§3.6). Answers the crop from a fixed list, in the same pass as the category, disease and pest heads. Scored in Tables 21 and 24. • YOLO26x-cls, fine-tuned. Carries the four-part head of §3.4, so one model answers crop, category, disease and pest together. Scored on crop in Table 21 and against DaViT-Base on all four heads in Tables 24 and 25. Two of the three sort an image into a fixed list of classes and one answers in free text. None draws a box around the problem, and none can name a crop outside its list. The “unknown” cut-off is a setting, not a learned value. Calibrating it against a target error rate is listed in §6.
3.4
Module M2: Disease and Pest Detection
3.4.1
Two Sub-Modules
Conditional Routing
3.4.3
Reference Implementation
The model in §5.6 does exactly this split. One shared backbone feeds four heads: • Crop, 110 choices, and category, 3 choices. The category answer is what picks between the next two heads, and it is made in the same pass. • Disease, 285 choices, filtered by a table of which diseases occur on which crop. It needs the crop first. • Pest, 92 choices, not filtered at all. This is one model with two separately gated heads, rather than two independent pipelines. In this reference model the four heads share a backbone and the pest head is not handed the crop label itself, so the optional input of Figure 6 is part of the proposed design and not of the model scored in §5.6.
3.5
Route A: Fine-Tuned Language Model
Vision-
Figure 7 shows the route: one model answers all three stages in one call. 3.5.1
Candidate Model
Qwen3-VL-4B, fine-tuned on data curated from the panel-labelled images (§4). Gemma-3-4B was considered alongside it; we chose Qwen3-VL-4B on the proof-ofconcept results. None of the Qwen fine-tune’s training images appears in the scoring set of §4.4, so its scores in §5 are out of sample. It also trained on almost the same images as Route B, but none of the images used to validate or test Route B were included, so the two routes are close to a matched comparison.
• Disease detection (M2.1). Input: image plus the crop label from M1. Output: disease name and confidence. Severity is out of scope here (§6). 6
Disease classifier
pathogen and viral, list narrowed to the crop
narrows the list
M1 crop label
M0-cleared image
goes to both heads
one input, no narrowing
Pest detector
insect, all 92 choices stay open
Figure 6: Proposed M2 routing. The crop label reaches both heads and narrows only disease. Table 10: Proposed M2 routing, replacing one Plantix call that returns crop and problem together. “Needs” means the route cannot run without the label. Route
Needs M1 crop label
Model type
Disease (pathogen and viral) Pest (insect)
✓ Optional
Per-crop classifier, unchanged Unfiltered 92-way detector that may read the label, YOLO candidate
The pest share is a row in Table 6; the disease branch is its pathogen and virus rows together.
Image
Structured output
Fine-tuned VLM
quality, crop, disease, pest confidence follow-up question when the image is insufficient
Qwen3-VL-4B, single forward pass
Figure 7: Route A. One model, one call, no hand-off between stages.
3.5.2
• Limitations. Needs a GPU, which makes it the costlier route to host (Table 22), and at the slow end of the throughput we assume the two GPUs only just clear the busiest minute we measured, so more farmer queries need more GPUs (§5.3); disease accuracy is confounded by the stored production answer sitting in its training labels (§5.5).
What Fine-Tuning Changes
A general-purpose model like Qwen3-VL-4B will attempt an answer even when the photograph does not provide enough information. The fine-tuned model, in contrast, can recognize when more information is needed and ask for a specific follow-up photograph. Figure 8 shows one example, and Table 11 gives five more. • It asks for one of five things: a better photograph, a closer look at the crop, a closer look at the damaged part, whether the subject is an animal, or plain advice rather than another question. • The request is specific to the crop and suspected problem: it tells the farmer what to photograph and what the new image should help confirm. This is more useful than simply asking for a clearer picture. For example, the model may ask for a close-up that helps distinguish between two insects that look similar but require different treatments (Table 11). • It also asks when it is already right: on many rows it named both the crop and the problem correctly and still asked for the photograph an agronomist would want before recommending treatment (§5.5). • Table 22 carries how often it asks, measured on the shared test rows of §4.4. 3.5.3
3.6
Route B: Computer Vision Model Orchestration
Figure 9 shows the route: one specialist model per stage, orchestrated. Each stage is versioned on its own, so a weak stage can be replaced without retraining the others. The price is a hand-off between each pair of stages. Table 12 lists the candidate per stage. • Advantages. Every threshold is explicit and independently tunable; a wrong answer traces to one stage; the models are small enough for CPU-only serving, with a mobile-viable quality gate; the lower hosting cost of the two routes (Table 22). • Limitations. Three models to version, monitor and keep in sync; a wrong M1 crop call misroutes M2, so error compounds down the chain; no reasoning layer and no follow-up question; more moving parts to build and maintain than one model.
Advantages and Limitations
• Advantages. Single pass with no hand-off between stages; asks for a better photograph instead of guessing; extracts unstructured detail alongside the structured fields; one model to deploy.
4
Data and Methodology
The labels are made by models, not by people labelling at scale. Plantix cannot supply them, because Plantix 7
{ "analysis_status": "full", "crop": { "present": true, "primary_name": "maize", "scientific_name": "Zea mays", "growth_stage": "vegetative", "other_crops_visible": [] }, "health_issue": { "present": true, "category": "pest", "specific_name": "fall armyworm", "severity": "moderate", "affected_area_percent_estimate": 30, "affected_parts": ["leaf", "whorl"], "symptoms_observed": [ "ragged elongated feeding hole in leaf", "dark frass-like exudate with reddish-brown coloration", "abundant pale granular frass deposits around feeding site", "tissue chewed and torn at the whorl" ], "co_occurring_issues": [] }, "image_quality": { "overall": "good", "issues": [], "diagnostic_usability": "yes", "authenticity": "real_photo" }, "context": { "setting": "open_field", "camera_distance": "close", "soil_visible": true, "soil_condition": null }, "recommended_followup_image": "close-up of the whorl interior to confirm presence of larvae and characteristic inverted-Y head capsule markings" }
Figure 8: One Route A output, as returned. Every field is machine-readable, including the follow-up it asks for. Table 11: Follow-up photographs the fine-tune asks for, taken from its own output. Crop
Diagnosis
The follow-up it asks for
Common bean Mango
Bean anthracnose White mango scale
Maize
Fall armyworm
Citrus
Mealybug
Tomato
Leaf miner
close-up of affected leaves to check for accompanying foliar lesions and confirm anthracnose close-up macro of the white encrustations to confirm individual scale insect bodies and check for crawlers close-up of the whorl interior to confirm presence of larvae and characteristic inverted-Y head capsule markings macro close-up of a single infested stem to confirm individual mealybug bodies and waxy filaments versus fungal growth close-up of a single affected leaf showing the serpentine mine trails to confirm leaf miner activity versus other foliar damage
Quoted as written by the model. Each names a part of the plant, a sign to look for, and the alternative it would rule out.
Image
M0
CNN quality gate usable / recoverable / unusable
M1
DaViT, YOLO26 classifier
M2
per-crop disease + pest head
Structured diagnosis
Figure 9: Route B. Three independently versioned models.
is one of the things being measured. Figure 10 shows how the labels are built: a council of models labels every image on its own, normalisation and a consensus vote turn those into one label per image, and human review feeds back into both training and the label list.
of the vote so that no model being tested helps write the label it is scored against. These are Plantix and the two models we fine-tune. Table 13 gives how often each model answered, and how often it agreed with the council, for crop and for disease.
4.1
4.3
Sources and Sampling
Eight models answer in free text, so the same plant can be described in eight different ways. Each answer is mapped onto the one label vocabulary of Table 6 by a fixed set of rules, run as a script so the mapping can be repeated. Table 14 gives the rules with real answers from the data. Two rules carry most of the work. Dropping filler (damage, infestation, feeding, complex) collapses the many ways a model describes the same insect, which is why thrips absorbs six written forms. Against that, a guard list of words that change the meaning (early, late, downy, powdery, bacterial, tree) stops the same rule merging two real classes. The final label is then a vote, with confidence and tiebreak rules. Disagreement, low confidence, new classes
The images come from Farmer.Chat production data. Each image has a diagnosis label from Plantix, our production diagnosis service. These labels are used to stratify the initial sample by crop and by healthy, disease, and pest categories (Table 15). Plantix data or models are not used for training. The training labels are generated by the model council (§4.2), with the production Plantix label used as the reference when the council does not reach agreement. Human review is the only non-model source of labels.
4.2
Normalisation and Consensus
Model Council
Eight models label every image on their own. Five of them vote on the final label. Three are kept out 8
Table 12: Route B candidates by stage. Stage
Candidate
Role in the route
M0, quality
MobileNetV3-small CNN
M1, crop (reference model)
DaViT-Base hierarchical crop head
M1, crop (next-generation candidate)
YOLO26x-cls, fine-tuned
M2, disease
DaViT-Base hierarchical head, crop-masked
M2, pest
Unfiltered pest head (§3.4.2), unmasked 92-way, inside the same reference model
The learned gate of §3.2.2, chosen on its agreement with the paid gate The crop head of the four-head model, which also serves M2 below, so one model covers both stages The same four-head design on a second backbone, scored against the reference model in §5.6 Picks the disease from the list allowed for the crop M1 predicted Names the insect from the full list, so a wrong M1 call cannot rule the right insect out
Sections 5.1, 5.2 and 5.6 score these candidates.
Sources
stratified samples
Model council 8 models label independently 5 vote, 3 held out
Normalization
Consensus
canonical label vocabulary
vote + confidence tie-break rules
Human review
disagreement, low confidence, new class, long tail
feeds training + taxonomy
Figure 10: Data curation path. Table 13: The eight-model label council. Model
Role
Sonnet 4.6 GPT-5.4 Opus 4.8 Kimi K2.5 Pixtral Large Qwen3-VL, off the shelf Plantix Gemma-3
panel panel panel panel panel held out held out held out
Crop answered
Crop agreement
Disease answered
Disease agreement
100.0% 99.8% 98.9% 90.9% 76.0% 99.5% 94.7% 65.0%
82.0% 81.2% 80.9% 80.5% 68.8% 78.4% 80.5% 72.1%
71.9% 55.1% 35.1% 59.8% 49.6% 75.1% 66.7% 54.6%
68.3% 65.7% 72.2% 77.9% 77.2% 69.8% 58.8% 61.4%
Crop columns are over the 114,721 rows with a council crop, disease columns over the 80,398 with a council disease. Whole labelled sample, not the test split. Agreement is a synonym match, the same map applied to the council label and to the answer as in §4.5, counted over the rows the model answered: Opus 4.8 names one disease in three and agrees on three quarters of those. Plantix answers crop as a candidate list; one candidate counts as an answer, several do not.
and rare classes go to a human reviewer. The vote often produces no diagnosis: more than half of all disease answers are refusals to name the problem, so many images have no council label to count.
4.4
apply the same synonym map to the reference and the prediction. The map merges names for one crop or problem that the label list holds twice, such as sugar beet and beet, so a system is not marked wrong for choosing the other spelling. 2. Count a non-answer as a wrong answer. Markers such as unspecified, and answers listing several candidates, become “no answer” and score as wrong. The row set is the same for every system, so no system can raise its score by answering less often. 3. Score twice. The strict column counts only an exact match. The containment column is more generous: it counts an answer when either label contains the other (“mosaic virus” against “cucumber mosaic virus”), which accommodates models that answer in their own words.
Dataset Splits
Table 15 gives the two modules trained here. Route A trains on its own draw from the same pool, and §5 names the row set behind every results table. Two experiments run side by side. Table 21 scores the single scoring dataset. The hierarchical test (Table 24) is a separate experiment with its own four-part label set, built to answer the disease-against-pest question that one diagnosis grade cannot. No number from one is compared with a number from the other.
4.5
Two kinds of row stay outside the diagnosis count for every system alike: references that are healthy or unspecified, and references that are Plantix’s own stored answer. Dropping the second kind stops Plantix grading itself. It does not do the same for GPT-5.4, which voted on the panel consensus (Table 13); §5.2 names
Scoring Rule
One rule scores every system, in three steps: 1. Clean both labels. Lower-case them, cut any free text after a dash, colon, comma or bracket, and 9
Table 14: Normalisation rules, with answers taken from the model council’s own output. Rule
Answers as written
Becomes
Plurals and spellings Filler words
tomatoes, tomato plant leaf miner damage, leaf miners, leaf miner feeding damage maize seedling, coffee plant huanglongbing, helicoverpa armigera panama disease, jassid, corn, paddy rice early blight, late blight; tree tomato, tomato No crop detected, not defined, unidentified crop cucurbit, legume, grass-family crop [bean, gram, soybean, pea]
tomato leaf miner
Growth stage Scientific name Regional or trade name Opposites kept apart Model declines to answer Family rather than crop Several candidates at once
maize, coffee citrus greening, cotton bollworm fusarium wilt, leafhopper, maize, rice left as four separate classes __unspecified__ kept, marked low specificity __multi__, each member recorded
Drawn from 922,976 answers (115,372 images by eight models). A refusal and a family-level answer are each kept as their own outcome, not discarded.
• Threshold rules alone trail every learned candidate (Table 19). A learned model is required, not only calibration. • No candidate reaches the 88% target of §3.2.2; the best is 1.1 points short. That may be close to the limit imposed by Gate A’s own inconsistency. • The four checks about what is in the picture are also the four on which the small model has the highest F1 (Table 20), led by Plant Detected and Dominant Content. Dominant Content is also the check that fails most often (Table 4).
Table 15: Images used to train, validate and test each module. Module
Train
Validation
Test
M0, quality gate M1, crop
19,994 82,991
3,000 9,220
4,999 16,275
The M0 splits never share an image identifier. The M1 split is the hierarchical manifest, so the same images train the M2 disease and pest heads.
that exposure where its column is read.
4.6
Baselines and Systems Compared
Table 16 defines the systems compared throughout the paper.
4.7
5.2
Table 21 is the headline comparison. Seven systems are scored on the same 10,335 images under the one rule of §4.5: the Route A fine-tune, the same model before fine-tuning, the two Route B backbones, the production baseline, a panel model (GPT-5.4) and a general-purpose baseline. Table 15 gives the row counts. Two systems are left out: the Gemma-3 fine-tune and the model it started from, both rejected candidates (§3.5). Scoring rules: a refusal counts as a miss, containment is the generous match of §4.5, and backbone rows are the mean of two seeds. GPT-5.4 voted on the panel labels (Table 13), so it helped write the reference that its own diagnosis column is scored against. Four readings of Table 21: • The set is every top-20 crop test row on which every scored system has a prediction on record. The production crop classifier determined the usable set, declining 12.5% of the 11,810 top-20 test rows and leaving the 10,335 scored here. • The two families fail differently. Neither backbone ever refuses to answer. The language models return no diagnosis on 15% to 45% of panel-labelled rows, and Plantix on 36%. Being generous about wording does not close the gap: DaViT-Base leads Plantix by 30.7 points on exact matches and by 23.9 when either label containing the other counts. • Route B names the crop best. The fine-tune is trained on rows disjoint from this test set and reaches 94.66%, below DaViT-Base. It learned from the same panel that wrote the reference labels, so part of its score may reflect agreement with that panel’s naming conventions.
Pipeline Configurations
Table 17 lists the configurations under test. Two comparisons follow from it: the existing gate against the proposed M0 (Table 19), and one model doing everything against the three-stage pipeline (§5.2, with Route B’s per-stage numbers in §5.6). Two more are listed in §6: M0 with and without the cleaning step, and training with and without human-reviewed labels.
4.8
Metrics by Stage
Table 18 lists the metrics reported per stage. Section 6 lists two further measurements that follow this paper: useful-image recall and false rejections at M0, and confidence calibration at M1.
5
Results
5.1
Component-Level Results: M0
Crop and Diagnosis on One Test Set
Table 19 and Figure 11 score the five Gate B candidates of §3.2.2 against Gate A’s answers on the 4,999-image held-out test of Table 15, on agreement, size and latency. Table 20 gives the per-check detail for the winner. • MobileNetV3-small wins at 86.91% F1 and 12 ms; the tiny CNN reaches 81.02% at 2 ms and is the only learned candidate that fits on a phone. Recommendation: MobileNetV3 for server deployment, the tiny CNN for on-device use. 10
Table 16: Baseline and proposed system definitions used throughout the paper. ID
Definition
Baseline 0 Baseline 1
Existing production pipeline (GPT-4o quality gate and Plantix diagnosis call) General-purpose VLM, no fine-tuning (Gemini 3.5 Flash reported; GPT-5.4, Claude and Kimi scored as labelling-panel members) Fine-tuned vision-language model, single call, all three stages Specialist computer-vision models, one per stage, orchestrated
Route A Route B
Table 17: Configurations under test. Configuration
M0
M1
M2
Baseline 0, existing Baseline 1, general VLM Route A, fine-tuned VLM Route B, CV orchestration
Existing VLM VLM CNN gate
Existing VLM VLM DaViT, YOLO26 cls
Existing VLM VLM Hierarchical disease and pest heads
Table 18: Metrics reported, one row per metric. Metric
What it measures
Stage
Agreement with the reference gate p50 latency Model size Crop accuracy Declined share Diagnosis accuracy Head accuracy Disease macro-F1 Monthly hosting cost Deployability
How often the small gate makes the same call as the paid gate, overall and per check Median time to judge one image File size, which decides whether the model fits on a phone Share of images whose crop is named correctly Share of images where the model names no crop Share of images whose problem is named correctly, exactly and with containment credit The same, one score per head: crop, category, disease, pest Every disease class scored equally, so rare ones count as much as common ones Cost of running the route at the demand we measured What hardware the route needs
M0 M0 M0 M1 M1 M2 M2 M2 System System
• The backbones answer from a fixed list of 110 crops and 285 diagnoses; the language models answer in their own words, which is what the containment column allows for (§4.5).
5.3
onion has only two test rows, so its accuracy is not informative.
5.5
Table 23 lists the crop pairs that the Route A fine-tune misclassifies most often, on the same 10,335 rows. • Nearly a third of its crop errors, 29.5%, name a family rather than a crop: cucurbit for cucumber, grass-family crop for wheat or maize. Those errors reflect an overlap in the label list, not a misreading of the image. • The fine-tune asks for a second photograph on 57.5% of the 10,335 rows, and on 1,288 of the 2,452 panellabelled rows where it had already named both the crop and the problem correctly.
Route A Against Route B
These costs are modelled, not billed: • Prices come from the AWS list published 10 August 2026 for ap-south-1, applied to the query volume we measured: 46,869 diagnosis calls in the month to 24 July 2026. • The CPU route’s speed comes from measured laptop figures. The GPU route’s speed, 0.5 to 5 images per second per GPU, is assumed with no measurement behind it. This is the least certain assumption in the $924 estimate. • Most of either bill is idle time: under 8% of the paid hours are used in every scenario, and even at the slow end of its measured range the CPU route provides 14.6× the capacity required for the busiest minute we measured, while the GPU route at its slowest assumed speed only just meets that demand.
5.4
Error Analysis
5.6
Backbone Benchmark: DaViT-Base Against YOLO26x-cls
A held-out test of 16,273 images. The disease filter uses the model’s own crop prediction, which is what a deployed system would have, not the true crop. Both backbones were trained the same way, 25 epochs on the hierarchical manifest of Table 15. Values are the mean of two seeds (42 and 1337), and the two seeds differ by less than a point on every head. • Mean head accuracy is the plain average of the four heads, which do not share a denominator. • DaViT-Base wins every head by 2.1 to 3.7 points (Figure 12), but the two differ by 3.0× in parameter count, so part of the gap is capacity rather than design. • YOLO26x-cls trains 2.3× faster (11.7 against 27.2 minutes per run, Table 25) and has about 1.8× higher
Country and Crop Analysis
• India has 45.0% of its submissions rejected before diagnosis, compared with 18.3% in Kenya (Table 7). A single global quality threshold under-serves one country or the other. • Wheat and maize are 41.8% of crop-named images, and the top 20 crops cover 89.2% (Figure 3). • Of the 20 reference crops, 17 reach 80% or better for the Route A fine-tune. Cotton (72.8% of 320 rows) and cucumber (76.0% of 495) are the exceptions; 11
Table 19: Gate B candidates on the 4,999-image held-out test, scored against Gate A. Candidate MobileNetV3-small CNN Hybrid (GBM + tiny CNN) Gradient-boosted trees Tiny custom CNN Calibrated threshold rules
F1
Accuracy
Size
p50 latency
86.91% 83.79% 83.08% 81.02% 76.58%
86.76% 83.88% 82.90% 82.52% 77.32%
2.42 MB 1.85 MB 1.29 MB 0.56 MB ∼0 MB
12 ms 78 ms 91 ms 2 ms 22 ms
Fits mobile (<1 MB) ✗ ✗ ✗ ✓ ✓
Green bold marks the best F1, accuracy and latency. A tick in the last column means the model fits inside a 1 MB on-device budget. F1
88%
MobileNetV3-small CNN F1 86.9% · 2.42 MB Hybrid, GBM + tiny CNN F1 83.8% · 1.85 MB
84%
80%
Gradient-boosted trees F1 83.1% · 1.29 MB
Tiny custom CNN F1 81.0% · 0.56 MB
Calibrated threshold rules F1 76.6% · ~0 MB
76%
0 ms 25 ms 50 ms p50 latency per image on CPU. Marker area is model size; hatched markers fit a mobile budget (under 1 MB). Vertical axis: F1 against the GPT-4o gate's decisions, 4,999 held-out images.
75 ms
100 ms
Figure 11: Gate B candidates, F1 against p50 latency on the held-out test. DaViT-Base, 87.4M params
YOLO26x-cls, 29.0M params
100%
90%
89.8
89.3
87.8
87.0 82.7
80%
79.8
77.9 74.1
73.7 70.2
70%
60%
Crop head
Category head
Disease head
Pest head
Mean of 4 heads
Held-out test, mean of 2 seeds. The disease filter uses the model's own crop prediction.
Figure 12: Head accuracy for the two backbones, mean of two seeds.
74.14% to 77.86%), which is the expected price of picking from all 92 pests with no help from the crop. • The labels come from the council’s vote, with the stored production answer standing in where the council did not agree, not from experts. So this table measures agreement with that consensus, the same caveat as §5.2.
Table 20: Per-check F1 for MobileNetV3-small, ranked. Check Plant Detected Dominant Content Orientation Obstruction Noise Color Balance Lighting Motion Blur Resolution Focus
F1 92.2% 88.2% 80.8% 80.2% 79.6% 79.0% 77.8% 76.8% 76.2% 74.8%
6
Impact and Future Work
Green bold marks the best per-check F1.
6.1
Impact
inference throughput, with one third as many parameters. • The disease result is not a 285-class classification. The crop filter cuts it to a median of 10 allowed diseases per crop (lowest 8, highest 43). • The pest head is never filtered by crop, by design. It scores below disease (70.21% to 73.66% against
No route runs on live farmer queries at the time of writing, so nothing below is an observed outcome. Table 26 compares the four systems that answer the crop and the diagnosis question (M1 and M2) on the three deployment dimensions of cost, accuracy and speed. Three readings of Table 26: • Cost and accuracy do not present a trade-off 12
Table 21: Seven systems on one test set, scored under one rule. System
Role
DaViT-Base, hierarchical Qwen3-VL-4B, fine-tuned YOLO26x-cls, hierarchical Plantix GPT-5.4 Gemini 3.5 Flash Qwen3-VL-4B, off the shelf
Route B reference model Route A, out of sample Route B next-generation candidate production baseline panel member general-purpose baseline Route A before fine-tuning
Crop
Diagnosis
Containment
No answer
95.41% 94.66% 93.78% 91.46% 88.25% 86.16% 73.34%
73.26% 33.04% 68.42% 42.61% 40.36% 39.05% 12.96%
77.43% 36.97% 72.63% 53.50% 45.86% 46.95% 16.03%
0% 41.30% 0% 35.99% 42.47% 14.67% 45.08%
Crop is over all 10,335 rows; diagnosis, containment and no-answer over the 7,636 whose reference came from the panel. Green bold marks the best per column. Table 22: Route A against Route B. Dimension
Route A, fine-tuned VLM
Route B, CV orchestration
Crop accuracy
94.66%, out of sample
Diagnosis on panel-labelled rows
33.04%, label-confounded
Monthly hosting cost, measured demand
$924 (two g6.xlarge, L4 GPU, one-year reservation) Asks for a second photograph on 57.5% of rows (§5.5) GPU required one wrong answer, no chain
95.41% reference model, 93.78% next-generation candidate 73.26% reference model, 68.42% next-generation candidate $100 (two c7g.xlarge, CPU only, one-year reservation) not available
Follow-up question Deployability Failure mode
CPU only; on-device quality gate possible a wrong M1 call misroutes M2
Crop and diagnosis from Table 21, cost from the hosting model below. Green bold marks the better route where the two are comparable.
2. Image enhancement. No cleaning method is evaluated. 3. Not modelled. Disease severity, multi-disease images, multi-crop images, disease progression over time. 4. AI-aided human review. Only a small set of independently reviewed images reaches training; reviewers will instead correct the pipeline’s own answer under written guidelines (§2.2.2). 5. Confidence calibration. M1 confidence will be calibrated, so a cut-off can be set from a target error rate.
Table 23: Crop pairs the Route A fine-tune misclassifies most often. Reference
Predicted
Cucumber Wheat Soybean Cotton Maize
Cucurbit Grass-family crop Common bean Okra Grass-family crop
Share of crop errors 13.9% 6.9% 3.6% 3.4% 3.4%
Shares are of the model’s 552 crop errors on the 10,335 scored rows, which are 5.3% of those rows.
here. The cheaper route is also the more accurate one on both questions, so the case for the GPU route rests on what it does besides classify: the follow-up question and the free-form reasoning (§3.5). • Speed is the weakest column and should not decide anything yet. The two backbones are timed batched on training hardware, the production path is timed per request, and the fine-tune has no number at all. A single benchmark using the deployment configuration for each system would resolve this. • The gap the table does not show is control. Every value in the Plantix row is fixed: no threshold to set, no crop to add, no confidence to read (Table 2). The two backbone rows provide control over all three. Beyond crops, the three questions the pipeline asks, whether an image can be read, what is in it, and what is wrong with it, carry to the condition of a farm animal, the grading of produce, and any image task where whether the photograph is usable is a separate question from what it shows. The photographs in this corpus that appear to show animals (Table 7, last column) are a natural next application.
6.2
7
Conclusion
• Diagnosing a crop from a field photograph is five separate decisions (§1), and the production system lets us adjust none of them. • Measuring M0, M1 and M2 separately means every threshold is ours to set and every failure can be traced to one stage, which a single closed call cannot do. • A small local MobileNetV3 matches the GPT-4o quality gate at 86.9% F1 and 12 ms, 1.1 points under the 88% target we set beforehand. • On one test set of 10,335 images, scored the same way for every system, the DaViT-Base backbone achieves the highest crop accuracy at 95.41%, 0.75 points above the fine-tuned VLM and 3.95 above Plantix, on images it never saw. It leads every system on diagnosis too: 73.26%, against 68.42% for the smaller YOLO backbone, 42.61% for Plantix and 33.04% for the fine-tune, at one ninth the VLM route’s modelled hosting cost. No diagnosis number is final until the stored production answer is out of every prompt and training set, and until disease and pest are split into separate label lists. • Route A and Route B sit at different points on ac-
Future Work
1. False-rejection rate at M0. How many usable images the gate rejects needs its own measurement, using a recall-versus-threshold sweep of Gate B. 13
Table 24: Four-head accuracy on the held-out test, mean of two seeds. Backbone
Params
DaViT-Base YOLO26x-cls
87.4M 29.0M
Mean head acc.
Crop
Category
82.65% 89.83% 79.78% 87.75%
Disease
Pest
Disease macro-F1
89.27% 77.86% 87.02% 74.14%
73.66% 70.21%
41.75% 39.03%
Crop (110 choices) and category (3) are over all 16,273 rows; disease (285, filtered by predicted crop) over 11,916; pest (92, unfiltered) over 4,276. Table 25: Training and inference cost of the same comparison. Backbone DaViT-Base YOLO26x-cls
Train img/s
Train wall (2 seeds)
Peak GPU memory
Eval img/s
1,430 3,699
54.4 min 23.3 min
18.97 GB 17.31 GB
655 1,197
Four L40S GPUs, distributed data parallel. Train wall is the measured elapsed time summed over the two seed runs. Green bold marks the better value per column. Table 26: Crop and diagnosis systems against cost, accuracy and speed. System
Monthly cost
Crop
Diagnosis
Measured speed
DaViT-Base, hierarchical YOLO26x-cls, hierarchical Qwen3-VL-4B, fine-tuned Plantix
$100, CPU $100, CPU $924, GPU paid per call
95.41% 93.78% 94.66% 91.46%
73.26% 68.42% 33.04% 42.61%
655 img/s, batched 1,197 img/s, batched not measured 2.8 to 4.4 s per image
Accuracy from Table 21, cost from Table 22. Backbone speed is batched evaluation on four L40S GPUs, not the CPU shape the cost assumes, and the Plantix figure is a single-request production trace (Table 2), so the last column compares three different measurements. Green bold marks the best per column.
curacy, cost and control (Table 22). Which one to deploy is a product decision, and this paper supplies the measurements it needs.
Dash. Fine-tuning and evaluating Conversational AI for agricultural advisory. arXiv preprint arXiv:2603.03294, 2026. [11] Chandrashekar M S, Vineet Singh, and Lakshmi Pedapudi. Benchmarking Automatic Speech Recognition for Indian Languages in agricultural contexts. arXiv preprint arXiv:2602.03868, 2026.
References [1] Namita Singh, Jacqueline Wang’ombe, Nereah Okanga, Tetyana Zelenska, Jona Repishti, Jayasankar G K, Sanjeev Mishra, Rajsekar Manokaran, Vineet Singh, Mohammed Irfan Rafiq, Rikin Gandhi, and Akshay Nambi. Farmer.chat: Scaling AI-powered agricultural services for smallholder farmers. arXiv preprint arXiv:2409.08916, 2024. [2] Sharada P. Mohanty, David P. Hughes, and Marcel Salathé. Using Deep Learning for image-based plant disease detection. Frontiers in Plant Science, 7:1419, 2016. [3] David P. Hughes and Marcel Salathé. An open access repository of images on plant health to enable the development of mobile disease diagnostics. arXiv preprint arXiv:1511.08060, 2015. [4] Davinder Singh, Naman Jain, Pranjali Jain, Pratik Kayal, Sudhakar Kumawat, and Nipun Batra. Plantdoc: A dataset for visual plant disease detection. In Proceedings of the 7th ACM IKDD CoDS and 25th COMAD, pages 249–253, 2020. [5] Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V. Le, and Hartwig Adam. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1314–1324, 2019. [6] Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan. Davit: Dual attention vision transformers. In European Conference on Computer Vision (ECCV), 2022. [7] Ultralytics. Ultralytics yolo26. https://docs.ultralytics. com/models/yolo26/, 2026. Software, released January 2026; variant used: YOLO26x-cls. [8] Shuai Bai et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. Qwen Team, Alibaba Cloud. [9] Gemma Team. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025. [10] Sanyam Singh, Naga Ganesh, Vineet Singh, Lakshmi Pedapudi, Ritesh Kumar, S S P Jyothi, Archana Karanam, Waseem Pasha, Ekta Kumari, C. Yashoda, Mettu Vijaya Rekha Reddy, Shesha Phani Debbesa, and Chandan
14