ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs Ruman Wang1 , Hangting Ye2 , 1
Liaoning University of Traditional Chinese Medicine; 2 School of Artificial Intelligence, Jilin University [email protected], [email protected]
arXiv:2607.28538v1 [cs.CV] 30 Jul 2026
Abstract Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local datagovernance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and featurelevel SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Refinement raises executability from 66.7% to 95.0% and the candidate evidence-pass rate from 70.0% to 91.7%, while final filtering ensures 100% coverage among retained features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.
1
Introduction
Keloids (KD) and hypertrophic scars (HS) are pathological responses to wound healing that can appear similar in clinical photographs but differ in growth pattern, prognosis, and treatment. Keloids extend beyond the original wound and commonly recur, whereas hypertrophic scars remain within the wound boundary and may regress over time (Bayat, McGrouther, and Ferguson 2003; Berman, Maderal, and Raphael 2017). Distinguishing them affects treatment planning. Because clinical scar assessment relies on observerscored properties, the original VSS and POSAS studies explicitly evaluate interrater reliability (Baryza and Baryza 1995; Draaijers et al. 2004). Reproducible image-based decision support could therefore provide a standardized second
opinion or triage aid when scar expertise is not immediately available; it is intended to assist, not replace, clinical evaluation. Building such a system presents two coupled challenges. First, scar representations must be data-efficient and generalize across clinical sites. Expert annotation is costly, and specialized scar cohorts are modest relative to the data typically used to learn image representations. Photographs also vary across hospitals in camera, illumination, viewpoint, skin tone, and anatomical site. Limited labeled cohorts and heterogeneous acquisition are persistent challenges in medical image learning (Litjens et al. 2017). The relevant difficulty is thus not an inability to collect data from multiple hospitals— our cohort spans three—but learning a stable decision rule from a limited number of specialist-labeled cases. Second, clinical knowledge in LLMs must be externalized without turning the LLM into an opaque image classifier. LLMs have demonstrated medical question answering, knowledge recall, and reasoning capabilities (Singhal et al. 2023; Nori et al. 2023). Direct multimodal inference, however, either sends photographs to a hosted service or requires a large model to be deployed locally. In both cases, the prediction remains tied to a black-box, potentially variable generation process; a pilot study of chat-based models on KD–HS images also found that direct diagnosis was not yet clinically adequate (Shiraishi et al. 2024). Generating code avoids direct diagnosis, but introduces its own reliability problem: a plausible program may fail to execute, measure the wrong visual property, or lack clinical support. A useful knowledge-transfer mechanism must therefore keep images local while producing executable, evidence-grounded, and auditable measurements. These challenges lead to our central question: Can an LLM contribute clinical knowledge to data-efficient medical image classification without observing patient images or making the final diagnosis? Scar assessment offers a natural interface between textual knowledge and image computation. Instruments such as the Vancouver Scar Scale (VSS) and Patient and Observer Scar Assessment Scale (POSAS) organize scar evaluation around named properties, including pigmentation and vascularity (Baryza and Baryza 1995; Draaijers et al. 2004). Although photographs cannot capture scale items that require palpation or patient report, their visually assessable concepts suggest a structured representation: convert observ-
able image patterns into explicit measurements, then learn a decision rule in that low-dimensional space. Building on this insight, we propose ScaFE (Scar Feature Engineering), a bounded program-synthesis framework that uses an LLM as a knowledge-driven feature engineer. A web-enabled LLM first retrieves clinical sources under a traceability contract, then generates candidate Python programs whose output dimensions have names, operational definitions, clinical rationales, and supporting evidence. The programs run locally under a restricted interface with no access to the network, labels, filenames, metadata, or patient records. A fixed Random Forest (RF) evaluates each representation on an inner validation split. ScaFE returns only aggregate execution errors, class-wise confusion counts, balanced accuracy, invalid or constant feature rates, and global SHAP importance to the LLM. Over several rounds, this feedback guides code repair and feature revision without exposing any image or patient-level output. Unsupported dimensions are removed by a final evidence gate, and the selected program and RF are frozen before evaluation on an unseen hospital. This design separates three roles that end-to-end and direct-VLM approaches conflate: the LLM supplies clinical priors, deterministic code performs local measurement, and a lightweight learner estimates the diagnostic boundary. The separation reduces dependence on labeled images, makes each feature inspectable, and permits a hosted LLM to assist without receiving clinical photographs. In leaveone-site-out evaluation on 600 images from three hospitals, ScaFE reaches 81.0% site-macro balanced accuracy, 10.0 points above the strongest baseline. Its advantage increases to 11.8 points when only 10% of the development cohort is available, and ablations verify the contributions of literature grounding and semantically aligned validation feedback. Our contributions are threefold: • We formulate LLM-assisted medical image learning as evidence-grounded feature-program search and introduce ScaFE, which transfers clinical knowledge into deterministic local measurements rather than direct image predictions. • We develop a validation-guided refinement loop that improves program executability and feature utility using only aggregate feedback, while retaining a traceable evidence record for every feature used by the final predictor. • We conduct patient-level leave-one-site-out evaluation across three hospitals, demonstrate gains over supervised, foundation-model, handcrafted, and direct-VLM baselines, and audit executability, evidence grounding, feature faithfulness, and cross-LLM stability.
2 2.1
Related Work
Scar Assessment and Image-Based Learning
Clinical scar assessment commonly uses structured instruments such as VSS and POSAS (Baryza and Baryza 1995; Draaijers et al. 2004). Their named dimensions make clinical reasoning explicit, but several items require palpation or patient report and cannot be recovered from an ordinary photo-
graph. Earlier computational pipelines translated observable color, texture, and morphology into handcrafted descriptors. Such features are inspectable but laborious to design and brittle under changes in illumination or acquisition. Deep networks instead learn image representations and have achieved strong results in dermatology and medical imaging (Esteva et al. 2017; Liu et al. 2020; Litjens et al. 2017). Recent dermatology foundation models and ontologyaligned vision–language resources extend this trend across broader tasks (Yan et al. 2025b,a). Their effectiveness nevertheless depends on task-relevant data and validation under the intended deployment shift. For KD–HS differentiation specifically, a pilot that repeatedly queried chat-based models with scar images reported low diagnostic accuracy and concluded that direct use was not clinically ready (Shiraishi et al. 2024). We evaluate supervised networks and frozen visual encoders under patient-level, leave-one-hospital-out testing.
2.2
Language and Vision-Language Models in Medicine
Medical LLMs encode clinical knowledge, while vision– language models extend it to biomedical images and report generation (Singhal et al. 2023; Nori et al. 2023; Zhou et al. 2024; Li et al. 2023; Rao et al. 2025; Sellergren et al. 2026). These systems map an image to a prediction, answer, or explanation; a recent evaluation further shows that multimodal medical predictions can rely strongly on accompanying text (Buckley et al. 2026). Hosted inference requires transmitting the image, whereas local deployment still couples the decision to a large, opaque model. ScaFE instead asks the LLM to produce an inspectable program without seeing a clinical image; the program then executes deterministically inside the hospital without an LLM at inference time.
2.3
Knowledge-Guided Structured Representations
Predictive performance depends on data representation and on avoiding common evaluation pitfalls (Bengio, Courville, and Vincent 2013; Domingos 2012). Concept bottleneck models learn named intermediate concepts (Koh et al. 2020), while knowledge-guided networks inject clinical priors into model training (Xie et al. 2019). Neurosymbolic work integrates neural learning with symbolic knowledge (Garcez and Lamb 2023); separately, prototype methods organize tabular representations around explicit global prototypes (Ye et al. 2024). ScaFE differs in where the representation comes from: it does not learn a concept layer from the same small image cohort or rely on a manually fixed feature set. A web-enabled LLM converts source-traceable clinical concepts into executable measurements, and held-out aggregate feedback iteratively improves the program. This separation provides dataefficient learning, local image processing, and feature-level auditability within one pipeline.
3
Problem Formulation H×W ×3
Scar image classification. Let X ⊂ R denote the space of RGB scar images, where H, W ∈ N are the image height and width, and let Y = {1, . . . , C} denote C diagnostic categories. We are given a labeled development set and a disjoint test set, +Nte Dtrain = {(Ii , yi )}N Dtest = {(Ii , yi )}N i=1 , i=N +1 , (1) where Ii ∈ X and yi ∈ Y. The objective is to learn a predictor from Dtrain that generalizes to Dtest . The test set is excluded from feature construction, classifier fitting, and model selection. Program-based representation. Instead of learning an image representation and a decision function jointly, we decompose the predictor into an executable feature program g and a classifier hg : RKg → Y: g : X → R Kg , fg (I) = hg (g(I)). (2) The program-dependent dimension Kg allows different candidates to encode different sets of visual measurements. We call g executable if it terminates under the prescribed resource limit and returns a deterministic, finite, fixed-length numeric vector for every valid input. Let Gexec denote the set of executable candidates. Validation-guided program search. We partition the development data into fitting and validation subsets, Dtrain = Dfit ∪˙ Dval . (3) Let TrainRF denote the fixed RF training procedure. For a candidate g, the downstream classifier is fitted only on the structured training pairs hg = TrainRF {(g(Ii ), yi ) : (Ii , yi ) ∈ Dfit } . (4) Given a scalar validation metric M, our goal is to search for g ∗ = arg max M(hg ◦ g; Dval ). (5) g∈Gexec
ScaFE performs this search by asking an LLM to generate and revise feature programs. The LLM observes online clinical evidence, candidate source code, and aggregate validation feedback, but not raw images or patient-level records.
4
The ScaFE Framework
ScaFE converts online clinical evidence into executable image features through three stages: web-grounded candidate generation, validation-guided local refinement, and final model construction with a Random Forest (RF). Figure 1 summarizes the separation between LLM knowledge transfer, local feature execution, and downstream prediction.
4.1
Overview
ScaFE runs T rounds. At round t ∈ {1, . . . , T }, the LLM retrieves public clinical evidence and proposes M feature programs. Each program executes locally, and a fixed RF trained on Dfit evaluates it on Dval . Only aggregate diagnostics return to the LLM to guide the next repair. After round T , ScaFE selects one program, refits the RF on all development data, and evaluates the frozen pipeline on Dtest . The LLM may search public literature, but generated code has no network, file-system, label, filename, or metadata access. Raw clinical images remain local throughout the search.
4.2
Web-Grounded Feature Program Generation
Autonomous literature search. Let LΘ denote a fixed LLM configuration indexed by Θ, and let B denote an online search operator that maps a set of queries to retrieved evidence records. ScaFE does not provide a predetermined reference collection. Instead, the prompt describes the diagnostic task and permits LΘ to formulate its own query set Qt . At round t, the retrieved evidence set is Rt = B(Qt ).
(6)
The search prioritizes peer-reviewed articles, official guidelines, and primary descriptions of criteria such as VSS and POSAS (Baryza and Baryza 1995; Draaijers et al. 2004). Each evidence record stores bibliographic identifiers, access time, a supporting passage, and the supported clinical concept. A record enters Rt only if its identifier resolves and its passage is recoverable. Queries, verification outcomes, and records are archived with the code. Listing 1 in Section B of the Supplementary Document gives the exact contract. Candidate program generation. Let Pt denote the prompt used at round t, and let St−1 denote the collection of aggregate feedback records from the previous round. The first round uses the generation prompt Pgen and evidence R1 ; later rounds additionally receive the previous programs and feedback. We write this process as (m) Gt = LΘ Pt , Rt , Gt−1 , St−1 , Gt = {gt }M m=1 , (7) where P1 = Pgen and G0 = S0 = ∅. Each candidate implements (m)
gt
(I) = [ϕt,m,1 (I), . . . , ϕt,m,Kt,m (I)]⊤ ∈ RKt,m . (8)
Here Kt,m = Kg(m) is the candidate’s output dimension and t ϕt,m,j is its jth scalar feature. Alongside the code, the LLM returns the name, definition, clinical rationale, and online source record for every output dimension. Listing 2 in Section B of the Supplementary Document gives the complete generation template and output contract.
4.3
Validation-Guided Iterative Refinement
Local execution check. Each program is parsed and run under fixed time and memory budgets. Valid code uses only approved libraries and returns a deterministic, finite, fixedlength vector. Failures yield a sanitized category and message; the triggering image and intermediate values remain local. Before evaluation, dimensions lacking a valid source record and recoverable passage are removed, and candidates with none are discarded. Thus every downstream feature passes the evidence gate. These checks audit reliability and traceability; held-out evaluation measures utility. Candidate evaluation. For A ∈ {fit, val, train}, let NA = |DA |. Applying an executable program g row-wise gives g XA = [g(Ii )⊤ ](Ii ,yi )∈DA ∈ RNA ×Kg , (9) YA = [yi ](Ii ,yi )∈DA ∈ Y NA .
context Use verified sources and link each feature to supporting evidence. … Measure observable color, texture, and boundary. … Return deterministic local code; no data leakage… question Generate executable features for keloid vs. hypertrophic scar classification.
User
Program
LLM
Prompt
Data (Image)
Local feature execution
hypertrop hic scar
Color
Texture
Morphological
Clinical Composite
…
…
…
…
keloid
or Features
Machine learning algorithm
Prediction
Figure 1: Overview of ScaFE. An LLM turns source-traceable clinical concepts into deterministic feature programs that process local images into interpretable features. A fixed RF evaluates candidates, and only aggregate diagnostics return for refinement. The selected program and RF are frozen for prediction; raw images never enter the LLM context. g We fit hg on (Xfit , Yfit ) and predict the validation samples. The class-wise results are summarized by X g Cab = 1[yi = a ∧ hg (g(Ii )) = b] , (10)
level feedback record is St (g) = {status, error, C g , BAccval (g), g g shape(Xfit ), shape(Xval ),
(13)
g
rinvalid , rconstant , ψ̄ }.
(Ii ,yi )∈Dval g where 1[·] is the indicator function and, for a, b ∈ Y, Cab
(11)
The round-level collection is St = {St (g) : g ∈ Gt }. Only these aggregate records are exposed to the LLM; patientlevel features, predictions, and SHAP values remain local. The machine-readable fields are reproduced in Listing 4 of the Supplementary Document.
Aggregate feature feedback. Classification counts alone do not indicate whether a program produces unusable or ignored features. ScaFE therefore also reports the featurematrix shapes, the fractions of non-finite and constant dimensions, and global feature importance. We compute SHAP values (Lundberg and Lee 2017) for the validation predictions and aggregate feature j as
Program refinement. The LLM repairs execution failures, preserves supported features with useful validation contributions, and revises weak features; it may search again for a missing clinical concept. Listing 3 of the Supplementary Document gives the exact template. We fix M , T , the split in Eq. (3), and the RF before search. Because Dval is queried repeatedly, it is an optimization set; only the sealed test set estimates generalization.
counts validation examples from class a predicted as class b. The scalar comparison metric in Eq. (5) is balanced accuracy, C
BAccval (g) =
ψ̄jg =
1
g Ccc 1 X . PC g C c=1 b=1 Ccb
N C val X X
Nval C i=1 c=1
g |ψijc |,
4.4 (12)
g where ψijc is the attribution of feature j ∈ {1, . . . , Kg } g C to class c for validation sample i. Let C g = (Cab )a,b=1 , K
g ψ̄ g = (ψ̄jg )j=1 , and let rinvalid and rconstant be the fractions of invalid and constant output dimensions. The candidate-
Random Forest as the Downstream Learner
Rationale. A feature program produces a low-dimensional table of heterogeneous color, texture, and morphology measurements rather than a spatial tensor. RF accommodates mixed scales without standardization, captures nonlinear thresholds and interactions, and reduces tree variance through bootstrap aggregation and feature subsampling (Breiman 2001).
Algorithm 1: Validation-Guided ScaFE Require: Dtrain , sealed Dtest , LLM LΘ , web tool B, rounds T , candidates M Ensure: Feature program g ∗ and RF classifier h∗ 1: Split Dtrain into Dfit , Dval ; set G0 , S0 ← ∅ 2: for t = 1 to T do 3: Qt , Rt ← Search(LΘ , B, Pt , St−1 ) 4: Gt ← Generate(Pt , Rt , Gt−1 , St−1 , M ) 5: St ← {Evaluate(g, Dfit , Dval ) : g ∈ Gt } 6: end for 7: Select g ∗ from GTexec by validation BAcc g∗ 8: Fit h∗ ← TrainRF(Xtrain , Ytrain ); evaluate once on Dtest ∗ ∗ 9: return g , h
Role in ScaFE. ScaFE fixes one RF configuration across candidates and rounds, so validation changes primarily reflect the representation rather than classifier tuning. RF serves as both search-time evaluator and final learner. A bounded search budget and sealed test set, not the RF itself, separate search from evaluation.
4.5
Final Model Construction and Inference
Let GTexec denote the executable candidates in the last round. We select
g ∗ = arg max BAccval (g), exec g∈GT
(14)
breaking exact ties by the smaller feature dimension. We then extract features from the complete development set and refit the RF: g∗ h∗ = TrainRF(Xtrain , Ytrain ). (15) For a new image I, ScaFE predicts ŷ = h∗ (g ∗ (I)). Both g ∗ and h∗ are frozen before evaluating Dtest . Algorithm 1 summarizes the complete procedure.
5
Experiments
Our evaluation is organized around four questions. RQ1 asks whether ScaFE generalizes to an unseen clinical site better than classical features, supervised transfer learning, visual foundation models, and direct VLM inference. RQ2 identifies which parts of the iterative search account for any gain and whether RF is an appropriate downstream learner. RQ3 examines data efficiency and robustness to acquisition shifts. RQ4 tests whether the generated programs are executable, evidence-grounded, stable, interpretable, and computationally practical.
5.1
Experimental Protocol
Multi-institutional cohort. We study binary classification of keloids (KD) and hypertrophic scars (HS) on 600 deidentified clinical photographs collected at three hospitals. Each site contributes 200 images, with 100 KD and 100 HS cases. The cohort contains 600 unique patients (200 per site; one image per patient). Duplicate and near-duplicate images are removed before any split is formed. Reference diagnoses are taken from the clinical record and independently verified by two board-certified plastic surgeons with 8 and 11 years of experience; disagreements are resolved by a third specialist
with 15 years of experience. Inter-rater agreement before adjudication is Cohen’s κ = 0.86. Images were acquired from January 2020 to December 2025 using smartphones (72%) and DSLR cameras (28%). Inclusion requires an adult patient, a record-confirmed KD or HS diagnosis, and a focused pre-treatment photograph; images with an obscured lesion, postoperative dressing, or unresolved duplicate are excluded. The retrospective, de-identified study was approved by the ethics committees of the participating institutions with a waiver of additional consent; identifying approval numbers are omitted for blind review. The clinicians establish reference labels only and do not design features, inspect search feedback, or participate in ScaFE. Table 1 in Section A.1 of the Supplementary Document gives the per-site composition and further reports exclusions, demographic coverage, missingness, and patient/duplicate leakage checks. None of these attributes is used as input. Nested cross-site evaluation. We use three leave-one-siteout folds. In each fold, one hospital (200 images) is sealed as the test set and the other two hospitals (400 images) form the development set. The latter is divided patient-wise into 80% fitting and 20% validation data, stratified by site and diagnosis. The fitting subset trains classifiers, whereas the validation subset supplies all program-refinement, hyperparameterselection, and early-stopping signals. ScaFE is rerun from scratch within every outer fold. The held-out site is opened only after its feature program, classifier, and all thresholds have been frozen. Every method uses the same partitions; neither site identifiers nor test outcomes are available to generated code or to the LLM. Metrics and statistical analysis. The primary endpoint is site-macro balanced accuracy (BAcc): BAcc is computed separately on Sites A–C and then averaged without weighting by site size. We also report macro-F1, AUROC, sensitivity, specificity, Brier score, and expected calibration error. Stochastic procedures are repeated with five prespecified seeds. We report their mean and a 95% confidence interval obtained from 10,000 paired, patient-level bootstrap resamples within each site. Pairwise improvements over all baselines use the same resamples and Holm correction. Sitelevel results and calibration curves are retained rather than reporting only a single cross-site average. Because there are only three hospitals, these intervals quantify patient sampling within the observed sites, not uncertainty over the wider population of hospitals. Implementation details. Images are stripped of metadata, orientation-corrected, and resized with aspect-ratiopreserving padding to 512×512 for ScaFE; neural models use their native input resolution. No manual crop, lesion mask, or clinical metadata is supplied. We generate M = 4 programs in each of T = 3 rounds with temperature 0.2 using gpt-4.1-2025-04-14 and web search (accessed July 15–20, 2026). Every candidate p uses the same RF: 500 trees, depth 8, minimum leaf size 3, Kg features per split for a Kg -dimensional program, class balancing, and a fixed seed. Table 2 in Section A.2 of the Supplementary Document lists the complete selection budgets, software, model revisions,
Table 1: Patient-level leave-one-site-out performance. Site A–C columns report held-out BAcc (%); Avg. is their unweighted mean. Macro-F1, AUROC, calibration, and confidence intervals appear in Section A.4 of the Supplementary Document. Method
Site A
Site B
Site C
Avg.
Handcrafted + RF ResNet-18 EfficientNet-B0 DINOv3 + linear probe Derm Foundation + probe BiomedCLIP + linear probe Local VLM-direct
62.0 65.0 66.0 70.0 71.5 72.5 67.0
60.5 63.0 64.0 67.5 68.5 69.5 65.5
61.5 64.5 65.5 68.5 69.0 71.0 66.0
61.3 64.2 65.2 68.7 69.7 71.0 66.2
ScaFE (ours)
82.5
80.0
80.5
81.0
augmentation policy, and archived run artifacts.
5.2
Baselines and Controls
We compare ScaFE with complementary, non-strawman alternatives. Clinical handcrafted+RF uses prespecified color histograms, GLCM/LBP texture, and threshold-based morphology motivated by VSS and POSAS, without an LLM. ResNet-18 (He et al. 2016) and EfficientNet-B0 (Tan and Le 2019) are initialized from ImageNet weights and fine-tuned with validation-based early stopping. Frozen DINOv3 (Siméoni et al. 2025), Derm Foundation (Kiraly et al. 2024), and BiomedCLIP (Zhang et al. 2025) encoders use ℓ2 -regularized linear probes. Local VLM-direct prompts MedGemma-1.5-4B-IT (Sellergren et al. 2026) inside the hospital environment and normalizes the sequence likelihoods of the two allowed labels. Each learned baseline receives at most M T = 12 validation configurations. Two negative controls test arbitrary frozen ResNet features and 1,000 within-site development-label permutations with preserved class counts. Both use held-out labels only for evaluation and cannot encode sample identity.
5.3
Cross-Site Generalization (RQ1)
Table 1 is the primary comparison. The three site columns make a result auditable: a high average cannot conceal failure at one hospital. Tables 6 and 7 in Section A.4 of the Supplementary Document supply the complete secondary metrics, confidence intervals, and leakage controls. Results. ScaFE obtains 81.0% site-macro BAcc, 10.0 points above the strongest baseline, BiomedCLIP (71.0%). The margins are 10.0, 10.5, and 9.5 points on Sites A–C, respectively, so the average does not conceal a site-specific failure. The paired patient-level bootstrap gives a 95% CI of 7.2–12.8 points for the improvement, with Holm-adjusted p < 0.001. ScaFE also reaches 80.7% macro-F1 and 87.9% AUROC, while reducing ECE to 0.043. The random-weight encoder and label-permutation controls remain near chance at 51.2% and 50.0% BAcc, respectively, arguing against an obvious sample-identity or split-leakage explanation for the gain.
5.4
What Produces the Gain? (RQ2)
Search-mechanism ablations. Figure 2(a) separates clinical grounding, mechanical code repair, prediction-level feedback, and feature-level attribution. The one-shot variant generates M candidates only at t = 1. “Runtime feedback only” passes syntax and contract failures but withholds confusion counts, BAcc, and SHAP. “No SHAP” retains confusion counts and BAcc. In the shuffled-feedback control, aggregate records are randomly reassigned among candidates within a round, preserving their marginal values but destroying their semantic link to the program being revised. Downstream learner. To test rather than assume the RF choice, we freeze each selected feature program and replace only its classifier. Logistic regression and RBF-SVM test linear and smooth nonlinear boundaries, a single decision tree exposes the variance of unbagged trees, and XGBoost (Chen and Guestrin 2016) provides a strong boosting comparator. All hyperparameters are selected on the inner validation set under the same budget; Table 5 in Section A.3 of the Supplementary Document gives all five classifier results. Results. Iterative refinement raises the executableprogram rate from 66.7% to 95.0% and BAcc from 73.8% to 81.0%. Removing online literature, prediction feedback, or SHAP loses 4.9, 6.6, and 2.8 points, respectively. Shuffling candidate feedback reduces BAcc to 72.9%, slightly below the one-shot variant, while retaining a high execution rate; this pattern isolates the value of semantically aligned feedback from that of additional LLM calls. RF has the highest BAcc and, across the same five synthesis trajectories, the lowest pipeline SD. XGBoost is close in mean BAcc (80.4%) but less stable (1.9 versus 1.1 points).
5.5
Data Efficiency and Robustness (RQ3)
We subsample 10%, 25%, 50%, and 75% of each development set while preserving site, class, and patient groups. Crucially, the entire ScaFE search is rerun inside each subsample; discovering a program with all 400 development images and only refitting its classifier on fewer images would leak information into the low-data curve. Figure 3 reports site-macro BAcc. We additionally evaluate frozen models under brightness and contrast shifts, JPEG compression, and mild Gaussian blur, with perturbation strengths fixed before test evaluation. Section A.5 of the Supplementary Document reports this acquisition-shift audit without elevating it to a separate robustness claim. Where demographic and body-site metadata are sufficiently complete, we audit coverage and missingness without using those attributes as inputs. Results. ScaFE obtains BAcc values of 72.0%, 75.8%, 78.4%, 79.8%, and 81.0% as the development fraction increases. Its lead over the strongest baseline is largest at 10% (11.8 points) and remains 10.0 points with all development data. Under the prespecified acquisition shifts, every model degrades; ScaFE retains the highest absolute BAcc, but its relative degradation is comparable to the baselines and we therefore make no separate robustness claim.
(a)
(b) 81.0
Rate / validation BAcc (%)
Full ScaFE 78.2
No SHAP 76.1
No literature 74.4
Runtime only
73.8
One-shot
90
80
Executable Contract-valid Evidence-verified Validation BAcc
70
72.9
Shuffled feedback
60 72
74
76
78
80
82
1
2
Held-out BAcc (%)
3
Refinement round
Figure 2: Search-loop analysis. (a) Held-out BAcc for ScaFE and its ablations. (b) Executability, contract validity, evidence verification, and validation BAcc across refinement rounds. Test BAcc is computed only after freezing the pipeline; exact values appear in Section A.3 of the Supplementary Document.
Site-macro BAcc (%)
+10.0 pp
80
minutes per fold, whereas local extraction takes 34.7 ms/image. A second LLM remains comparable at 79.6 ± 1.5% BAcc (Supplementary Section A.6).
+11.8 pp
70
6
60
50 10
25
50
75
100
Development data used per outer fold (%) ScaFE BiomedCLIP Derm Foundation
DINOv3 ResNet-18 Handcrafted+RF
Figure 3: Data efficiency under stratified subsampling. Fractions 10/25/50/75/100 use approximately 40/100/200/300/400 development images per fold; shading marks ScaFE’s gain over the strongest baseline, BiomedCLIP.
5.6
Reliability, Interpretability, and Cost (RQ4)
RQ4 audits execution, source traceability, feature faithfulness, and cross-trajectory stability. Domain experts assess the retained source–feature mappings and visual proxies and confirm their clinical reasonableness. Supplementary Section A.6 reports this audit, second-LLM replication, and cost. Results. Executability rises from 66.7% to 95.0%; candidate evidence passes at 91.7%, while final coverage is 100%. Top-three SHAP deletion loses 8.6 points versus 2.1 for random deletion (p < 0.001), and trajectories agree on 91.4 ± 1.8% of predictions. Search takes 12 calls and 18.6
Conclusion
ScaFE converts LLM clinical knowledge into locally executed, evidence-grounded feature programs. On 600 photographs from three hospitals, it achieves 81.0% site-macro BAcc—10.0 points above the strongest baseline and 11.8 points ahead with only 10% of the development data. Refinement raises executability from 66.7% to 95.0% and the candidate evidence-pass rate from 70.0% to 91.7%; final filtering ensures 100% coverage. Ablations attribute the gains to grounding, semantic feedback, and feature importance. ScaFE thus separates knowledge transfer from clinical-data access while retaining a data-efficient, auditable, and reproducible prediction pipeline.
7
Limitations
ScaFE supports, not replaces, diagnosis. This retrospective two-class, three-hospital study requires prospective clinician validation and omits nonvisual findings. Deployment requires pinned artifacts, sandboxing, and governance.
References Baryza, M. J.; and Baryza, G. A. 1995. The Vancouver Scar Scale: an administration tool and its interrater reliability. Journal of Burn Care & Rehabilitation, 16(5): 535–538. Bayat, A.; McGrouther, D. A.; and Ferguson, M. W. J. 2003. Skin scarring. BMJ, 326(7380): 88–92. Bengio, Y.; Courville, A.; and Vincent, P. 2013. Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8): 1798–1828.
Berman, B.; Maderal, A.; and Raphael, B. 2017. Keloids and hypertrophic scars: pathophysiology, classification, and treatment. Dermatologic Surgery, 43: S3–S18. Breiman, L. 2001. Random forests. Machine learning, 45(1): 5–32. Buckley, T. A.; Diao, J. A.; Srivastava, C. N.; Brodeur, P. G.; Rajpurkar, P.; Rodman, A.; and Manrai, A. K. 2026. Multimodal Foundation Models Exploit Text to Make Medical Image Predictions. Nature Communications, 17(1): 7475. Chen, T.; and Guestrin, C. 2016. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. Domingos, P. 2012. A few useful things to know about machine learning. Communications of the ACM, 55(10): 78–87. Draaijers, L. J.; Tempelman, F. R. H.; Botman, Y. A. M.; Tuinebreijer, W. E.; Middelkoop, E.; Kreis, R. W.; and van Zuijlen, P. P. M. 2004. The Patient and Observer Scar Assessment Scale: a reliable and feasible tool for scar evaluation. Plastic and Reconstructive Surgery, 113(7): 1960–1965. Esteva, A.; Kuprel, B.; Novoa, R. A.; Ko, J.; Swetter, S. M.; Blau, H. M.; and Thrun, S. 2017. Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639): 115–118. Garcez, A. d.; and Lamb, L. C. 2023. Neurosymbolic AI: The 3rd wave. Artificial Intelligence Review, 56: 12387–12406. He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770–778. Kiraly, A. P.; Baur, S.; Philbrick, K.; Mahvar, F.; Yatziv, L.; Chen, T.; Sterling, B.; George, N.; Jamil, F.; Tang, J.; Bailey, K.; Ahmed, F.; Goel, A.; Ward, A.; Yang, L.; Sellergren, A.; Matias, Y.; Hassidim, A.; Shetty, S.; Golden, D.; Azizi, S.; Steiner, D. F.; Liu, Y.; Thelin, T.; Pilgrim, R.; and Kirmizibayrak, C. 2024. Health AI Developer Foundations. arXiv preprint arXiv:2411.15128. Koh, P. W.; Nguyen, T.; Tang, Y. S.; Mussmann, S.; Pierson, E.; Kim, B.; and Liang, P. 2020. Concept bottleneck models. In International Conference on Machine Learning, 5338– 5348. Li, C.; Wong, C.; Zhang, S.; Usuyama, N.; Liu, H.; Yang, J.; Naumann, T.; Poon, H.; and Gao, J. 2023. LLaVAMed: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. Advances in Neural Information Processing Systems, 36: 28541–28564. Litjens, G.; Kooi, T.; Bejnordi, B. E.; Setio, A. A. A.; Ciompi, F.; Ghafoorian, M.; van der Laak, J. A. W. M.; van Ginneken, B.; and Sánchez, C. I. 2017. A survey on deep learning in medical image analysis. Medical image analysis, 42: 60–88. Liu, Y.; Jain, A.; Eng, C.; Way, D. H.; Lee, K.; Bui, P.; Kanada, K.; de Oliveira Marinho, G.; Gallegos, J.; Gabriele, S.; Gupta, V.; Singh, N.; Natarajan, V.; Hofmann-Wellenhof, R.; Corrado, G. S.; Peng, L. H.; Webster, D. R.; Ai, D.; Huang, S. J.; Liu, Y.; Dunn, R. C.; and Coz, D. 2020. A deep
learning system for differential diagnosis of skin diseases. Nature Medicine, 26(6): 900–908. Lundberg, S. M.; and Lee, S.-I. 2017. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc. Nori, H.; King, N.; McKinney, S. M.; Carignan, D.; and Horvitz, E. 2023. Capabilities of GPT-4 on medical challenge problems. arXiv preprint arXiv:2303.13375. Rao, V. M.; Hla, M.; Moor, M.; Adithan, S.; Kwak, S.; Topol, E. J.; and Rajpurkar, P. 2025. Multimodal Generative AI for Medical Image Interpretation. Nature, 639(8056): 888–896. Sellergren, A.; Gao, C.; Mahvar, F.; Kohlberger, T.; Jamil, F.; Traverse, M.; Tono, A.; Sadjad, B.; Yang, L.; Lau, C.; Yatziv, L.; Chen, T.; Sterling, B.; Philbrick, K.; Tiwari, R.; Liu, Y.; Jajoo, M.; Sankarapu, C.; Vispute, S.; Purandare, H.; Mishra, A. B.; Schmidgall, S.; Tu, T.; Palepu, A.; Park, C.; Strother, T.; Thapa, R.; Cheng, Y.; Singh, P.; Black, K.; Matias, Y.; Chou, K.; Hassidim, A.; Goel, K.; Barral, J.; Warkentin, T.; Shetty, S.; Webster, D.; Virmani, S.; Steiner, D. F.; Kirmizibayrak, C.; and Golden, D. 2026. MedGemma 1.5 Technical Report. arXiv preprint arXiv:2604.05081. Shiraishi, M.; Miyamoto, S.; Takeishi, H.; Kurita, D.; Furuse, K.; Ohba, J.; Moriwaki, Y.; Fujisawa, K.; and Okazaki, M. 2024. The potential of chat-based artificial intelligence models in differentiating between keloid and hypertrophic scars: a pilot study. Aesthetic Plastic Surgery, 48(24): 5367–5372. Siméoni, O.; Vo, H. V.; Seitzer, M.; Baldassarre, F.; Oquab, M.; Jose, C.; Khalidov, V.; Szafraniec, M.; Yi, S.; Ramamonjisoa, M.; Massa, F.; Haziza, D.; Wehrstedt, L.; Wang, J.; Darcet, T.; Moutakanni, T.; Sentana, L.; Roberts, C.; Vedaldi, A.; Tolan, J.; Brandt, J.; Couprie, C.; Mairal, J.; Jégou, H.; Labatut, P.; and Bojanowski, P. 2025. DINOv3. arXiv preprint arXiv:2508.10104. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S. S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; Payne, P.; Seneviratne, M.; Gamble, P.; Kelly, C.; Babiker, A.; Schärli, N.; Chowdhery, A.; Mansfield, P.; Demner-Fushman, D.; Agüera y Arcas, B.; Webster, D.; Corrado, G. S.; Matias, Y.; Chou, K.; Gottweis, J.; Tomašev, N.; Liu, Y.; Rajkomar, A.; Barral, J.; Semturs, C.; Karthikesalingam, A.; and Natarajan, V. 2023. Large language models encode clinical knowledge. Nature, 620(7972): 172–180. Tan, M.; and Le, Q. V. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning, 6105–6114. Xie, Y.; Xia, Y.; Zhang, J.; Song, Y.; Feng, D.; Fulham, M.; and Cai, W. 2019. Knowledge-based collaborative deep learning for benign-malignant lung nodule classification on chest CT. IEEE Transactions on Medical Imaging, 38(4): 991–1004. Yan, S.; Hu, M.; Jiang, Y.; Li, X.; Fei, H.; Tschandl, P.; Kittler, H.; and Ge, Z. 2025a. Derm1M: A Million-Scale Vision– Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 12681–12690.
Yan, S.; Yu, Z.; Primiero, C.; Vico-Alonso, C.; Wang, Z.; Yang, L.; Tschandl, P.; Hu, M.; Ju, L.; Tan, G.; Tang, V.; Ng, A. B.; Powell, D.; Bonnington, P.; See, S.; Magnaterra, E.; Ferguson, P.; Nguyen, J.; Guitera, P.; Banuls, J.; Janda, M.; Mar, V.; Kittler, H.; Soyer, H. P.; and Ge, Z. 2025b. A Multimodal Vision Foundation Model for Clinical Dermatology. Nature Medicine, 31(8): 2691–2702. Ye, H.; Fan, W.; Song, X.; Zheng, S.; Zhao, H.; Guo, D.; and Chang, Y. 2024. PTaRL: Prototype-based Tabular Representation Learning via Space Calibration. In The Twelfth International Conference on Learning Representations. Zhang, S.; Xu, Y.; Usuyama, N.; Xu, H.; Bagga, J.; Tinn, R.; Preston, S.; Rao, R.; Wei, M.; Valluri, N.; Wong, C.; Tupini, A.; Wang, Y.; Mazzola, M.; Shukla, S.; Liden, L.; Gao, J.; Crabtree, A.; Piening, B.; Bifulco, C.; Lungren, M. P.; Naumann, T.; Wang, S.; and Poon, H. 2025. A Multimodal Biomedical Foundation Model Trained from Fifteen Million Image–Text Pairs. NEJM AI, 2(1). Zhou, J.; He, X.; Sun, L.; Xu, J.; Chen, X.; Chu, Y.; Zhou, L.; Liao, X.; Zhang, B.; Afvari, S.; and Gao, X. 2024. Pre-trained multimodal large language model enhances dermatological diagnosis using SkinGPT-4. Nature Communications, 15(1): 5649.