Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?
arXiv:2605.27082v1 [cs.AI] 26 May 2026
Qingyuan Zeng1 Ziyang Chen1 Pengxiang Cai1 Zixin Guan2 1 1 1 Anglin Liu Lang Qin Xinyao LAI Jintai Chen1 1 The Hong Kong University of Science and Technology (Guangzhou) 2 Guangzhou University of Chinese Medicine
Abstract Biomedical discovery often requires reconciling broad biomedical knowledge with specific experimental or clinical data. While background knowledge suggests useful biological mechanisms, it is typically too general to map directly onto concrete dataset variables. Conversely, data-driven patterns are often dataset-specific and lack mechanistic insight, creating a gap between abstract principles and concrete evidence. We formulate the bridging of this missing link as knowledge contextualization: transforming broad biomedical knowledge into evidence-supported, scenario-grounded propositions, with the goal of allowing domain experts to efficiently inspect, replay, and validate candidate hypotheses. To achieve this, we propose SCENE, a bi-level multi-agent framework that implements knowledge contextualization as an iterative search process. The upper level translates broad knowledge into search directions and maps them onto the dataset schema. The lower level then executes these grounded directions using multi-objective optimization to find concrete propositions that balance evidential strength with sufficient data support. A feedback loop between the two levels progressively refines these search directions. We evaluate SCENE in two distinct settings: discovering patient subgroups with heterogeneous treatment benefits in clinical trial scenarios, and identifying context-specific biological responses in LINCS L1000 studies. In clinical trial scenarios, SCENE outperforms existing baselines by discovering highly specific, strongly supported subgroups. Furthermore, SCENE uniquely enables the discovery of perturbational contexts with strong target-response matching and high positive rates in LINCS L1000 studies. These results demonstrate that SCENE effectively bridges the gap between broad knowledge and scenario-specific evidence, providing traceable and inspectable hypotheses for follow-up validation.
1
Introduction
Biomedical discovery often requires reconciling broad biomedical knowledge with specific experimental or clinical datasets. While background knowledge suggests useful biological mechanisms, it is typically too generic to directly map onto the variables and constraints of a concrete setting. Conversely, the evidence collected in a specific dataset is rich in context but rarely explains how it connects to reusable biomedical principles [1, 2, 3]. This mismatch creates a practical obstacle: it becomes difficult to prioritize which data patterns deserve follow-up, because statistically strong patterns may lack biological rationale, while plausible biomedical ideas may remain too abstract to test in the available data. Existing methods address parts of this problem but typically approach it from a single direction, failing to bridge the gap between broad biomedical knowledge and concrete data. On the knowledge side, integration methods inject external knowledge into prompts, representations, or retrieval modules [4, 5]. However, this approach leaves the knowledge largely static; it does not specify how broad Preprint.
biomedical knowledge should be grounded under the specific variables, endpoints, and validity constraints of a given dataset. On the evidence side, data-driven algorithms, such as tabular models and subgroup discovery, search for statistical patterns within a single dataset [6, 7, 8, 9, 10, 11, 2, 3]. These discovered patterns are often weakly tied to prior biomedical knowledge, making them difficult to validate or reuse as traceable candidate hypotheses. Consequently, there remains a lack of an explicit conversion mechanism that translates broad knowledge into data-grounded, scenario-specific propositions. We therefore formulate this missing conversion step as knowledge contextualization: transforming broad biomedical knowledge into evidence-supported, scenario-grounded propositions. The goal is to produce traceable candidate hypotheses that domain experts can efficiently inspect, replay, and validate in follow-up analysis [12, 13]. This problem is challenging for three reasons. First, broad biomedical concepts do not directly match the concrete variables and measurements available in a specific dataset. For example, a concept such as metabolic dysregulation may need to be instantiated using available measurements such as BMI, fasting glucose, or triglycerides, depending on the dataset schema. Second, the same biomedical principle requires different grounding across distinct settings; for example, it may manifest as a patient subgroup in a clinical trial, or as a context-specific cellular response in a perturbational assay [6, 3]. Third, even after mapping these concepts to concrete variables, the space of valid candidate propositions is immense, making exhaustive search computationally impractical [14, 10]. Addressing these challenges requires an explicit contextualization process that can ground biomedical priors, evaluate them against available evidence, and return interpretable results; Figure 1 illustrates the intended form of such scenario-grounded propositions. To address these challenges, we propose SCENE, a bi-level multi-agent framework that implements knowledge contextualization as an iterative search process [15, 16, 17]. At the upper level, SCENE translates broad biomedical knowledge into testable search directions and formulates a schemagrounding plan. At the lower level, it executes this plan using multi-objective optimization to find concrete propositions that balance evidential strength with sufficient data support [18]. Crucially, the two levels form a bidirectional closed loop: the upper level constrains the search space based on biomedical priors, while the scenario-specific evidence from the lower level guides the upper level to refine or pivot the search directions. By explicitly linking knowledge-guided generation with data-driven evaluation, SCENE provides a traceable mechanism to turn abstract knowledge into scenario-grounded propositions. • We formulate knowledge contextualization as a standalone problem in biomedical discovery: transforming broad biomedical knowledge into evidence-supported, scenario-grounded propositions to produce traceable, inspectable candidate hypotheses. • We propose SCENE, a bi-level multi-agent framework that forms a bidirectional closed loop between upper-level direction planning and lower-level multi-objective evolutionary search: the upper level guides the lower level by constraining and prioritizing its executable search space, while the lower level feeds scenario-specific evidence back to the upper level to refine subsequent directions. • We demonstrate SCENE’s generality across two distinct biomedical settings: clinical-trial treatment-benefit subgroup discovery and LINCS L1000 perturbational context discovery. In clinical trials, SCENE outperforms existing baselines by discovering subgroups with higher risk reduction and stronger data support. In L1000 studies, it uniquely enables the discovery of perturbational contexts with strong target-response matching and high positive rates. We further show that SCENE propositions can improve downstream few-shot classification.
2
Related Work
Knowledge integration in biomedicine. A line of work incorporates external biomedical knowledge into machine learning systems through knowledge graphs, structured retrieval, or knowledge injection. Recent efforts such as HypKG [5] contextualize biomedical knowledge graphs with patientspecific EHR information, while benchmarks such as FEDMEKI [4] study how medical knowledge can be injected into foundation models under privacy constraints. However, these approaches primarily operate at the level of representation learning or model training, aiming to improve downstream prediction or generation, rather than directly producing interpretable, scenario-grounded propositions in the current setting. 2
Executable subgroup rule
Direction Rule Broad prior knowledge SCENE contextualizes
B. L1000 proposition
A. Clinical-trial proposition
Broad prior + Scenario evidence
Context-specific perturbation response Direction
inflammation heterogeneity IF high inflammation + poor glucose control
MEK-pathway response
IF lung epithelial + MEK inhibitor + high dose / 24 h
Finding
THEN stronger MEK-like expression response
THEN treatment benefit appears larger
Evidence
Matched patients show a larger observed treatment-benefit contrast.
Evidence
Matched L1000 profiles respond more strongly than comparable background profiles.
Why it matters
helps identify patients who may benefit more from a given treatment.
Why it matters
helps identify assay settings where target-directed responses are stronger.
Perturbation context space all patients
Scenario data + schema
matched subgroup
bad outcome: control bad outcome: treatment
MEK inhibitor PI3K inhibitor EGFR inhibitor ...
Lung low high epithelial 48 h 24 h
Neuron
risk reduction
Immune
lower is better
...
low
high
48 h 24 h low
high
48 h 24 h
low
high
48 h 24 h low
high
48 h 24 h low
high
48 h 24 h
high
...
48 h 24 h
...
low
low
high
...
48 h 24 h
...
low
high
...
48 h 24 h
...
Matched profiles
Comparable background
Target-response match
Stronger Target-response pattern
Matched Background
Figure 1: Examples of scenario-grounded propositions. SCENE contextualizes a broad biomedical prior with scenario evidence and schema constraints to produce an executable, evidence-supported rule or finding. In clinical trials, the proposition identifies a subgroup in which treatment benefit appears larger; in LINCS L1000, it identifies a cell-perturbagen-dose-time context with a stronger target-response pattern than comparable background profiles.
Single-scenario subgroup discovery and treatment effect heterogeneity. Another related literature studies treatment effect heterogeneity and subgroup discovery, especially in clinical trials. Classical methods such as Virtual Twins [6] and SIDES [7] search for treatment-responsive subgroups through predicted individual responses or recursive partitioning, while honest trees [8] and causal forests [9] extend this line through tree- or forest-based estimation of heterogeneous treatment effects. More recent approaches, such as CURLS [10] and Learning Subgroups with Maximum Treatment Effects without Causal Heuristics [11] cast subgroup discovery as rule learning or structured optimization. Despite their differences, these methods share a common perspective: they treat the problem as data-only effect estimation, partitioning, or rule search within a given dataset. They neither use broad biomedical knowledge as an explicit hypothesis space, nor address how such knowledge can be converted into scenario-grounded propositions. Our position. SCENE is not defined around subgroup discovery, treatment effect estimation, or any single downstream task. Instead, we study knowledge contextualization itself: how broad biomedical knowledge can be proposed, grounded, tested, and refined in a concrete scenario. Under this view, subgroup rules in clinical trials are only one form of scenario-grounded proposition, alongside contextbounded findings in perturbational transcriptomic studies. SCENE therefore addresses a broader problem than existing single-scenario methods, while still producing interpretable, evidence-grounded rules or findings in each task.
3
Methodology
3.1
Task and Overview
For a biomedical scenario s, SCENE receives three inputs: biomedical prior knowledge K(s) , scenario evidence D(s) = (X (s) , Y (s) , C (s) ), and a machine-readable schema S (s) . Here X (s) denotes observable features, Y (s) denotes outcomes or response variables used for evidence evaluation, and C (s) denotes task context. The schema S (s) specifies observable variables, admissible windows or operators, forbidden fields for candidate construction, and validity constraints. Biomedical prior knowledge K(s) is treated as a finite set of pre-discovery records, such as statements derived from publications, trial records, or study resources. These records provide source-traceable starting points for search directions. Detailed record formats, manifest fields, and provenance requirements are given in Appendices A.1 and A.7. 3
Bi-level Multi-Agent System
Inputs
Upper-Level Direction Planning Reusable
Biomedical Prior Knowledge
Planner State
H1
Task
Prior
H2
Feedback update
Memory
Schema Match
Direction Proposal
Concept
H3
Search Guidance
Feature
Target preference
Window
Bounds restriction Grounded Rules / Findings
Composite
Control update
Gap
Specific
Scenario Evidence
Lower-Level Knowledge Discovery Feedback
A Rule Population
Trade-off
Evolution Loop
E Execution Log x₂
···
xₚ
Rule 2
Evidence / Provenance Records
C Feature Proposal Dynamic
D Evaluation Kept
Rule 1
B Mutation ×
Data Schema
x₁
Searching
Crossover Tuning Injection
x1 > t x2 ≤ t x1 & x3
Failure
Scenario-Grounded Propositions
Discarded
f2
Virtual
Pareto Select f1: effect f2: support Pareto Front
f1
Figure 2: SCENE (Scenario-Contextualized Evidence and Knowledge Engine) artifact flow. Biomedical prior knowledge, scenario evidence, and a data schema define the discovery scenario. The upper level converts broad knowledge into bounded search directions and schema-grounding plans that constrain and prioritize the lower-level executable search space. The lower level searches candidate rules or findings with multi-objective effect/support optimization and returns scenario-specific feedback to refine later directions. Exported outputs are scenario-grounded propositions, each pairing a direction, an executable rule or finding, and an evidence record with provenance. The target output is a set of scenario-grounded propositions P (s) = {pi }m i=1 ,
pi = (di , qi , ωi ).
(1)
Here di is a search direction, qi is a grounded rule or finding in the current scenario, and ωi is an evidence record containing effect, support, diagnostics, source links, and provenance. A search direction specifies the biomedical concept or relation to be explored. The grounded object qi is the executable scenario-specific instance produced by search: in clinical subgroup discovery it may be a subgroup rule, while in L1000 perturbational discovery it may be a cell–perturbagen–dose– time context. The evidence record ωi makes the proposition inspectable and replayable. Thus, the reportable unit is the entire proposition (di , qi , ωi ), not a rule or finding alone. SCENE implements knowledge contextualization as a closed-loop search process. The upper level translates prior knowledge and task context into bounded search directions and schema-grounding plans, which constrain and prioritize the executable search space for the lower level. The lower level searches admissible candidate rules or findings with multi-objective optimization over effect and support, and returns scenario-specific feedback for later direction refinement. The final outputs are proposition-level artifacts that combine a direction, an executable grounded object, and an evidence record. Figure 2 summarizes the operational flow. Given an active direction at round t, the schema-grounding plan Gt induces an admissible candidate family Qt . A scenario adapter A(s) specifies the executable candidate language, validity checks, support and effect functions, and reporting boundary. Details of the candidate grammar, grounding operators, reporting checks, and scenario-specific estimators are deferred to Appendices A.6, A.7, and A.8. The round-level protocol is shown in Algorithm 1. 3.2
Upper-Level Direction Planning
The upper level decides what should be searched next. It receives prior records, task context, schema summaries, and discovery-visible feedback from previous rounds. Its output is not a biomedical 4
Algorithm 1 SCENE round-level search protocol Require: prior records K(s) , scenario evidence D(s) , schema S (s) , adapter A(s) , round budget T Ensure: reportable propositions P (s) 1: initialize compact state m0 , feedback f0 , controls η10 , and P (s) ← ∅ 2: for t ← 1 to T do 3: (dt , ηt ) ← P LAN(K(s) , C (s) , S (s) , mt−1 , ft−1 , ηt0 ) ▷ propose a bounded search direction 4: Gt ← G ROUND(dt , S (s) , ηt ) ▷ map the direction to schema-visible objects 5: (Gt , ηt ) ← VALIDATE(Gt , ηt , A(s) ) ▷ remove invalid or leakage-risk objects 6: Ct,0 ← S EED(Gt , ηt ) ▷ initialize candidate rules or findings 7: (Ft , λt ) ← E VOLVE(dt , Gt , Ct,0 , D(s) , A(s) , ηt ) ▷ search and retain a Pareto frontier 8: (ft , mt , η̄t+1 ) ← U PDATE F EEDBACK(Ft , λt , dt , mt−1 ) ▷ update next-round search using discovery-visible logs 9: P (s) ← P (s) ∪ R EPORT(Ft , dt , D(s) , A(s) ) ▷ export propositions passing reporting checks 0 10: ηt+1 ← η̄t+1 11: end for 12: return P (s) claim, but a bounded search direction together with search guidance: (dt , ηt ) = Planθ (K(s) , C (s) , S (s) , mt−1 , ft−1 , ηt0 ),
ηt = (πt , Ht , ρt , βt ).
(2)
Here πt is the operator mixture, Ht is the generation budget, ρt is the support floor, and βt contains bounded preferences such as target feature families, admissible ranges, avoid lists, or virtual/dynamic feature toggles. The symbol θ denotes fixed role instructions, decoding settings, controller bounds, and replay policy, all specified before the discovery run. Operationally, P LAN performs three functions. First, it selects or revises a search direction from prior knowledge, task context, and prior feedback. Second, it grounds the direction to schema-visible feature families, windows, or admissible composites. Third, it converts the grounded direction into search guidance for the lower level. All role outputs, including directions and grounding plans, are treated as structured proposals and must pass deterministic validation before entering executable search. Detailed role contracts and controller policies are given in Appendices A.2 and A.3. 3.3
Lower-Level Knowledge Discovery
Conditioned on the active direction and its grounding plan, the lower level searches over an admissible candidate family Qt . Candidates are executable scenario-specific objects: subgroup rules in clinical settings and context-bounded findings in L1000 settings. The search starts from grounded seeds and diversity-oriented admissible candidates, then applies variation operators such as mutation, crossover, tuning, and injection. When raw observables are insufficient to instantiate a direction, virtual or dynamic features may be proposed, but they enter the search space only after deterministic construction, provenance recording, and validity checks. (s)
Each candidate q ∈ Qt is evaluated by the scenario adapter, which returns an effect objective Jeff (q), (s) a support objective Jsup (q), and discovery-visible diagnostics V (s) (q). The effect objective measures the configured scenario contrast, such as a treatment-benefit contrast in clinical subgroup discovery or a target-response contrast in L1000 discovery. The support objective prevents the search from favoring highly specific but poorly supported fragments. These objectives define discovery-side ranking under the frozen split and manifest. Pareto dominance and frontier membership use only effect and support: (s) (s) Jeff (q ′ ) ≥ Jeff (q), (s) ′ (s) Ft = q ∈ Qt : ∄q ′ ∈ Qt such that Jsup . (3) (q ) ≥ J (q), sup and at least one inequality is strict The retained frontier is both a candidate set and a feedback object. A coherent frontier indicates that the current direction admits repeated grounded instantiations, whereas empty, low-support, invalid, or redundant frontiers expose failure modes. Diagnostics can annotate instability or break exact ties after Pareto ranking and diversity filtering, but they cannot create or remove Pareto 5
Table 1: Held-out clinical subgroup discovery. Frozen Best-1 rules are replayed over 50 paired splits. Task columns report held-out bad-outcome ARR, B/P denote breast-cancer/PCOS tasks and BL/TR denote baseline/trajectory views. P-ARR is population-adjusted ARR, U-Supp. is usable support, and D-Cons. is train-to-holdout direction consistency. Higher is better for all displayed metrics. Method
B-D-BL B-D-TR P-M-BL P-M-TR P-A-FPG P-A-AUC P-ARR ↑ U-Supp. ↑ D-Cons. ↑
Virtual Twins SIDES Causal Forest Honest Causal Tree CURLS MaxTE Subgroups SCENE (ours)
0.023 0.007 0.066 0.057 0.015 0.076 0.086
-0.027 0.008 0.005 0.037 0.003 0.012 0.095
0.036 0.195 0.104 0.232 0.025 0.234 0.441
0.197 -0.005 0.035 0.042 0.210 0.121 0.439
0.078 0.006 0.060 0.053 0.118 0.117 0.269
0.004 -0.058 0.130 0.087 0.103 0.132 0.270
0.014 0.010 0.016 0.032 0.010 0.054 0.093
58.7 26.7 56.4 70.0 26.3 38.5 84.7
51.5 74.3 51.4 71.7 45.7 68.6 100.0
Method
Conn. ↑
Strong-Pos. ↑
AUPRC ↑
U-Supp. ↑
Gap ↓
Minimal w/o Dynamic w/o Pareto w/o Upper w/o Virtual SCENE (ours)
0.166 0.171 0.184 0.177 0.170 0.188
27.6 28.5 18.8 35.1 24.2 41.5
0.260 0.273 0.273 0.271 0.275 0.268
33.5 34.2 30.0 29.5 34.3 31.4
0.111 0.068 0.081 0.106 0.087 0.063
High Benefit Rule Rate 66%
SCENE
40%
CURLs MaxTE Subgroups
37%
Virtual Twins
36%
Causal Forest
33%
Honest Causal Tree
32% 26%
SIDES 0
10
20
30
40
50
60
70
Rules with >=10 pp benefit effect (%)
Figure 3: Fraction of held-out Best1 clinical rules whose treatmentbenefit effect is at least 10 percentage points.
Table 2: L1000 component ablations. Minimal is the jointablation baseline. Conn. is connectivity margin, Strong-Pos. is the strong-response rate, U-Supp. is usable support, and Gap is the normalized train-to-holdout connectivity gap.
dominance. Evolutionary operators and default hyperparameters are detailed in Appendix A.5; executable grounding operators are described in Appendix A.6. 3.4
Closed-Loop Interaction and Reporting Boundary
The upper and lower levels interact through validated artifacts and frontier logs rather than unconstrained dialogue. At each round, the upper level sends a direction, a grounding plan, and bounded search guidance to shape the lower-level search; the lower level returns a retained frontier Ft and an execution log λt for the next planning step. The feedback update summarizes failure modes, support–effect trade-offs, alignment gaps, diversity, and diagnostic contrast: (ft , mt , η̄t+1 ) = Φfb (Ft , λt , dt , mt−1 ).
(4)
These signals help later rounds preserve, narrow, broaden, or replace search directions. Export is stricter than feedback. A candidate may influence later search if it is discovery-visible, but it becomes a reported proposition only if it lies on the retained frontier and passes validity, support, provenance, leakage, and diagnostic checks. Final held-out or test evidence is used only for postdiscovery audit and is never fed back into the same run to revise directions, rules, findings, or search controls. This boundary distinguishes SCENE from a pipeline that simply combines language-model suggestions with evolutionary rule mining: SCENE constrains language-model proposals through schema validation, evaluates grounded candidates through explicit effect/support objectives, and exports only proposition-level outputs with evidence records and provenance. Detailed split-lineage and reporting schemas are given in Appendix A.7.
4
Experiments
4.1
Experimental Setup
We evaluate SCENE on two discovery settings: clinical subgroup discovery and perturbational response discovery, and one downstream knowledge-reuse setting. The clinical benchmark uses patientlevel tables linked to two registered trials, with NCT identifiers used only for registry and protocol 6
Rule: BMI > 25.4 · C-peptide > 1.05 · HOMA-beta > 156.3
80
prior methods
50
75
full cohort
Adiposity
No screening
waist >= 80 cm
Apply SCENE rule
Benefit gap (pp)
SCENE
60
IR HOMA-IR >= 2.5
40
Triglycerides
20
TG >= 1.7 mmol/L
0
BMI
+ C-pep
0
+ HOMA-beta
25
100
Subgroup members satisfying risk criterion (%)
Rule condition applied
Figure 4: Concrete clinical meaning of a SCENE subgroup rule. (Left) The displayed SCENE rule increases the favorable outcome treatment–control gap. (Right) Rule-level clinical-risk audit. Gray densities summarize valid top-ranked rules from prior methods replayed on the same cohort; blue densities summarize sensitivity-only local threshold perturbations around the SCENE rule. linkage rather than patient-level rows: NCT00174655 from Project Data Sphere and NCT02491333 from the Dryad release [19, 20, 21, 22]. We selected the trials by pre-analysis feasibility and coverage criteria: randomized arms, patient-level covariates, interpretable endpoints, sufficient support for repeated subgroup replay, and complementary biomedical settings. We compare against subgroup and heterogeneous treatment-effect baselines [6, 7, 8, 9, 10, 11]. Across 50 paired splits, each clinical run searches the training split, exports one prioritized rule, and replays it on held-out patients for six task frames: B-D-BL, B-D-TR, P-M-BL, P-M-TR, P-A-FPG, P-A-AUC (Appendix A.10). The primary clinical metric is held-out bad-outcome ARR (absolute risk reduction, larger is better). P-ARR records population-adjusted benefit, U-Supp. records whether held-out replay has sufficient usable support, and D-Cons. records train-to-holdout benefit-direction preservation, definitions are in Appendix A.11. In the perturbational benchmark, based on CMap/LINCS L1000 perturbational signatures [2, 3, 23, 24], each episode fixes a target mechanism and performs discovery on the training split only. We report Conn. as the held-out connectivity margin between rule-covered signatures and background, Strong-Pos. as the strong response rate, AUPRC [25, 26] as target mechanism retrieval, U-Supp. as usable held-out support, and Gap as the normalized train–holdout connectivity gap; Appendix A.12 gives the replay contract. Finally, the reuse setting compares downstream few-shot classification with C0 (few-shot row serialization), C1 (C0 plus generic knowledge context), and C2 (C1 plus SCENE knowledge cards), using final-test AUPRC, MacroF1 [27], and accuracy to measure ranking, class-balanced F1, and correctness. Held-out outcomes and connectivity are never used to select rules or revise prompts; downstream test labels are never used to choose SCENE cards. 4.2
Clinical Trial Subgroup Discovery
Table 1 evaluates held-out ARRbad after the exported Best-1 rule is frozen. SCENE is best on all six clinical frames, with gains spanning both table regimes: on baseline static frames (B-D-BL, P-M-BL), it reaches 0.086/0.441 versus 0.076/0.234 for existing SOTA methods; on trajectoryenhanced dynamic frames (B-D-TR, P-M-TR), it reaches 0.095/0.439 versus 0.037/0.210. SCENE also leads on the P-A endpoint frames, indicating that the advantage is not tied to one endpoint or table view. SCENE also outperforms existing SOTA methods in aggregate diagnostics : improves P-ARR, U-Supp., and D-Cons. from 0.054, 70.0%, and 74.3% to 0.093, 84.7%, and 100.0%. Thus, SCENE handles both static and dynamic clinical tables while discovering support-aware propositions whose treatment-benefit direction is preserved after export to held-out patients. 4.3
L1000 Perturbational Context Discovery
Table 2 evaluates SCENE in the L1000 mechanism discovery setting by using component ablations to examine how different modules affect perturbational context discovery. SCENE (ours) achieves the most favorable overall trade-off across the target response effect, the strong positive rate, and stability between the training and holdout sets. Removing Pareto selection (w/o Pareto) primarily collapses the high-response tail. Removing upper-level planning (w/o Upper) produces a different failure mode: the strong positive rate falls and the gap increases, suggesting weaker transfer from discovery to the holdout set. Furthermore, removing virtual features (w/o Virtual) trades effect strength for coverage. These results indicate that the components are complementary rather than redundant. Specifically, 7
Search Improves Benefit Discovery 70
Benefit Effect (pp)
60
Trajectory Best-so-far Average
Scenario
Mode
AUPRC
MacroF1
ACC
Clinical Trial
C0 C1 C2
0.5877 0.6140 0.6747
0.4408 0.4745 0.5291
0.4898 0.5510 0.6939
L1000
C0 C1 C2
0.3582 0.3725 0.4113
0.5536 0.5914 0.6528
0.5600 0.6100 0.7200
50 40 30 20 10 0 2.5
5.0
7.5
10.0
12.5
15.0
17.5
20.0
Generation
Figure 5: Discovery-side SCENE search trajectory showing median best-so-far and mean candidate benefit across generations.
Table 3: Few-shot reuse of SCENE propositions. C0: rowserialized few-shot baseline; C1: C0 + generic knowledge context; C2: C1 + validation-selected, row-conditioned SCENE cards. Metrics are means over 10 exemplar seeds.
upper-level direction planning stabilizes discovery across the boundary between training and holdout sets, Pareto selection preserves strong holdout responders, and virtual or dynamic grounding expands the search space without replacing the need for evidence-based frontier selection. 4.4
Contextual Knowledge Reuse
Table 3 evaluates whether propositions discovered by SCENE can be reused as auxiliary knowledge for downstream few-shot tabular classification. We convert the exported propositions from the discovery runs into compact knowledge cards and inject them into an in-context prediction setting, while keeping the row serialization, exemplar budget, data split, and LLM decoding policy fixed across conditions. The comparison includes ordinary few-shot row serialization (C0), C0 plus generic scene context (C1), and C1 plus row-conditioned SCENE knowledge cards (C2). In the clinical scene, the downstream label is a row-level proxy for treatment-benefit evidence under the task endpoint; in the L1000 scene, the downstream label is the target-mechanism positive class defined for the selected perturbational task. C2 is best on all final test metrics in both scenarios, improving clinical accuracy from 0.5510 to 0.6939 and L1000 accuracy from 0.6100 to 0.7200 over C1. These results suggest that SCENE propositions provide reusable scenario-conditioned information beyond generic task descriptions, while serving as downstream utility evidence rather than independent biomedical validation. Seed-level summaries and prompt artifacts are audited in Appendix A.13. 4.5
Search, Replay, and Efficiency Diagnostics
To complement the aggregate results, we further examine representative propositions, replay behavior, and efficiency diagnostics. Figure 4 shows that a metabolic P-A-AUC rule raises the favorable outcome treatment–control gap and remains enriched for adiposity, insulin resistance, and elevated triglycerides relative to the rule distributions induced by prior methods in Table 1, suggesting a stable, task-relevant endocrine–metabolic phenotype rather than a clinically unstructured output. Figure 6 replays exported rules for RPS6, MYC, and AURKB. Rule-covered profiles consistently shift toward higher Distil SS and TAS (L1000 signature-strength and transcriptional-activity diagnostics [3]), indicating that SCENE identifies coherent response contexts rather than isolated score gains. These cases illustrate SCENE’s output boundary: schema-valid contexts for follow-up analysis, not standalone clinical rules or mechanism validation. Figure 3 shows that 66% of SCENE’s reported clinical rules exceed a 10-percentage-point benefit threshold. Figure 5 shows that both the median best-so-far and the average discovery-side benefit among candidates improve across generations. Together with Table 3, these analyses support the view that SCENE outputs compact, scenario-grounded propositions that remain reusable outside the original discovery run. Figure 7 compares five third-party API backends on the P-M-TR task; benefit is held-out bad-outcome ARR (pp), averaged over the top-10 discovery-ranked rules from each run. Qwen3.5-9B provides the most favorable runtime–benefit trade-off, achieving the lowest runtime while maintaining high held-out benefit. 4.6
Further Analysis
Appendix B provides additional post-hoc audits beyond aggregate benchmark scores. Appendix B.1 analyzes clinical benefit–support–variability geometry, showing that SCENE remains high-benefit, 8
RPS6
MYC
AURKB
0.8
TAS
0.6
0.4
0.2 5
10
15
Distil SS
5
10
Distil SS
15
5
10
Distil SS
15
Figure 6: Held-out L1000 rule replay on 3 target mechanisms. Gray hexbins summarize the held-out background, colored points denote signatures satisfying the selected rule, and arrows indicate the median shift from the background to rule-covered signatures.
Held-out benefit (pp)
18 DeepSeek-V3.2
Backend
17
Qwen3.5-9B [28] GLM-4.5-Air [29] Qwen3-8B [30] Qwen3.5-27B [31] DeepSeek-V3.2 [32]
Qwen3.5-27B
16 Qwen3.5-9B
15
GLM-4.5-Air
14
Qwen3-8B
13 6
7
8
9
10
11
Avg. run (min) Avg. gen. (s) 6.2 6.4 7.8 8.6 12.0
19.8 26.9 25.8 31.2 43.7
12
Wall-Clock Time (min)
Figure 7: Benefit–runtime trade-off and runtime cost across LLM backends. (Left) Each point is one third-party API backend under the same discovery protocol. (Right) Avg. run is mean wall-clock time per discovery run; Avg. gen. is mean wall-clock time per evolutionary generation within a run.
high-support, and low-variability. Appendix B.2 compares reported rules on predefined clinical axes, showing stronger alignment with task-relevant phenotypes than prior-method rule vocabularies. Appendix B.3 audits L1000 inspectability, showing improved response-quality diagnostics and consistent TAS–Distil-SS displacement. Appendix B.4 evaluates replay reliability, showing that benefit directions are largely preserved and magnitude stability improves with stricter support floors. Appendix B.5 analyzes rule composition and provenance, showing that SCENE combines interpretable feature families rather than returning opaque rule strings. Appendix B.6 expands the L1000 ablation, showing complementary contributions from upper-level planning, Pareto selection, and virtual or dynamic grounding. Appendix B.7 tracks generated features through evidence gates, showing that generated features are validated before entering reported propositions. Appendix B.8 audits grounding traces, showing a visible path from broad knowledge directions to executable propositions. Together, these analyses support the view that SCENE outputs robust, inspectable, and provenance-aware scenario-grounded propositions rather than isolated high-scoring rules.
5
Conclusion
In this work, we studied knowledge contextualization as a biomedical discovery problem: converting broad prior knowledge into evidence-supported, scenario-grounded propositions. We proposed SCENE, a bi-level multi-agent framework that couples direction planning, schema grounding, and multi-objective evidence search to produce traceable propositions rather than isolated data patterns. Across clinical-trial subgroup discovery and LINCS L1000 perturbational discovery, SCENE consistently identifies stronger, better-supported, and inspectable propositions than prior baselines, and further improves downstream few-shot classification. These results suggest that broad biomedical knowledge can be operationalized as auditable, scenario-grounded hypotheses for follow-up validation. 9
References [1] Deborah A. Zarin, Tony Tse, Rebecca J. Williams, Robert M. Califf, and Nicholas C. Ide. The ClinicalTrials.gov results database–update and key issues. The New England Journal of Medicine, 364(9):852–860, 2011. [2] Justin Lamb, Emily D. Crawford, David Peck, Joshua W. Modell, Irene C. Blat, Matthew J. Wrobel, Jim Lerner, Jean-Philippe Brunet, Aravind Subramanian, Kenneth N. Ross, Michael Reich, Haley Hieronymus, Guo Wei, Scott A. Armstrong, Stephen J. Haggarty, Paul A. Clemons, Ru Wei, Steven A. Carr, Eric S. Lander, and Todd R. Golub. The connectivity map: Using geneexpression signatures to connect small molecules, genes, and disease. Science, 313(5795):1929– 1935, 2006. [3] Aravind Subramanian, Rajiv Narayan, Steven M. Corsello, David D. Peck, Ted E. Natoli, Xiaodong Lu, Joshua Gould, John F. Davis, Andrew A. Tubelli, Jacob K. Asiedu, David L. Lahr, Jodi E. Hirschman, Zihan Liu, Melanie Donahue, Bina Julian, Mariya Khan, David Wadden, Ian C. Smith, Daniel Lam, Arthur Liberzon, Courtney Toder, Mukta Bagul, Marek Orzechowski, Oana M. Enache, Federica Piccioni, Sarah A. Johnson, Nicholas J. Lyons, Alice H. Berger, Alykhan F. Shamji, Angela N. Brooks, Anita Vrcic, Corey Flynn, Julien Rosains, David Y. Takeda, Roger Hu, David Davison, Justin Lamb, Kristin Ardlie, Larson Hogstrom, Peyton Greenside, Nathanael S. Gray, Paul A. Clemons, Serena Silver, Xiaoyun Wu, Wen-Ning Zhao, Wilfried Read-Button, Xiaohua Wu, Stephen J. Haggarty, Lucienne V. Ronco, Jesse S. Boehm, Stuart L. Schreiber, John G. Doench, Joshua A. Bittker, David E. Root, Bang Wong, and Todd R. Golub. A next generation connectivity map: L1000 platform and the first 1,000,000 profiles. Cell, 171(6):1437–1452.e17, 2017. [4] Jiaqi Wang, Xiaochen Wang, Lingjuan Lyu, Jinghui Chen, and Fenglong Ma. FEDMEKI: A benchmark for scaling medical foundation models via federated knowledge injection. In Proceedings of the Conference on Neural Information Processing Systems, 2024. Datasets and Benchmarks Track. [5] Yuzhang Xie, Xu Han, Ran Xu, Xiao Hu, Jiaying Lu, and Carl Yang. HypKG: Hypergraph-based knowledge graph contextualization for precision healthcare. In Proceedings of the International Semantic Web Conference, 2025. [6] Jared C. Foster, Jeremy M. G. Taylor, and Stephen J. Ruberg. Subgroup identification from randomized clinical trial data. Statistics in Medicine, 30(24):2867–2880, 2011. [7] Ilya Lipkovich, Alex Dmitrienko, Jonathan Denne, and Gregory Enas. Subgroup identification based on differential effect search—a recursive partitioning method for establishing response to treatment in patient subpopulations. Statistics in Medicine, 30(21):2601–2621, 2011. [8] Susan Athey and Guido Imbens. Recursive partitioning for heterogeneous causal effects. Proceedings of the National Academy of Sciences, 113(27):7353–7360, 2016. [9] Susan Athey, Julie Tibshirani, and Stefan Wager. Generalized random forests. The Annals of Statistics, 47(2):1148–1178, 2019. [10] Jiehui Zhou, Linxiao Yang, Xingyu Liu, Xinyue Gu, Liang Sun, and Wei Chen. CURLS: Causal rule learning for subgroups with significant treatment effect. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 4619–4630, 2024. [11] Lincen Yang, Zhong Li, Matthijs van Leeuwen, and Saber Salehkaleybar. Learning subgroups with maximum treatment effects without causal heuristics. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 27565–27573, 2026. [12] Payal Chandak, Kexin Huang, and Marinka Zitnik. Building a knowledge graph to enable precision medicine. Scientific Data, 10(1):67, 2023. [13] Luc Moreau, Paolo Missier, Khalid Belhajjame, Reza B’Far, James Cheney, Sam Coppens, Stephen Cresswell, Yolanda Gil, Paul Groth, Graham Klyne, Timothy Lebo, James McCusker, Simon Miles, Jim Myers, Satya Sahoo, and Curt Tilmes. The PROV data model and abstract syntax notation. W3C recommendation, World Wide Web Consortium (W3C), 2013. 10
[14] Ilya Lipkovich, David Svensson, Bohdana Ratitch, and Alex Dmitrienko. Modern approaches for evaluating treatment effect heterogeneity from clinical trials and observational data. Statistics in Medicine, 43(22):4388–4436, 2024. [15] Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. MetaGPT: Meta programming for a multiagent collaborative framework. In Proceedings of the International Conference on Learning Representations, 2024. [16] Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. AutoGen: Enabling next-gen LLM applications via multi-agent conversations. In Proceedings of the Conference on Language Modeling, 2024. [17] Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, Bingnan Zheng, Bang Liu, Yuyu Luo, and Chenglin Wu. AFlow: Automating agentic workflow generation. In Proceedings of the International Conference on Learning Representations, 2025. [18] Kalyanmoy Deb, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. A fast and elitist multiobjective genetic algorithm: NSGA-II. IEEE Transactions on Evolutionary Computation, 6(2):182–197, 2002. [19] Project Data Sphere. Project Data Sphere: Data access. https://data.projectdatasphe re.org/projectdatasphere/html/access. Accessed 2026-05-05. [20] Angela K. Green, Katherine E. Reeder-Hayes, Robert W. Corty, Ethan Basch, Matthew I. Milowsky, Stacie B. Dusetzina, Antonia V. Bennett, and William A. Wood. The project data sphere initiative: Accelerating cancer research by sharing data. The Oncologist, 20(5):464–e20, 2015. [21] Qidan Wen, Min Hu, Maohua Lai, Juan Li, Zhenxing Hu, Kewei Quan, Jia Liu, Hua Liu, Yanbing Meng, Suling Wang, Xiaohui Wen, Chuyi Yu, Shuna Li, Shiya Huang, Yanhua Zheng, Han Lin, Xingyan Liang, Lingjing Lu, Zhefen Mai, Chunren Zhang, Taixiang Wu, Ernest H. Y. Ng, Elisabet Stener-Victorin, and Hongxia Ma. Data: Effect of acupuncture and metformin on insulin sensitivity in women with polycystic ovary syndrome and insulin resistance: A three-armed randomized controlled trial. Dryad Dataset, 2021. [22] Qidan Wen, Min Hu, Maohua Lai, Juan Li, Zhenxing Hu, Kewei Quan, Jia Liu, Hua Liu, Yanbing Meng, Suling Wang, Xiaohui Wen, Chuyi Yu, Shuna Li, Shiya Huang, Yanhua Zheng, Han Lin, Xingyan Liang, Lingjing Lu, Zhefen Mai, Chunren Zhang, Taixiang Wu, Ernest H. Y. Ng, Elisabet Stener-Victorin, and Hongxia Ma. Effect of acupuncture and metformin on insulin sensitivity in women with polycystic ovary syndrome and insulin resistance: A three-armed randomized controlled trial. Human Reproduction, 37(3):542–552, 2022. [23] Alexandra B. Keenan, Sherry L. Jenkins, Kathleen M. Jagodnik, Simon Koplev, Edward He, Denis Torre, Zichen Wang, Anders B. Dohlman, Moshe C. Silverstein, Alexander Lachmann, Maxim V. Kuleshov, Avi Ma’ayan, Vasileios Stathias, Raymond Terryn, Daniel Cooper, Michele Forlin, Amar Koleti, Dusica Vidovic, Caty Chung, Stephan C. Schürer, et al. The library of integrated network-based cellular signatures NIH program: System-level cataloging of human cells response to perturbations. Cell Systems, 6(1):13–24, 2018. [24] Amar Koleti, Raymond Terryn, Vasileios Stathias, Caty Chung, Daniel J. Cooper, John P. Turner, Dusica Vidovic, Michele Forlin, Tanya T. Kelley, Alessandro D’Urso, et al. Data portal for the library of integrated network-based cellular signatures (LINCS) program: Integrated access to diverse large-scale cellular perturbation response data. Nucleic Acids Research, 46(D1):D558–D566, 2018. [25] Jesse Davis and Mark Goadrich. The relationship between precision-recall and ROC curves. In Proceedings of the International Conference on Machine Learning, pages 233–240, 2006. 11
[26] Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3):e0118432, 2015. [27] Marina Sokolova and Guy Lapalme. A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4):427–437, 2009. [28] Qwen Team. Qwen/Qwen3.5-9B model card. https://huggingface.co/Qwen/Qwen3.5-9 B, 2026. Accessed 2026-05-07. [29] Z.AI. GLM-4.5-Air model card. https://huggingface.co/zai-org/GLM-4.5-Air, 2025. Accessed 2026-05-07. [30] Qwen Team. Qwen3: Think deeper, act faster. https://qwenlm.github.io/blog/qwen3/, 2025. Accessed 2026-05-07. [31] Qwen Team. Qwen/Qwen3.5-27B model card. https://huggingface.co/Qwen/Qwen3. 5-27B, 2026. Accessed 2026-05-07. [32] DeepSeek-AI. DeepSeek-V3.2 release. https://api-docs.deepseek.com/news/news25 1201, 2025. Accessed 2026-05-07. [33] ClinicalTrials.gov. ClinicalTrials.gov record for NCT00174655. https://clinicaltria ls.gov/study/NCT00174655, 2011. ClinicalTrials.gov identifier NCT00174655; accessed 2026-04-28. [34] ClinicalTrials.gov. ClinicalTrials.gov record for NCT02491333. https://clinicaltria ls.gov/study/NCT02491333, 2022. ClinicalTrials.gov identifier NCT02491333; accessed 2026-04-28. [35] Peter M. Rothwell. Subgroup analysis in randomised controlled trials: Importance, indications, and interpretation. The Lancet, 365(9454):176–186, 2005. [36] David M. Kent, Jessica K. Paulus, David van Klaveren, Ralph D’Agostino, Steven Goodman, Rodney Hayward, John P. A. Ioannidis, Bray Patrick-Lake, Sally Morton, Michael Pencina, Gowri Raman, Joseph S. Ross, Harry P. Selker, Ravi Varadhan, Andrew Vickers, John B. Wong, and Ewout W. Steyerberg. The predictive approaches to treatment effect heterogeneity (PATH) statement. Annals of Internal Medicine, 172(1):35–45, 2020. [37] Arthur Liberzon, Chet Birger, Helga Thorvaldsdóttir, Mahmoud Ghandi, Jill P. Mesirov, and Pablo Tamayo. The molecular signatures database hallmark gene set collection. Cell Systems, 1(6):417–425, 2015. [38] Qwen Team. Qwen3.5: Towards native multimodal agents. https://qwen.ai/blog?id=qw en3.5, 2026. Accessed 2026-05-03. [39] Tom Fawcett. An introduction to ROC analysis. Pattern Recognition Letters, 27(8):861–874, 2006. [40] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Proceedings of the Conference on Neural Information Processing Systems, pages 1877–1901, 2020.
12
Overview In this appendix, we provide additional details and supplementary analyses for SCENE. The appendix is organized as follows. Appendix A presents extended methodological details, including scenario formalization, explicit LLM role contracts, the evolutionary grounding procedure, comprehensive experimental setups, and detailed implementation protocols for all main tables. Appendix B provides supplementary post-hoc audits (e.g., clinical effect-support geometry, L1000 inspectability, and generated feature lifecycles). Finally, Appendix C and Appendix D discuss limitations and broader impacts, respectively.
Reproducibility To support reproducibility, we provide the concrete settings used in all reported experiments, including the scenario-specific data schemas, LLM role configurations, multi-objective evolutionary hyperparameters, threshold grids, and validation split boundaries. The source code, processed datasets (for both clinical-trial cohorts and L1000 perturbational signatures), and supplementary agent artifacts are included in the supplemental material submitted with this paper. We will also release the code publicly upon acceptance. In addition, we give explicit definitions of the operational boundaries used by SCENE, such as the role instruction contracts, deterministic reporting predicates, and scenario-specific estimators, so that the bi-level knowledge contextualization procedure can be reproduced as faithfully as possible.
Contents 1
Introduction
1
2
Related Work
2
3
Methodology
3
3.1
Task and Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
3
3.2
Upper-Level Direction Planning . . . . . . . . . . . . . . . . . . . . . . . . . . .
4
3.3
Lower-Level Knowledge Discovery . . . . . . . . . . . . . . . . . . . . . . . . .
5
3.4
Closed-Loop Interaction and Reporting Boundary . . . . . . . . . . . . . . . . . .
6
4
5
Experiments
6
4.1
Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
6
4.2
Clinical Trial Subgroup Discovery . . . . . . . . . . . . . . . . . . . . . . . . . .
7
4.3
L1000 Perturbational Context Discovery . . . . . . . . . . . . . . . . . . . . . . .
7
4.4
Contextual Knowledge Reuse . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
8
4.5
Search, Replay, and Efficiency Diagnostics . . . . . . . . . . . . . . . . . . . . .
8
4.6
Further Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
8
Conclusion
9
A Additional Method Details
15
A.1 Notation and Scenario Formalization . . . . . . . . . . . . . . . . . . . . . . . . .
15
A.2 Upper-Level Controller and Role Artifacts . . . . . . . . . . . . . . . . . . . . . .
16
A.3 Role Instruction Contracts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
17
13
A.4 Role Artifact Templates and One-Round Trace . . . . . . . . . . . . . . . . . . . .
19
A.5 Evolutionary Grounding Procedure . . . . . . . . . . . . . . . . . . . . . . . . . .
20
A.6 Grounding Operators . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
20
A.7 Reporting, Provenance, and Split Boundaries . . . . . . . . . . . . . . . . . . . .
21
A.8 Scenario-Specific Estimators . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
23
A.9 Experiment Matrix and Reporting Contract . . . . . . . . . . . . . . . . . . . . .
24
A.10 Clinical Task Definitions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
25
A.11 Implementation Details for Table 1 . . . . . . . . . . . . . . . . . . . . . . . . . .
25
A.12 Implementation Details for Table 2 . . . . . . . . . . . . . . . . . . . . . . . . . .
29
A.13 Implementation Details for Table 3 . . . . . . . . . . . . . . . . . . . . . . . . . .
33
A.14 LLM Backend Runtime Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . .
36
B Additional Experiments and Analyses
37
B.1 Clinical Effect–Support–Uncertainty Geometry . . . . . . . . . . . . . . . . . . .
37
B.2 Clinical Contextualization Advantage over Prior Rule Vocabularies . . . . . . . . .
41
B.3 L1000 Proposition Inspectability . . . . . . . . . . . . . . . . . . . . . . . . . . .
42
B.4 Holdout Reliability and Support–Gap Diagnostics . . . . . . . . . . . . . . . . . .
43
B.5 Semantic Rule Composition and Feature Provenance . . . . . . . . . . . . . . . .
44
B.6 L1000 Ablation and Holdout Evidence Diagnostics . . . . . . . . . . . . . . . . .
46
B.7 Evidence-Gated Lifecycle of Generated Features . . . . . . . . . . . . . . . . . .
47
B.8 Grounding Trace from Prior Knowledge to Executable Propositions . . . . . . . .
47
C Limitations
48
D Broader Impacts
49
14
A
Additional Method Details
A.1
Notation and Scenario Formalization
This appendix expands the compressed methodological description in the main text. The main text defines SCENE at the level of the discovery protocol, while this section makes explicit the formal objects that are passed through the protocol. SCENE reports scenario-grounded, hypothesisgenerating propositions with provenance and diagnostics; it does not report confirmed clinical recommendations, externally validated mechanisms, or claims certified by a language model. For a biomedical scenario s, the scenario evidence is written as D(s) = (X (s) , Y (s) , C (s) ),
(5)
where X (s) denotes observable features, Y (s) denotes outcomes or response variables used for evidence evaluation, and C (s) denotes task context. Biomedical prior knowledge is represented as a finite manifest of source-documented records available before discovery: s K(s) = {κa }A a=1 ,
κa = (ida , ua , σa , χa , ψa ).
(6)
Here As is the number of knowledge records, ida is a stable record identifier, ua is a compact knowledge statement, σa records the source descriptor, χa records the scenario scope, and ψa records the transform that made the record available to SCENE. These records provide source-traceable search cues, but they are not treated as conclusions about the current dataset. The exported output is contextualized knowledge, represented as a set of scenario-grounded propositions: P (s) = {pi }m pi = (di , qi , ωi ). (7) i=1 , Here di is a bounded search direction, qi is an executable grounded rule or finding in the current scenario, and ωi is an evidence record containing effect, support, diagnostics, source links, split information, and provenance. The reportable unit is therefore the entire proposition pi , not a rule or finding alone. This convention also defines the output boundary used by Figure 2: grounded rules or findings and evidence records are components of one scenario-grounded proposition rather than two independent outputs. Table 4: Notation used by SCENE. Symbol
Meaning
s D (s) X (s) Y (s) C (s) K(s) κa S (s)
Biomedical scenario or task instance. Scenario evidence, equal to (X (s) , Y (s) , C (s) ). Observable features or measurements available under the scenario schema. Outcomes, responses, or configured evidence variables used for evaluation. Task context, including scenario description, endpoint definition, and discovery constraints. Source-documented prior knowledge records available before discovery. One prior record with identifier, statement, source descriptor, scope, and transform metadata. Machine-readable schema with variable types, admissible windows, forbidden fields, missingness rules, and composition constraints. Active direction and bounded search guidance at round t. Direction-conditioned vocabulary and grounding plan. Candidate rule/finding and admissible candidate family. Scenario-defined support and effect objectives. Pareto frontier and execution log returned by the lower level. Feedback packet and compact working memory. Exported set of scenario-grounded propositions. One reportable proposition: direction, executable grounded object, and evidence record. Evidence and provenance record attached to a reported proposition.
dt , ηt Zt , Gt q, Qt (s) (s) Jsup , Jeff Ft , λt ft , mt P (s) pi = (di , qi , ωi ) ωi
The main text keeps candidate search at the protocol level. Executably, the active direction and (s) validated grounding plan induce a vocabulary Zt and validity function Γt , giving (s)
(s) Qt = {q ∈ H(Zt ) : |q| ≤ L, Γt (q) = 1, Jsup (q) ≥ ρt }.
(8)
In the tabular instantiations used here, the default object in H(Zt ) is a bounded conjunction q(x) =
ℓ ^
I[zkj ▷◁j τj ],
j=1
15
ℓ ≤ L,
(9)
Table 5: Scenario formalization under the common SCENE adapter interface. Component
Clinical-trial scenario
Evidence
Subjects with admissible subgroup covariates, treatment arm, Perturbational profiles or signatures with cell, perturbagen, endpoint, and protocol context. Treatment and outcome fields dose/time, response or connectivity score, and optional are evaluation fields, not subgroup literals. signature-derived attributes. Bounded subgroup conjunction over baseline or explicitly Bounded context conjunction over admissible perturbation, landmarked pre-decision summaries. cell, dose/time, compound, or signature attributes. Type-correct literals, no treatment/outcome/post-outcome Compatible context literals, admissible dose/time windows, leakage, admissible windows, nonempty arms, and support no response-derived features, matched or recurrent support, floor. and support floor. min{ntreat (q), nctrl (q)}. Number of matched or recurrent admissible evidence instances. Configured within-subgroup treatment-benefit contrast. Configured connectivity or response contrast, blocked when a valid blocking design is available. Scenario-grounded proposition containing a subgroup rule, Scenario-grounded proposition containing a perturbational evidence record, and provenance. context finding, evidence record, and provenance.
Candidate q Validity
Support Effect Output
LINCS L1000 perturbational scenario
where literals are drawn from raw schema-visible observables or approved virtual/dynamic operators. In non-clinical settings, the same notation denotes an executable context-bounded finding rather than a clinical subgroup rule. The scenario adapter used by Algorithm 1 can be written as (s) (s) A(s) = O(s) , Γ(s) , Jsup , Jeff , T (s) , Ω(s) , (s)
(10) (s)
where O(s) specifies observable families, Γ(s) validity, Jsup support, Jeff effect, T (s) optional grounding operators, and Ω(s) split, provenance, diagnostic, and output schemas. A.2
Upper-Level Controller and Role Artifacts
The upper level is implemented through typed role artifacts rather than requiring physically separate models. The same language model may instantiate multiple roles, but each role consumes and emits a different typed object. This design is intended to make the contextualization process replayable and auditable: language-model calls may propose directions, grounding maps, or critiques, but they do not directly validate biomedical claims or bypass deterministic controllers. The accepted information flow is direction planning
semantic grounding
(K(s) , C (s) , mt−1 , ft−1 ) −−−−−−−−−→ (dt , ηt ) −−−−−−−−−−→ Gt search
feedback
(11)
−−−→ (Ft , λt ) −−−−→ (ft , mt , η̄t+1 ). The upper level emits search contracts, not final biomedical conclusions. These contracts define what should be grounded next, which schema-visible objects may be used, and how the lower-level search should be bounded. The U PDATE F EEDBACK call in Algorithm 1 is implemented by the deterministic priority rule in Algorithm 2. This rule maps invalidity, support failures, diversity collapse, diagnostic contrast, alignment gaps, and stagnation into bounded next-round actions. The compact feedback statement in the main text is therefore instantiated by a deterministic priority order. Default thresholds are recorded in the manifest, e.g., diagnostic-contrast high/low cutoffs, diversity floor, invalid-rate ceiling, low-support ceiling, and upper-level stagnation patience. These values may be tuned per scenario before discovery, but not by final held-out or test metrics. Language-model calls are validated as JSON-producing proposal functions. A strategist artifact contains direction, knowledge_ids, grounding_cues, target_preference, control_adjustment, and rationale. A proposer artifact contains semantic_matches, grounding_map, candidate_literals, invalid_combinations, seed_candidates, and optional feature_proposals. Critic/refiner artifacts contain frontier summaries, failure modes, stability notes, next-direction hints, memory updates, and bounded control deltas. Extra top-level fields, unknown feature identifiers, unregistered source IDs, non-finite numbers, and claims of validation success are rejected. 16
Table 6: Upper-level role artifacts and fail-closed behavior. Artifact
Required content
Validation or fallback
Direction dt
Search intent, supporting knowledge IDs or model_suggested, grounding cues, target preferences, rationale. Operator mixture, generation budget, support floor, bounds restrictions, avoid list, dynamic-feature toggles. Matched features/windows/composites, valid literal keys, invalid combinations, seed candidates, grounding rationale. Frontier summary, alignment gaps, failure modes, support-effect trade-offs, stability notes. Retained/rejected cues, recent directions, stable families, failure modes, next-round hint.
Unknown source IDs are rejected or marked model_suggested; unsupported cues cannot be reported as prior-supported.
Guidance ηt
Grounding plan Gt
Feedback ft
Memory mt
Numeric controls are clipped to manifest ranges; invalid toggles are ignored and logged.
Unmaterialized or leakage-risk literals are dropped before search.
Missing diagnostics trigger conservative defaults.
Memory is bounded and summarized; it is not converted into external prior knowledge without a new manifest record.
Algorithm 2 UpdateFeedback / FrontierFeedbackUpdate Require: frontier Ft , execution log λt , direction dt , memory mt−1 , manifest thresholds τ Ensure: feedback ft , memory mt , bounded controls η̄t+1 1: compute invalid_rate, low_support_rate, diversity, diagnostic_contrast, alignment_gap, and stagnation from (Ft , λt ) 2: if leakage_flag or schema_violation or invalid_rate > τinvalid then 3: action ← replace 4: else if Ft = ∅ then 5: action ← broaden if one manifest-permitted retry remains; otherwise replace 6: else if low_support_rate > τlow_support or alignment_gap is high then 7: action ← broaden 8: else if diversity < τdiversity or stagnation exceeds patience then 9: action ← narrow 10: else if diagnostic_contrast is high and support is adequate then 11: action ← preserve with conservative control deltas 12: else 13: action ← preserve 14: end if 15: construct ft with the action, frontier summary, failure modes, alignment gaps, and stability notes 16: update mt with retained/rejected cues, stable families, and recent failure modes 17: set η̄t+1 by clipping any control deltas to manifest bounds 18: return (ft , mt , η̄t+1 )
A.3
Role Instruction Contracts
SCENE specifies role instructions as contracts rather than as unconstrained prompts. Each languagemodel call is treated as a JSON-producing proposal function: the model may suggest a direction, grounding, critique, or refinement, but its output has no executable effect until parsed, checked against the frozen manifest, and validated by deterministic controllers. The controller supplies only discoveryvisible context, including source records, schema summaries, frontier feedback, split-visibility flags, and bounded control ranges. It rejects malformed JSON, unknown source identifiers, unknown feature keys, forbidden fields, non-finite values, out-of-range controls, and unsupported biomedical assertions. The following compact contracts summarize the role interfaces used by SCENE; scenario-specific prompts instantiate the placeholders with manifest-visible context, but the executable boundary is the schema and controller rather than the wording of the prompt.
17
Strategist contract role: strategist input: { "scenario_context": "<task and discovery-visible split context>", "knowledge_records": [{"id": "kappa_a", "cue": "..."}], "schema_summary": {"feature_families": ["..."], "forbidden": ["..."]}, "memory": "<bounded prior-round summary>", "frontier_summary": "<recent frontier status and trade-offs>", "control_bounds": "<manifest ranges>" } instructions: - Propose a search direction, not a biomedical conclusion. - Use supplied knowledge IDs only; otherwise mark "model_suggested". - Map concepts to schema-visible grounding cues or mark them unresolved. - Propose a replacement or revised direction when frontier feedback indicates failure. - Do not report effect size, statistical significance, or validation status. output_json: { "direction": "...", "status": "source_supported|model_suggested|feedback_derived", "knowledge_ids": ["kappa_a"], "grounding_cues": ["..."], "target_preference": {"prefer": ["..."], "avoid": ["..."]}, "control_adjustment": [{"param": "...", "delta": 0.0, "reason": "..."}], "rationale": "short evidence-bounded rationale" }
Proposer-grounder contract role: proposer_grounder input: { "direction": "<selected direction>", "schema_summary": "<feature keys, types, windows, grammar>", "allowed_operators": ["mutation", "crossover", "tuning", "injection"], "forbidden_fields": ["..."], "feature_statistics": "<discovery-visible summaries only>" } instructions: - Use only supplied feature keys, windows, and operators. - Do not use treatment, outcome, post-outcome, response-derived, or validation-derived fields as literals. - Translate each cue into matched features, admissible windows, candidate literals, or an explicit invalid/unresolved entry. - Do not claim effect, significance, or validation success. output_json: { "semantic_matches": [{"cue": "...", "features": ["..."], "windows": ["..."]}], "grounding_map": {"cue_id": {"literal_keys": ["..."], "status": "matched"}}, "candidate_literals": [{"feature": "...", "op": ">|<=", "value_source": "grid"}], "seed_candidates": [["literal_key_a", "literal_key_b"]], "invalid_combinations": [{"cue": "...", "reason": "not_schema_visible"}], "feature_proposals": [{"request": "...", "reason": "..."}], "safety": {"uses_forbidden_field": false} }
18
Critic-refiner contract role: critic_refiner input: { "direction": "<current direction>", "frontier_summary": "<Pareto set, support/effect, diversity>", "execution_log": "<invalid, low-support, weak-alignment cases>", "current_controls": "<bounded runtime controls>" } instructions: - Use only frontier and execution-log evidence; do not invent experience. - Summarize failure_modes and stability_notes separately. - Choose one next action: preserve, narrow, broaden, or replace. - Keep control_delta bounded, conservative, and next-round only. - Do not convert feedback into external prior knowledge. output_json: { "frontier_summary": {"status": "coherent|empty|unstable|low_support"}, "failure_modes": ["..."], "stability_notes": ["..."], "action": "preserve|narrow|broaden|replace", "next_direction_hint": "...", "control_delta": [{"param": "...", "delta": 0.0, "reason": "..."}], "memory_update": "bounded summary" }
When feature_proposals are present, the scenario adapter may request one virtual or dynamic feature plan using only registered tables, allowed variables, admissible windows, and leakage-safe operators; invalid or unmaterialized plans are dropped and logged. If executable synthesis is used for an accepted feature plan, it is separately sandboxed and validated before generated features enter search. A.4
Role Artifact Templates and One-Round Trace
The following schematic artifact illustrates the role interface using a simplified metformin-response example. Field values are illustrative rather than experimental observations, and the object is intended to show the replay contract rather than measured support or effect. The example emphasizes that a reported output is a proposition-level object containing a direction, an executable grounded object, and an evidence record. Metformin-response example {
"knowledge_record": { "id": "kappa_metabolic_01", "cue": "metabolic dysregulation may modify metformin response", "scope": "clinical subgroup discovery" }, "direction": { "text": "search for an insulin-resistance or metabolic-risk subgroup", "knowledge_ids": ["kappa_metabolic_01"], "status": "source_supported" }, "semantic_match": { "features": ["baseline_BMI", "fasting_glucose", "HOMA_IR_if_available"], "windows": ["baseline"], "dropped": ["post-treatment glucose"] }, "search_guidance": { "support_floor": "manifest_min_balanced_support", "operators": ["mutation", "crossover", "tuning"],
19
}
"avoid": ["treatment", "outcome", "post-outcome"] }, "grounded_object": { "candidate_key": "baseline_BMI>tau_1 AND fasting_glucose>tau_2", "threshold_source": "predeclared discovery grid" }, "frontier_log": { "frontier_status": "coherent", "failure_modes": ["HOMA_IR unavailable in this schema"] }, "feedback": { "action": "preserve", "memory_update": "metabolic baseline features remain productive" }, "reported_proposition": { "direction": "search for an insulin-resistance or metabolic-risk subgroup", "grounded_object": "baseline_BMI>tau_1 AND fasting_glucose>tau_2", "evidence_record": { "support": "...", "effect": "...", "source_id": "kappa_metabolic_01", "NoLeak": true, "split_id": "...", "diagnostics": "..." } }
A one-round trace is therefore an artifact chain rather than a transcript of agent conversation: κa → dt → (Gt , ηt ) → (Ft , λt ) → ft → pi .
(12)
The knowledge record supplies a provenance-bearing cue; the direction states a bounded search intent; schema match and guidance translate that intent into executable schema objects; the frontier and log expose what the lower level could or could not ground; feedback updates memory and controls; and the final arrow emits a proposition-level object pi = (di , qi , ωi ) only if the candidate also satisfies the reporting predicate in Appendix A.7. This trace is illustrative of the interface contract and does not encode a measured effect size, support count, or empirical success claim. A.5
Evolutionary Grounding Procedure
The grounding loop executes the proposer’s search space under the current direction. It maintains a population Ct,g , an archive Rt , and an append-only log λt . The elements of this population are executable candidates: clinical subgroup rules in clinical tasks and context-bounded findings in (s) (s) L1000 tasks. Pareto dominance uses only (Jeff , Jsup ); pseudo-validation can diagnose instability or break exact ties after Pareto rank and diversity, but it cannot create dominance. Scenario manifests may set stricter floors or disable operators, but final validation metrics cannot alter search defaults for the same run. A.6
Grounding Operators
Some directions refer to concepts that are only partially visible in raw observables. This is the main operational challenge of knowledge contextualization: broad biomedical ideas are rarely expressed in the same vocabulary as the dataset schema. SCENE can therefore expand the grounding vocabulary from Z (0) to Zt = Z (0) ∪ Z (virt) ∪ Z (dyn) .
(13)
Here Z (0) contains raw schema-visible observables, Z (virt) contains approved virtual composites, and Z (dyn) contains approved dynamic or windowed summaries. The executable grammar is fixed 20
Algorithm 3 Direction-conditioned grounding at round t (s)
Require: direction dt , contract Gt = (Zt , Γt , Ct,0 ), evidence D(s) , adapter A(s) , controls ηt = (πt , Ht , ρt , βt ) Ensure: frontier Ft and execution log λt 1: freeze replay configuration ξt ▷ seed, split, adapter, grammar, controls 2: for g ← 0 to Ht do (s) 3: Cˆt,g ← VALIDATE(Ct,g , Γt , ρt ) ▷ deduplicate and filter (s) (s) ˆ 4: Rt,g ← E VALUATE(Ct,g , D , A ) ▷ effect/support diagnostics 5: update archive Rt and log λt with kept, rejected, and invalid cases 6: if g < Ht then 7: Ut,g ← VARY(Cˆt,g , dt , Gt , πt ) ▷ mutation, crossover, tuning, injection ˆ 8: Ct,g+1 ← S ELECT(Ct,g ∪ Ut,g , Rt , ηt ) ▷ rank, diversity, key order 9: end if 10: end for 11: return Ft ← Pareto(Rt ), λt Table 7: Executable grounding controls recorded in ξt and λt . Component
Operation
Initialization
Mix proposer seeds with admissible diversity-oriented candi- Seed source, validity status, duplicates, support failures. dates. Add, drop, or modify literals, thresholds, bins, dose/time Parent key, edited field, grid identifier, validity reason. windows, or landmark windows. Combine compatible partial structures from frontier or archive Parent keys, compatibility checks, resulting semantic key. candidates. Move thresholds/windows along a predeclared grid around Old/new values, grid step, support/effect change. supported candidates. Introduce fresh candidates from grounding cues, feature pro- Trigger, source cue, source record or model_suggested tag. posals, or refiner hints. Apply validity, deduplication, support floor, Pareto rank, di- Full ordering key and retained/rejected status. versity, deterministic key order. Freeze seed stream, split ID, adapter version, candidate gram- Configuration hash and role-output hash when available. mar, population size, archive cap.
Mutation Crossover Tuning Injection Selection Replay
Logged state
by the scenario manifest before discovery: q ::= ℓ1 ∧ · · · ∧ ℓr , 1 ≤ r ≤ L, ℓ ::= z ▷◁ τ | z ∈ B | missing(z) = b, z ::= xk | virtr (xa1 , . . . , xam ) | dynr (xa1 , . . . , xam ; w), ▷◁ ∈ {<, ≤, >, ≥}.
(14)
Here xk must be a schema-visible observable, B and τ must come from predeclared grids or manifestapproved category sets, w is an admissible temporal or perturbational window, and every operator instance must return a typed feature with provenance. Mutation, crossover, tuning, and injection can edit only objects generated by this grammar; invalid literals are dropped before scoring and recorded in λt . Virtual operators derive scalar or categorical features from existing tables, such as thresholded composites, ratios, bins, or feature-family summaries. Dynamic operators summarize auxiliary longitudinal or perturbational traces, such as pre/post changes, event counts, window aggregates, dose-time summaries, or perturbational signature summaries. Both return typed features with construction provenance, missingness indicators, admissibility tags, and fail-closed behavior. A proposed feature that cannot be materialized from available columns, violates the schema, or uses forbidden response/validation information is excluded and logged. Thus, generated features expand the grounding space but do not relax the executable boundary. A.7
Reporting, Provenance, and Split Boundaries
A candidate may guide feedback even when it is not reportable, provided it was produced from discovery-visible information. Export is stricter than feedback: a candidate becomes a reported 21
Table 8: Default executable values used when a scenario manifest does not specify stricter settings. Field
Default
Role in execution
Population Nt
Clinical: 240; L1000: 160; smoke: 40
Generations Ht
Clinical: 20; L1000: 14; smoke: 2
Rule length and support Operator mix π0
L = 3; clinical ρt ≥ 30 balanced-arm support; L1000 ρt ≥ 60 instances Mutation 0.60, crossover 0.25, injection 0.15
Retry/archive
Retry cap 5Nt ; archive cap min(1000, 5Nt )
Stagnation
5 generations lower-level; 3 rounds upper-level
Controller thresholds
∆high = 0.65, ∆low = 0.25, diversity floor 0.20, invalid-rate ceiling 0.50, low-support-rate ceiling 0.60
Candidate population and offspring budget per generation after clipping to [40, 320]. Full runs are clipped to [5, 40]; smoke mode is flagged as non-evaluation. Scenario adapters may raise floors; broadening cannot lower them below the manifest floor. Each coordinate is clipped to [0.05, 0.80] and renormalized. Invalid, duplicate, or low-support offspring are retried; archives are pruned by Pareto rank, diversity, then key order. Triggers early stopping below or direction replacement above. Used for preserve/narrow/broaden/replace hints before final validation is visible.
Table 9: Examples of how broad biomedical cues can be grounded under different scenario schemas. These examples illustrate the grounding interface; they are not additional rules or reported findings. Broad cue
Clinical-trial grounding
Metabolic dysregulation
BMI, fasting glucose, triglycerides, HOMA- Pathway or signature summaries related to Schema-visible IR, insulin-related summaries when avail- metabolic programs when materialized from variables only; no able. registered profiles. outcome or validation-derived fields. Baseline or landmarked pre-decision sub- Context literals over cell, perturbagen, dose, Clinical group literals, with treatment and endpoint time, and admissible signature attributes. treatment/outcome reserved for evaluation. fields and L1000 response-derived fields cannot become literals. Predeclared longitudinal or trajectory sum- Dose-time summaries or perturbational sig- Window must be maries within admissible windows. nature summaries within allowed windows. declared before search and replayable on holdout. Available only when represented by baseline, Cell-line context, perturbagen class, pathway Comparison pools and laboratory, pathology, or approved derived summaries, and target-compatible signature target-derived labels measurements. attributes. cannot be constructed post hoc.
Treatment-response heterogeneity
Dynamic response context
Mechanistic assay context
L1000 grounding
Gate
scenario-grounded proposition only if it has frontier membership, validity, support, provenance, leakage checks, and diagnostics. We write the reporting predicate as (s)
(s) Report(q, dt , t) = 1{q ∈ Ft , Γt (q) = 1, Jsup (q) ≥ ρt , Prov(q, dt , t), NoLeak(s) (q), Diag(q, t)}. (15) When this predicate is satisfied, SCENE exports a proposition
pi = (dt , q, ωi ),
(16)
where ωi stores the evidence and provenance fields. If the predicate fails, the candidate may remain in the execution log for feedback, but it is not exported as contextualized knowledge. Split lineage is frozen before discovery. Source-documented priors may be consumed by strategist and proposer; within-run memory and logs may guide the same discovery episode; discovery-visible pseudo-validation may guide conservative bounded control update; final held-out or test archives are post-discovery audit artifacts only. Importing a completed validation summary into a later run requires a new frozen knowledge record with source run, split ID, timestamp, evidence-status label, and allowed consumer roles, so that validation-derived knowledge is explicit rather than silently recycled. A run manifest records scenario ID, freeze time, schema hash, split ID, controller thresholds, control bounds, allowed feature families, forbidden fields, enabled operators, role-output hashes when available, adapter version, and final-validation boundary. Source-free suggestions may be logged 22
Table 10: Field-level checks required before a candidate can be exported as a scenario-grounded proposition. Check
Required fields
Fail-closed behavior
Provenance
Candidate, round, direction, manifest, schema, split, replay-config IDs; knowledge-record IDs or model_suggested; grounding map; operator provenance. Forbidden-field list, literal keys, operator input columns, split lineage, L1000 comparison-pool lineage.
Missing IDs, hashes, grounding entries, or operator provenance prevent export.
No leakage
Diagnostics
Effect/support, frontier membership, Pareto rank or archive status, diagnostic contrast, pseudo-validation status or missingness, blocking status where relevant, invalid/low-support counts.
Clinical treatment/outcome/post-outcome literals fail. L1000 response/connectivity-derived, validation-derived, comparison-pool, or post hoc block features fail. Missing diagnostics prevent export, although the case may remain in λt for feedback.
as model_suggested only for data-grounded search hypotheses; they cannot be cited as external biomedical knowledge or prior-supported mechanisms. Minimal manifest fields.
An executable run stores a compact JSON-like manifest before discovery:
Manifest contract {
}
"scenario_id": "clinical_trial_or_l1000_task", "freeze_time_utc": "...", "schema_hash": "sha256:...", "split_id": "discovery_split_seed", "knowledge_records": [{"record_id": "...", "source_type": "..."}], "controller_policy": { "action_priority": ["replace", "broaden", "narrow", "preserve"], "delta_high": 0.65, "diversity_floor": 0.20, "invalid_rate_ceiling": 0.50 }, "control_bounds": { "population_size": [40, 320], "generations": [5, 40], "operator_probability": [0.05, 0.80], "max_rule_len": 3 }, "final_validation_boundary": { "visible_to_discovery": false, "metrics_feed_back": false }
Actual manifests additionally store scenario-specific paths, hashes, replay seeds, adapter versions, enabled operators, forbidden fields, and role-output hashes when role outputs are cached. A.8
Scenario-Specific Estimators
The scenario adapter defines the concrete effect and support functions used by lower-level search. (s) (s) The main text presents these functions abstractly as Jeff and Jsup ; this section records the two main instantiations used in the paper. For clinical treatment-benefit discovery and a fixed admissible rule q, the configured effect is a descriptive bad-outcome reduction: (trial)
(trial) Jsup (q) = min{n1 (q), n0 (q)}. (17) This sign convention matches the main clinical endpoint: larger values indicate a larger reduction in unfavorable outcomes under the treatment arm relative to the control arm. Because SCENE
Jeff
(q) = Pr(Ybad = 1 | q(X) = 1, T = 0)−Pr(Ybad = 1 | q(X) = 1, T = 1),
23
adaptively searches over many candidates and refines later directions using discovery feedback, reported subgroup effects are exploratory, selection-conditioned estimates. Unless a separate preregistered or multiplicity-adjusted analysis is performed, SCENE does not report valid post-selection p-values, treatment-labeling claims, or external clinical recommendations. Rules using post-baseline summaries are marked as predictive or associational unless a valid landmark interpretation is declared by the scenario. For LINCS L1000, the configured effect is a context-versus-background target-response or connectivity contrast. The adapter freezes blocking variables, comparison-pool eligibility, exclusion lists, threshold grids, and feature-construction rules before search. A candidate may determine selected-instance membership through its admissible literals, but it may not determine controls after observing its effect. Let ci be the response or connectivity score. For blocking strata B(q) with both selected and comparison evidence, Mb (q) = {i ∈ M(q) : b(i) = b}, M0b (q) = {i ∈ M0 (q) : b(i) = b}, P with weights wb (q) = |Mb (q)|/ b′ ∈B(q) |Mb′ (q)|. The blocked contrast is (l1000,blk)
Jeff
(q) =
X
wb (q) c̄Mb (q) − c̄M0b (q) .
(18)
(19)
b∈B(q)
If no admissible blocking design is available, the adapter uses the predeclared unblocked fallback c̄M(q) − c̄M0 (q) when support conditions are met and records blocking_status=unblocked. L1000 findings are descriptive perturbational associations within the configured assay and matching design, not causal compound effects or validated mechanisms. A.9
Experiment Matrix and Reporting Contract
This subsection links the method objects defined above to the three empirical reporting settings. It does not introduce additional search rules; detailed task construction and metric definitions are deferred to Appendices A.10–A.13. SCENE is evaluated in three settings: (i) clinical subgroup discovery, (ii) perturbational mechanism discovery, and (iii) downstream few-shot contextual knowledge reuse. These settings share the same high-level discovery backbone but differ in unit of analysis, target variable, candidate language, and evaluation functional. Clinical subgroup discovery. The unit is one patient. Each run operates on one clinical task frame, one paired train–holdout split, and one discovery seed. Each method returns an ordered rule list, but only the first committed rule is used in the paper-facing best-1 benchmark. Holdout outcomes are never used to choose or reorder rules. Under the proposition notation above, the grounded object qi is a subgroup rule and ωi stores the held-out replay evidence, support, diagnostics, and provenance. Perturbational mechanism discovery. The unit is one L1000 signature instance. Each run fixes one target mechanism, one split, one search seed, and one ablation profile. Discovery is executed on the training split only; all reported perturbational metrics are computed by replaying frozen rules on the held-out split. Here qi is a context-bounded perturbational finding, and ωi records target-response evidence, support, replay diagnostics, and split lineage. Contextual knowledge reuse. The unit is one downstream tabular row, either a patient or a signature instance. Each run fixes one scene, one outer split, one few-shot exemplar set, one prompt condition, and one deterministic LLM decoding policy. Discovered propositions are converted into compact knowledge cards only from train-side discovery and selected only on validation data. Test rows are used only once for final prediction and metric computation. This setting evaluates whether proposition-level outputs can be reused as auxiliary context; it is not an independent validation of every biomedical mechanism encoded in a card. Across all three settings, we maintain strict train/validation/test separation: held-out outcomes, heldout connectivity, and final downstream labels are not used to revise prompts, re-select rules, or change knowledge-card content within the same evaluation episode. 24
Table 11: Clinical task definitions used in the treatment-benefit subgroup-discovery benchmark. BL/TR denote baseline/static and trajectory-enhanced feature views. Header
Trial and clinical setting
B-D-BL
NCT00174655; node-positive operable breast cancer.
B-D-TR P-M-BL
P-M-TR
Treatment question and feature view
Doxorubicin-containing chemotherapyschedule comparison using baseline/static covariates. NCT00174655; node-positive operable Same doxorubicin-schedule comparibreast cancer. son using trajectory-enhanced features. NCT02491333; polycystic ovary syn- Metformin versus placebo using basedrome (PCOS) under sham acupunc- line/static covariates. ture. NCT02491333; PCOS under sham Metformin versus placebo using acupuncture. trajectory-enhanced features.
P-A-FPG
NCT02491333; PCOS under placebo metformin.
Acupuncture-response subgroup discovery for fasting-plasma-glucose improvement.
P-A-AUC
NCT02491333; PCOS under placebo metformin.
Acupuncture-response subgroup discovery for OGTT glucose-AUC improvement.
A.10
Response or unfavorable-outcome definition Bad outcome is recurrence or death by the preprocessed disease-free-survival (DFS) status. Bad outcome is recurrence or death by the preprocessed DFS status. Good response is month-4 HOMA-IR reduction of at least 25% from baseline; bad outcome is failing this threshold. Good response is month-4 HOMA-IR reduction of at least 25% from baseline; bad outcome is failing this threshold. Good response is month-4 fasting plasma glucose reduction of at least 0.2 mmol/L from baseline; bad outcome is failing this threshold. Good response is month-4 OGTT glucose-AUC reduction of at least 1.0 from baseline; bad outcome is failing this threshold.
Clinical Task Definitions
Table 11 expands the short clinical task headers used in Table 1. The six tasks are constructed from two de-identified patient-level trial tables associated with the registered trials in the table; ClinicalTrials.gov provides identifiers and protocol context, not patient-level modeling rows [33, 34, 1]. The trials were selected by pre-analysis feasibility and coverage criteria: randomized arms, patient-level covariates, interpretable treatment/endpoints, enough support for repeated replay, and coverage of more than one biomedical setting. No trial was added or removed after inspecting SCENE held-out performance. The NCT00174655 oncology table is accessed through Project Data Sphere [19, 20]; the NCT02491333 PCOS table is derived from the Dryad release of Wen et al. [21, 22]. These are treatment-benefit subgroup-discovery tasks: each asks whether a prioritized subgroup rule identifies patients for whom treatment reduces a task-specific unfavorable clinical outcome, rather than estimating a global average treatment effect. These definitions determine the task-specific held-out endpoint used by the main clinical endpoint, ARRbad . Treatment, outcome, post-outcome, identifier, and bookkeeping fields are excluded from candidate subgroup features before fitting any method. A.11
Implementation Details for Table 1
Table 1 is implemented as a best-1 rule-selection benchmark, following the convention that exploratory subgroup findings should be fixed before independent assessment [35, 36]. The unit of evaluation is not an unranked pool of discovered subgroups, but the first rule that a method commits to before held-out outcomes are inspected. For clinical task frame d, split identifier b, and method m, the discovery code exports an ordered rule list (1) (2) Rm,d,b = qm,d,b , qm,d,b , . . . . (20) The reported rule for that method–task–split is (1)
⋆ qm,d,b = qm,d,b .
(21)
⋆ The held-out partition is used only after qm,d,b has been fixed. We do not choose the rule with the best test-set risk difference, and we do not average multiple rules from a single method–split run. This implementation matches the scientific use case of SCENE: the system should return prioritized, interpretable subgroup hypotheses whose held-out replay can be audited, rather than producing a large catalogue from which favorable test-set rules are selected post hoc.
25
Task frames and endpoint construction. Table 1 uses the six paper-facing clinical reporting frames B-D-BL, B-D-TR, P-M-BL, P-M-TR, P-A-FPG, and P-A-AUC. These headers follow Table 11: B-D-BL/B-D-TR are baseline/static and trajectory-enhanced breast-cancer doxorubicin-schedule frames; P-M-BL/P-M-TR are baseline/static and trajectory-enhanced PCOS metformin-response frames; P-A-FPG and P-A-AUC are PCOS acupuncture-response frames for fasting-plasma-glucose and OGTT-AUC improvement. Each modeling table contains one row per patient, so row-level splitting is patient-level splitting. The treatment indicator is the task-specific randomized arm encoded in the corresponding modeling frame. The favorable-response indicator is constructed as follows. Let DDFS = {B-D-BL, B-D-TR} and DHOMA = {P-M-BL, P-M-TR}. Let FiDFS denote the preprocessed disease-free-survival indicator, with FiDFS = 1 indicating no recurrence or death. We write riHOMA for the month-4 relative HOMA-IR reduction, and aFPG and aAUC i i for the corresponding absolute reductions in fasting plasma glucose and OGTT glucose AUC. The task-specific favorable endpoint is DFS F , d ∈ DDFS , i HOMA 1{ri ≥ 0.25}, d ∈ DHOMA , Gi,d = (22) FPG 1{a ≥ 0.2 mmol/L}, d = P-A-FPG, i 1{aAUC ≥ 1.0}, d = P-A-AUC. i The bad-outcome endpoint used by the Table 1 risk-difference estimator is bad Yi,d = 1 − Gi,d .
(23)
Treatment, outcome, death, post-outcome, identifier, and bookkeeping columns are excluded from the feature set before fitting any method. Numeric imputers, categorical encoders, binarization maps, and threshold grids are fit on the training partition and replayed on the held-out partition. The response thresholds are fixed before method comparison and are used only to define task endpoints; they are not tuned using SCENE or baseline held-out performance. Split manifest and pairing. The reporting run repeats the complete rule-discovery and best-1 replay pipeline over B = 50 patient-level split identifiers. Each split uses a 75/25 train/test partition, i.e. test_size=0.25, stratified by treatment arm and bad-outcome status when feasible. The benchmark runner supports either an explicit comma-separated seed list or automatic seed generation sb = s0 + (b − 1)∆s ,
b = 1, . . . , B.
(24)
For the Table 1 reporting run, we use the manifest seed list induced by base_seed=58 and the runner seed policy recorded in the output manifest. When the default benchmark policy is used, this gives sb = 58 + 10(b − 1); if an explicit seed list is supplied, that list overrides –num-runs. The renderer materializes a split identifier as split_id=seed_<seed>: for baselines the seed is parsed from run_<id>_seed_<seed>, and for SCENE it is read from the archived config_snapshot.json. The final paper rendering requires the paired equality Im,d = Id
for all methods m on task frame d,
(25)
where Im,d is the set of used split identifiers for method m on task d. A rendered table is considered reportable only when the input manifest contains the expected Bm,d = 50 used best-1 artifacts for every method–task row and the split sets satisfy this equality. Partial diagnostic renderings with smaller, uneven, or unpaired Bm,d are useful for debugging, but the final paper rendering is configured to abort rather than silently skip or impute rows. SCENE validation artifacts. SCENE entries are read from post-discovery replay files rather than from training-side search logs. For B-D-BL and B-D-TR, the renderer uses the baseline/static and trajectory-enhanced validation outputs under the NCT00174655 SCENE root. For P-M-BL, P-M-TR, P-A-FPG, and P-A-AUC, it uses the corresponding NCT02491333 validation roots. The SCENE columns consumed by the renderer are test_subgroup_size_pct, test_bad_rate_treat, test_bad_rate_ctrl, test_rd_bad, test_rd_bad_ci_low, and test_rd_bad_ci_high. The SCENE clinical implementation uses population size 240, 20 generations, maximum rule length 3, minimum discovery-side subgroup size 30, evolutionary-loop V2, and loop-stage virtual feature construction, with the dynamic-feature switch enabled for B-D-TR, P-M-TR, P-A-FPG, and P-A-AUC, and disabled for B-D-BL and P-M-BL. These settings are fixed before held-out replay. The validation files are held-out replay artifacts: they are not fed back into the same discovery episode and are not used to revise the rule, direction, feature proposals, or operator mixture for that split. 26
Table 12: Artifact contract used by the Table 1 renderer. The contract is designed to make the best-1 boundary and the paired-split accounting auditable from files, not from prose alone. Source
Best-1 rule source
Required audit fields
Baselines
First row of the exported rule CSV after training-side ranking. First validation row after SCENE discovery-side Pareto/export ordering.
Method, task, run index, seed, source path, required test columns, used/skipped status. Task header, feature-view/dynamic flag, run index, seed, source path, required test columns, used/skipped status. Bm,d , split identifier coverage, metric columns, skipped-file reasons.
SCENE
Renderer outputs
The run-level file is the single source for summary aggregation.
Table 13: Method-specific implementation details for the non-SCENE Table 1 baselines. The values shown are fixed before held-out replay and are not tuned using Table 1 outcomes. Method
Wrapper implementation
Main fixed knobs
Rule ordering before replay
Virtual Twins
Separate treated/control randomforest classifiers for bad outcome, followed by a regression-tree surrogate over predicted benefit.
Predicted benefit, training support, and lower training-side RDbad .
SIDES
Local SIDES-style beam search over encoded categorical literals and numeric thresholds.
Honest causal tree
Transformed-outcome tree with a within-training structure/ranking split.
Causal forest surrogate
Transformed-outcome random forest distilled into a shallow interpretable tree.
CURLS
Wrapper around the released CURLS implementation.
MaxTE Subgroups
Strict R wrapper using rpart, pruning by minimum crossvalidation error, and rule extraction.
RF estimators = 300, RF leaf size = 5, surrogate max depth from rule cap, surrogate min leaf = 30, minimum arm support = 8. Beam width = 12, max feature pool = 60, max thresholds per feature = 5, min node size max(30, 2× min arm support). Honest ranking fraction = 0.5, tree max depth from rule cap, min leaf tied to arm support and leaf floor. RF regressors = 300, RF leaf size max(2, 30/2), surrogate max depth from rule cap, surrogate min leaf = 30. Thresholds = 4, negated literals enabled, max selected features = 60, variance weight = 0, minimum coverage = 0.1, rule length from cap. Numeric-only feature interface, within-training structure/ranking split = 0.5, Rscript path fixed, rule-length cap from runner.
Training-side SIDES score combining benefit, separation, support, and two-proportion evidence. Leaf effects estimated on the honest ranking split, then support and lower RDbad . Surrogate CATE score, training support, and lower training-side RDbad . CURLS exported order is preserved; no held-out reranking is applied.
Ranking-split treatment-benefit estimate from the R helper, followed by the Python rule-length cap.
Agent and compute resources. The non-SCENE Table 1 baselines are CPU-side tabular learners or rule-mining procedures. For SCENE, the agent-facing steps use third-party LLM API routes rather than local GPU inference; the logged configuration records the provider, model identifier, decoding settings, retry policy, and returned role artifacts. Rule replay, risk-difference estimation, paired-split checks, and final table rendering are performed from cached artifacts on CPU workers. CPU-side aggregation was run on a single CPU workstation (AMD Ryzen 7 7840H, 16 GB RAM), no local GPU is required once the archived SCENE artifacts are available. The Table 1 CPU-side replay and rendering jobs complete within a few CPU-hours after the processed modeling tables are available, excluding provider-dependent LLM API latency. Baseline wrappers and fixed knobs. The non-SCENE baselines cover classical subgroup discovery and recent rule-learning baselines: Virtual Twins [6], SIDES [7], honest causal trees [8], causal forests [9], CURLS [10], and MaxTE Subgroups [11]. All wrappers receive the same task list, endpoint definition, leakage exclusions, split identifier, test fraction, minimum arm-support policy when available, and rule-complexity budget. Each wrapper converts the underlying method into the same exported object: a ranked list of readable subgroup rules and held-out replay columns for those rules. The runner writes baseline files under runs/<method>/run_<index>_seed_<seed>/. The command below is the reporting form of the baseline execution; a shorter command with fewer runs is used only for sanity checks.
27
Baseline benchmark command & ".\.venv_curls_repro\Scripts\python.exe" ".\analysis\quantitative_benchmark_runner.py" ` --methods "virtual_twins,sides,honest_causal_tree,causal_forest,curls,MaxTE_subgroups" ` --datasets "B-D-BL,B-D-TR,P-M-BL,P-M-TR,P-A-FPG,P-A-AUC" ` --num-runs 50 ` --base-seed 58 ` --seed-step 10 ` --test-size 0.25 ` --rule-feature-cap 5 ` --analysis-dir ".\analysis" ` --out-dir $Out ` --rscript-path "$RscriptExe" ` --curls-root ".\CURLS\osfstorage-archive"
If a method produces no valid rule or a matched file lacks required columns, the file is recorded in table1_input_manifest.csv or table1_skipped_files.csv with a reason and does not silently contribute imputed metrics. With –expected-runs 50 –require-paired-splits, the renderer aborts final Table 1 generation if any method–task row has fewer than 50 used artifacts or if the split-id sets differ across methods within a task. Metric definitions.
test Let Dd,b be the held-out patient set for task frame d and split b, and let test S(q) = {i ∈ Dd,b : q(xi ) = 1}
be the subgroup selected by rule q. For treatment arm t ∈ {0, 1}, P bad X i∈S(q) 1{Ti = t}Yi . nt (q) = 1{Ti = t}, p̂t,bad (q) = nt (q)
(26)
i∈S(q)
The current main-text Table 1 reports the benefit-oriented bad-outcome absolute risk reduction, dbad (q) = p̂1,bad (q)− p̂0,bad (q), dbad (q) = p̂0,bad (q)− p̂1,bad (q). (27) [ bad (q) = −RD RD ARR [ bad is better: positive values mean that the committed subgroup has a lower bad-outcome Higher ARR rate under the treatment arm than under the control arm. The six task columns B-D-BL, B-D-TR, [ bad values for the method’s P-M-BL, P-M-TR, P-A-FPG, and P-A-AUC are the mean held-out ARR Best-1 rules in the corresponding clinical frame. The three rightmost columns are aggregate diagnostics over the same Best-1 replay records. Let |S(q)| cov(q) = (28) test | |Dd,b be held-out subgroup coverage. Population-adjusted ARR is a descriptive impact summary, [ bad (q) cov(q), PARR(q) = ARR (29) and the displayed P-ARR entry averages this quantity over the six task frames. It is included as a coverage-adjusted summary of the displayed ARR effects, not as an independent causal endpoint. Usable support is an outcome-free support-validity indicator, U (q) = 1 {0.10 ≤ cov(q) ≤ 0.60 and min{n0 (q), n1 (q)} ≥ 5} , (30) ⋆ and U-Supp. is 100 times the mean of U (qm,d,b ), averaged over the six task frames. This statistic penalizes both tiny subgroups and near-population rules, while requiring minimal arm-specific support. Direction consistency compares the discovery-side treatment-benefit direction with the held-out direction. If Atrain m,d,b denotes the training-side benefit sign recorded for the exported Best-1 test [ bad (q ⋆ rule and Am,d,b = ARR m,d,b ), then P train test d,b 1{sign(Am,d,b ) = sign(Am,d,b )} DConsm = 100 P . (31) train test d,b 1{Am,d,b and Am,d,b are nonmissing} Arm-specific bad-outcome rates, raw subgroup coverage, smaller-arm support, and risk-difference intervals remain in the run-level audit files but are not displayed in the compact main-text table. 28
Aggregation. For a scalar run-level metric z used by Table 1, the task-level entry for method m and task frame d is the mean over used split artifacts, X 1 z̄m,d = zm,d,b , (32) Bm,d b∈Im,d
where Im,d is the set of split identifiers marked used in the manifest and Bm,d = |Im,d |. The six [ bad (q ⋆ ARR columns use zm,d,b = ARR m,d,b ). The P-ARR and U-Supp. columns first form the ⋆ ⋆ task-level means of PARR(qm,d,b ) and 100U (qm,d,b ), respectively, and then average those task-level means across the six clinical frames so that each frame receives equal weight. D-Cons. is computed from the nonmissing run-level direction labels across the six frames. No missing metric is imputed; files with missing required columns or invalid Best-1 rules are recorded in the skipped-file manifest. Single-split Wald intervals for two proportions are still stored in the run-level artifacts for audit, s p̂1,bad (q)(1 − p̂1,bad (q)) p̂0,bad (q)(1 − p̂0,bad (q)) c SE(q) = + , (33) n1 (q) n0 (q) but these per-rule intervals are not used to choose rules and are not part of the compact main-text Table 1. Interpretation boundary. This implementation is intended to make the Table 1 comparison about rule discovery rather than about leakage, outcome redefinition, or test-set rule search. The shared endpoint and exclusion policy align the covariate boundary across methods; the best-1 boundary prevents a method from searching the held-out set over many exported rules; and the manifest makes the repeated-split accounting auditable. These design choices support the paper’s claim at the level of exploratory, prioritized clinical subgroup hypotheses. They do not turn the resulting rules into confirmatory treatment-effect estimates, valid post-selection p-values, or clinical recommendations. A.12
Implementation Details for Table 2
Table 2 is implemented as a paired L1000 second-scene module ablation. The unit of analysis is one complete discovery episode for a fixed target mechanism g, one train/holdout split and search seed s, and one ablation profile a. Each episode freezes the target mechanism, the split, and the search seed before holdout evaluation, and then executes the full second-scene chain: L1000 table holdout knowledge-proposition → SCENE discovery → → construction rule replay export. The paper experiment uses three target mechanisms and ten split/search seeds, so the reporting contract for each ablation row is |G| × |S| = 3 × 10 = 30 paired target-seed episodes. The convenience runner run_l1000_ablation_table2.py can execute and collect runs for a single prepared modeling table. For the manuscript table, we repeat that runner over the three target-specific modeling tables and then use report_table2_table3_c ontracts.py to enforce the full target-seed-profile matrix before aggregation. A Table 2 row is reportable only if this checker finds exactly one run for every expected target-seed cell and every reported metric is finite in all 30 cells; otherwise it writes table2_pairing_errors.csv or table2_metric_errors.csv and aborts. L1000 table construction. The reported experiments use a fixed raw-to-modeling conversion script that constructs one row per L1000 signature instance. Each row in the resulting modeling table corresponds to one L1000 signature instance indexed by sig_id. Static fields describe the cell and perturbagen context, while dynamic fields summarize dose, time, pathway, and landmark-gene response programs. The default paper-facing builder uses LINCS/L1000 Level-5 COMPZ signatures and metadata resources [2, 3, 23, 24]; pathway summaries use curated molecular-signature programs [37]. It keeps compound perturbations for the modeling population, uses genetic perturbations as target-mechanism references, applies tas_min=0.2, keeps 6h/24h/96h signatures, retains the top 20 cell contexts after a minimum 200 signatures-per-cell filter, caps the modeling subset at 50,000 29
signatures, caps the reference pool at 5,000 signatures, and reads the GCTX matrix in column batches of 2,048. For a target mechanism g, the builder constructs a target centroid from reference perturbation signatures when at least min_target_ref=20 references are available. Cell-specific target centroids are used when at least min_cell_target_ref=3 references exist for that cell; otherwise the global target centroid is used. The continuous endpoint used by Table 2 is the configured connectivity score ci , recorded as connectivity_score and related target-margin fields. The binary outcome field is retained for compatibility with shared rule-mining utilities, but Table 2 reports holdout connectivity contrasts rather than treating this compatibility label as the primary endpoint. Leakage controls. Discovery excludes identifiers, split keys, target-derived connectivity fields, outcome columns, and known assay-density shortcuts from rule literals. In L1000, perturbagenfrequency features are disabled by default, and tokens such as distil_, tas, pct_self_rank_q25, and dose_time_exposure are hard-blocked. The preferred holdout split is group-aware by pert_iname; the split key is sig_id when unique and otherwise the internal row id. The default holdout fraction is 0.25. The split specification is written to l1000_split_spec.json and reused by the evaluator, so validation metrics cannot alter the rule search in the same episode. Ablation matrix. All profiles share the same L1000 adapter, target construction, split policy, rule grammar, and evaluator. They differ only in which SCENE modules are enabled. The internal profile names are kept in commands and manifests, while the paper-facing labels match Table 2: • full: enables all modules and is reported as SCENE (ours). • no_top_layer: removes the upper strategic layer while retaining lower-level search, and is reported as w/o Upper. • no_dynamic: removes dynamic feature planning and executable dynamic feature replay, and is reported as w/o Dynamic. • no_virtual: removes virtual feature construction and refinement, and is reported as w/o Virtual. • no_pareto: disables Pareto multi-objective selection and keeps scalar-fitness selection, and is reported as w/o Pareto. • minimal_baseline: removes the ablated modules jointly and is reported as Minimal. Unless overridden by an ablation profile, the shared mining knobs are population size 160, 14 generations, maximum rule length 3, minimum subgroup size 60, initial feature sample size 32, support metric transfer_support, generalization group pert_iname, minimum group coverage 12, minimum group coverage ratio 0.01, and support-size weight 0.22. LLM and feature-generation routes. The L1000 discovery script uses separate third-party LLM API routes for structured JSON decisions, free-form text synthesis, dynamic feature planning, and dynamic code generation; it does not require local GPU inference. The paper configuration records the provider, model name, retry policy, timeout, and enable_thinking setting in l1000_run_ma nifest.json. The JSON route defaults to Qwen/Qwen3.5-9B [38] with JSON response format and thinking disabled; text synthesis uses the configured text route; dynamic planning and code generation use their dedicated route configs. CPU-side aggregation was run on a single CPU workstation (AMD Ryzen 7 7840H, 16 GB RAM). Individual target–seed–profile episodes typically require tens of minutes to a few CPU-hours after the modeling table is prepared. Exact bitwise replay of live LLM calls requires the archived prompts and returned artifacts; the numerical Table 2 aggregation is based on the saved run artifacts, not on re-querying the LLM. Holdout evaluation. Rules are ranked before holdout replay by discovery-side order: Pareto rank when enabled, fitness, training-side connectivity contrast, support, p-value, and generation. The evaluator replays the top K = 20 rules on the frozen holdout table; the quality summary also uses K = 20 unless the run manifest records an override. Virtual features are fit on the discovery-training side and replayed on holdout. Dynamic features are exported as executable plans and replayed before validation. If a rule depends on a virtual or dynamic feature that cannot be replayed on holdout, that rule is marked invalid and contributes to the valid-rule-rate denominator. 30
Table 14: Artifact contract for Table 2. These files make the target-seed pairing, split boundary, ablation state, and metric source auditable. Artifact
Role
Required audit fields
l1000_build_summary.json
Modeling-table provenance.
l1000_run_manifest.json
Single-run manifest.
l1000_split_spec.json
Frozen discovery/holdout split.
l1000_agentic_results.csv
Discovery-side ranked rules.
test_rule_eval.csv and l1000_quality_eval.json
Holdout rule replay.
l1000_table2_metrics.json
Run-level Table 2 source.
l1000_knowledge_propositions.json
Downstream knowledge export.
Raw L1000 path, target perturbagen, reference counts, filters, row counts, output paths. Target mechanism, ablation profile, seed, enabled components, split-spec path, train/holdout sizes, LLM configs, blocked feature tokens. Split key, train/test keys, split group column, holdout fraction, random seed. Rule text, fitness, Pareto rank, discovery contrast, support, transfer support, group coverage, p-value, generation. Rule validity, subgroup size, support ratio, ∆Conn, PosRate90, p-values, rank stability, generalization gap. Compact metric set consumed by the Table 2 collector. Selected validated rules, target mechanism, support/effect evidence, feature tokens, source labels.
test Let Dg,s be the frozen holdout set for target g and seed s, and let ci be the configured connectivity score. For a rule q, test S(q) = {i ∈ Dg,s : q(xi ) = 1},
test B(q) = Dg,s \ S(q).
The mean holdout connectivity contrast is ∆Conntest (q) =
X X 1 1 ci − ci . |S(q)| |B(q)| i∈S(q)
i∈B(q)
The median version replaces the two means by medians. The strong-connectivity threshold τ is fixed inside the evaluator. In auto mode, the evaluator uses a CLUE-like τ = 90 threshold when the score scale supports it and otherwise uses the holdout 90th percentile. Thus X PosRate90(q) = |S(q)|−1 1{ci ≥ τ }, i∈S(q)
and ∆PosRate90(q) subtracts the corresponding background rate over B(q). Metrics and aggregation. For one run, Table 2 summarizes a valid top-K replay subset Qa,g,s for ablation profile a, target g, and seed s. Quality metrics are computed from valid replayed rules, with one support focus used by the evaluator. If at least max(5, ⌈K/2⌉) valid rules have support_ratio_test >= 0.02, the mean contrast, PosRate90, p-value rate, and gap diagnostics are computed on this support-focused subset and the run records quality_focus_mode_topk=support_ge_0.02; otherwise the evaluator uses all valid top-K rules and records quality_focus_mode_topk=all. In both cases, Valid Rule Rate uses the raw top-K replayed rules as its denominator, including rules with parsing errors or empty subgroups. The focus mode and focus rule count are preserved in l1000_table2_metrics.json and in the reporting-contract manifest. The Conn. column is the run-level mean of the held-out connectivity contrast, X 1 Conna,g,s = ∆Conntest (q), |Qa,g,s |
(34)
q∈Qa,g,s
where larger values indicate stronger target-directed connectivity inside the rule-covered signatures than in the held-out background. The Strong-Pos. column is 100 times the corresponding mean PosRate90(q) over Qa,g,s , so it measures the fraction of rule-covered signatures falling into the strong-response tail. The significance-rate diagnostic uses effective holdout p < 0.05, with permutation p-values automatically refined from 1,000 to at most 5,000 permutations when floor saturation 31
is detected. These p-values are retained in the audit files but are post-selection diagnostics, not confirmatory inference for adaptively searched rules. Target retrieval is evaluated by a rule-ensemble score on holdout signatures. Each top discovery rule votes for matching signatures with positive discovery-side contrast when available, otherwise a unit vote. The preferred positive label is connectivity_target_margin > 0; if that label is unavailable, the evaluator falls back to the strong-connectivity threshold. AUPRC is the area under the precision–recall curve of this ensemble score, following standard PR evaluation conventions for imbalanced binary retrieval [25, 26]; AUROC, Hit@K, recall@K, precision@K, and target-rank fields are also stored in the audit files following PR/ROC conventions [39]. When perturbagen names and a target perturbagen are available, perturbagen-level median scores are used for target rank and Hit@K. The current compact table reports U-Supp. rather than raw support ratio. For a replayed L1000 rule, |S(q)| UL1000 (q) = 1 0.10 ≤ ≤ 0.60 and |S(q)| ≥ 5 , (35) test | |Dg,s P and U-Supp. is 100|Qa,g,s |−1 q∈Qa,g,s UL1000 (q), averaged across episodes. This makes the reported support diagnostic a usability rate rather than a monotone reward for larger subgroups. The Gap column is the normalized train–holdout contrast gap, Gapa,g,s =
∆Conntrain,a,g,s − ∆Conntest,a,g,s max{10−8 , |∆Conntrain,a,g,s |}
,
(36)
so smaller values indicate better transfer from discovery to holdout. For scalar metric z, the Table 2 entry for ablation profile a is 1 XX z̄a = za,g,s . 30 g∈G s∈S
The reporting artifacts store the mean and empirical standard deviation over the same 30 target-seed episodes; the compact main-text table reports the means only. No missing numeric metric is imputed. A run is nonreportable if discovery produces no rule export, if holdout replay fails, if required features cannot be materialized, or if l1000_table2_metrics.json is missing. Interpretation boundary. Table 2 evaluates whether SCENE modules improve replayable L1000 mechanism-connectivity discovery under fixed targets and splits. The output should be interpreted as context-bounded perturbational association evidence: a rule identifies cell, perturbagen, dose/time, or signature contexts with stronger held-out target connectivity. It is not a clinical recommendation, an externally validated biological mechanism, or a causal compound-effect estimate. Reproduction command. The paper-facing execution uses run_l1000_ablation_table2.py; the script default contains a convenience preset, while the paper run overrides it with ten seeds and repeats the runner for the three target mechanisms. Final manuscript aggregation is performed by report_table2_table3_contracts.py, which checks exact one-run-per-cell coverage and per-metric completeness for the three-target by ten-seed matrix before writing table2_paired_ru n_manifest.csv and table2_paired_summary.csv. A reproducible shell form is: Reproduction pipeline $ProjectRoot = "<absolute path to the project root containing analysis/>" $PythonExe = "<absolute path to the Python executable>" $OutRoot = "<absolute path to the Table 2 output root>" $AnalysisDir = Join-Path $ProjectRoot "analysis" $Targets = @("<target_1>", "<target_2>", "<target_3>") $Seeds = "44,45,46,47,48,49,50,51,52,53" $Profiles = "full,no_top_layer,no_dynamic,no_virtual,no_pareto,minimal_baseline" foreach ($target in $Targets) { $tableDir = Join-Path (Join-Path $OutRoot "tables") $target
32
& $PythonExe (Join-Path $AnalysisDir "prepare_l1000_tables.py") ` --target-pert-iname $target ` --out-dir $tableDir
}
$runRoot = Join-Path (Join-Path $OutRoot "runs") $target & $PythonExe (Join-Path $AnalysisDir "run_l1000_ablation_table2.py") ` --modeling-csv (Join-Path $tableDir "l1000_modeling_table.csv") ` --result-root $runRoot ` --study-config paper_table2_v1 ` --profiles $Profiles ` --seeds $Seeds ` --run ` --generations 14 ` --population-size 160 ` --min-subgroup-size 60 ` --eval-top-k 20 ` --eval-quality-top-k 20 ` --out-dir (Join-Path $runRoot "summary")
A.13
Implementation Details for Table 3
Table 3 evaluates whether contextualized knowledge discovered by SCENE can be reused as auxiliary knowledge for downstream few-shot tabular classification under an in-context learning protocol [40]. This is a post-discovery utility experiment: Table 1 and Table 2 produce rules and knowledge propositions; Table 3 converts them into compact knowledge cards and asks whether they improve few-shot predictions under the same row serialization, exemplar budget, split, and LLM decoding policy. No Table 3 prediction, prompt, or metric is fed back into the upstream discovery runs. Scenes and labels. The core runner is r u n _ k n o w l e d g e _ a u g m e n t e d _ f e w s h o t . py, with task context supplied by k n o w l e d g e _ b a s e . p y. The active scene is selected by CODE_CONFIG["active_scene"] or by the corresponding command-line arguments. Clinical tasks use patient-level records. In the paper-facing treatment-benefit mode, the positive class is evidence consistent with a positive treatment-over-control signal: a treatment-arm record with a good outcome or a control-arm record with a bad outcome. This is a row-level downstream classification proxy for treatment-benefit evidence, not an individual treatment-effect estimate. For L1000, the row is a signature instance. The preferred paper label is an external target-mechanism label when available. When the L1000 run is configured with l1000_label_mode=target_margin, the positive class is the target-mechanism proxy connectivity_target_margin > 0. If strong_connectivity is used, it is reported as a secondary stress test. Label columns, target-derived connectivity fields, outcome, sig_id, pert_id, distil_id, internal row ids, SMILES/InChI keys, and manifest hard-block tokens are excluded from row serialization and from rule-card literals unless an explicit stress-test mode enables them. Biological context fields such as cell_id, dose/time fields, and pert_iname may remain in the serialized row; when pert_iname is available, the downstream split is group-aware by pert_iname and the split manifest records train-test group overlap. Split and exemplar protocol. Each run fixes a scene, one outer split, one few-shot exemplar set, one prompt condition, and one deterministic LLM decoding setting. The default train/validation/test fractions are 0.6/0.2/0.2 with split seed 44, unless a source split is supplied. Clinical runs can bind the final test partition to upstream source test identifiers. L1000 runs bind to the selected source l100 0_split_spec.json when –use-source-test-ids is enabled; with –require-group-split, train/validation/test splitting is group-aware by pert_iname and the split manifest records train-test group overlap. Few-shot exemplars are sampled only from the train partition, while knowledge-card selection is scored only on the validation partition. The final test partition is used only for Table 3 prediction and metric computation. Table 3 uses three prompt conditions and ten exemplar seeds per condition in the paper configuration. The checked source of truth is the command line and the resulting run_config.json; the smaller 33
seed list in CODE_CONFIG is a convenience/debug default and is not used for the manuscript table. C0 is the ordinary few-shot baseline with row serialization and labeled exemplars. C1 adds generic scene/task context to the same few-shot protocol. C2 uses the SCENE discovered-card protocol: task-specific scene context, task knowledge context when available, row-conditioned discovered rule cards, and the same few-shot exemplars. The runner writes these same condition names in all metric files; report_table2_table3_contracts.py checks that all three conditions exist for each scene/shot/source directory, checks the exact exemplar seed set and duplicate-free condition-seed rows, checks L1000 split/leakage invariants when pert_iname is serialized, and writes the manuscriptfacing table3_mapped_metrics_summary.csv and table3_mapped_metrics_agg.csv. Knowledge-card construction. Candidate cards are loaded from the upstream rule artifacts of the corresponding source run. The loader parses rule expressions, canonicalizes semantic duplicates, removes rules using prohibited or blacklisted fields, and scores the remaining rules on the validation split only. For ordinary binary classification, the validation card effect is ∆val (q) = Pr(Y = 1 | q(X) = 1, Dval ) − Pr(Y = 1 | Dval ), p and the ranking score is ∆val (q) nval (q). For clinical treatment-benefit mode, the validation effect is the within-rule treatment-minus-control good-outcome contrast, ∆benefit (q) = Pr(G = 1 | q(X) = 1, T = 1, Dval ) − Pr(G = 1 | q(X) = 1, T = 0, Dval ), val and support is min(ntreat , nctrl ) inside the rule. The paper discovered-card pool keeps validationdirectionally helpful rules with min_card_val_effect_delta=0.0, applies the minimum validation support threshold, and selects the top top_k_cards=5 cards unless a run manifest records an override. Each selected card stores a card id, rule condition, consequence text, validation effect, validation support, and source-run evidence. Cards are treated as soft priors, not deterministic labels. With row_conditioned_cards=True, a query prompt receives only selected cards whose rule condition matches the query row. If no selected card matches, the prompt explicitly states that no selected rule card matches the current row. For L1000, use_l1000_knowledge_json_cards=False is the default in this few-shot runner to avoid reusing cards selected by the same upstream holdout when that holdout is also bound as the final downstream test split. Agent and compute resources. Table 3 uses third-party LLM API inference for the few-shot prediction calls in all prompt conditions. C0, C1, and C2 share the same model route, decoding policy, exemplar budget, and parsing rules; only the supplied knowledge context differs. The CPU-side aggregation was run on a single CPU workstation (AMD Ryzen 7 7840H, 16 GB RAM). No local GPU is required unless one chooses to replace the API route with local LLM inference. The CPU-side reporting work is lightweight once prompts and responses are cached; end-to-end wall-clock time is dominated by provider-dependent API latency for the 2 scenes, 3 prompt conditions, and 10 exemplar seeds. Prompt and inference protocol. Rows and exemplars are serialized as feature-value statements using the same selected row feature list for all conditions in a scene. The selector first keeps configured canonical features, then scene-priority features, then features used by selected cards, and finally fills remaining slots up to max_row_features=24; blacklist tokens such as follow-up, recurrence, death, outcome, and event fields are applied before serialization. Few-shot sampling is class-balanced; in clinical label_group mode it also balances treatment arm when possible to reduce treatment-label shortcut prompts. The paper run uses real LLM inference, temperature=0.0, thinking disabled when supported by the provider, prompt saving enabled, and llm_parse_fallback=error. The configured model and provider are recorded in run_config.json; our reporting command fixes Qwen/Qwen3.5-9B [38] for the few-shot runner unless a run manifest explicitly records another paper-approved model. Model outputs are parsed as yes/no answers from <answer> tags and task-specific positive/negative hints. If the first parse fails, the script performs one repair call that asks only for <answer>yes</answer> or <answer>no</answer>. Under the paper fail-closed policy, unresolved parse failures abort the run rather than being replaced by mock or majority predictions. 34
Table 15: Artifact contract for Table 3. These outputs make the few-shot split, card selection, prompts, model outputs, and parsed predictions auditable. Artifact
Role
Required audit fields
run_config.json
Resolved downstream configuration.
split_indices.json
Downstream split manifest.
candidate_rules_scored.c sv selected_discovered_cards. json prompts/*.json
Validation-side card ranking.
Full prompt replay record.
fewshot_predictions.csv
Row-level predictions.
fewshot_metrics_summary. csv and fewshot_metrics_agg.csv table3_mapped_metrics_su mmary.csv and table3_mappe d_metrics_agg.csv
Table 3 metric sources.
Scene, task id, source run directory, data/rule/knowledge paths, label source, conditions, shots, exemplar seeds, row features, LLM fields. Train/validation/test indices, split seed, group column, source-test binding, fixed-split information. Rule, semantic key, validation support, validation effect, validation score, leakage-filter status. Card ids, rules, validation effects, supports, card group, source evidence. System prompt, user prompt, exemplar ids, active card ids, raw output, repair output, parsed answer, gold label. Condition, shot, seed, query row, true label, parsed prediction, parsed probability, active cards, parse-failure flag. Per-condition/per-seed metrics and condition-level mean/std summaries.
Paper C2 card source.
Final paper-label reporting files.
C0/C1/C2 completeness, exact seed-set and duplicate checks, L1000 leakage checks, final Table 3 aggregate metrics.
Metrics and aggregation. For condition c, shot count k, exemplar seed e, and final test set Dtest , the script records parsed binary predictions ŷi and parsed binary scores p̂i . Table 3 reports three standard binary-classification metrics. AUPRC is average precision, i.e. the area under the precision–recall curve induced by the parsed scores [25, 26], AUPRCc,k,e = AP {(yi , p̂i ) : i ∈ Dtest } . MacroF1 is the unweighted mean of the class-specific F1 scores, 1 MacroF1c,k,e = (F10,c,k,e + F11,c,k,e ) , 2 where each class is treated once as the positive class. Accuracy is X 1 1{ŷi = yi }. ACCc,k,e = test |D | test i∈D
Because the prompted model is constrained to a yes/no answer, AUPRC has limited score resolution and is interpreted together with these thresholded metrics. Additional audit fields include AUROC, balanced accuracy, majority-class accuracy, accuracy minus majority accuracy, sensitivity, specificity, precision, test positive rate, predicted positive rate, and parse-failure count [39]. For scalar metric z, 10 1 X zc,k,e , z̄c,k = 10 e=1 and fewshot_metrics_agg.csv reports the empirical standard deviation across the ten exemplar seeds. The compact main-text Table 3 displays the means only. The final Table 3 renderer should consume table3_mapped_metrics_agg.csv, not the runner’s intermediate aggregate, so that condition completeness, seed identity, and leakage checks are always applied. Failure handling and interpretation. A Table 3 run is nonreportable if the label has a single class, source split binding fails, the test split is empty, candidate rules cannot be parsed, all candidates are removed by leakage filters, no discovered card survives validation scoring for paper C2, the required real LLM is unavailable, or an answer cannot be parsed under the fail-closed policy. If max_test_samples is used to control LLM cost, the subsampling mode, sampled class prevalence, and sampled row count are written to the metric files and run config. Missing cards are not replaced with generic biomedical text. Table 3 therefore measures the downstream prompt utility of SCENEderived contextualized knowledge, not the independent biological or clinical truth of every card and not a causal validation of the upstream rules. Reproduction command.
The paper run invokes the three manuscript conditions directly:
35
Reproduction pipeline $ProjectRoot = "<absolute path to the project root containing analysis/>" $PythonExe = "<absolute path to the Python executable>" $OutRoot = "<absolute path to the Table 3 output root>" $env:PYTHONPATH = $ProjectRoot $env:SILICONFLOW_API_KEY = "<provider API key>" $env:SILICONFLOW_BASE_URL = "<OpenAI-compatible base URL>" & $PythonExe -m analysis.run_knowledge_augmented_fewshot ` --scene clinical ` --task-id <clinical_task_id> ` --source-run-dir "<absolute path to the Table 1 source run directory>" ` --source-scenario both ` --source-run-id run_01 ` --conditions C0,C1,C2 ` --shots 4 ` --seeds 0,1,2,3,4,5,6,7,8,9 ` --top-k-cards 5 ` --min-card-val-effect-delta 0.0 ` --row-conditioned-cards ` --fewshot-sampling-mode label_group ` --temperature 0.0 ` --llm-model Qwen/Qwen3.5-9B ` --no-llm-enable-thinking ` --save-prompts ` --strict-source-binding ` --require-group-split ` --require-real-llm ` --llm-parse-fallback error ` --out-dir (Join-Path $OutRoot "clinical") & $PythonExe -m analysis.run_knowledge_augmented_fewshot ` --scene l1000 ` --task-id <l1000_task_id> ` --source-run-dir "<absolute path to the Table 2 full-profile source run directory>" ` --use-source-test-ids ` --l1000-label-mode target_margin ` --conditions C0,C1,C2 ` --shots 4 ` --seeds 0,1,2,3,4,5,6,7,8,9 ` --top-k-cards 5 ` --min-card-val-effect-delta 0.0 ` --row-conditioned-cards ` --temperature 0.0 ` --llm-model Qwen/Qwen3.5-9B ` --no-llm-enable-thinking ` --save-prompts ` --strict-source-binding ` --require-group-split ` --require-real-llm ` --llm-parse-fallback error ` --out-dir (Join-Path $OutRoot "l1000")
A.14
LLM Backend Runtime Protocol
The backend-runtime comparison changes only the third-party LLM API backend while keeping the discovery scenario, split, seed, schema, population size, generation budget, support floor, and reporting predicate fixed. Total runtime is measured as wall-clock time for the full frozen discovery run, including API calls, parsing, controller validation, evolutionary scoring, and logging. Generation latency is the mean wall-clock time per evolutionary generation within that run. Because all backends are accessed through third-party APIs, these measurements are indicative runtime costs rather than hardware-normalized throughput; provider load and network conditions can change absolute values. The sweep is reported as a runtime sensitivity analysis and is not used to select rules or revise the 36
main benchmark protocols. Figure 7 fixes the P-M-TR task, split, seed, schema, controller settings, and discovery budget, and changes only the third-party API backend. Held-out benefit is computed after discovery as bad-outcome ARR in percentage points, averaged over the top-10 discovery-ranked rules exported by each run; held-out outcomes are not used to rank or select these rules. Runtime includes API calls, parsing, controller validation, evolutionary scoring, and logging, and should be interpreted as indicative API wall-clock cost rather than hardware-normalized throughput.
B
Additional Experiments and Analyses
This section provides implementation details and post-hoc evidence audits for SCENE. All supplementary analyses reuse completed clinical and L1000 outputs; no additional LLM prompting, rule discovery, or holdout-driven model selection is performed. The goal is to make the empirical evidence more legible along axes that are central to SCENE: effect–support geometry, contextualization of reported rules, L1000 inspectability, held-out replay reliability, semantic provenance, component behavior, generated-feature gating, and grounding traces. These audits should therefore be read as descriptive robustness and interpretation analyses rather than as new model-selection procedures. For the clinical analyses, we use a benefit-oriented sign convention. The source summaries report RD(Bad) as treatment bad-outcome rate minus control bad-outcome rate. We therefore define Benefit(R) = Pr(Ybad = 1 | R, T = 0) − Pr(Ybad = 1 | R, T = 1) = −RD(Bad).
(37)
Larger values indicate a larger reduction in bad outcomes under treatment. Unless otherwise stated, the clinical support axis denotes subgroup support, and the rule-level reliability analysis uses the smaller treatment-arm count inside the rule as the held-out support measure. This convention keeps the appendix aligned with the multi-objective evidence view used by SCENE: rules should have both a favorable scenario effect and enough support to be interpretable. B.1
Clinical Effect–Support–Uncertainty Geometry
The main clinical comparison table reports selected subgroup summaries, but a table alone does not show how treatment benefit is positioned relative to the support-validity criterion used by the primary benchmark. Raw support is difficult to interpret as a monotone objective: a rule covering nearly everyone can be uninformative, while a very small rule can be unstable. We therefore visualize the clinical comparison with the same usable-support diagnostic reported in Table 1, and then complements it with uncertainty and variability diagnostics. For Figure 8, we use the method-level usable support column reported in Table 1. Usable support measures how often a method’s committed Best-1 rule has enough held-out coverage to be evaluated as a subgroup without degenerating into either a tiny fragment or a near-population rule. It is computed by first forming the task-level usability rate X 1 ut,m = 100 · I{10% ≤ sr ≤ 60%, min(nr,treat , nr,control ) ≥ 5}, |Rt,m | r∈Rt,m
where Rt,m is the set of Best-1 records available for method m in clinical frame t, sr is the held-out subgroup coverage percentage for record r, and nr,treat and nr,control are the held-out treated and control counts inside the subgroup. The plotted support coordinate is then 6
Um =
1X ut,m , 6 t=1
which is the U-Supp. entry in Table 1. This binary usability flag penalizes both very small subgroups, which can be unstable, and near-population rules, which are weak subgroup discoveries, while also requiring minimal arm-specific support. The statistic is outcome-free: it does not use treatment effects or confidence intervals, and it asks only whether the committed Best-1 rule has a support profile suitable for held-out subgroup evaluation. Across the 42 primary method–frame summaries used in this audit, the plotted benefit values reproduce the Table 1 ARR entries. SCENE has a median benefit effect of 0.269, compared with 0.055 for the non-SCENE baselines. On the Table 1 U-Supp. axis, SCENE reaches 84.7%, whereas 37
Figure 8: Clinical benefit–support landscape. Each panel shows one clinical frame. The x-axis is usable-support value from Table 1, the y-axis is the corresponding benefit effect for that frame. Marker shape and color denote the method. SCENE is consistently separated on benefit while also retaining the highest usable support.
the strongest non-SCENE baseline reaches 70.0%. At the per-frame level, SCENE has the largest benefit-effect point estimate in all six clinical frames. The correct interpretation is therefore not that SCENE always maximizes raw subgroup size, but that its high-benefit rules also satisfy the support-validity criterion used by the primary comparison. Figure 8 places SCENE and the clinical baselines into this benefit–usable-support landscape. Because the x-axis is the method-level Table 1 U-Supp. diagnostic, each method keeps the same support coordinate across panels, while the y-coordinate changes with the clinical frame’s ARR. SCENE lies in the upper-right region across the six panels: it combines the largest frame-specific benefit values with the highest Table 1 usable support. This pattern is consistent with the role of SCENE’s lower-level multi-objective search: it should not simply maximize subgroup size, nor should it collapse to the smallest high-effect subgroup. The clinical comparison in Table 1 reports Best-1 subgroup summaries for each method and task. The table is designed to make the primary comparison compact, but it does not directly show whether a high benefit estimate is accompanied by unstable behavior across repeated Best-1 records. We therefore add a complementary diagnostic that removes confidence intervals from the visualization and instead summarizes empirical variability of the Best-1 benefit effect. This audit is descriptive rather than inferential: it is not used for model selection, and it is not intended to replace the patientlevel confidence intervals in the main quantitative comparison. Its purpose is to check whether the favorable clinical effect estimates of SCENE are accompanied by excessive run-to-run variability. For this analysis, the benefit effect is the bad-outcome reduction, ∆ = ControlBadRate − TreatmentBadRate. Larger values are better. The input records are aligned with the Table 1 reporting protocol: when a method exports a ranked rule list, only the prioritized rule is used as its Best-1 rule; archived replay-audit records contribute one Best-1 replay record per row. For each method and clinical frame, we first compute a task-level center and a task-level variability statistic over these Best-1 records. We then average those task-level summaries across the six Table 1 clinical frames. The resulting figures therefore show one point per method, not per-run points, and do not encode or reveal the number of available records for any method. 38
Mean variability (SD; lower is better)
SCENE
Mean median absolute deviation (lower is better)
favorable high benefit, low uncertainty
0.10
0.12
0.14
SIDES Honest Causal Tree
0.16
MaxTE Subgroups Virtual Twins
0.18
Causal Forest
0.20
0.22
0.24 0.00
CURLS
0.05
0.10
0.15
0.20
SCENE
MaxTE Subgroups SIDES 0.08 Honest Causal Tree
0.10 Virtual Twins
0.12
CURLS
Causal Forest 0.14 0.00
0.25
favorable high benefit, low uncertainty
0.06
0.05
0.10
0.15
0.20
0.25
Mean task-level median Best-1 benefit
Mean Best-1 benefit effect
Figure 9: Clinical benefit–variability profiles. (Left) Mean benefit with task-level SD. (Right) Median benefit with task-level MAD. In both panels, moving right/up indicates higher benefit and lower empirical variability.
We use three variability definitions. The first uses the mean benefit as the center and the standard deviation (SD) as the variability statistic. The second uses the median benefit as the center and the median absolute deviation (MAD) around that median as a robust variability statistic. The third again uses the median benefit as the center but uses the interquartile range (IQR) as the variability statistic. In each plot, the horizontal axis is the benefit center, so moving right is better. The vertical axis is the empirical variability estimate, so moving upward is better because the axis is drawn with lower variability at the top. The upper-right region is therefore the favorable region: high benefit and low variability. Figure 9 gives the most conventional variability view. SCENE has the largest mean benefit effect (0.267), while the strongest baseline by mean benefit is MaxTE Subgroups (0.112). The same figure also shows that SCENE does not obtain this larger benefit by becoming more variable. Its average task-level SD is 0.097, lower than every baseline in this audit. The closest baselines on variability are SIDES (0.140), Honest Causal Tree (0.155), and MaxTE Subgroups (0.165), but all three are substantially left of SCENE on the benefit axis. Conversely, CURLS and Causal Forest are not only lower-benefit but also more variable under this summary. Thus the mean-SD view supports the main table’s interpretation: SCENE is not merely selecting a high-effect subgroup in a noisy manner; it is separated from the baselines in the favorable direction on both axes. The median-MAD view in Figure 9 asks whether the conclusion is driven by a small number of extreme Best-1 records. The answer is no. SCENE again has the largest center statistic, with a mean task-level median benefit of 0.265. Its average MAD is 0.062, which is also the smallest among all methods. The strongest baseline after applying the same robust summary is Honest Causal Tree by the stability-adjusted read-out, but its median benefit is only 0.114 and its MAD is 0.089. MaxTE Subgroups is close to Honest Causal Tree under this robust view, with median benefit 0.098 and MAD 0.073, but it remains far below SCENE in benefit. This is important because MAD is less sensitive to outlying records than SD. The fact that SCENE remains in the favorable upper-right region under MAD suggests that the result is not simply a mean-based artifact. The compact audit table in Figure 10 uses the same task-level scorecard as the benefit–support landscape. The raw Benefit effect is the bad-outcome reduction in percentage points. The Effect/CI signal is the benefit effect divided by the width of the corresponding confidence interval, so it rewards effects that are large relative to their uncertainty band. The Benefit score and Signal score are task-wise min–max normalized versions of the benefit effect and Effect/CI signal, respectively. The Composite score is a benefit-weighted scorecard read-out, computed as 0.50 times the Benefit score plus 0.25 times the normalized support score and 0.25 times the Signal score. These score rows are therefore relative geometry diagnostics within each clinical frame, not new clinical endpoints. Figure 10 uses IQR to focus on the middle half of the Best-1 records. This gives a second robust dispersion check that is different from MAD. The pattern remains stable. SCENE has a mean task-level median benefit of 0.265 and an average IQR of 0.089. The best baseline median benefit 39
Mean IQR of Best-1 benefit (lower is better)
0.075
SCENE
favorable high benefit, low uncertainty
0.100
0.125
Task-averaged geometry audit 0.150 SIDES
MaxTE Subgroups
0.175 Virtual Twins Honest Causal Tree 0.200
Read-out
SCENE
Best base.
Margin
Benefit effect Effect/CI signal Benefit score Signal score Composite score
26.7 pp 0.234 0.963 0.904 0.829
11.6 pp 0.127 0.430 0.541 0.520
+15.1 pp +0.107 +0.534 +0.363 +0.309
0.225 Causal Forest CURLS
0.250
0.00
0.05
0.10
0.15
0.20
0.25
Mean task-level median Best-1 benefit
Figure 10: Robust clinical geometry and aggregate read-outs. (Left) Median benefit with task-level IQR. (Right) Method-level averages over the six clinical frames; the margin compares SCENE with the strongest baseline for each read-out.
Table 16: Summary of non-CI benefit–variability diagnostics. The values are method-level averages over the six Table 1 clinical frames. “Adjusted” is center minus variability and is reported only as a numerical read-out; it is not used as an additional figure axis. Diagnostic Mean–SD Median–MAD Median–IQR
SCENE Center
SCENE variability
SCENE adjusted
Best baseline adjusted
0.267 0.265 0.265
0.097 0.062 0.089
0.170 0.203 0.176
-0.053 0.025 -0.066
remains much smaller, and the baseline IQR values are all larger: SIDES and MaxTE Subgroups are around 0.163–0.164, Honest Causal Tree and Virtual Twins are around 0.193, and Causal Forest and CURLS are around 0.242–0.250. The table next to the IQR panel adds a complementary taskaveraged audit from the effect–support–uncertainty scorecard. SCENE has the highest average benefit effect, effect-to-CI signal, normalized benefit score, normalized signal score, and benefit-weighted composite score across the six clinical frames, with consistent margins over the strongest baseline for each read-out. Thus the IQR view and the compact table make distinct points: the figure audits empirical dispersion, while the table summarizes method-level separation on the clinical geometry scorecard without reusing the variability axes. Table 16 summarizes the same observation numerically. If one subtracts the variability estimate from the corresponding benefit center, SCENE remains the top method under all three non-CI diagnostics. The margin is largest in the mean-SD and median-IQR views, where all baseline adjusted values are negative. In the median-MAD view, the best baselines obtain small positive adjusted values, but they remain far below SCENE. This makes the qualitative conclusion robust to the choice of variability statistic. These diagnostics should be interpreted carefully. SD, MAD, and IQR measure empirical variability across archived Best-1 records, not uncertainty in the formal statistical sense and not a replacement for held-out clinical validation. They are useful here because they avoid relying on potentially inconsistent CI endpoints from heterogeneous baseline output files and because they do not expose the number of replay records for each method. The consistent placement of SCENE across mean-SD, median-MAD, and median-IQR views provides a simple robustness check for the main clinical table: the method’s advantage is not limited to a single confidence-interval rendering or a single dispersion statistic. Instead, SCENE remains the method with the highest clinical benefit and the lowest empirical variability across all three descriptive views. 40
Clinical-criterion rates in rule-selected subgroups SCENE median
B-D-BL
prior methods
breast, baseline
Poor grade
LH/FSH <=1.68
ER-negative
Hirsutism <=6
PR-negative
full cohort
P-M-TR
metformin, trajectory
0
50 100 acupuncture, fasting glucose
0
50 acupuncture, OGTT AUC
HbA1c >5.1 0
B-D-TR
50 breast, trajectory
100
P-A-FPG
Poor grade
TG <=1.0
ER-negative
C-peptide low
PR-negative
FPG <=5.5 0
P-M-BL
50 metformin, baseline
100
P-A-AUC
Hirsutism <=6
Adiposity
Insulin AUC low
IR
C-peptide low
100
Triglycerides 0
50
100
0
50
100
Subgroup members satisfying task-specific clinical criterion (%) Figure 11: Task-specific clinical-context audit of reported top rules. Each task block contains three endpoint-relevant non-outcome clinical criteria. Dark vertical markers show the median across actual SCENE Best-1 rules outputs of 50 runs; gray densities summarize replayable Best-1 rules of 50 runs from prior methods; the dashed line is the full-cohort criterion rate. The horizontal axis is the percentage of subgroup members satisfying the displayed criterion. B.2
Clinical Contextualization Advantage over Prior Rule Vocabularies
Table 1 evaluates whether each method’s exported rule improves the held-out clinical endpoint. The present audit asks a different question: when a method commits to its top-ranked rules, do the induced subgroups also concentrate on clinically meaningful, task-relevant axes rather than arbitrary threshold combinations? For each clinical task frame, we define three non-outcome clinical criteria before plotting. The breast-cancer frames use standard prognostic axes from the trial setting: poor histologic grade, estrogen receptor negativity, and progesterone receptor negativity. The PCOS frames use endpoint-matched endocrine, glycemic, lipid, and insulin-resistance axes. For metformin-response tasks these include hirsutism/gonadotropin phenotype and insulin-exposure markers; for the fastingglucose acupuncture task these include fasting-glucose, triglyceride, and C-peptide criteria. For the OGTT AUC acupuncture task, we intentionally reuse the same three metabolic-risk criteria shown in the main-text clinical-risk audit: central adiposity (waist ≥ 80 cm), insulin resistance (HOMA-IR ≥ 2.5), and elevated triglycerides (TG ≥ 1.7 mmol/L). These criteria are used only for post-hoc contextualization auditing and are never used to select or revise the rules. Figure 11 compares actual SCENE Best-1 rule outputs with replayable Best-1 rules from prior methods for each task. For each task-specific clinical criterion, the gray density summarizes the distribution over prior-method top-rule occurrences, while the dark vertical marker reports the median over SCENE top-rule occurrences. We visualize SCENE by its median rather than by a kernel density because the number of SCENE reported top-rule records is smaller than the overall prior-method rule set. The dashed vertical line is the corresponding full-cohort rate. A rightward SCENE median therefore means that SCENE’s reported top rules more consistently identify subgroups concentrated on that clinical axis. Table 17 expands the task-level summary into a full axis-level audit over the predefined non-outcome axes used in Figure 11. This larger table is included for transparency rather than for additional model selection: it shows the cohort rate, the prior-method rule-vocabulary median, and the SCENE 41
Table 17: Axis-level clinical contextualization matrix. Values are percentages of subgroup members satisfying the displayed non-outcome clinical criterion. “Prior” is the median over replayable topranked prior-method rules, and “SCENE” is the median over SCENE top-rule outputs. The axes are fixed contextualization probes and are not used to select or revise rules. This is a post-hoc contextualization audit, not causal feature attribution. Frame
Clinical axis
Cohort
Prior
SCENE
SCENE–prior
SCENE–cohort
B-D-BL B-D-BL B-D-BL
ER-negative PR-negative Poor grade
29.2 36.3 43.7
29.3 36.2 43.7
46.9 47.1 100.0
+17.6 +10.8 +56.3
+17.6 +10.7 +56.3
B-D-TR B-D-TR B-D-TR
ER-negative PR-negative Poor grade
29.2 36.3 43.7
29.5 36.4 44.3
45.6 46.7 100.0
+16.0 +10.2 +55.7
+16.3 +10.4 +56.3
P-M-BL P-M-BL P-M-BL
C-peptide low Hirsutism ≤ 6 Insulin AUC low
50.0 77.3 50.0
49.2 80.0 50.0
63.2 88.7 62.3
+13.9 +8.7 +12.3
+13.2 +11.4 +12.3
P-M-TR P-M-TR P-M-TR
HbA1c > 5.1 Hirsutism ≤ 6 LH/FSH ≤ 1.68
66.2 80.0 45.6
69.0 83.1 47.2
77.1 100.0 73.2
+8.1 +16.9 +25.9
+11.0 +20.0 +27.5
P-A-FPG P-A-FPG P-A-FPG
C-peptide low FPG ≤ 5.5 TG ≤ 1.0
50.3 77.4 31.3
54.1 74.3 31.2
71.7 84.1 48.7
+17.7 +9.8 +17.5
+21.5 +6.7 +17.4
P-A-AUC P-A-AUC P-A-AUC
Adiposity HOMA-IR ≥ 2.5 TG ≥ 1.7
70.3 83.3 29.2
80.2 89.4 33.3
95.6 100.0 42.0
+15.5 +10.6 +8.6
+25.3 +16.7 +12.8
median for every predefined clinical axis used in Figure 11. Reporting the complete matrix makes the contextualization claim less dependent on a single aggregate read-out. Across all six clinical frames, the SCENE median is shifted toward clinically interpretable taskspecific phenotypes relative to the prior-method rule vocabulary. The expanded matrix makes the consistency of this shift visible: all 18 predefined axes have positive SCENE–prior margins. The smallest margin is still positive (HbA1c in the metformin trajectory frame, +8.1 percentage points), while the largest margins occur on the breast-cancer histologic-grade axis, where SCENE subgroups concentrate poor-grade tumors much more strongly than the replayed prior-method rule vocabulary. The PCOS frames show the same pattern on different clinical axes: metformin tasks shift toward endocrine and insulin-exposure phenotypes, whereas acupuncture glucose and OGTT tasks shift toward glycemic, insulin-resistance, adiposity, and triglyceride axes. Thus the table adds information beyond the task-level summary: it shows that the contextualization advantage is not carried by one favorable axis or one disease frame. This does not establish causal feature attribution or clinical treatment guidance. It supports the more limited claim that SCENE’s reported rule vocabulary is more contextualized: the rules that obtain held-out benefit also map to recognizable clinical axes in the corresponding task frame, and this mapping remains visible when the audit is expanded from task averages to the individual clinical probes. B.3
L1000 Proposition Inspectability
SCENE is designed to return executable propositions rather than only a scalar score or an uninterpreted high-performing subset. In the L1000 setting, this means that a discovered rule should be replayable on held-out perturbational signatures and should identify a concrete region of the assay space that can be inspected by standard LINCS/L1000 quality and response measurements. This appendix evaluates that property from two complementary views: a proposition-level selected-versus-background audit and a target-level replay of top-ranked rule sets. Figure 12 reports the selected-versus-background audit, while Figure 13 visualizes the replayed rule unions in L1000 assay space. The selected-versus-background audit in Figure 12 tests whether exported propositions remain meaningful after held-out replay. The unit of the boxplot is a replayed proposition: for each rule, we compute the median metric value among selected signatures and the median value among its 42
SCENE propositions select higher-quality L1000 signatures than background
Up landmark genes
Down landmark genes
median shift +0.29
median shift +102.50
median shift +122.75
0.75 0.60 0.45
SCENE
down landmark genes
Replicate concordance up landmark genes
Distil CC q75
0.90
background
240 160 80
Background
SCENE
240 160 80
Background
0
SCENE
Background
up + down landmark genes
SCENE
Modulated landmark genes 600
median shift +221.75
400
200
SCENE
Background
Figure 12: Matched L1000 replay audit. SCENE-selected signatures are compared with each proposition’s held-out background using complementary L1000 audit metrics. Panels report replicate concordance, up-regulated landmark genes, down-regulated landmark genes, and total modulated landmark genes. Top-ranked SCENE rules show positive median shifts in L1000 assay space RPS6 top-10 union: n=2053
MYC top-10 union: n=677
AURKB top-10 union: n=495
0.8
TAS
0.6
0.4
Median shift: TAS +0.44, SS +9.2
0.2 5
10
Distil SS
15
Median shift: TAS +0.35, SS +4.6 5
10
Distil SS
15
Median shift: TAS +0.39, SS +6.3 5
10
Distil SS
15
Figure 13: Union of Best-1 full-profile SCENE rules across ten target-seed runs. Colored points are signatures covered by at least one selected SCENE rule; gray hexagons show held-out density and arrows show median displacement. corresponding background signatures. The panels use audit metrics that do not duplicate the TAS– Distil SS replay axes: replicate concordance, up-regulated landmark genes, down-regulated landmark genes, and total modulated landmark genes. Across the propositions pooled over RPS6, MYC, and AURKB, median selected-background differences are positive in all four views: +0.29 Distil CC q75, +102.5 up-regulated landmark genes, +122.8 down-regulated landmark genes, and +221.8 total modulated landmark genes. The background boxes are narrow because each background is a large, overlapping held-out complement; this stability makes the selected-background separation more interpretable rather than less informative. Figure 13 complements the proposition-level audit by showing where the selected signatures lie in the assay space. For each target, we take the Best-1 full-profile rule from each of the ten Table 2 target-seed runs and visualize the union of held-out signatures covered by at least one selected rule. The top-rule unions cover 2053 RPS6 signatures, 677 MYC signatures, and 495 AURKB signatures. They shift the covered-signature median by +0.44 TAS and +9.2 Distil SS for RPS6, +0.35 TAS and +4.6 Distil SS for MYC, and +0.39 TAS and +6.3 Distil SS for AURKB. Visually, these median-displacement arrows move toward the high-activity, high-strength region of the L1000 assay space. These positive shifts support the practical inspectability claim: SCENE outputs rules that can be replayed, audited, and interpreted as concrete perturbational contexts. B.4
Holdout Reliability and Support–Gap Diagnostics
A central failure mode of subgroup discovery is rule overfitting: a rule may have a strong discoveryside contrast but fail when replayed on held-out patients. This section therefore audits rule-level replay records instead of only reporting selected task-level summaries. This audit is especially important for SCENE because the lower-level search explores many candidate rules; a credible discovery system must expose where the archive is stable and where it is support-limited. 43
Table 18: Support-floor sensitivity of held-out replay reliability. The table reports cumulative minimum smaller-arm support floors. Retained archive is shown as a percentage of eligible replay records. The gap is |∆train − ∆holdout | and is shown in percentage points. Support floor
Retained (%)
Direction preserved (%)
Held-out positive (%)
Gap median/P75 (pp)
All replayed mmin ≥ 10 mmin ≥ 20 mmin ≥ 50
100.0 67.4 47.9 21.5
84.9 91.8 92.0 98.4
91.3 93.8 94.2 100.0
7.1 / 17.5 5.0 / 8.4 3.9 / 6.9 2.0 / 3.2
Figure 14: Reliability by held-out support bin. The left panel shows the median absolute train– holdout gap by smaller-arm held-out support. The right panel shows the direction-preservation rate. Magnitude reliability improves sharply with support, while direction reliability is generally higher for supported rules but not perfectly monotonic across intermediate bins. Figure 14 shows the held-out effect direction is preserved in 84.9% of evaluations. Direction preservation is 86.4% for full longitudinal contexts and 83.4% for static-only contexts. The median absolute train–holdout gap is 0.071, and the mean absolute gap is 0.175. These values indicate that direction is more reliable than magnitude: many rules preserve whether treatment appears beneficial, but the exact effect size is often attenuated on held-out patients. Table 18 complements the binned figure from a different angle: it treats held-out smaller-arm support as a reporting boundary and asks how the replay audit changes under increasingly conservative minimum-support floors. The rows are cumulative thresholds. This table is therefore a sensitivity check on how much of the replay archive remains interpretable under stricter support requirements, whereas Figure 14 shows where the support–gap relationship is located. Unlike Table 1 D-Cons., which is computed only over committed Best-1 rules, this audit includes the broader eligible replay archive, namely the Top-5 rules from each run. The sensitivity pattern is useful because it separates two claims that can otherwise be conflated. Directional replay is already high when all eligible rules are considered, but magnitude agreement is support-sensitive. As the support floor becomes stricter, a smaller part of the archive remains, while direction preservation and positive held-out replay increase and the train–holdout gap decreases. This supports a conservative reporting practice: SCENE’s replay archive should be interpreted with support-aware uncertainty, rather than treating every discovered rule as equally precise. B.5
Semantic Rule Composition and Feature Provenance
SCENE is designed to produce scenario-grounded rules rather than opaque prediction scores. This section asks what kinds of clinical evidence the discovered rules actually use. We parse clinical rule text into semantic feature families and canonicalize rule predicates to reduce duplicate syntax. This analysis is a syntax and provenance audit, not a causal feature-importance analysis. A rule may contribute to multiple feature families when it combines predicates from multiple evidence sources. 44
Feature-family alluvial map for unique canonical rules Clinical task
Feature family
Rule composition 1 family 3.5%
B-D-BL 18.4%
Baseline 22.7%
2 families 46.9%
B-D-TR 21.5%
Lab 37.3%
P-M-BL 15.8% P-M-TR 15.3%
Pathology 10.7% Dynamic 1.6%
3+ families 49.6%
P-A-FPG 11.1%
Generated 20.0%
P-A-AUC 18.0%
Other 7.6%
Figure 15: Feature-family assignments from clinical rules. The alluvial map links the six clinical tasks, rule-composition breadth, and feature family. Percentages are computed over rule-family assignments derived from the SCENE selected rules used in Table 1; multi-family rules contribute once to each family they use, and rules appearing in multiple tasks contribute once to each observed task. Table 19: Held-out evidence among recurring semantic archetypes. Rows summarize canonical rules in the displayed family combination for which held-out effects are available. Benefit is the bad-outcome reduction in percentage points; positive held-out is the share of those rules with benefit greater than zero. Semantic archetype
Median benefit
Positive held-out
Baseline + Lab
+19.0 pp
94.7%
Lab + Generated
+16.9 pp
94.1%
Baseline + Lab + Generated
+12.6 pp
100.0%
Lab only reference
+14.3 pp
72.7%
Interpretation Patient baseline context combined with direct clinical measurements. Measured variables augmented by derived scenario-specific summaries. Three-source contextualization linking patient state, measurement, and generated summaries. Single-family laboratory thresholds, included as a shortcut reference.
Figure 15 summarizes the SCENE selected rules used in Table 1 as percentages over rule-family assignment flow units. The middle column now shows rule-composition breadth: 3.5% of the displayed assignments come from single-family rules, 46.9% from two-family rules, and 49.6% from rules spanning three or more feature families. On the feature-family side, laboratory measurements are the largest component (37.3%), followed by baseline clinical history (22.7%), generated bins/formulas (20.0%), tumor/pathology descriptors (10.7%), other or unmapped features (7.6%), and longitudinal dynamics (1.6%). Dynamic predicates are therefore present but sparse; the appropriate interpretation is that longitudinal evidence acts as an additional context layer rather than dominating the rule vocabulary. Because multi-family rules contribute once to each used family, the displayed widths are assignment-weighted rather than unique-rule counts. Other includes unmapped or rare feature families. Table 19 adds a rule-internal evidence audit that is not visible in the alluvial map. The alluvial figure answers which feature families appear in the archive; the table asks whether recurring semantic archetypes with held-out replay evidence also retain positive clinical effect estimates. We report only aggregate read-outs for archetypes, not raw archive counts. The rows should be interpreted as syntax/provenance summaries of executable rules, not as causal attribution to individual feature families. 45
Holdout effect versus generalization gap Holdout DeltaConn_test (higher is better)
0.20
favorable: high DeltaConn, low gap
0.15
Below SCENE (ours) SCENE (ours) Minimal w/o Dynamic w/o Pareto w/o Upper w/o Virtual SCENE gap
0.10
0.05
0.00 0.06
0.07
0.08
0.09
0.10
Generalization gap (lower is better)
0.11
Figure 16: L1000 holdout effect versus generalization gap. Each point is a run-aggregated ablation variant. Higher ∆Conn indicates stronger holdout connectivity evidence, whereas lower gap indicates better train–holdout stability. The shaded upper-left region marks the favorable regime. SCENE (ours) combines high holdout effect with the smallest generalization gap among the displayed variants. The dashed vertical line marks the SCENE (ours) gap. Table 20: Minimal-baseline-centered L1000 effect–stability gains. Relative changes use the minimal baseline as the reference. Higher values are better in all three columns; for the gap column, a positive value means that the train–holdout gap is reduced. The table is descriptive and does not introduce a composite score. Profile
Effect lift ↑
Strong-tail lift ↑
Minimal w/o dynamic w/o virtual w/o Pareto w/o upper SCENE (ours)
0.0% +3.0% +2.4% +10.8% +6.6% +13.3%
0.0% +3.3% -12.3% -31.9% +27.2% +50.4%
B.6
Gap reduction ↑ Read-out 0.0% +38.7% +21.6% +27.0% +4.5% +43.2%
Reference discovery regime. Small effect and tail gains; gap decreases. Small effect lift with weaker strong-response tail. Effect lift without strong-tail concentration. Below SCENE (ours); limited gap reduction. Largest balanced gains beyond minimal.
L1000 Ablation and Holdout Evidence Diagnostics
The L1000 scenario evaluates whether the same SCENE framework can produce context-bounded target-response findings outside the clinical-trial setting. Here, the relevant outcome is not clinical efficacy but holdout connectivity evidence, target retrieval, support, significance, and generalization. The following diagnostics expand the main ablation table using the same paired L1000 outputs. They are descriptive because they summarize completed, frozen ablation artifacts and do not change the profile selection or replay protocol. Figure 16 separates the two most important robustness axes: holdout connectivity effect and generalization gap. The favorable region is high DeltaConn and low gap. SCENE (ours) lies closest to this region in the current run summaries, whereas several ablations retain positive DeltaConn but incur larger generalization gaps. Table 20 gives a minimal-baseline-centered view of the same L1000 ablation evidence. Instead of asking whether each profile has one favorable metric in isolation, it asks which profiles move the system beyond the minimal discovery regime along three core axes: held-out connectivity effect, the strong-positive response tail, and train–holdout stability. Table 20 clarifies the contribution of upper-level planning in the L1000 setting. Without the upper layer, the system remains below SCENE (ours) and provides only a small train–holdout gap reduction, suggesting that lower-level search alone is less stable across the holdout boundary. SCENE (ours) 46
Evidence-gated lifecycle of generated features Proposed
Executable
Valid
Search injected
Gate kept
In validated rules
100%
100%
98.6%
87.9%
11.6%
6.5%
Progressive evidence filtering before validated-rule use
Figure 17: Evidence-gated lifecycle of SCENE-generated features. Percentages are measured relative to all proposed generated run-features. Generated features are first required to be executable and non-degenerate, then injected into search, filtered by an evidence gate, and finally used only if they appear in validated rules. The large drop after search injection indicates that SCENE treats generated features as candidate contextualizations rather than automatically reportable discoveries.
exhibits a different profile: it achieves the largest effect lift, the strongest high-response tail, and the largest reduction in the train–holdout gap relative to the minimal baseline. These coupled gains indicate that upper-level planning helps coordinate contextual proposal generation, search, and validation so that the exported propositions preserve target-response signal across the holdout boundary. B.7
Evidence-Gated Lifecycle of Generated Features
SCENE allows the discovery process to introduce scenario-specific generated features, such as longitudinal summaries or perturbational signature summaries. This flexibility is useful for knowledge contextualization, but it also requires a strict execution boundary: generated features should not be treated as free-form model suggestions. They must be materialized from registered scenario tables, pass deterministic validity checks, enter the search only as executable feature objects, and survive evidence-based filtering before they can appear in reported rules. Figure 17 audits this lifecycle over the completed discovery records. The denominator is the set of proposed generated run-features. Nearly all proposals are executable and non-degenerate after deterministic construction, and most are injected into the search frontier. The sharp reduction occurs at the evidence gate: only a small subset is retained after the search evaluates whether the generated feature provides useful frontier evidence. An even smaller subset appears in the final validated rules. This pattern is desirable for SCENE. It shows that the framework uses generated features to expand the grounding space, but does not allow this expanded space to directly determine reported propositions. Instead, generated features must first become valid scenario objects and then earn their place through the same evidence-gated search process as ordinary literals. This audit supports the main claim that SCENE performs controlled knowledge contextualization rather than unconstrained feature invention. Broad biomedical cues can motivate new scenariospecific summaries, but the executable boundary is governed by the schema, construction grammar, and evidence gate. The final reported rules therefore reflect a conservative subset of generated features: features that are materialized, validated, searched, and retained with supporting evidence. The figure should be read as a protocol audit rather than an additional model-selection step; all quantities are computed from completed runs and are not used to revise the discovery process. B.8
Grounding Trace from Prior Knowledge to Executable Propositions
The main quantitative results evaluate whether SCENE discovers useful rules or findings, but they do not by themselves show how a broad biomedical direction becomes a concrete scenario-level proposition. Figure 18 provides this missing trace. Each card follows one selected SCENE output from a knowledge direction, through grounded feature families, into an executable rule and an associated held-out evidence record. Trajectory examples use predeclared landmarked or trajectory-enhanced features and should be read as predictive subgroup hypotheses, not baseline-only treatment-decision rules. 47
From broad concept to scenario-grounded rule Knowledge direction
Scenario-grounded rule
Grounded feature families
Held-out evidence
C1 Breast cancer trial, longitudinal subgroup Early treatment response should be interpreted with endocrine tumor status and baseline...
Baseline
Generated
Dynamic
Pathology
early platelet change remains low; estrogen-receptor positive status is absent; and no injury or poisoning history.
Risk-difference benefit: +6.0 pp held-out min-arm n: 27 Gap: 0.7 pp · direction preserved Train Held-out
C2 Breast cancer trial, laboratory-pathology subgroup Renal and hepatic laboratory context can refine pathology-defined clinical benefit strata.
maximum creatinine during cycles 1-4 is higher; poor or undifferentiated tumor grade is present; and alkaline phosphatase is not elevated.
Dynamic
Lab
Pathology
Risk-difference benefit: +13.3 pp held-out support: 15 Gap: 1.6 pp · direction preserved Train Held-out
C3 PCOS metformin trial, endocrine subgroup Endocrine phenotype markers can identify a metformin-response subgroup in the PCOS...
Baseline
baseline LH bin is low; baseline hirsutism score is lower; and total testosterone is observed.
Generated Lab
Risk-difference benefit: +12.6 pp held-out support: 15 Gap: 18.7 pp · direction preserved Train Held-out
C4 L1000 perturbation extension Pathway state and transcriptional signature strength ground targetconnectivity enrichment.
Cell context
non-CD34 cell context; dose-range signature span is high; and DNArepair-like program is reduced.
Pathway program Signature magnitude
Connectivity delta: +0.249 held-out subgroup n: 292 Gap: 0.029 · direction preserved Train Held-out
Figure 18: Grounding trace from broad biomedical directions to executable scenario-grounded propositions. Each card links a knowledge direction to grounded feature families, a readable executable rule, and held-out evidence. Clinical cards report risk-difference benefit in percentage points; the L1000 card reports a held-out connectivity delta. The cards are selected examples for inspecting SCENE’s reporting path, not an additional model-selection procedure. The four cards were selected to be illustrative but evidence-bearing rather than purely anecdotal. The first three cards come from clinical-trial subgroup discovery and show breast-cancer and PCOS examples in which baseline descriptors, laboratory measurements, longitudinal summaries, generated feature bins, pathology descriptors, or endocrine markers are combined into executable subgroup rules. The last card comes from the L1000 perturbational setting and shows the same grounding idea in a different schema: a target-connectivity direction is instantiated through cell context, pathwayprogram state, and transcriptional signature magnitude. In all four cases, the output is not just a natural-language explanation. The rule is an executable predicate over scenario-visible features, and the rightmost column reports post-discovery evidence on held-out or validation data. This trace is important for SCENE’s claim of knowledge contextualization. A static knowledgeinjection approach can state that a biomedical concept is relevant, and a data-only rule miner can return a high-scoring rule, but neither necessarily records how the concept is grounded into the current schema. SCENE explicitly preserves this chain: a direction proposes the semantic intent, grounding maps it to feature families, search turns those families into a concrete rule. The clinical cards show positive risk-difference benefit estimates after replay, while the L1000 card shows a positive connectivity delta for a context-specific perturbational subgroup. These examples therefore make the reported propositions inspectable. The figure is not intended as an aggregate performance comparison and should not be interpreted as independent biomedical validation of the displayed mechanisms. Rather, it is a qualitative audit of the reporting boundary. It demonstrates that SCENE exports propositions with a visible grounding path, rather than reporting isolated rule strings or post-hoc textual rationales.
C
Limitations
SCENE is designed for biomedical hypothesis generation. The clinical analyses identify scenariogrounded subgroups with held-out outcome contrasts, but these propositions are not diagnostic criteria, treatment recommendations, or patient-level decision rules without independent validation. Live language-model calls can propose different directions or grounding cues across providers, model versions, and decoding settings. Similarly, L1000 replay provides perturbational assay evidence for prioritizing follow-up, not wet-lab mechanism confirmation or drug-efficacy evidence. The experiments cover the clinical task frames and L1000 target programs studied in this paper; applications to new endpoints, cohorts, assays, or disease areas should be accompanied by taskspecific validation and domain review. These results should therefore be interpreted as evidence 48
that broad biomedical knowledge can be contextualized into auditable propositions, rather than as a substitute for prospective clinical or experimental studies.
D
Broader Impacts
SCENE may improve biomedical discovery workflows by making hypothesis generation more transparent: outputs are expressed as executable rules with source-traceable search directions, support/effect summaries, and replay diagnostics that can be inspected by domain experts. This can help researchers organize large clinical or perturbational datasets and prioritize candidates for follow-up analysis. More broadly, SCENE points toward auditable auto-research workflows in which agents propose executable hypotheses rather than ungrounded claims; however, such automation can also create automation bias if concise rules are trusted without expert review. The main risk is overinterpretation of concise rules as actionable medical or therapeutic claims. Responsible use therefore requires expert review, independent validation before downstream decisions, and appropriate governance for data privacy, security, and fairness when patient data are involved. Clear reporting of validation scope and uncertainty diagnostics is important to keep the resulting propositions in their intended exploratory role. Finally, multi-agent discovery has computational and operational costs. Live LLM calls, evolutionary search, and repeated validation can increase compute use and complicate reproducibility. The implementation choices in this paper emphasize typed role outputs, deterministic validation, and reusable manifests so that future users can limit unnecessary calls, replay completed runs, and audit the provenance of reported propositions.
49