ConceptioArchivearXiv CS
arXiv CSopen access

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports Yufan Wang∗ , Anit Kumar Sahu¶ , Yan Fei Ng† , Daniel Kang∗ , Shayan Vassef∗ , Soorya Ram Shimgekar∗ , Koustuv Saha§ , Piyum Zonooz∗ , Navin Kumar∗ , Chee Leong Cheng†‡ , and Li Yan Khor†‡ ∗ Nimblemind.ai, USA † Department of Anatomical Pathology, Singapore General Hospital, Singapore ‡ Duke–NUS Medical School, Singapore § University of Illinois Urbana-Champaign

arXiv:2607.06435v1 [cs.AI] 7 Jul 2026

¶ Independent Researcher

Abstract—Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, and its eradication is key to gastric cancer prevention. However, evidence supporting H. pylori positivity and H. pylori-associated gastritis may be distributed across heterogeneous coded and free-text report fields and may require contextual interpretation of assertion and negation, limiting keyword search, and making manual review difficult to scale. We conducted a retrospective pilot evaluation of the Nimblemind Multi-Agent System (nMAS), a field-name-driven, evidence-linked extraction workflow, using 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. Four clinician-scoped binary fields were evaluated: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Across 216 feature-case decisions, nMAS correctly classified 213, corresponding to 98.61% overall accuracy. A separately implemented UMA-style MiniMax M2.5 comparator produced similar aggregate and per-field classification metrics. Although predictive performance was similar, nMAS maintained unified report-level outputs with supporting source sentences; the demonstrated contribution is therefore workflow integration and traceability rather than predictive superiority. Under an illustrative, unmeasured scenario, reviewing 1,000 reports at five minutes per manual review versus five seconds per evidence-linked verification would reduce review time from 83.3 to 1.4 staff-hours, corresponding to 81.9 staff-hours and about USD 6,100 in potential staff-time value. Larger multiinstitutional studies should evaluate evidence-span correctness, clinician verification time, and generalizability. Index Terms—clinical NLP, pathology reports, large language models, multi-agent systems, information extraction, Helicobacter pylori, gastric biopsy, H. pylori-associated gastritis

I. I NTRODUCTION About 31% of the Singapore population was estimated to have evidence of Helicobacter pylori infection, with prevalence increasing with age and varying across ethnic groups [1]. Persistent H. pylori infection is associated with chronic active gastritis and peptic ulcer disease, while eradication is central to gastric cancer prevention [2]–[4]. Reliable identification of biopsy-confirmed H. pylori-positive reports is therefore important for treatment review, eradication follow-up, clini-

cal audit, research cohort assembly, and quality-improvement workflows. Pathology-based case finding is not a simple keywordsearch task. Identical organism terms may appear in affirmative, negated, historical, or ancillary-stain contexts; for example, “Helicobacter organisms are identified” and “No Helicobacter organisms are identified” contain the same target terms but require opposite labels. Relevant evidence may also be distributed across coded specimen fields, specimen labels, diagnoses, microscopic descriptions, and ancillary-test comments, making manual review difficult to scale. Slow manual review can delay treatment review, eradication followup, and quality-improvement action, which may limit timely patient care [5]–[7]. At five minutes per report, screening 1,000 candidate pathology reports would require about 83 staff-hours before downstream review or analysis could begin. Rule-based and supervised extraction systems can improve classification but often require task-specific annotation, maintenance, or retraining as report templates and target schemas change. Large language models provide greater configurability, but extracted labels remain difficult to use for clinical audit or research unless they are linked to supporting source text. These limitations motivate a reusable extraction workflow that combines configurable field definitions, source-text grounding, validation, and clinician-reviewable outputs [8]–[11]. To address this gap, we evaluate the Nimblemind Multi-Agent System (nMAS), a modular, field-name-driven, evidence-linked workflow configured for H. pylori-related feature extraction from gastric biopsy pathology reports. nMAS routes user-specified target fields through tiered extraction modules and returns structured labels with supporting sourcetext evidence for clinician verification. In this pilot study, we applied nMAS to 54 de-identified gastric biopsy pathology reports from a large healthcare system in Singapore. We evaluated four clinician-scoped binary fields: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. We assessed whether the workflow could support initial H. pylori case finding while

preserving source-linked evidence for clinician verification. II. R ELATED W ORK Baseline Clinical and Pathology Information Extraction. Manual review remains the reference standard when interpretation depends on section context, negation, ambiguous wording, and local reporting conventions, but it is difficult to scale. In breast pathology abstraction, Wieneke et al. found that an NLP system still flagged 49.1% of reports for manual review and incorrectly coded 30.8% of reports [12]. Keyword-based retrieval can identify reports containing predefined terms but does not by itself resolve assertion status, section location, or diagnostic context [8], [13]. Rule-based NLP systems extend keyword matching through dictionaries, regular expressions, section parsing, assertion detection, and negation handling. For example, the Mayo Clinical Text Analysis and Knowledge Extraction System (cTAKES) demonstrated modular clinical text processing across electronic health record applications [13], and rule-based methods have also been applied to pathology report extraction [8], [14]. Learning-Based and LLM-Based Medical Abstraction. Supervised and transformer-based methods can model clinical context beyond exact keyword matching. Domain-specific language models such as BioBERT have improved performance across biomedical named-entity recognition, relation extraction, and question-answering tasks [15]. However, supervised clinical information extraction systems typically depend on task-specific annotated data and model development, which can limit their flexibility when extraction targets or output schemas change [8], [16]. This limitation is relevant when clinicians need to define new variables for different audits, research cohorts, or disease-specific workflows. Large language models (LLMs) have enabled more flexible zero-shot and prompt-based abstraction. UniMedAbstractor (UMA) uses configurable prompt templates and frontier models to structure multiple clinical attributes from real-world data [16]. In pathology, Truhn et al. evaluated GPT-4 for zero-shot extraction from colorectal cancer and glioblastoma histopathology reports [9], while Balasubramanian et al. compared multiple LLMs for structured extraction from breast cancer pathology reports [10]. These studies demonstrate that LLM-based abstraction can be configured for different extraction schemas. However, clinical use still requires reproducible schemas, reliable handling of uncertainty and negation, and mechanisms for identifying unsupported outputs [17], [18]. Current Solutions in Singapore and the Region. Clinical natural language processing (NLP) has also been evaluated in Singapore and other Asian healthcare settings, but directly comparable work on H. pylori gastric biopsy pathology extraction remains limited. In Singapore, Hardjojo et al. developed the Clinical History Extractor for Syndromic Surveillance (CHESS), a rule-based system that extracted infectiousdisease symptoms from free-text primary-care electronic medical records and distinguished affirmed, negated, and suspected assertions [19]. Tay et al. evaluated NLP approaches to infer metastatic disease sites from radiology reports in a Singapore

oncology setting [20]. These studies support the feasibility of structuring local clinical free text, but they address primarycare surveillance and oncology radiology rather than gastric biopsy pathology. Regional gastrointestinal studies provide closer comparisons to the present task. In Korea, Song et al. developed an NLP pipeline for extracting gastric disease information from unstructured esophagogastroduodenoscopy reports and linked pathology reports, including categories related to H. pyloriassociated gastritis [21]. Bae et al. developed a related pipeline for extracting quality indicators from free-text colonoscopy and pathology reports, demonstrating how automated abstraction can support large-scale gastrointestinal quality monitoring [22]. These studies show that gastrointestinal report extraction is feasible in Asian healthcare settings. However, they primarily evaluate task-specific pipelines and do not address a configurable, field-name driven, evidence-linked multi-agent workflow for Singapore gastric biopsy reports. Collectively, prior studies show that clinical free-text abstraction and gastrointestinal report extraction are feasible in local and regional healthcare settings [19]–[22]. However, they do not directly address the present operational use case: configurable extraction of H. pylori-related features from Singapore gastric biopsy pathology reports while preserving source-text evidence for clinician verification. This gap matters because biopsy-confirmed H. pylori case finding supports eradication follow-up, treatment-outcome audits, research cohort assembly, and quality-improvement review [5]–[7]. The present study addresses this gap by evaluating nMAS across four clinician-scoped target features: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. III. DATA Dataset Overview. The evaluation dataset consists of 54 de-identified gastric biopsy pathology report records from the Department of Anatomical Pathology, Singapore General Hospital (SGH). SGH is a large tertiary hospital serving a multi-ethnic population [23], [24]. Each report-level record contains de-identified administrative metadata, coded specimen and diagnosis fields, ordering and procedure information, sign-out information, and the full pathology report text. In the source table, these fields include specimen identifier, receive date, de-identified patient fields, TCode and MCode fields, diagnosis report text, order location, procedure code, sign-out field, reference target labels, and reference evidence spans. Direct patient identifiers were removed or replaced with de-identified placeholders before analysis, while clinically relevant wording needed for extraction was preserved. The reports represent heterogeneous free-text reporting styles across different pathologists, including variation in specimen labels, diagnostic phrasing, organism descriptions, negation patterns, microscopic descriptions, and ancillary test wording. A typical record may include both coded tissue fields, such as ‘T57010 (Gastric biopsy),” and diagnostic text, such as ‘Stomach; biopsy. Severe chronic acute antral and body

gastritis with mild colonization by Helicobacter pylori.” This illustrates why the extraction task requires both structured and unstructured evidence: coded fields help identify the gastric biopsy specimen, while diagnostic text may support H. pylori positivity and H. pylori-associated gastritis. Table I shows representative de-identified examples from the evaluation dataset with their corresponding reference labels. These examples are included to illustrate the evaluationdata format and gold-label definitions. Prompt-level examples, field definitions, and guardrails used to guide extraction are described separately in the Method section. Target Features. Four binary target features were evaluated. These features were selected through a clinician-informed scoping process. An initial broader set of potential gastric biopsy extraction targets was reviewed with four clinicians, including senior pathologists, to identify the minimum structured information needed for practical H. pylori-related case finding. Gastric/stomach biopsy and biopsy status were included as cohort-definition fields to confirm the anatomical site and specimen type. H. pylori positivity and H. pylori-associated gastritis were included as disease-relevance fields. Together, the four features distinguish relevant gastric biopsy specimens from non-target records and separate organism detection from an explicit diagnostic association between H. pylori and gastritis. 1) Gastric/Stomach Biopsy: whether the report indicates a gastric or stomach biopsy specimen. 2) Biopsy: whether the case is a biopsy specimen rather than a non-biopsy specimen type. 3) H. pylori Positive: whether H. pylori, Helicobacter organisms, or Helicobacter-like organisms are identified. 4) H. pylori Gastritis: whether the report supports H. pylori-associated gastritis, such as active chronic gastritis explicitly associated with H. pylori. Reference Standard and Evidence Annotation. Clinicianreviewed reference labels were defined for each report and each of the four target features. Reference labels were determined using the complete de-identified report record, including coded fields, specimen labels, final diagnosis text, microscopic descriptions, organism-related statements, and ancillary stain comments where available. When initial labeling disagreements occurred, they were resolved through adjudication by an additional pathologist, and the final adjudicated labels were used as the reference standard. Reference evidence spans were recorded, where applicable, to document the source text supporting each reference label. IV. M ETHOD We evaluated the Nimblemind Multi-Agent System (nMAS), a modular clinical information extraction workflow configured for field-name driven extraction from de-identified gastric biopsy pathology reports. The workflow was applied to the four target features defined in Section III-B: gastric/stomach biopsy, biopsy status, H. pylori positivity, and H. pylori-associated gastritis. Fig. 1 summarizes the implemented workflow.

Workflow Overview. As shown in Fig. 1, nMAS processes an input pathology report and a set of user-specified target fields through four stages: request assessment and parsing, complexity-based feature extraction, field-level result aggregation, and source-grounded output validation. The Achievability Agent and Query Parser prepare supported field requests for tiered extraction; FE-MUX then combines the resulting fieldlevel outputs, which are validated against the source report before the final structured output is returned. Input Standardization and Field Preparation. Before extraction, each report record was converted into a consistent textual representation. Formatting differences involving line breaks, spacing, section separators, and repeated administrative headers were normalized where appropriate. Clinically relevant content was preserved, including coded specimen information, specimen labels, final diagnosis text, microscopic descriptions, organism-related statements, ancillary test results, and negation. Each target field was paired with an entry in the configurable FIELD_LIBRARY. A field entry specified the target meaning, expected output format, clinical context, in-context demonstrations, and field-specific guardrails. Clinical input from four clinicians, including senior pathologists, informed the target-field definitions and the clinically relevant distinctions encoded in the demonstrations and guardrails. These components guided interpretation of the requested field at inference time. The four target fields were configured according to the definitions described in Section III. Achievability Agent and Query Parsing. The Achievability Agent assessed whether the input contained readable pathology content and whether each requested field was represented in the configured extraction schema. Requests lacking usable report text or a valid field definition were returned as unsupported rather than being assigned an inferred value. For supported requests, the Query Parser constructed a field-specific extraction request containing the report text, target field, corresponding field definition, and expected output schema. It then assigned the request to an extraction tier according to the configured linguistic and reasoning requirements of the field. Routing did not use the reference labels. Tiered Feature Extraction and Prompt Design. After query parsing, each supported field request was assigned to one of three extraction routes according to its configured linguistic and reasoning requirements. Tier 1, Named Entity Recognition Feature Extraction, was used for fields supported primarily by direct lexical or coded evidence. This route consisted of a pre-processor, a natural language processing extraction component, and a post-processor that converted the result into the required format. Tier 2, Small Language Model Feature Extraction, and Tier 3, Large Language Model Feature Extraction, were used for fields requiring contextual interpretation. Model assignments were determined before the present evaluation through engineering experiments using the same tiered nMAS workflow on a separate pathology information-extraction task. The broader candidate screening included the models ultimately

TABLE I R EPRESENTATIVE DE - IDENTIFIED GASTRIC BIOPSY PATHOLOGY EXCERPTS AND THEIR REFERENCE TARGET LABELS .

Case

Report Excerpt

Reference Target Labels

A

“Stomach; biopsy. Severe chronic acute antral and body gastritis with mild colonization by Helicobacter pylori.”

Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: Y; H. pylori gastritis: Y.

B

“Gastric biopsy: Chronic gastritis. No Helicobacter organisms are identified. Helicobacter organisms are absent.”

Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: N; H. pylori gastritis: N.

C

“Gastric biopsy: Helicobacter pylori-associated active chronic gastritis. Helicobacter organisms are identified.”

Gastric/stomach biopsy: Y; Biopsy: Y; H. pylori positive: Y; H. pylori gastritis: Y.

User Request Input Document

Achievability Agent

Field Names

Query Parser

Tier 1

Tier 2

Tier 3

Named Entity Recognition (NER) Feature Extraction

Small Language Model (SLM) Feature Extraction

Large Language Model (LLM) Feature Extraction

Pre-Processor

Pre-Processor

Pre-Processor

Natural Language Processing (NLP)

Small Language Model (SLM)

Large Language Model (LLM)

Post-Processor

Post-Processor

Post-Processor

Feature Extraction Multiplexer

Output Validation

Extraction Output

Fig. 1. Overview of the nMAS field-name driven extraction workflow. Target fields are routed through complexity-ranked extraction tiers, merged, validated against source-text evidence, and returned as structured outputs.

selected for deployment, Qwen2.5-7B-Instruct and DeepSeekV4-Flash, together with additional candidate models including

GLM-5, Gemini 3.1 Pro, DeepSeek V3.2, Kimi K2 Thinking, and Qwen3-Next-80B [25]–[31]. Models were compared

based on field-level extraction performance, structured-output reliability, source-text grounding, negation handling, and runtime efficiency. Following this broader screening, DeepSeek-V4-Flash [26] was evaluated in the deployment workflow and achieved the best overall performance across these criteria. It was therefore selected for Tier 3, whereas Qwen2.5-7B-Instruct [25] was selected for Tier 2. Neither the 54 reports in the present study nor their reference labels were used for model selection or training. Tier 2 was used for bounded contextual extraction, whereas Tier 3 was used for fields requiring broader contextual interpretation, including assertion-status assessment, negation handling, and the diagnostic association between H. pylori and gastritis. Each language-model route used a pre-processor to combine the report text with the relevant field definition, a model component to extract the requested information, and a post-processor to normalize the result into the required schema. The extraction prompt used a shared pathology informationextraction instruction instantiated with the report text and requested target field. Field-specific interpretation was supplied through the corresponding FIELD_LIBRARY entry, while the detailed demonstrations and guardrails are described in the following subsection. The model was required to return structured JSON with the extracted value and a verbatim supporting source sentence. Prompt 1: Core extraction instruction Use only information explicitly present in the report. Do not infer missing facts. Interpret the target feature using the field library, including its context, examples, and guardrails. If the requested value is absent or unclear, omit the feature. Every retained extraction must include a verbatim supporting source sentence, a confidence score, and explanatory notes.

These global guardrails applied to all target fields. They prohibited unsupported inference, required direct source-text evidence for every retained value, preserved conflicting supported mentions for review, and enforced the predefined JSON schema. When evidence appeared in multiple sections, definitive diagnostic interpretation was prioritized over less authoritative content such as clinical history or administrative metadata. Prompt 2: H. pylori-specific field guidance For h_pylori_positive, assign Y only when the report affirmatively identifies H. pylori, Helicobacter organisms, or Helicobacter-like organisms in the gastric specimen. Assign N when the report explicitly states absence or nonidentification, such as ‘‘No Helicobacter organisms are identified’’ or a negative H. pylori immunostain. Negation overrides keyword presence. For h_pylori_gastritis, assign Y only when the report explicitly links gastritis to H. pylori or Helicobacter organisms, such as ‘‘H. pylori-associated active chronic gastritis.’’ Gastritis alone is insufficient, and organism positivity alone is insufficient unless the report makes the diagnostic association clear. Assign N when gastritis is present but H. pylori is absent or not diagnostically associated with the gastritis.

Field-Specific In-Context Demonstrations and Guardrails.

Each field-specific demonstration paired a short pathologystyle excerpt with an expected field value and a guardrail defining the relevant extraction boundary. As illustrated in Table II, the demonstrations addressed gastric and stomach terminology, common biopsy abbreviations, affirmative and negated organism findings, ancillary-stain evidence, and the requirement for an explicit diagnostic association before assigning H. pylori-associated gastritis. The complete set of demonstrations is provided in Appendix Table V. Feature Extraction Multiplexing. Outputs from the active extraction routes were passed to FE-MUX, which assembled the field-level results into a single report-level structure. The merged representation retained each predicted value together with its source sentence, date, confidence score, and explanatory notes where available. Schema compliance and sourcegrounding checks were performed by the subsequent validation stage. Output Validation and Generation. The validation stage assessed both output structure and source grounding. It checked that each result followed the required schema, that the reported source sentence appeared verbatim in the input report, and that the sentence supported the predicted value in context. Positive labels required affirmative source evidence. Negative labels required explicit absence or negation rather than lack of mention. For H. pylori-associated gastritis, evidence of organism positivity alone was insufficient without diagnostic evidence linking H. pylori to gastritis. Missing, unsupported, or internally inconsistent outputs were flagged rather than accepted as fully grounded predictions. The generation stage then assembled the validated fields into the final report-level output for clinician review. External UMA-Style Benchmark. To provide a relevant external comparator, we implemented a Universal Abstraction (UMA)-style benchmark based on the one-attribute prompting strategy described by Wong et al. [16]. UMA was selected because it is a validated, schema-conditioned clinical abstraction framework designed for configurable attribute extraction without task-specific model training. This made it a closer comparator to the present field-name driven pathology task than a keyword search, a disease-specific rule set, or a supervised model requiring newly annotated training data. The benchmark used the same 54 report records and the same four target features as the nMAS evaluation. For each report-feature pair, the model received only the Diagnosis (full report) text, the target attribute name, a concise target definition, fieldspecific guidance, and positive and negative examples derived from the feature definitions in this manuscript. The prompt deliberately omitted reference labels, nMAS outputs, row-level correctness indicators, and any other gold-label information so that the benchmark measured independent extraction rather than agreement with known answers or imitation of nMAS behavior. Each UMA call requested one binary value and source evidence in a compact JSON object:

TABLE II R EPRESENTATIVE H. pylori-S PECIFIC ICL D EMONSTRATIONS AND F IELD -M APPING G UARDRAILS Target Feature

Pathology-Style Demonstration Excerpt

Expected Field Mapping

Gastric Biopsy

“Stomach; biopsy.”

gastric_biopsy = Y when the specimen is Accept stomach and gastric explicitly from stomach or gastric tissue. as equivalent site terms.

Biopsy

“GASTRIC BX.”

biopsy = Y when biopsy is expressed using Accept common biopsy aban accepted abbreviation such as “BX.” breviations.

H. pylori Positive

“No Helicobacter organisms are identified.”

h_pylori_positive = N when the organ- Negation overrides keyword ism statement is explicitly negated. presence.

H. pylori Gastritis

“Helicobacter pylori associated active chronic gastritis.”

h_pylori_gastritis = Y when gastritis is Require explicit diagnostic explicitly associated with H. pylori. association.

{ "attribute": "h_pylori_positive", "value": "Y|N|null", "evidence": "verbatim quote", "confidence": 0.0, "notes": "short reason" }

The benchmark was run using MiniMax M2.5 [32], which was selected as a practical strong-LLM comparator because it was available through the deployment channel used for the broader benchmarking workflow and provided reliable structured JSON responses in preliminary checks. Reference Standard and Evaluation Metrics. Extraction performance was evaluated using the clinician-reviewed reference labels and evidence annotations described in Section III-C. Each of the 54 reports contributed one prediction for each of the four target fields, yielding 216 feature-case evaluations for nMAS and 216 feature-case evaluations for the UMA-style MiniMax benchmark. A prediction was counted as correct when its binary value matched the corresponding reference label. Missing, invalid, unsupported, or non-parseable predictions were counted as incorrect rather than being converted into negative labels. Accuracy was calculated separately for each field and across all feature-case evaluations. Positive-class precision, recall, and F1 score were also calculated using the clinician-reviewed labels as the reference standard. Tier-3 Prompt-Component Ablation. To examine the contribution of field-specific prompt guidance, we conducted a paired ablation experiment on the two Tier 3 fields: H. pylori positivity and H. pylori-associated gastritis. These fields were selected for ablation because they were determined by clinical experts to be the most clinically relevant context-dependent targets for H. pylori-related case finding. Unlike the specimenrelated fields, they require interpretation of assertion status, negation, ancillary-stain findings, and the explicit diagnostic association between H. pylori and gastritis. The full condition used the shared Tier 3 extraction prompt together with the complete field-specific FIELD_LIBRARY entries, whereas the ablated condition omitted the field-specific definitions, aliases, clinical context, guardrails, and in-context demonstrations. For this ablation analysis, reports containing exact textual overlap with the in-context demonstrations were excluded to prevent demonstration wording from affecting the comparison

Guardrail Demonstrated

between conditions. Both conditions were therefore evaluated on the same remaining 41 reports, yielding 82 paired featurecase decisions per condition. V. R ESULTS The nMAS workflow correctly classified 213 of 216 featurecase decisions across 54 de-identified gastric biopsy pathology reports, corresponding to an overall feature-case accuracy of 98.61% (Table III). Gastric/stomach biopsy identification and biopsy status each achieved 100.00% accuracy. Accuracy was 98.15% for H. pylori positivity and 96.30% for H. pyloriassociated gastritis. The three observed errors occurred in these two disease-relevance fields. The external UMA-style MiniMax M2.5 baseline produced the similar aggregate and per-field performance as nMAS. Both methods correctly classified all gastric/stomach biopsy and biopsy-status decisions and produced the same three errors in the two H. pylori-related fields. Representative Evidence-Linked Output. In addition to binary feature labels, nMAS returned source-text evidence and supporting metadata for field-level review. The following representative output illustrates the structure returned for an affirmative H. pylori organism finding: { "h_pylori_positive": { "value": "Y", "source_sentence": "MILD COLONIZATION BY HELICOBACTER PYLORI.", "date": null, "confidence_score": 5, "notes": "Organism positivity is explicitly stated." } }

This output format allowed each predicted label to be reviewed together with the source sentence supporting the extraction. Tier-3 Field-Library Ablation. Table IV summarizes the controlled field-library ablation on the 41 reports remaining after exclusion of exact textual overlap with the ICL demonstrations. The same report set was used for both the full and ablated conditions. Across the 82 feature-case decisions per condition, both settings achieved high label-level performance. The full condition correctly classified 81 of 82 decisions, corresponding to 98.78% accuracy and 98.88% positive-class F1. The ablated condition produced the same label-level performance after

TABLE III N MAS AND E XTERNAL UMA-S TYLE M INI M AX M2.5 P ERFORMANCE ACROSS 54 G ASTRIC B IOPSY PATHOLOGY R EPORTS

Method

Feature

N

Correct

Acc.

Prec.

Rec.

F1

nMAS nMAS nMAS nMAS

Gastric/stomach biopsy Biopsy H. pylori positive H. pylori gastritis

54 54 54 54

54 54 53 52

100.00% 100.00% 98.15% 96.30%

100.00% 100.00% 100.00% 96.30%

100.00% 100.00% 96.43% 96.30%

100.00% 100.00% 98.18% 96.30%

UMA + MiniMax M2.5 UMA + MiniMax M2.5 UMA + MiniMax M2.5 UMA + MiniMax M2.5

Gastric/stomach biopsy Biopsy H. pylori positive H. pylori gastritis

54 54 54 54

54 54 53 52

100.00% 100.00% 98.15% 96.30%

100.00% 100.00% 100.00% 96.30%

100.00% 100.00% 96.43% 96.30%

100.00% 100.00% 98.18% 96.30%

nMAS overall UMA + MiniMax M2.5 overall

Feature-case level Feature-case level

216 216

213 213

98.61% 98.61%

99.38% 99.38%

98.77% 98.77%

99.08% 99.08%

TABLE IV T IER -3 PROMPT- COMPONENT ABLATION AFTER EXCLUDING REPORTS WITH EXACT TEXTUAL OVERLAP WITH THE ICL DEMONSTRATIONS . Cond.

Feature

Full Full Full

H. pylori positive 41 H. pylori gastritis 41 Overall 82

N Corr. 41 40 81

100.00% 100.00% 100.00% 97.56% 100.00% 97.78% 98.78% 100.00% 98.88%

Acc.

Prec.

F1

Ablated H. pylori positive 41 Ablated H. pylori gastritis 41 Ablated Overall 82

41 40 81

100.00% 100.00% 100.00% 97.56% 100.00% 97.78% 98.78% 100.00% 98.88%

excluding ICL-overlap reports. Performance for H. pylori positivity was perfect in both conditions. The only remaining error in both settings occurred for H. pylori-associated gastritis in the same report. Removing field-specific guidance did not change label-level performance in this leakage-controlled pilot ablation; both conditions produced the same overall accuracy, F1 score, and remaining H. pylori-associated gastritis error. Error Analysis. The three observed errors were confined to the two context-dependent disease-relevance fields. They comprised one false positive for H. pylori-associated gastritis and two false negatives from a single report: one for H. pylori positivity and one for H. pylori-associated gastritis. No errors occurred in either gastric/stomach biopsy identification or biopsy status. The same three errors were produced by nMAS and the external UMA-style MiniMax M2.5 baseline. VI. D ISCUSSION We highlight the following core findings. First, specimenrelated fields were extracted correctly in all 54 reports, reflecting their reliance on explicit specimen labels or coded evidence. Second, all three errors occurred in the more contextdependent H. pylori positivity and H. pylori-associated gastritis fields, indicating that assertion status, negation, ancillarystain findings, and explicit diagnostic association remain the main technical challenges.

These error patterns also explain why label-level performance alone is insufficient for this use case. In pathology case finding, a reviewer must be able to verify whether a predicted label is supported by the relevant sentence, especially when the same organism terms can appear in affirmative or negated contexts. Similar predictive performance therefore does not establish operational equivalence. UMA outputs required subsequent assembly into report-level results, whereas nMAS maintained one reviewer-facing contract across configuration, extraction, validation, and delivery. Field definitions, validation rules, or extraction models can therefore be revised without changing the clinician-facing output. The contribution is workflow integration and traceability rather than superior performance, consistent with the need for reusable clinical AI infrastructure [11]. The ablation analysis showed identical label-level performance with and without field-specific guidance. This finding should be interpreted in the context of the pilot task: the two Tier 3 H. pylori fields were clinically important but relatively constrained, and many reports contained direct diagnostic wording or explicit negation. The result therefore does not imply that field-specific guardrails and demonstrations are unnecessary in general. Rather, it suggests that the baseline prompt was already sufficient for this small and comparatively straightforward extraction task. For future requests involving larger schemas, more ambiguous clinical concepts, cross-sentence temporal reasoning, unsupported labels, or institution-specific terminology, field-specific guidance is expected to be more important for maintaining reliable and auditable extraction behavior. To illustrate the possible operational effect in a Singapore context, consider screening 1,000 reports. At five minutes per manual review versus five seconds to verify an evidence-linked output, review time would fall from 83.3 to 1.4 staff-hours, saving 81.9 hours. Using a rounded representative hourly rate of USD 75, informed by Singapore salary benchmarks and the SGD–USD exchange rate, this corresponds to about

USD 6,100 in potential staff-time value [33]–[36]. This sensitivity estimate is not a measured saving and excludes variation in implementation costs, report complexity, reviewer seniority, interface design, and the proportion of outputs requiring complete report review. This pilot study was limited to 54 de-identified reports from a single institution and four binary target fields. It did not prospectively measure latency, implementation effort, clinician review time, or the accuracy and completeness of the returned evidence spans. Performance may also differ across other pathology templates, ancillary-test conventions, multilingual records, scanned documents, addenda, historical findings, and more complex target schemas. Future work should evaluate nMAS on larger, multi-institutional datasets, assess label and evidence-span correctness separately, and prospectively measure clinician review time, usability, and inter-reviewer agreement [18]. Reproducible deployment will additionally require versioned prompts and field definitions, execution logging, and monitoring for unsupported labels, missing evidence, negation errors, parse failures, and model or prompt drift. Until such validation is completed, nMAS should remain an evidence-linked case-finding and reviewsupport workflow rather than a replacement for pathologist review or clinical judgment. VII. C ONCLUSION This pilot study evaluated nMAS for evidence-linked extraction of four H. pylori-related fields from 54 de-identified gastric biopsy pathology reports. nMAS achieved 98.61% overall feature-case accuracy and returned structured predictions with supporting source-text evidence. An external UMA-style MiniMax M2.5 comparator achieved the same classification performance, indicating that the present contribution lies in workflow integration, report-level aggregation, validation, and traceability rather than predictive superiority. Larger multiinstitutional evaluations, independent evidence-span assessment, prospective clinician-review studies, and componentlevel ablations are needed before operational use. A PPENDIX A R EPRESENTATIVE D E - IDENTIFIED G ASTRIC B IOPSY R EPORT The following de-identified report illustrates the sourcerecord structure used for extraction, including coded specimen fields, diagnostic text, and gross-description content. Specimen ID: 22:AB00001 Receive Date: 2022-01-11 Patient Name: Patient 1 ID: 0000001 Race: A Sex: B

Diagnosis (full report): CPOE CPOE MESSAGE RECEIVED: 11/01/22 1214 HISTOPATHOLOGY: Y VETTED & ORDER FORM COMPLETED (DR ONLY): Y SPECIMEN LABEL COMPLETED: Y ORDERSET: Routine Specimen: 1. Rectal sigmoid polyp 2. Gastric BX CLINICAL DIAGNOSIS: anaemia TIME: 11:56 DIAGNOSIS (1) Rectosigmoid polyp TUBULAR ADENOMA WITH LOW GRADE DYSPLASIA. (2) Stomach; biopsy SEVERE CHRONIC ACUTE ANTRAL AND BODY GASTRITIS WITH MILD COLONIZATION BY HELICOBACTER PYLORI. THERE IS NO INTESTINAL METAPLASIA, DYSPLASIA OR MALIGNANCY. GROSS DESCRIPTION The specimens are received in formalin, labelled with patient’s data and designated as follows. (A) Rectal sigmoid polyp It consists of a piece of tissue measuring 0.4 cm in greatest dimension. (A1-inked blue; no reserve) (B) Gastric biopsy It consists of 3 pieces of tissue measuring 0.1 cm to 0.3 cm in greatest dimension. (B1-inked yellow; no reserve) The specimens were fixed in formalin for 6--72 hours. Order Location: S42A Procedure: HT.SPECIALBX2 HP.HE EMBED Signout: AP-ABC

A PPENDIX B C OMPLETE H. PYLORI S PECIFIC ICL D EMONSTRATIONS Table V shows the complete set of field-specific ICL demonstrations used to define the four target features. The demonstrations include multiple examples per feature and cover positive wording, negative wording, abbreviations, protocolbased biopsy descriptions, organism identification, ancillary stain evidence, and diagnostic association with gastritis. A PPENDIX C T IER -3 A BLATION P ROMPT T EMPLATES Prompt 3: Ablated Tier-3 baseline prompt

TCode: T57010 (Gastric biopsy) T57010 (Gastric biopsy) T59604 (Rectal biopsy) MCode: GA003 (Severe) M43000|T57010 (CHRONIC GASTRITIS) M82110 (Tubular adenoma, NOS)

You are a clinical NLP assistant specialised in parsing Singapore hospital gastrointestinal pathology reports. Extract the following two H. pylori features from the report: h_pylori_pos and h_pylori_gastritis. Use only the pathology report text. Answer each feature with Y or N. Return only valid JSON with the keys: h_pylori_pos, verbatim_h_pylori_pos, h_pylori_gastritis, verbatim_h_pylori_gastritis, and reasoning.

TABLE V C OMPLETE H. pylori-S PECIFIC ICL D EMONSTRATIONS AND F IELD -M APPING G UARDRAILS Target Feature

Pathology-Style Demonstration Excerpt

Expected Field Mapping

Gastric Biopsy

“Stomach; biopsy.”

gastric_biopsy = Y when the specimen is Accept stomach and gastric explicitly from stomach or gastric tissue. as equivalent site terms.

Gastric Biopsy

“GASTRIC ANTRUM BIOPSY.”

gastric_biopsy = Y when the biopsy site Gastric subsite wording supis antrum, body, cardia, fundus, or another gastric ports gastric biopsy. subsite.

Gastric Biopsy

“Stomach, Sydney protocol biopsy.”

gastric_biopsy = Y when Sydney proto- Recognize protocol-based col biopsy is described as a stomach/gastric gastric sampling. biopsy.

Biopsy

“Gastric biopsy. It consists of 3 pieces of tissue measuring biopsy = Y when the specimen is explicitly Use specimen type and gross 0.1 cm to 0.3 cm.” described as biopsy tissue. description.

Biopsy

”GASTRIC BX.”

Biopsy

“Random biopsies taken (Updated Sydney Protocol) antrum, biopsy = Y when the procedure states that Procedure text can support incisura, and body.” biopsy samples were taken. biopsy status.

H. pylori Positive

“MILD COLONIZATION BY HELICOBACTER PYLORI.” h_pylori_positive = Y when the report Affirmative organism eviaffirmatively states colonization by H. pylori. dence supports positivity.

H. pylori Positive

“Helicobacter organisms are identified.”

H. pylori Positive

“Helicobacter pylori-like organisms are highlighted by im- h_pylori_positive = Y when ancillary Positive IHC evidence supmunohistochemistry.” staining highlights H. pylori-like organisms. ports positivity.

H. pylori Positive

“No Helicobacter organisms are identified.”

h_pylori_positive = N when the organ- Negation overrides keyword ism statement is explicitly negated. presence.

H. pylori Positive

“No definite Helicobacter pylori identified.”

h_pylori_positive = N when the report Do not infer positivity from states that definite organisms are not identified. uncertain or negative wording.

H. pylori Positive

“No conspicuous Helicobacter pylori organisms are identified, h_pylori_positive = N when negative Negative IHC supports a negcorroborated by negative H. pylori immunostain.” morphology is supported by negative immunos- ative value. tain.

H. pylori Gastritis

“Helicobacter pylori associated active chronic gastritis.”

H. pylori Gastritis

“Moderate to marked H. pylori associated active chronic h_pylori_gastritis = Y when severity is Preserve the association even gastritis.” described together with H. pylori-associated gas- with severity modifiers. tritis.

H. pylori Gastritis

“Helicobacter pylori-associated moderate chronic gastritis h_pylori_gastritis = Y when the diag- Final diagnosis can support with focal activity.” nosis links H. pylori to chronic gastritis. the association.

H. pylori Gastritis

“Mild chronic gastritis. No Helicobacter organisms are iden- h_pylori_gastritis = N when gastritis is Gastritis alone is insufficient. tified.” present but organisms are explicitly absent.

H. pylori Gastritis

“Mild-to-moderate chronic gastritis. No definite Helicobacter h_pylori_gastritis = N when chronic Do not label non-associated pylori identified.” gastritis is not attributed to H. pylori. gastritis as H. pylori gastritis.

H. pylori Gastritis

“Helicobacter organisms are absent.”

Prompt 4: Full Tier-3 prompt with field-specific guidance Use the same baseline Tier-3 prompt shown in Prompt 3. In addition, apply the field-specific FIELD_LIBRARY entries for h_pylori_positive and h_pylori_gastritis, including their clinical context, guardrails, and in-context demonstrations. The output key h_pylori_pos corresponds to the field-library entry h_pylori_positive.

R EFERENCES [1] K. M. Fock, “Helicobacter pylori infection: current status in singapore,” Annals of the Academy of Medicine, Singapore, vol. 26, no. 5, pp. 637– 641, 1997. [2] P. Malfertheiner, F. Megraud, T. Rokkas, J. P. Gisbert, J.-M. Liou, C. Schulz, A. Gasbarrini, R. H. Hunt, M. Leja, C. O’Morain et al., “Management of Helicobacter pylori infection: the Maastricht VI/Florence consensus report,” Gut, vol. 71, no. 9, pp. 1724–1762, 2022.

Guardrail Demonstrated

biopsy = Y when biopsy is expressed using Accept common biopsy aban accepted abbreviation such as ‘BX.” breviations.

h_pylori_positive = Y when Helicobac- Direct organism identificater organisms are explicitly identified. tion supports positivity.

h_pylori_gastritis = Y when gastritis is Require explicit diagnostic explicitly associated with H. pylori. association.

h_pylori_gastritis = N when organism Organism absence rules absence prevents attribution of gastritis to H. against H. pylori-associated pylori. gastritis.

[3] W. D. Chey, C. W. Howden, S. F. Moss, D. R. Morgan, K. B. Greer, S. Grover, and S. C. Shah, “ACG clinical guideline: Treatment of Helicobacter pylori infection,” American Journal of Gastroenterology, vol. 119, no. 9, pp. 1730–1753, 2024. [4] Y.-C. Lee, T.-H. Chiang, C.-K. Chou, Y.-K. Tu, W.-C. Liao, M.-S. Wu, and D. Y. Graham, “Association between Helicobacter pylori eradication and gastric cancer incidence: A systematic review and meta-analysis,” Gastroenterology, vol. 150, no. 5, pp. 1113–1124.e5, 2016. [5] C. A. Z. Chew, T. F. Lye, D. Ang, and T. L. Ang, “The diagnosis and management of h. pylori infection in singapore,” Singapore Medical Journal, vol. 58, no. 5, pp. 234–240, 2017. [6] T. L. Ang and D. Ang, “Helicobacter pylori treatment strategies in singapore,” Gut and Liver, vol. 15, no. 1, pp. 13–18, 2021. [7] T. L. Ang, K. W. Lim, D. Ang, Y. J. Wong, M. Tan, and A. S. Y. Wong, “Clinical audit of current Helicobacter pylori treatment outcomes in singapore,” Singapore Medical Journal, vol. 63, no. 9, pp. 503–508, 2022.

[8] Y. Wang, L. Wang, M. Rastegar-Mojarad, S. Moon, F. Shen, N. Afzal, S. Liu, Y. Zeng, S. Mehrabi, S. Sohn et al., “Clinical information extraction applications: A literature review,” Journal of Biomedical Informatics, vol. 77, pp. 34–49, 2018. [9] D. Truhn, C. M. L. Loeffler, G. Müller-Franzes, S. Nebelung, K. J. Hewitt, S. Brandner, K. K. Bressem, S. Foersch, and J. N. Kather, “Extracting structured information from unstructured histopathology reports using generative pre-trained transformer 4 (GPT-4),” The Journal of Pathology, vol. 262, no. 3, pp. 310–319, 2024. [10] J. B. Balasubramanian, D. Adams, I. Roxanis, A. Berrington de Gonzalez, P. Coulson, J. S. Almeida, and M. Garcı́a-Closas, “Leveraging large language models for structured information extraction from pathology reports,” Journal of Pathology Informatics, vol. 19, p. 100521, 2025. [11] J. Bian, M. Afshar, C. M. Scifres, E. Webber, D. Burton, D. Vawdrey, F. Wang, G. B. Melton, N. Shah, R. E. Patzer, and Y. Zhang, “The bottleneck was never data or algorithms: building a learning utility for ai-enabled learning health systems,” npj Health Systems, vol. 3, no. 43, 2026. [Online]. Available: https://www.nature.com/articles/s44401-026-00107-x [12] A. E. Wieneke, E. J. A. Bowles, D. Cronkite, K. J. Wernli, H. Gao, D. Carrell, and D. S. M. Buist, “Validation of natural language processing to extract breast cancer pathology procedures and results,” Journal of Pathology Informatics, vol. 6, p. 38, 2015. [13] G. K. Savova, J. J. Masanz, P. V. Ogren, J. Zheng, S. Sohn, K. C. KipperSchuler, and C. G. Chute, “Mayo clinical text analysis and knowledge extraction system (cTAKES): architecture, component evaluation and applications,” Journal of the American Medical Informatics Association, vol. 17, no. 5, pp. 507–513, 2010. [14] O. J. Achilonu, E. Singh, G. Nimako, R. M. J. C. Eijkemans, and E. Musenge, “Rule-based information extraction from free-text pathology reports reveals trends in south african female breast cancer molecular subtypes and ki67 expression,” BioMed Research International, vol. 2022, p. 6157861, 2022. [15] J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang, “BioBERT: a pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics, vol. 36, no. 4, pp. 1234–1240, 2020. [16] C. Wong, S. Preston, Q. Liu, Z. Gero, J. Bagga, S. Zhang, S. Jain, T. Zhao, Y. Gu, Y. Xu, S. Kiblawi, R. Weerasinghe, R. Leidner, K. Young, B. Piening, C. Bifulco, T. Naumann, M. Wei, and H. Poon, “Universal abstraction: Harnessing frontier models to structure realworld data at scale,” 2025. [17] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023. [18] B. Vasey, M. Nagendran, B. Campbell, D. A. Clifton, G. S. Collins, S. Denaxas, A. K. Denniston, L. Faes, B. Geerts, M. Ibrahim, X. Liu, B. A. Mateen, P. Mathur, M. D. McCradden, L. Morgan, J. Ordish, C. Rogers, S. Saria, D. S. W. Ting, P. Watkinson, W. Weber, P. Wheatstone, P. McCulloch, and DECIDE-AI expert group, “Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI,” Nature Medicine, vol. 28, no. 5, pp. 924–933, 2022. [19] A. Hardjojo, E. H. Goh, F. Sundram, L. W. Tan, D. Y. T. Goh, M. C. Phoon, V. J. Lee, M. Hartman, and A. R. Cook, “Validation of a natural language processing algorithm for detecting infectious disease symptoms in primary care electronic medical records in singapore,” JMIR Medical Informatics, vol. 6, no. 2, p. e36, 2018. [20] S. B. Tay, G. H. Low, G. J. E. Wong, H. J. Tey, F. L. Leong, C. Li, M. L. K. Chua, D. S. W. Tan, C. H. Thng, I. B. H. Tan, and R. S. Y. C. Tan, “Use of natural language processing to infer sites of metastatic disease from radiology reports at scale,” JCO Clinical Cancer Informatics, vol. 8, p. e2300122, 2024. [21] G. Song, S. J. Chung, J. Y. Seo, S. Y. Yang, E. H. Jin, G. E. Chung, S. R. Shim, S. Sa, M. S. Hong, K. H. Kim, E. Jang, C. W. Lee, J. H. Bae, and H. W. Han, “Natural language processing for information extraction of gastric diseases and its application in large-scale clinical research,” Journal of Clinical Medicine, vol. 11, no. 11, p. 2967, 2022. [22] J. H. Bae, H. W. Han, S. Y. Yang, G. Song, S. Sa, G. E. Chung, J. Y. Seo, E. H. Jin, H. Kim, and D. An, “Natural language processing for assessing quality indicators in free-text colonoscopy and pathology reports: Development and usability study,” JMIR Medical Informatics, vol. 10, no. 4, p. e35257, 2022.

[23] Singapore General Hospital, “Who We Are: Singapore General Hospital,” 2026, [Online]. Available: https://www.sgh.com.sg/about-sgh/ who-we-are. Accessed: Jun. 2, 2026. [24] Singapore Department of Statistics, “Census of Population 2020: Statistical Release 1 – Demographic Characteristics, Education, Language and Religion,” 2021, singapore resident population ethnic composition reported as 74.3% Chinese, 13.5% Malays, 9.0% Indians, and 3.2% Others. [25] A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei et al., “Qwen2.5 Technical Report,” arXiv preprint arXiv:2412.15115, 2024. [26] DeepSeek-AI, “DeepSeek-V4: Towards highly efficient million-token context intelligence,” arXiv preprint arXiv:2606.19348, 2026. [27] GLM-5 Team, “GLM-5: From vibe coding to agentic engineering,” arXiv preprint arXiv:2602.15763, 2026. [28] Google, “Gemini 3.1 Pro Preview,” Google AI for Developers, 2026, accessed: Jun. 30, 2026. [Online]. Available: https://ai.google.dev/ gemini-api/docs/models/gemini-3.1-pro-preview [29] DeepSeek-AI, “DeepSeek-V3.2: Pushing the frontier of open large language models,” arXiv preprint arXiv:2512.02556, 2025. [30] Moonshot AI, “Introducing Kimi K2 Thinking,” Official model documentation, 2025, accessed: Jun. 30, 2026. [Online]. Available: https://moonshotai.github.io/Kimi-K2/thinking.html [31] Qwen Team, “Qwen3-Next-80B-A3B,” Official Qwen model release, 2025, accessed: Jun. 30, 2026. [Online]. Available: https://qwen.ai/blog?from=research.latest-advancements-list&id= 4074cca80393150c248e508aa62983f9cb7d27cd [32] MiniMax, “MiniMax M2.5: Built for Real-World Productivity,” Feb. 2026, accessed: June 28, 2026. [Online]. Available: https: //www.minimax.io/news/minimax-m25 [33] JobStreet, “Clinical research coordinator salary in singapore,” 2026, accessed: 2026-06-05. Reported average monthly salary range: SGD 3,700–4,200. [Online]. Available: https://sg.jobstreet.com/career-advice/ role/clinical-research-coordinator/salary [34] Ministry of Health Singapore, “Average and median salaries earned by public sector healthcare workers from 2021 to 2025,” 2025, accessed: 2026-06-05. Reported median monthly base salary of public-sector doctors in 2024: approximately SGD 14,400. [35] ERI SalaryExpert, “Pathologist salary in singapore,” 2026, accessed: 2026-06-05. Reported average hourly rate: SGD 143.64. [36] Wise, “Singapore dollar to us dollars exchange rate history,” 2026, accessed: 2026-06-05. Mid-market exchange rate on June 5, 2026: 1 SGD = USD 0.7745.

Record · ID 346551 · SHA-256 21d9c93d31a7dbc6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.