Conceptio › Archive › arXiv CS
arXiv CSopen access

woma: a real-time foundation model and its fine-tuned models for endoscopy

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

1

woma: a real-time foundation model and its fine-tuned models for endoscopy Thang Tran∗

Lan Dang

CloudKites AI Lab New South Wales, Australia [email protected]

Monash Business School, Monash University Victoria, Australia [email protected]

arXiv:2609.15130v1 [cs.SE] 14 Sep 2026

September 2026

Abstract woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which task models are fine-tuned. We contribute a systematic design for production. Requirements and pass marks were fixed before any run, eight candidates screened under pre-registered rules, self-supervised training taken to a stopping rule, then fine-tuning and deployment optimisation, all on one self-contained library, numbat. We also contribute woma itself with two fine-tuned models, every outcome reported met or missed. Our colonoscopy model finds and outlines polyps, names which colon segment is in view, suggests polyp type and grades bowel preparation. Our gastroscopy model names a station out of 22 protocol sites, flags and outlines lesions, and names one of seven findings. Every number was read on data never seen in training, and shipped weights were chosen on that record. In colonoscopy, 96 % of polyps in a six-hospital PolypGen set are found at precision ≥0.85, and 19 of 19 polyps across fifteen full REAL-Colon videos at 1.6 false alarms per procedure. In gastroscopy, landmark region is named correctly on 92 % of frames from unseen patients, and 37 of 39 held-out neoplasia frames are flagged at specificity 0.91. On one workstation GPU every task runs over 1080p video at about 100 frames per second, faster than PyTorch, ONNX Runtime and TensorRT in all four precision regimes tested. TensorRT comes closest: one pass of our foundation model takes it 3 to 27 % longer than ours, and we deliver 6 to 31 % more frames per second from frame to results. A second build links no vendor library at all — our own kernels over Vulkan — so a site deploys two files and needs no toolkit, no cuDNN and no framework; in f32 it beats the CUDA build on the same card.

Keywords: foundation model; self-supervised learning; real-time inference; colonoscopy; gastroscopy

1

Introduction

Computer-aided polyp detection has crossed from research into practice: a randomised trial showed a real-time detector raising the adenoma detection rate [1], a meta-analysis confirmed the effect [2], and a first device is sold with its design and validation described in print [3]. In the stomach, a system that watches for unexamined blind spots cut the blind-spot rate from 22 % to 6 % in a randomised trial [4], and a convolutional network found gastric cancer with a per-lesion sensitivity of 92 % [5]. What such reports rarely give is the whole path: how the network at the centre was chosen, what trained each head, what each number was read on, how the shipped weights were picked, and how fast the system runs when measured rather than estimated. This paper gives that path for one system, and it contributes two things. ∗

Corresponding author: [email protected]

woma: real-time endoscopy foundation model

2

1. A systematic design for production, from the first day. Requirements and pass marks were written before anything ran; eight candidate architectures were screened under pre-registered rules; data pipeline, trainer, label-free training to a stopping rule, fine-tuning of task models and deployment optimisation all ran on one self-contained library, numbat [6, 7]; and what came out is benchmarked against PyTorch, ONNX Runtime and TensorRT on one machine, for our foundation model and for both fine-tuned models end to end. 2. A real-time foundation model and its fine-tuned models for endoscopy, with transparent outcomes. woma, with colon and upper-GI models built on it, is reported against every pass mark met or missed, on data training never saw, alongside rival architectures fine-tuned on identical recipes. Section 2 covers requirements, data, design experiment, fine-tuning recipes and evaluation protocol. Section 3 reports screening outcome, held-out results and speed. Section 4 weighs what both contributions are worth, why held-out evaluation matters, and what locally collected data would add. Every number here is retrospective, on public data.

2

Methods

2.1

Requirements, targets and criteria

The foundation model is the part of the network that turns a video frame into features; every task head reads those features (Figure 1). At 60 frames per second a frame lasts 16.7 milliseconds (ms), and the foundation model and its feature-pyramid neck may spend at most about half of that on a 640×640 pixel input, on one workstation graphics processing unit (GPU). Sixty is the demanding end of what endoscopy processors emit, and a stack that keeps up with it keeps up with the 25, 30 and 50 frames per second that others produce, so one budget serves every site. It is a requirement on deployment rather than a property of our material — nothing here was recorded at 60, and the two clips the delivered program is timed on run at 15 and 30 frames per second (Table S2) — so speed is measured against a fixed input under the protocol of Section 2.5, never inferred from the recordings the accuracy numbers come from. Beyond speed, it must serve four kinds of head at once (finding lesions, outlining them, naming the anatomical site and grading quality), it must do so in hospitals it never saw, and it must stay quiet on clean mucosa, because an alert that fires on nothing is what makes clinicians switch a system off. Pass marks came from published anchors, fixed before any run (Table 1).

2.2

Data

Figure 2 shows the data: about 1.5 million unlabelled frames for pretraining, labelled sets for the heads, and two sets walled off from training from the first day. PolypGen [14] holds polyp images from six hospitals and was built to test the unseen-hospital question; REAL-Colon [15] holds sixty full-length colonoscopy videos with histology for every polyp. Both are read only. Everything else is split by video or by patient, because consecutive endoscopy frames are near-identical: a split by frame once inflated our detection score from 0.67 to a fictitious 0.91. The colon and upper-GI training sets, their exclusions and their routing rules are given in Figures S1 and S2.

2.3

Designing the foundation model

Learning without labels, and three routes. A network can learn from unlabelled frames by solving a task that needs no labels, such as filling in a masked part of the image or agreeing with itself across two views of the same image. Figure 4 sorts objectives by three properties that matter here: whether the recipe works on a convolutional network, the kind that fits a 16.7 ms budget on a small GPU; whether it teaches every location of the feature map, which detection and outlining need; and whether it guards against collapse, the failure in which every feature drifts to the same value. Lesions are recognised by texture, which favours objectives that must rebuild texture over

woma: real-time endoscopy foundation model

3

detection head 60 fps video

preprocess crop, letterbox

foundation model (woma)

feature pyramid neck

P2 P3 P4 P5 + pooled

frame budget 16.7 ms: model+neck ≤ 8

segmentation region-gated

temporal filter

overlay

station / quality heads

heads ≈ 5

temporal ≈ 2

margin ≈ 1.7

Figure 1: Where the foundation model sits. Every head reads its features, and the time budget of one video frame is split so that the foundation model and neck may spend at most about half of it. GastroNet-5M — 1.0M held of 4.8M, 8 centres

Own corpus — ∼0.5M frames, growing

HyperKvasir — 99k unlabelled + 10.6k labelled

pretraining pool ∼1.5M frames held — self-supervision and distillation targets

screening probes frozen detection + frozen stations, at fixed intervals

Internal labels — detection; EGD stations

GastroHUN — 8.8k images, 22 stations

fine-tuning detection, segmentation and station heads at 6402

Kvasir-SEG — 1k segmentation masks

SUN-SEG — 158k video frames

PolypGen — 8k+ images, 6 centres

REAL-Colon — 60 full videos, 2.7M frames

temporal evaluation video segmentation consistency

held-out evaluation — never trained on unseen-centre accuracy; false alarms per procedure

Figure 2: Data pool: every dataset, its size, and the single role it plays. The bottom lane is walled off — PolypGen and REAL-Colon feed evaluation only, so the unseen-centre and false-alarm numbers mean what they say. Sources: GastroNet-5M [8], HyperKvasir [9], GastroHUN [10], Kvasir-SEG [11], SUN-SEG [12, 13], PolypGen [14], REAL-Colon [15].

objectives that learn to ignore it. There are three routes to a pretrained network (Figure 5): from a random start on the unlabelled pool, from public weights continued on the pool, or by distilling a large foundation model into a small one. The selection experiment carried each route on at least two candidates, so that a route failing could be told apart from a candidate failing.

woma: real-time endoscopy foundation model

4

Design study the problem, the criteria, a review of the field, eight candidates, the plan

Stage 0: measure latency of every deployable form · training throughput · probe calibration

Builds eight candidates in the training harness, the same probes on all the eight candidates 1 R34 · 2 RVGG · 3 CNXV2 4 FMS · 5 CSPR · 6 CSPR-A 7 RVT · 8 CSPR-S

Stage 1: screening aspects I to IV on all eight: equal-clock twelve-hour runs, matched-images re-probe, route-optimal wave, adaptability fine-tune

ranked ladder: rungs 1–2 advance

a run below the floor, a rank collapse, or a station probe at chance: killed

Stage 2: confirmation 72-hour runs, both finalists in parallel, plus a fine-tune smoke test

winner by pre-registered rank

all fail: fall back to the strongest measured deploy stack, trained with labels only

Production foundation model the winner’s route, run to its stopping rule on the full pool

Multi-task fine-tune detection, outline, station and findings heads on the labelled sets

Packaging one model file per model; latency verified against the budget

Evaluation on held-out data PolypGen unseen centres · REAL-Colon false alarms per procedure · GastroHUN stations · Kvasir-SEG

Report this paper

Figure 3: Whole project from the design study to the evaluated models, with its two gates. Every exit is a decision that was written down before anything ran; the two “killed” branches are outcomes we were prepared to take.

woma: real-time endoscopy foundation model

5

Table 1: Pass marks, fixed before any run, and the published anchor each was set from. One row per mark; the corpora they are read on are named in Section 2.5. instrument

mark

anchor

colonoscopy PolypGen sensitivity, precision ≥0.85 REAL-Colon per-polyp sensitivity outline Dice, three other-hospital sets colon segment macro recall, ten segments colon segment accuracy, six merged classes polyp type macro recall bowel preparation adequacy F1

≥0.90 ≥0.93 ≥0.82 ≥0.70 0.666 ≥0.80 ≥0.90

cross-centre transfer [14] per-procedure reporting [15] PraNet protocol [16] single-frame best [17] the same, at parity [17] reported, no published parity Boston scale [18]

gastroscopy station landmark-region accuracy station 23-code macro recall lesion gate sensitivity, specificity ≥0.90 findings macro recall, validation findings accuracy, UGIAD coarse-5 findings accuracy, GastroVision half

≥0.92 ≥0.85 ≥0.93 ≥0.85 ≥0.90 ≥0.85

0.887, the dataset paper [10] the same, reported 0.922 per image [5] 0.916–0.939, FastUGI-Net [19] the same the same

speed, one workstation GPU whole stack, per frame foundation model and neck, per frame

≤16.7 ms ≤8 ms

≈10 ms, commercial [3] the same

The detection marks are read on PolypGen’s six hospitals [14, 20] and on REAL-Colon study 4 at no more than two sustained false alarms per procedure [15, 21]; the outline mark is the mean over CVC-ColonDB [22], ETIS [23] and CVC-300 [24]; the lesion gate is 39 neoplasia frames against 1,347 clean frames from another hospital. A mark with no published parity is still fixed in advance and reported either way.

ways to learn without labels

masked reconstruction (texture-hungry)

view-invariance (texture-averse)

DINO [25] self-distillation convnets: yes per-location: no

DINOv2 [26] patch-level loss convnets: via distilled students per-location: yes

MAE [27] patch reconstruction transformer-only convnets: no

SparK [28] sparse masked conv convnets: yes per-location: yes

statistical regularisation

ConvNeXt-V2 FCMAE [29] GRN guard convnets: yes per-location: yes

VICReg [30] variance term, local convnets: yes per-location: yes

Figure 4: Families of label-free training objectives, keyed to what matters here: does the recipe work on a convolutional network, does it teach every location of the feature map (“per-location”), and does it have a built-in guard against the features collapsing to the same value. The masked-reconstruction family with that guard is the one woma uses.

woma: real-time endoscopy foundation model

6

Table 2: Eight candidates and a Stage-0 latency gate (batch 1, 6402 , f16, one forward of the network alone on an RTX 3090 as the engine proxy; the pass mark is 8 ms for network and neck). Route: A from scratch, B public weights then adaptation, C distillation. Candidate 8’s kernel path was deferred by design, so it was screened for representation quality only. #

name

strategy

1

R34

2

RVGG

3

CNXV2

4

FMS

5

CSPR

6

CSPR-A

7

RVT

8

CSPR-S

ResNet-34-D, a plain residual network deployed as itself RepVGG-A2, trained with branches, folded to a plain 3×3 stack ConvNeXt-V2-Nano, a masked-autoencoder convnet CSP-R-13 distilled from an in-domain foundation teacher CSP-R-22, a detection trunk made safe for self-supervision candidate 4’s body with two attention blocks at the deepest scale RepViT-M1.1, a reparameterised mobile hybrid candidate 4’s body with a state-space layer at the deepest scale

route

params

ms

gate

B

21.8M

2.22

pass

B

25.5M

1.79

pass

B

15.6M

3.54

C

12.7M

1.90

pass, flagged for its depthwise convolutions pass

A

20.7M

2.42

pass, at the line

C

19.3M

1.97

pass

B

8.2M

4.14

C

14M

—

pass, the mobile-regime trade-off made visible deferred

Route A — from scratch

Route B — initialise, adapt

Route C — distil a foundation

random start

public weights (supervised or label-free)

foundation teacher (in-domain if the licence fits)

long label-free run on the held pool

short in-domain continuation

distil into the deploy-size foundation model on the pool

the slowest route; viable at our pool size

converges far faster than a random start [31]

how current real-time detectors are built

Figure 5: Three routes to a pretrained foundation model. They compose: Route B can prepare the teacher that Route C distils. Every route was carried by at least two of the eight candidates, so a route failing could be told apart from a candidate failing.

Design principles. Our review compressed into eight rules, and those rules generated our candidate set. Measure cost on target hardware, never from operation counts. Prefer public weights or distillation for speed, keeping training from scratch as a control our pool size allows [31]. Pretrain a convolutional network only with an objective that teaches every location and guards against collapse [28–30]. Favour reconstruction objectives for texture-defined lesions. Hold an unbroken additive identity path through every stage [32]. Anti-alias each downsampling step [33], since aliasing causes flicker and flicker causes false alarms. Give a pooled endpoint real width. And keep our operator set boring, so a network compiles to any standard engine. Candidate set. Table 2 and Figure 6 list the eight candidates. They are one pick per family of a 2025–26 landscape our survey found deployable within our latency requirement: plain residual networks, plain networks that fold into one branch at deployment, cross-stage-partial trunks that real-time detectors are built from, modernised convnets, and a mobile hybrid. Hierarchical vision transformers and state-space encoders were surveyed and excluded, transformers for unproven latency and quantisation at this resolution, state-space encoders for needing a custom kernel our engine rule forbids. This set samples every axis selection must price at least twice: topology, pretraining route, context at the deepest scale (a matched triplet 4, 6 and 8 differs in exactly one thing there) and deployment regime.

woma: real-time endoscopy foundation model

7

input 2242 (pretrain) / 6402 (deploy)

input 2242 / 6402

deep stem: 3×3 s2 (32) + 3×3 (32) + 3×3 (64); blur-pool s2

stem: 1× RepVGG block — 64 ch — stride 2

S1: 3× basic residual — 64 ch — stride 4

P2

S2: blur; 4× basic residual — 128 ch — stride 8

P3

S1: blur; 2× RepVGG — 96 ch — stride 4

P2

S2: blur; 4× RepVGG — 192 ch — stride 8

P3

S3: blur; 6× basic residual — 256 ch — stride 16

P4 + probe 1×1→256

S3: blur; 14× RepVGG — 384 ch — stride 16

P4 + probe 1×1→256

S4: blur; 3× basic residual — 512 ch — stride 32

P5

S4: blur; 1× RepVGG — 1408 ch — stride 32

P5

1×1: 512 → 2048 — pretraining only

1×1: 1408 → 2048 — pretraining only

global average pool → 2048-d

global average pool → 2048-d

(a) Candidate 1, R34: ResNet-34-D with a deep stem and blur-pooled downsampling, deployed as itself.

(b) Candidate 2, RVGG: RepVGG-A2, trained with parallel branches and folded into plain 3×3 convolutions for deployment.

input 2242 / 6402

input 2242 / 6402

patchify stem: 4×4 conv, stride 4 — 80 ch

stem: 3×3 s2 (48); blur; 3×3 s2 — 96 ch

S1: 2× ConvNeXt-V2 block — 80 ch — stride 4

P2

S1: 2× residual-CSP unit — 96 ch — stride 4

P2

S2: LN; down 2×2; 2× blocks — 160 ch — stride 8

P3

S2: blur; down; 3× units — 192 ch — stride 8

P3

S3: LN; down; 8× blocks — 320 ch — stride 16

P4 + probe 1×1→256

S3: blur; down; 5× units — 384 ch — stride 16

P4 + probe 1×1→256

S4: LN; down; 2× blocks — 640 ch — stride 32

P5

S4: blur; down; 2× units — 640 ch — stride 32

P5

1×1: 640 → 2048 — pretraining only

1×1: 640 → 2048 — pretraining only

global average pool → 2048-d

global average pool → 2048-d

(c) Candidate 3, CNXV2: ConvNeXt-V2-Nano with the FCMAE pretraining endpoint; became woma.

input 2242 / 6402

input 2242 / 6402 stem: 3×3 s2 (48); blur; 3×3 s2 — 96 ch S1: 2× CSP units — 96 ch — stride 4

(d) Candidate 4, FMS: the CSP-R-13 student of a foundation-model teacher.

CSP-R-13 body: stem + S1–S3 as in Figure 6d (strides 4–16)

P2 P3 P4 + probe

P2

S2: blur; down; 4× units — 192 ch — stride 8

P3

S3: blur; down; 8× units — 384 ch — stride 16

P4 + probe 1×1→256

S4: blur; down; 3× units — 768 ch — stride 32

P5

S4: blur; down; 2× CSP units — 640 ch — stride 32 2× multi-head self-attention blocks — dim 640, 8 heads — on the S/32 grid

P5 (attended)

1×1: 640 → 2048 — pretraining only 1×1: 768 → 2048 — pretraining only

global average pool of the attended map → 2048-d

global average pool → 2048-d

(e) Candidate 5, CSPR: CSP-R-22, the detector-family residual cross-stage-partial trunk; the runner-up.

(f) Candidate 6, CSPR-A: CSP-R-13 with two selfattention blocks at the stride-32 stage.

input 2242 / 6402 stem: two 3×3 s2 convs — 64 ch S1: 2× RepViT block (rep-dw 3×3 + SE alt.) — 64 ch — stride 4 S2: down; 2× blocks — 128 ch — stride 8

input 2242 / 6402 P3

S3: down; 12× blocks — 256 ch — stride 16

P4 + probe 1×1→256

S4: down; 2× blocks — 512 ch — stride 32

P5

global average pool → 512-d

(g) Candidate 7, RVT: RepViT-M1.1, a reparameterised mobile hybrid.

CSP-R-13 body: stem + S1–S4 as in Figure 6d (strides 4–32, 96–640 ch) 2× state-space (SSM) token-mixer blocks — dim 640 — linear-cost scan over the S/32 grid

P2 P3 P4 + probe

P5 (mixed)

global average pool of the mixed map → 640-d

(h) Candidate 8, CSPR-S: the state-space (selective-scan) sibling of candidate 5.

Figure 6: Eight candidate foundation models of the selection study. Every candidate exposes the same three feature maps (strides 8, 16 and 32, called P3 to P5 ) for the heads; the dashed 1×1 layer and the pooled endpoint exist only for pretraining. Candidate 3 became woma; candidate 5, the runner-up, was fine-tuned on both models for comparison.

woma: real-time endoscopy foundation model

8

Plan, and its rules. Figure 3 shows this project whole, with its two gates; a guiding idea is to spend measurements before GPU-hours and GPU-hours before weeks. Stage 0 measures before anything trains and can already kill a candidate, for a latency miss in its deployable form or training throughput too low for a twelve-hour run to mean anything. Stage 1 trains every candidate for the same wall-clock budget and probes the frozen network at fixed intervals with two tests that need no fine-tuning: a linear station classifier on GastroHUN [10] (macro F1 over 22 stations) and a linear polyp-outline head on Kvasir-SEG [11] (Dice), each read against its untrained floor (0.056 and the image-independent null of 0.390). A third instrument, the effective rank of the dense feature map, watches for collapse, with a warning line at 64. Four views of each run (equal clock; matched images seen; a route-optimal second wave; and a short adaptability fine-tune of a whole network through one outline head, twelve epochs for every candidate) combine into a single ranked ladder. Our advance rule was fixed in advance: top two rungs go on, provided each clears three times its random-network floor, ends above that effective-rank warning line, and beats twice chance on our station probe. Stage 2 gives both finalists 72 hours each plus a short supervised fine-tune; whichever wins on detection and station together goes on, ties broken by measured latency. Winner’s route then runs to its stopping rule on the full pool: every twenty minutes we probe that frozen network as above, and a run stops once both probes have gained less than 0.005 over their last 3,000 steps, which is less than the spread between two random seeds.

2.4

Fine-tuned models

One forward of the foundation model serves every head (Table S1). Detection uses a light featurepyramid neck [34] under an anchor-free head of the YOLOv8 kind [35], with a single class in both models: a two-class upper-GI variant (Barrett’s segment against a focal lesion) lost, because a focal lesion often sits inside a Barrett’s field and one detector cannot serve a field and a focus at once. Polyp type and findings are read per frame in this version. Figures S1 and S2 give each model’s complete recipe: training sets, optimiser and schedule, augmentation, what is measured while a run trains, what is measured on held-out data for every snapshot, and the decision rule. Three choices deserve a sentence. Each training step draws one task by weight and runs one batch of it, so our foundation model sees every head’s gradient in proportion to those weights and no head’s batch statistics leak into another’s. The foundation model is pretrained while the heads start from random weights, so heads learn at ten times its rate. Weights we evaluate and save are a running average of live ones. Each model’s shipped snapshot is chosen on its held-out record: lesion gate filters first, then most pass marks met wins. Training curves are never consulted. Every arm reported here was fine-tuned on one machine at one setting: three cards, four images each, effective batch twelve. The first pass was not — candidates 1 and 7 ran on the second box while the first was busy, and two cards reach twelve only at six images each. Normalisation statistics inside a trunk are estimated per replica and never synchronised, so a replica sees four samples or six, and for a trunk built of convolution and batch normalisation that is no detail. Both arms were re-run at four and it is the re-runs that Table 4 reports; the earlier pair is kept, and what it now measures is the split. Candidate 1, a residual network with batch normalisation throughout, reads PolypGen 0.860–0.889 at six across all ten snapshots and 0.912–0.970 at four: two bands that do not meet at any epoch, one wholly under the pass mark and one wholly over it, while every other instrument moves by less than 0.03. Candidate 7, a transformer with layer normalisation, does not move — 0.953–0.974 against 0.944–0.973. Machine and split changed together, on the same model of card and byte-identical data, and a pair cannot separate them; the matched re-runs are the answer to that, not an argument that it does not matter. Repeats of one setting settle a different question: which instruments tolerate a single run at all. Across two seeds PolypGen sensitivity moves by at most 0.015 and in-domain detection by 0.023, against 0.18 for held-out segment recall — as large as the 0.19 that separates candidate 2’s two batch settings, so we do not read that instrument as ordered between single arms. Detection comparisons in this paper therefore rest on ground the seeds support; segment-recall differences smaller than about 0.2 do not, and we read them as indistinguishable rather than ordered.

woma: real-time endoscopy foundation model

2.5

9

Evaluation

Held-out instruments and metrics. Every result is read on data training never saw: for our colon model PolypGen [14], REAL-Colon study 4 [15] (15 full videos), 15 held-out CAS-Colon videos [36], HyperKvasir’s landmark folders [9], KUMC’s held-out patients [37] and three other-hospital PraNet outline sets; for our upper-GI model GastroHUN’s validation and test patients, 39 EDD2020 [38] neoplasia frames against 1,347 clean frames from another hospital, UGIAD’s published test split [19] and half of GastroVision [39]. Sensitivity is the share of true lesions found, specificity the share of clean frames left alone, macro recall that share averaged over classes, and Dice the overlap of a predicted outline with a true one. Every candidate carried to fine-tuning ran both recipes unchanged and was scored on these instruments at every second epoch. Speed protocol. All four engines were timed in one sitting on one machine, on a card at its factory 370 W limit with no display attached; nothing else ran, and both cards were idle before and after. Foundation-model rows ran PyTorch first and frame-to-results rows numbat first, so neither system always met a cold box. Each engine runs 300 synchronised iterations after 30 warm-ups on the same layers: the same ConvNeXt-V2-Nano from timm for the foundation-model rows, and for the fine-tuned rows our own layers end to end — foundation model, neck, every head with its class counts, the box decode and non-maximum suppression (NMS) — from a frame in host memory to results in host memory. Every cell was measured three times and we report a median of three medians. What those repeats show matters more than their number. Each numbat repeat is a separate launch, starting its allocator and selecting its kernels afresh, while framework repeats run inside one process after their warm-ups — friendlier of two arrangements, and not ours. Across all 31 timed cells, widest spread between repeats is 0.067 ms and median spread 0.014 ms, and in all 31 engine-to-engine comparisons numbat’s slowest repeat is faster than the fastest repeat of the engine it is set against: no pair of distributions overlaps anywhere in the table. The closest case is the bf16 foundation model, where numbat’s worst reading of 2.49 ms still beats TensorRT’s best of 2.57. PyTorch is version 2.14.0 with CUDA 13 in eager mode, channels-last for its 16-bit rows; ONNX Runtime 1.29.0 and TensorRT 11.3.0 read one exported graph of those layers, TensorRT as an engine built at optimisation level 5; every engine’s outputs were checked against PyTorch’s on a fixed frame before anything was timed. numbat is the 0.9.14 release build, timed as the program’s own stages with a device sync after each stage, over 151-frame clips.

3

Results

3.1

Screening outcome

Table 3 gives our Stage-1 field. Candidate 3 ranked first: highest frozen station probe (0.772), an outline probe within 0.01 of best, effective rank well clear of its warning line, and highest fine-tuned Dice (0.869) in that adaptability test. Candidate 2 ranked second. Candidate 5, the only network trained from a random start, sat fifth on the frozen probes but showed the largest movement under fine-tuning (+0.268), which is what a detection-style trunk trained from scratch is expected to do, and it was carried forward as the from-scratch control. Three distillation candidates (4, 6 and 8) went unranked: their features collapsed to an effective rank near 6, and the cause was traced to the teacher’s dense target itself, which at matched positions carries an effective rank of 20.7, three points above what an untrained student already has. A whitening term repaired the collapse (effective rank 66 at step 1,000), but too late for the ladder. Candidate 3 then won the confirmation on both detection and station, and its route (public pretrain-only weights of ConvNeXt-V2-Nano, self-supervised on ImageNet photographs [40, 41], continued with the same masked autoencoder on the one-million-frame GastroNet-5M holding [8] plus in-house video) stopped under its rule at step 11,500. That checkpoint is woma. Its shape is that of ConvNeXt-V2-Nano (Figure 6c): a stem that turns each 4×4 patch into 80 features, four stages of 2, 2, 8 and 2 blocks at 80, 160, 320 and

woma: real-time endoscopy foundation model

10

Table 3: Stage-1 screening, all eight candidates at the exit of their equal-clock run: the two frozen probes (station macro F1 on GastroHUN, linear Dice on Kvasir-SEG), the effective rank of the dense map, the Dice after the twelve-epoch adaptability fine-tune with an identical head, and the outcome. Floors: station 0.056, Dice 0.390 (image-independent null), effective-rank warning line 64. candidate

route

station

Dice

eff. rank

fine-tuned Dice

3 CNXV2

B, masked autoencoder

0.772

0.731

101.6

0.869

2 RVGG 1 R34 7 RVT 5 CSPR

B, ImageNet weights B, ImageNet weights B, ImageNet weights A, from scratch

0.761 0.698 0.713 0.735

0.717 0.739 0.735 0.524

105.3 114.8 107.5 77.4

0.862 0.853 0.833 0.792 (+0.268)

4 FMS

C, distilled

0.459

0.081

6.2

—

6 CSPR-A 8 CSPR-S

C, distilled C, distilled

0.517 0.240

0.204 0.053

— —

— —

outcome rung 1; won the confirmation; became woma rung 2; finalist ranked third ranked fourth fifth; carried as the from-scratch control and fine-tuned on both models unranked: collapsed on the teacher’s target unranked: collapsed unranked: collapsed

640 channels, 15.6 million parameters; each block is a 7×7 depthwise convolution, a normalisation over channels, two 1×1 convolutions with a GELU (Gaussian error linear unit) between them, and the global response normalisation (GRN) that keeps channels diverse. Heads read stage outputs at strides 8, 16 and 32, called P3 to P5 .

3.2

Other candidates, same recipes

A screening is a claim about what a trunk is worth downstream, and the only way to test it is to spend the fine-tunes. Table 4 puts every candidate through both product recipes on the same data, schedule, augmentation, evaluation, pass marks and machine, four images per replica throughout (Section 2.4); Figures 7 and 8 plot each arm per snapshot on every held-out instrument. Finding lesions barely separates them. Every candidate’s gastroscopy arm clears the lesion gate, several above woma (candidate 1 at 0.974, candidate 2 at 1.000, against 0.949), and every colonoscopy arm clears its detection gate at every snapshot but one — woma’s own epoch 20, at 0.898, where candidate 5 still reads 0.934, having matched woma at the peak (0.969 against 0.970) and faded less. A detector head on a reasonable trunk finds lesions. Describing a frame separates them completely, and screening predicted that order. Stage 1’s frozen station probe ranked all five candidates 3 > 2 > 5 > 7 > 1 (0.772, 0.761, 0.735, 0.713, 0.698, Table 3). Their landmark-region accuracies after a full gastroscopy fine-tune fall in the same order, with no inversion: 0.923, 0.738, 0.662, 0.634, 0.575. A twelve-hour probe on frozen features ordered five trunks exactly as two weeks of fine-tuning did, which is the case for spending measurements before GPU-hours. woma is also the only candidate whose region accuracy clears its bar at all, by a margin no other comes within 0.15 of. One class can hide that. Candidate 5’s caecum recall on another centre climbs from 0.46 to 0.95 through training while its retroflexed-rectum recall falls from 0.42 to 0.11 and its ten-segment macro declines from 0.43 to 0.35: it answers caecum more often rather than recognising more segments, where woma’s macro holds between 0.48 and 0.52. Read on caecum alone it would look like a better station model, which is why we fixed that pass mark on macro recall before any run. Candidate 1 reads that difference most clearly, being no from-scratch control but a supervised ImageNet residual network, most conventional trunk in our set. Its gastroscopy arm flags lesions as well as anything here — lesion gate 0.974, above woma — yet names landmark region on 57.5 % of frames where woma names it on 92.3 %, reads 23-code macro at 0.415 against 0.822, and meets two of six bars against five. Its colonoscopy arm is the counter-case and we report it as one: at the matched setting it clears the detection gate at every snapshot, meets the same two bars woma meets, and is not separated from woma by more than 0.08 on any of the four colonoscopy instruments. An earlier run of this arm at six images per replica missed the gate at every snapshot; that was the split,

woma: real-time endoscopy foundation model

11

Table 4: Every candidate carried through both product recipes, identical in data, schedule, augmentation, evaluation and pass marks; only the foundation model differs. Each row is the snapshot the pre-registered rule selects — pass the anatomy’s gate, then meet the most bars. ⋆ marks a bar met; sel. is the snapshot, met counts the gated bars. A dash is an instrument not yet scored for that snapshot. colonoscopy arm

sel.

PolypGen REAL-Colon ⋆

⋆

Dice

segment

type

prep

met

0.521 0.397 0.449 0.436 0.446

0.634 0.527 0.629 0.583 0.671

0.737 0.561 0.666 0.715 0.702

2 2 2 2 2

woma (3 CNXV2) 5 CSPR 1 R34 2 RVGG 7 RVT

6 8 8 2 2

0.960 0.959⋆ 0.954⋆ 0.943⋆ 0.955⋆

1.000 1.000⋆ 1.000⋆ 0.947⋆ 0.947⋆

0.807 0.732 0.749 0.616 0.708

gastroscopy arm

sel.

gate

region

23-code findings UGIAD GastroV.

woma (3 CNXV2) 5 CSPR 1 R34 2 RVGG 7 RVT

asm10c2 asm14c4 asm18c2 asm12c4 asm2c4

0.949⋆ 0.974⋆ 0.974⋆ 0.949⋆ 0.974⋆

0.923⋆ 0.662 0.575 0.738 0.634

0.822 0.481 0.415 0.624 0.359

0.877⋆ 0.755 0.837 0.830 0.770

0.932⋆ 0.693 0.904⋆ 0.856 0.741

0.861⋆ 0.786 0.837 0.835 0.811

met 5 1 2 1 1

Colonoscopy bars: PolypGen sensitivity ≥0.90 at precision ≥0.85; REAL-Colon per-polyp sensitivity ≥0.93 at ≤2 false alarms per procedure; other-hospital Dice ≥0.82; segment recall ≥0.70. Gastroscopy bars: lesion gate sensitivity ≥0.93 at specificity ≥0.90; landmark-region accuracy ≥0.92; 23-code macro ≥0.85; findings macro ≥0.85; UGIAD coarse-5 ≥0.90; GastroVision-half ≥0.85.

not the trunk (Section 2.4). What supervised ImageNet features do not substitute for is therefore narrower than a whole trunk’s worth, and sharper for being narrow: not finding a lesion, but naming where the frame was taken and what else is in it. Candidate 2 clears the gate and still loses on everything the gate does not test. Its colonoscopy arm meets both detection marks at epoch 2 — 0.943 of PolypGen polyps found at precision ≥0.85, and 0.947 of REAL-Colon polyps at two false alarms per procedure — and meets two bars, the count woma meets. Every descriptive head reads lower: other-hospital outline Dice 0.616 against 0.807, ten-segment recall 0.436 against 0.521, polyp type 0.583 against 0.634. Nor does it hold. By epoch 20 segment recall is 0.101 and Dice 0.305 while detection still reads 0.930, so the rule selects its second snapshot because that is the last one worth having. A trunk can keep what a detector needs while shedding what every other head reads, and this is the arm that shows it plainly. The comparison also exposed a defect in our trainer that only a network with batch normalisation can show; it was fixed and verified before the candidate-5 runs were repeated, and no woma number changed.

woma: real-time endoscopy foundation model

sensitivity

sensitivity

0.95 0.9 0.85 2

6

10

14

18

1 0.95 0.9 0.85 0.8

2

Dice

0.7 0.6 2

6

10

14

18

6 10 14 18 polyp type (KUMC)

0.8 0.6 0.4

2

6

epoch woma (candidate 3) candidate 2

10

14

0.6 0.4 2 6 10 14 18 HyperKvasir landmarks, macro recall

macro recall

macro recall

cross-dataset Dice 0.8

colon segment, macro recall

REAL-Colon per-polyp sensitivity

1

macro recall

PolypGen sensitivity

12

0.8 0.6 0.4

18

epoch candidate 5 candidate 7

2

6

10

14

18

epoch

candidate 1 bar

Figure 7: Colonoscopy fine-tunes on one recipe, scored on held-out data at every second epoch: woma (candidate 3), candidate 5, candidate 1, candidate 2, candidate 7. Line dash and marker identify the arm. Dotted lines are the bars of Figure S1 (PolypGen sensitivity at precision ≥0.85; REAL-Colon per-polyp sensitivity at ≤2 false activations per procedure over 19 polyps; cross-dataset Dice over CVC-ColonDB, ETIS and CVC-300; polyp type on the KUMC held-out split). The last panel is HyperKvasir, another centre: the macro over its two well-supported landmarks, the caecum (1,009 frames) and the retroflexed rectum (391), which a head cannot raise by answering caecum more often; its ileum carries nine frames and is reported in Table 5 rather than plotted. Kept model is epoch 6 of woma.

station region accuracy

2

0.8

sensitivity

0.6 0.4

0.6 0.4 2

6 10 14 18 findings validation macro

6 10 14 18 UGIAD coarse-5 accuracy

0.8 0.7 0.6

1 0.95 0.9 0.85 0.8

2

6 10 14 18 GastroVision-half accuracy

0.9

0.9 accuracy

macro recall

macro recall

0.8

lesion gate sensitivity

station 23-code macro

accuracy

accuracy

1

0.8 0.6 0.4

2

6

10

14

18

2

6

epoch woma (candidate 3) candidate 2

10

14

18

epoch candidate 5 candidate 7

0.8 0.7 0.6 0.5

2

candidate 1 bar

6

10

14

18

epoch

Figure 8: Gastroscopy fine-tunes on one root and recipe, scored on held-out data at every second epoch: woma (candidate 3), candidate 5, candidate 1, candidate 2, candidate 7. Every arm clears the lesion gate and they separate on everything read from pooled features; at the snapshots the rule selects (Table 4) that separation runs in the order the frozen probe of Table 3 put them in. Bars of Figure S2 dotted; the shipped model is the assembled epoch-10 snapshot of Section 2.4.

woma: real-time endoscopy foundation model

3.3

13

Held-out results of the two models

Tables 5 and 6 give every held-out number against its pass mark. In both models the run’s own validation kept improving through the final learning-rate decay while every other-hospital number fell: in the colon run, PolypGen sensitivity peaked at epoch 4 (0.970) and fell to 0.898 by epoch 20 while the in-domain detection score plateaued, and an internal composite selector would have picked epoch 9. Epoch 6 is the colon keep, the only snapshot that passes both detection marks with margin. Upper-GI heads reach their marks at different epochs, so what ships is an assembly: epoch 10’s foundation model, detector, outline and station heads with a findings head retrained for two epochs on that frozen network; heads retrained the same way on the epoch-4, 6 and 8 networks clear every mark too. Table 5: Colon model, epoch 6 as shipped, against its pass marks, every number read on data the training never saw. 95 % confidence intervals are Wilson score intervals from the frame or polyp counts [42]. instrument

value

95 % CI

mark

met

0.960 0.897 1.000 17 of 19

0.948–0.969 — 0.83–1.00 —

≥0.90 — ≥0.93 —

yes reported yes reported

polyp outline, Dice Kvasir-SEG, in domain CVC-ClinicDB [43], in domain CVC-ColonDB, another hospital ETIS, another hospital CVC-300, another hospital mean of the three other-hospital sets

0.907 0.900 0.748 0.752 0.922 0.807

— — — — — —

— — — — — ≥0.82

reported reported reported reported reported no, by 0.013

colon segment, 15 held-out videos ten segments, macro recall ten segments, with an 8-second memory six merged classes, accuracy (Saito protocol) HyperKvasir, another hospital: caecum recall HyperKvasir: retroflexed-rectum recall HyperKvasir: terminal-ileum recall (9 frames)

0.521 0.574 0.655 0.737 0.529 0.220

— — — — — —

≥0.70 — 0.666 — — —

no reported parity reported reported reported

4 of 15 2 of 2 0 of 2 0.617 0.635

— — — — —

— — — ≥0.80 ≥0.80

no — — no no

0.737 0.845 0.985

— — —

— — —

polyp detection PolypGen sensitivity, precision ≥0.85 PolypGen sensitivity, half the clean frames flagged REAL-Colon per-polyp, 1.6 alarms/procedure REAL-Colon per-polyp, 0.2 alarms/procedure

polyp type per lesion, REAL-Colon: adenoma per lesion, REAL-Colon: hyperplastic per lesion, REAL-Colon: serrated per frame, REAL-Colon: macro recall per frame, KUMC held-out patients: macro recall bowel preparation macro recall over the four Boston classes weighted kappa grouped accuracy, adequate against not

reported reported reported

PolypGen is 1,347 polyp frames and 193 clean frames from six hospitals: at the operating point above, 1,293 polyps are found and 54 missed, of which only 20 carry no overlapping box at all (Section 4). REAL-Colon study 4 is 15 full videos holding 19 polyps, scored per polyp with a two-second persistence rule; the mark allows two sustained false alarms per procedure and the shipped model runs at 1.6.

Both detection marks are met. Read that per-frame number with its operating point: at the loosest threshold that holds precision at 0.85, nearly all of PolypGen’s 193 clean frames receive some box, which is why per-procedure counting is what a clinician should weigh and why our program ships at stricter confidence. Outlines miss their other-hospital mark by 0.013, on the two sets with the smallest and flattest polyps. Segment recall reads parity with published single-frame results on a merged six-class protocol, and 0.52 across ten classes; its errors sit in adjacent tube segments one frame cannot separate, and an 8-second memory adds 0.05. Polyp type falls far short on a class that matters, trained on 256 serrated frames. Our upper-GI model meets five of five gated marks in one set of weights, and region accuracy holds up on separate test patients; its 23-code

woma: real-time endoscopy foundation model

14

Table 6: Upper-GI model, assembled epoch-10 snapshot as shipped, against its pass marks, on data the training never saw. 95 % CI as in Table 5. The tiers are the order the selection rule reads the marks in: a snapshot must clear the lesion gate first. instrument

value

95 % CI

mark

met

station, unseen patients — tier 1 validation, 793 frames: landmark-region accuracy validation: landmark-region macro recall validation: 23-code macro recall validation: 23-code accuracy test, 803 frames from 59 other patients, read once test: 23-code macro recall

0.923 0.896 0.822 0.836 0.917 0.834

0.902–0.940 — — — 0.895–0.934 —

≥0.92 — ≥0.85 — ≥0.92 —

yes reported no reported yes reported

lesion gate and findings — tier 2 sensitivity, 39 neoplasia frames specificity, 1,347 clean frames, another hospital findings, validation macro recall over seven classes findings, UGIAD test: coarse-5 accuracy findings, UGIAD test: fine accuracy findings, UGIAD test: fine macro recall

0.949 0.912 0.877 0.932 0.866 0.828

0.83–0.99 0.895–0.926 — 0.910–0.950 — —

≥0.93 ≥0.90 ≥0.85 ≥0.90 — —

yes yes yes yes reported reported

other-hospital findings — tier 3 GastroVision half, two other hospitals: accuracy

0.861

0.840–0.880

≥0.85

yes

outline EDD2020 validation, 52 masks: Dice

0.769

—

—

reported

The station corpus is GastroHUN’s patient split [10]; the landmark region merges the 23 codes into the seven anatomical regions a report names. The lesion gate is EDD2020’s neoplasia frames [38] against lesion-free landmark frames from HyperKvasir [9], a hit counted at a quarter overlap; 37 of the 39 are flagged. UGIAD is 606 frames of its published test split [19] and GastroVision half is 1,152 frames [39]. The outline head is research-use-only because EDD2020 is.

macro is reported against its line, since codes it misses are neighbouring views of one landmark that annotators themselves disagreed on.

3.4

Speed

Table 7 sets numbat against three programs such a system would otherwise be built on: PyTorch, ONNX Runtime and TensorRT. Each was handed identical layers and identical weights on one card and asked to do two jobs — one pass of our foundation model, and a whole frame from video memory to finished results. Four precisions appear, meaning how many bits each number carries: ordinary 32-bit arithmetic (f32), NVIDIA’s faster 32-bit mode (TF32), and two 16-bit formats (bf16, f16). Fewer bits means quicker arithmetic and less traffic to memory; what that costs in accuracy we measured separately (Section 4). numbat is fastest in every row. One pass of our foundation model takes TensorRT, strongest of the three, 12 % longer in ordinary 32-bit, 19 % in TF32, 3 % in bf16 and 27 % in f16, and takes ONNX Runtime and PyTorch 1.6 to 3.0 times longer. Frame to finished results — every head run, detections turned into boxes, overlapping boxes dropped — our colon model delivers 248 frames a second in f16 where TensorRT delivers 215, and 160 against 129 in ordinary 32-bit; our gastroscopy model 243 against 213, and 161 against 129. Two rounds of work got there, each timing our own code step by step beside whichever program was beating us. Round one deleted work rather than speeding it up. A convolution block used to write its result to memory, read it back to add a bias, read it again to apply an activation, and again for a residual and a normalisation. Doing all of that inside the pass that already holds the numbers removes four journeys to memory and changes no answer — outputs stayed identical to the last digit. Round one also taught every part of the program, not our foundation model alone, to work in 16-bit numbers from end to end. Round two compared single operations with TensorRT’s. Ours were already quicker at 1×1 convolutions, commonest operation in this network. Its 7×7 depthwise convolution — one filter per channel, which is what gives a ConvNeXt block its wide view of a frame — ran five times faster than ours, and that one operation was the whole remaining gap. Rewriting it closed that gap: a row

woma: real-time endoscopy foundation model

15

Table 7: numbat against the engines a deployment is measured against, on the same GPU, the same layers, in one sitting under the protocol of Section 2.5. Each entry is the median of three repeats with their standard deviation, so every number in this table is a distribution rather than one run. The foundation model is one ConvNeXt-V2-Nano forward at 1×3×640×640 with the input already on the card, in milliseconds; each fine-tuned model is a frame in host memory to boxes, mask, station and findings in host memory, in frames per second. Best in each row in bold. regime

numbat

PyTorch

ONNX Runtime

TensorRT

foundation model, milliseconds, lower is better strict f32 4.30 ± 0.01 f32 with TF32 3.74 ± 0.01 bf16 2.49 ± 0.00 f16 2.42 ± 0.01

9.86 ± 0.03 6.21 ± 0.00 3.88 ± 0.01 3.82 ± 0.02

11.94 ± 0.04 9.87 ± 0.00 — 7.15 ± 0.01

4.83 ± 0.01 4.46 ± 0.00 2.57 ± 0.00 3.08 ± 0.00

upper-GI model, frames per second strict f32 161.3 ± 0.4 f32 with TF32 175.4 ± 0.5 bf16 237.0 ± 0.9 f16 243.3 ± 0.0

68.5 ± 0.0 96.9 ± 0.1 — 155.3 ± 0.3

61.7 ± 0.1 75.4 ± 0.1 — 100.2 ± 0.1

128.9 ± 0.1 135.6 ± 0.1 222.7 ± 1.4 213.2 ± 0.9

colon model, frames per second strict f32 159.7 ± 0.1 f32 with TF32 177.9 ± 0.2 bf16 243.3 ± 0.0 f16 247.5 ± 0.6

68.4 ± 0.0 96.9 ± 0.0 — 154.7 ± 0.2

61.7 ± 0.1 75.3 ± 0.1 — 99.8 ± 0.2

129.0 ± 0.1 135.9 ± 0.1 224.0 ± 0.3 215.3 ± 0.4

TF32 is the tensor-core mode every framework turns on by default for convolutions. A dash is a regime the engine has no path for: ONNX Runtime’s CUDA provider does not take a bf16 graph, and PyTorch’s bf16 rows are not run end to end. numbat 0.9.14; PyTorch 2.14.0 with CUDA 13 in eager mode, channels-last for its 16-bit rows, with torch.compile at max-autotune measured beside it (3.59 ms on the foundation model in bf16, its best compiled result, against 3.57 eager); ONNX Runtime 1.29.0 on its CUDA execution provider; TensorRT 11.3.0, engines built from the same ONNX graph at optimisation level 5. Every engine’s outputs were checked against PyTorch’s on a fixed frame before anything was timed. Repeats were taken in alternating order, numbat’s each from a separate launch of the program; in every row numbat’s slowest repeat is faster than the fastest repeat of every engine it is compared with.

of output now reads its strip of the frame once into the small fast memory beside the arithmetic units, one thread produces four neighbouring outputs, and 16-bit inputs stay 16-bit instead of being widened to 32. Two smaller changes followed. One normalisation step is folded into the weights of the layer after it, so it costs nothing while the program runs. And the pass that writes an activation now hands back the totals the next step needs, instead of that step reading the whole activation a second time to work them out. f16 is what the program runs by default. Ordinary 32-bit stays available as a switch and is what every held-out number in this paper was read with, and the optimised build reproduces the previous release’s 32-bit output on every frame of both test clips. A build that needs nothing installed. Everything above reaches the card through NVIDIA’s libraries — CUDA, cuBLAS, cuDNN — which a workstation must have installed, and kept in step with its graphics driver. numbat can run the same weights a second way. We write each GPU operation once and compile it to SPIR-V, a portable format for GPU programs that modern drivers accept, then send it to the card through Vulkan, which ships inside the driver itself. Nothing from any vendor is built in. What a site installs is then two files, one program and one model, and on Linux that program asks the system for its C library and loader and for nothing else. No toolkit, no cuDNN version to match, no PyTorch, no Python. Those two files run on any card whose driver can run Vulkan, which today means NVIDIA, AMD and Intel, and will mean Apple silicon once we have tested that route. Portability usually costs speed. Table 8 runs one program both ways, on one card and on the same two clips, timed exactly as Table 7 timed the four engines — so the two tables can be read against each other. In 32-bit the portable build is the faster of the two, 185 frames a second against 162 on our colon model and 180 against 161 on our gastroscopy model, because our own code puts a 32-bit convolution on the card’s matrix-multiply units where cuDNN takes a slower route. In 16-bit

woma: real-time endoscopy foundation model

16

it is behind by 12 and 16 %, 217 against 246 and 207 against 245, and cuDNN’s 7×7 depthwise convolution is most of that difference. It still clears 200 frames a second, eight times what a screen shows, and it finds the same things: identical box counts in every cell where both planes run. A department with no CUDA, or with a card that is not an NVIDIA one, gives up nothing in what this system finds and, in 32-bit, gains speed. Table 8: numbat’s two GPU planes, measured under the protocol and in the units of Table 7: same clips, three repeats, and the same quantity — a whole frame from video memory to finished results, in frames per second. Both columns were timed together on numbat 0.9.16, and this CUDA column agrees with Table 7’s numbat column to within 1.5 %, which is the spread between sittings. What each plane finds is identical: 113 boxes in 32-bit and 112 in 16-bit on the colonoscopy clip, 2 on the gastroscopy clip, in every cell where both planes run. Faster of the two in bold. colon model, FPS

upper-GI model, FPS

regime

CUDA

portable

CUDA

portable

strict f32 f32 with TF32 bf16 f16

162.1 179.9 243.9 245.7

184.8 — — 217.4

161.0 178.9 241.0 244.5

180.2 — — 206.6

CUDA is NVIDIA’s libraries (cuBLAS, cuDNN); portable is our own kernels through Vulkan, with no vendor library linked. A dash is a regime the portable plane has no path for: TF32 is a mode of NVIDIA’s tensor cores and has no counterpart, so its f32 row stands for both; bf16 kernels are not written for it yet and the program stops with UnsupportedType rather than guessing. One pass of the foundation model alone, the first block of Table 7, is not quoted for the portable plane: it queues work and returns before the card has run it, so a per-stage stopwatch measures the queueing, not the arithmetic. The whole-frame figure is sound on both planes because reading detections back forces the card to finish.

Table S2 gives our program as delivered: what ships is one file per model (66.6 MB, holding foundation model, every head, class names, thresholds and a fingerprint of every source weight) and one program on that same library, with video decoding and encoding compiled in and no framework at runtime. It runs as a three-stage pipeline (decode and letterbox; foundation model, heads and overlay; encode) whose numbers are identical to a serial program on every frame, and it draws nothing that obstructs the mucosa: four corner marks per detection, contours only for outlines, a station checklist that ticks a site only after a one-second vote, and a badge with the delivered rate.

4

Discussion

What the systematic design bought. Fixing requirements, candidates and decision rules before any run changed what the experiment could find. That funnel caught a whole route failing — three distilled candidates collapsing on a teacher target that was smooth rather than rich — and could say so because two candidates carried every route, keeping route failure distinguishable from candidate failure. It also caught what training curves hide: in both models in-domain validation kept rising through the final decay while every other-hospital number fell, so a selector reading those curves would have shipped a worse model. Running probes, trainer, evaluators and runtime on one library, with identical operators in training and inference, is what lets a shipped file reproduce its acceptance record to three decimals, and what made the speed work possible — profiling against TensorRT said which kernel was behind, and we could rewrite it that day. Re-scored in f16 both files move 6 of 141 upper-GI record fields and 5 of 41 colon fields, each by at most one frame; in bf16, 27 and 12. No pass mark changes status. What the models are worth. woma is a 15.6-million-parameter network serving nine heads across two products inside a 16.7 ms budget. Three marks go unmet: colon segment (a single-frame ceiling, where an 8-second memory adds 0.05 and 0.7 needs a model with memory), polyp type (4 of 15 adenomas per lesion, on 256 serrated training frames) and other-hospital outline, short by 0.013 and 0.03 Dice under the best published Kvasir-SEG numbers [44] from single-task models several times as expensive. Other candidates make our choice of foundation model legible: on identical

woma: real-time endoscopy foundation model

17

recipes they find lesions as well as woma yet describe a frame worse, in the order a frozen probe had already put them in. A supervised ImageNet residual network, most conventional among them, makes that point sharpest — it flags lesions as well as anything here and names a landmark region on 57.5 % of frames where woma names it on 92.3 %. Where they fail. Of 1,347 PolypGen polyp frames the colon detector misses 54 at the half-overlap its sensitivity is read at, but only 20 carry no overlapping box at all: two thirds of the misses are a box on the lesion that is not tight enough, which costs a number and would not cost a polyp. All 19 REAL-Colon lesions are found while up to two sustained false activations per procedure are allowed; tightened to 0.2, two lesions are held only by boxes under that threshold. The upper-GI gate flags 37 of 39 held-out neoplasia frames while leaving 91 % of 1,347 clean frames from another hospital alone; the two it misses carry no box at a quarter overlap. The findings head fails where its training data is thinnest and in one direction: its largest single UGIAD error is neoplasia read as gastric inflammation (23 of 153 frames), and it recovers 7 of 32 GastroVision gastric polyps. Both classes are carried in the hundreds against tens of thousands of normal frames, and both err towards the benign reading — the direction that matters most in a screening tool, and what a local fine-tune is aimed at. What held-out evaluation does not settle. Every number here was read on other hospitals, other patients or other videos, and shipped weights were chosen on that record. That is one step, not a destination: deciding samples are small. REAL-Colon study 4 holds 19 polyps, so 19 of 19 carries a 95 % lower bound of 0.83 [42], and the upper-GI gate has 39 positive frames, mostly oesophageal stills. Neither gate sees blur, bubbles, instruments or gastric cancers in motion, no endoscopist has used the overlay during a procedure, and the GPU is a workstation card rather than a procedure-room box. What is missing is a full-procedure gastroscopy set with per-lesion truth, and none is public. Data from our own institutions would serve twice — as unlabelled pretraining frames, and as the held-out set that finally measures a full procedure on local equipment — and everything here was built so it enters without a change of method: same splits by patient, same pass marks, same held-out selection.

5

Conclusion

We set out to build a real-time foundation model for endoscopy as one builds a product: requirements first, every decision written down before it was taken. Out came woma, chosen from eight candidates by a pre-registered experiment and trained without labels to a stopping rule, and two fine-tuned models on it that meet most of their pass marks on data training never saw and state the rest. Every step, from screening probes to shipped file, ran on one self-contained library, and both foundation model and fine-tuned models run faster than PyTorch, ONNX Runtime and TensorRT on the same GPU in every precision regime. What ships is one file per model and one program, and on a second build that program carries no vendor library at all.

Declarations Ethics. This study used only publicly available, de-identified datasets under their published licences, named with their roles in Figures 2, S1 and S2; the authors collected no data from patients, and no ethical approval was required. Datasets licensed for research use only (EDD2020, UGIAD, SAGE [45], GastroEndoNet [46], Kvasir v2 [47] and GastroVision) trained heads that are marked as research heads and are not offered for clinical use. The models are research artefacts and not medical devices; no endoscopist has used them during a procedure. Data and code availability. Every held-out score is a record written by one scorer, and every table and chart regenerates from those records. Records, both model files, our program and its

woma: real-time endoscopy foundation model

18

benchmark scripts are available from the corresponding author. Our library, numbat, is described in Tran and Dang [6] and Tran and Dang [7]. Funding. This work received no grant or funding from any agency in the public, commercial or not-for-profit sectors. Declaration of competing interests. T.T. develops numbat and is affiliated with CloudKites AI Lab, which may commercialise the library and the models described here. L.D. declares no competing interests. CRediT author contributions. Thang Tran: conceptualisation, methodology, software, investigation, data curation, validation, visualisation, writing – original draft. Lan Dang: conceptualisation, methodology, project administration, writing – review and editing. Acknowledgement. AI assistance was used under our direction to perform literature search, and to draft and edit text and figure code.

References [1] Alessandro Repici, Matteo Badalamenti, Roberta Maselli, Loredana Correale, Franco Radaelli, Emanuele Rondonotti, Elisa Ferrara, Marco Spadaccini, Asma Alkandari, Alessandro Fugazza, Andrea Anderloni, Piera Alessia Galtieri, Gaia Pellegatta, Silvia Carrara, Milena Di Leo, Vincenzo Craviotto, Laura Lamonaca, Roberto Lorenzetti, Alida Andrealli, Giulio Antonelli, Michael Wallace, Prateek Sharma, Thomas Rosch, and Cesare Hassan. Efficacy of real-time computer-aided detection of colorectal neoplasia in a randomized trial. Gastroenterology, 159 (2):512–520.e7, 2020. doi: 10.1053/j.gastro.2020.04.062. [2] Cesare Hassan, Marco Spadaccini, Andrea Iannone, Roberta Maselli, Manol Jovani, Viveksandeep Thoguluva Chandrasekar, Giulio Antonelli, Honggang Yu, Miguel Areia, Mario DinisRibeiro, Pradeep Bhandari, Prateek Sharma, Douglas K. Rex, Thomas Rösch, Michael Wallace, and Alessandro Repici. Performance of artificial intelligence in colonoscopy for adenoma and polyp detection: a systematic review and meta-analysis. Gastrointestinal Endoscopy, 93(1): 77–85.e6, 2021. doi: 10.1016/j.gie.2020.06.059. [3] Andrea Cherubini and Nhan Ngo Dinh. A review of the technology, training, and assessment methods for the first real-time AI-enhanced medical device for endoscopy. Bioengineering, 10 (4):404, 2023. doi: 10.3390/bioengineering10040404. [4] Lianlian Wu, Jun Zhang, Wei Zhou, Ping An, Lei Shen, Jun Liu, Xiaoda Jiang, Xu Huang, Ganggang Mu, Xinyue Wan, Xiaoguang Lv, Juan Gao, Ning Cui, Shan Hu, Yiyun Chen, Xiao Hu, Jiangjie Li, Di Chen, Dexin Gong, Xinqi He, Qianshan Ding, Xiaoyun Zhu, Suqin Li, Xiao Wei, Xia Li, Xuemei Wang, Jie Zhou, Mengjiao Zhang, and Hong Gang Yu. Randomised controlled trial of WISENSE, a real-time quality improving system for monitoring blind spots during esophagogastroduodenoscopy. Gut, 68(12):2161–2169, 2019. doi: 10.1136/gutjnl-2018-317366. [5] Toshiaki Hirasawa, Kazuharu Aoyama, Tetsuya Tanimoto, Soichiro Ishihara, Satoki Shichijo, Tsuyoshi Ozawa, Tatsuya Ohnishi, Mitsuhiro Fujishiro, Keigo Matsuo, Junko Fujisaki, and Tomohiro Tada. Application of artificial intelligence using a convolutional neural network for detecting gastric cancer in endoscopic images. Gastric Cancer, 21(4):653–660, 2018. doi: 10.1007/s10120-018-0793-2. [6] Thang Tran and Lan Dang. Numbat: Building and verifying a self-contained machine-learning stack. arXiv:2609.10632 [cs.SE], 2026.

woma: real-time endoscopy foundation model

19

[7] Thang Tran and Lan Dang. Cross-stack validation of language-model training: A clinical fine-tuning case study. arXiv:2608.24267 [cs.SE], 2026. [8] Tim G. W. Boers, Kiki N. Fockens, Joost A. van der Putten, Tim J. M. Jaspers, Carolus H. J. Kusters, Jelmer B. Jukema, Martijn R. Jong, Maarten R. Struyvenberg, Jeroen de Groof, Jacques J. Bergman, Peter H. N. de With, and Fons van der Sommen. Foundation models in gastrointestinal endoscopic AI: Impact of architecture, pre-training approach and data efficiency. Medical Image Analysis, 98:103298, 2024. doi: 10.1016/j.media.2024.103298. [9] Hanna Borgli, Vajira Thambawita, Pia H. Smedsrud, Steven Hicks, Debesh Jha, Sigrun L. Eskeland, Kristin Ranheim Randel, Konstantin Pogorelov, Mathias Lux, Duc Tien Dang Nguyen, Dag Johansen, Carsten Griwodz, Håkon K. Stensland, Enrique Garcia-Ceja, Peter T. Schmidt, Hugo L. Hammer, Michael A. Riegler, Pål Halvorsen, and Thomas de Lange. HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy. Scientific Data, 7:283, 2020. doi: 10.1038/s41597-020-00622-y. [10] Diego Bravo, Juan Frias, Felipe Vera, Juan Trejos, Carlos Martínez, Martín Gómez, Fabio González, and Eduardo Romero. GastroHUN an endoscopy dataset of complete systematic screening protocol for the stomach. Scientific Data, 12:102, 2025. doi: 10.1038/s41597-025-04401-5. [11] Debesh Jha, Pia H. Smedsrud, Michael A. Riegler, Pål Halvorsen, Thomas de Lange, Dag Johansen, and Håvard D. Johansen. Kvasir-SEG: A segmented polyp dataset. In MultiMedia Modeling (MMM 2020), volume 11962 of Lecture Notes in Computer Science, pages 451–462. Springer, 2020. doi: 10.1007/978-3-030-37734-2_37. [12] Masashi Misawa, Shin-ei Kudo, Yuichi Mori, Kinichi Hotta, Kazuo Ohtsuka, Takahisa Matsuda, Shoichi Saito, Toyoki Kudo, Toshiyuki Baba, Fumio Ishida, Hayato Itoh, Masahiro Oda, and Kensaku Mori. Development of a computer-aided detection system for colonoscopy and a publicly accessible large colonoscopy video database (with video). Gastrointestinal Endoscopy, 93(4):960–967.e3, 2021. doi: 10.1016/j.gie.2020.07.060. [13] Ge-Peng Ji, Guobao Xiao, Yu-Cheng Chou, Deng-Ping Fan, Kai Zhao, Geng Chen, and Luc Van Gool. Video polyp segmentation: A deep learning perspective. Machine Intelligence Research, 19(6):531–549, 2022. doi: 10.1007/s11633-022-1371-y. [14] Sharib Ali, Debesh Jha, Noha Ghatwary, Stefano Realdon, Renato Cannizzaro, Osama E. Salem, Dominique Lamarque, Christian Daul, Michael A. Riegler, Kim V. Anonsen, Andreas Petlund, Pål Halvorsen, Jens Rittscher, Thomas de Lange, and James E. East. A multi-centre polyp detection and segmentation dataset for generalisability assessment. Scientific Data, 10:75, 2023. doi: 10.1038/s41597-023-01981-y. [15] Carlo Biffi, Giulio Antonelli, Sebastian Bernhofer, Cesare Hassan, Daizen Hirata, Mineo Iwatate, Andreas Maieron, Pietro Salvagnini, and Andrea Cherubini. REAL-Colon: A dataset for developing real-world AI applications in colonoscopy. Scientific Data, 11:539, 2024. doi: 10.1038/s41597-024-03359-0. [16] Deng-Ping Fan, Ge-Peng Ji, Tao Zhou, Geng Chen, Huazhu Fu, Jianbing Shen, and Ling Shao. PraNet: Parallel reverse attention network for polyp segmentation. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2020, volume 12266 of Lecture Notes in Computer Science, pages 263–273. Springer, 2020. doi: 10.1007/978-3-030-59725-2_26. [17] Hiroaki Saito, Tetsuya Tanimoto, Tsuyoshi Ozawa, Soichiro Ishihara, Mitsuhiro Fujishiro, Satoki Shichijo, Dai Hirasawa, Tomoki Matsuda, Yuma Endo, and Tomohiro Tada. Automatic anatomical classification of colonoscopic images using deep convolutional neural networks. Gastroenterology Report, 9(3):226–233, 2021. doi: 10.1093/gastro/goaa078.

woma: real-time endoscopy foundation model

20

[18] Edwin J. Lai, Audrey H. Calderwood, Gheorghe Doros, Oren K. Fix, and Brian C. Jacobson. The Boston bowel preparation scale: a valid and reliable instrument for colonoscopy-oriented research. Gastrointestinal Endoscopy, 69(3, Part 2):620–625, 2009. doi: 10.1016/j.gie.2008.05.057. [19] In Neng Chan, Pak Kin Wong, Tao Yan, Yanyan Hu, Chon In Chan, Peixuan Ge, Zheng Li, Ying Hu, Shan Gao, and Hon Ho Yu. FastUGI-Net: Enhanced real-time endoscopic diagnosis with efficient multi-task learning. Expert Systems with Applications, 280:127444, 2025. doi: 10.1016/j.eswa.2025.127444. [20] Sharib Ali, Noha Ghatwary, Debesh Jha, Ece Isik-Polat, Gorkem Polat, Chen Yang, Wuyang Li, Adrian Galdran, et al. Assessing generalisability of deep learning-based polyp detection and segmentation methods through a computer vision challenge. Scientific Reports, 14:2032, 2024. doi: 10.1038/s41598-024-52063-x. [21] Erik A. Holzwanger, Mohammad Bilal, Jeremy R. Glissen Brown, Shailendra Singh, Aymeric Becq, Kenneth Ernest-Suarez, and Tyler M. Berzin. Benchmarking definitions of false-positive alerts during computer-aided polyp detection in colonoscopy. Endoscopy, 53(9):937–940, 2021. doi: 10.1055/a-1302-2942. [22] Jorge Bernal, Javier Sánchez, and Fernando Vilariño. Towards automatic polyp detection with a polyp appearance model. Pattern Recognition, 45(9):3166–3182, 2012. doi: 10.1016/j.patcog. 2012.03.002. [23] Juan Silva, Aymeric Histace, Olivier Romain, Xavier Dray, and Bertrand Granado. Toward embedded detection of polyps in WCE images for early diagnosis of colorectal cancer. International Journal of Computer Assisted Radiology and Surgery, 9:283–293, 2014. doi: 10.1007/s11548-013-0926-3. [24] David Vázquez, Jorge Bernal, F. Javier Sánchez, Gloria Fernández-Esparrach, Antonio M. López, Adriana Romero, Michal Drozdzal, and Aaron Courville. A benchmark for endoluminal scene segmentation of colonoscopy images. Journal of Healthcare Engineering, 2017:4037190, 2017. doi: 10.1155/2017/4037190. [25] Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9650–9660, 2021. arXiv:2104.14294. [26] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2024. arXiv:2304.07193; OpenReview a68SUt6zFt. [27] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16000–16009, 2022. doi: 10.1109/ CVPR52688.2022.01553. arXiv:2111.06377. [28] Keyu Tian, Yi Jiang, Qishuai Diao, Chen Lin, Liwei Wang, and Zehuan Yuan. Designing BERT for convolutional networks: Sparse and hierarchical masked modeling. In International Conference on Learning Representations (ICLR), 2023. arXiv:2301.03580; ICLR 2023 spotlight, OpenReview NRxydtWup1S. [29] Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. ConvNeXt V2: Co-designing and scaling ConvNets with masked autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16133–16142, 2023. arXiv:2301.00808.

woma: real-time endoscopy foundation model

21

[30] Adrien Bardes, Jean Ponce, and Yann LeCun. VICReg: Variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations (ICLR), 2022. arXiv:2105.04906; ICLR 2022 poster. [31] Colorado J. Reed, Xiangyu Yue, Ani Nrusimha, Sayna Ebrahimi, Vivek Vijaykumar, Richard Mao, Bo Li, Shanghang Zhang, Devin Guillory, Sean Metzger, Kurt Keutzer, and Trevor Darrell. Self-supervised pretraining improves self-supervised pretraining. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2584–2594, 2022. doi: 10.1109/WACV51458.2022.00112. arXiv:2103.12718. [32] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. doi: 10.1109/CVPR.2016.90. [33] Richard Zhang. Making convolutional networks shift-invariant again. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pages 7324–7334, 2019. URL https://proceedings.mlr.press/ v97/zhang19a.html. arXiv:1904.11486. [34] Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 936–944, 2017. doi: 10.1109/CVPR. 2017.106. [35] Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ultralytics YOLOv8, 2023. URL https: //github.com/ultralytics/ultralytics. AGPL-3.0. [36] Yiming Song, Zhengjie Zhang, Ruilan Wang, Ling Zhong, Crystal Cai, Jinnan Chen, Yujie Zhou, Xinyuan Wang, Zhao Li, Liuyi Yang, Zeyu Li, Hao Yan, Qingwei Zhang, Dahong Qian, and Xiaobo Li. CAS-Colon: A comprehensive colonoscopy anatomical segmentation dataset for artificial intelligence development. Scientific Data, 12:1382, 2025. doi: 10.1038/ s41597-025-05588-3. [37] Kaidong Li, Mohammad I. Fathan, Krushi Patel, Tianxiao Zhang, Cuncong Zhong, Ajay Bansal, Amit Rastogi, Jean S. Wang, and Guanghui Wang. Colonoscopy polyp detection and classification: Dataset creation and comparative evaluations. PLOS ONE, 16(8):e0255809, 2021. doi: 10.1371/journal.pone.0255809. [38] Sharib Ali, Noha Ghatwary, Barbara Braden, Dominique Lamarque, Adam Bailey, Stefano Realdon, Renato Cannizzaro, Jens Rittscher, Christian Daul, and James East. Endoscopy disease detection challenge 2020, 2020. arXiv:2003.03376. [39] Debesh Jha, Vanshali Sharma, Neethi Dasu, Nikhil Kumar Tomar, Steven Hicks, M. K. Bhuyan, Pradip K. Das, Michael A. Riegler, Pål Halvorsen, Ulas Bagci, and Thomas de Lange. GastroVision: A multi-class endoscopy image dataset for computer aided gastrointestinal disease detection. In Machine Learning for Multimodal Healthcare Data (ICML 2023 Workshop ML4MHD), volume 14315 of Lecture Notes in Computer Science. Springer, 2024. doi: 10.1007/ 978-3-031-47679-2_10. arXiv:2307.08140. [40] Ross Wightman. PyTorch image models, 2019. URL https://github.com/huggingface/ pytorch-image-models. [41] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y.

woma: real-time endoscopy foundation model

22

[42] Edwin B. Wilson. Probable inference, the law of succession, and statistical inference. Journal of the American Statistical Association, 22(158):209–212, 1927. doi: 10.1080/01621459.1927. 10502953. [43] Jorge Bernal, F. Javier Sánchez, Gloria Fernández-Esparrach, Debora Gil, Cristina Rodríguez, and Fernando Vilariño. WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians. Computerized Medical Imaging and Graphics, 43: 99–111, 2015. doi: 10.1016/j.compmedimag.2015.02.007. [44] Peng Li, Jianhua Ding, and Chia S. Lim. VMDU-net: a dual encoder multi-scale fusion network for polyp segmentation with Vision Mamba and Cross-Shape Transformer integration. Frontiers in Artificial Intelligence, 8:1557508, 2025. doi: 10.3389/frai.2025.1557508. [45] Niyoj Oli, Sachin Acharya, Sandesh Pokhrel, Sanjay Bhandari, Ramesh Rana, Nikesh Mani Shrestha, Ram Bahadur Gurung, Yash Raj Shrestha, Prashnna K. Gyawali, and Binod Bhattarai. SAGE: An expert-annotated south asian GI endoscopy dataset for multimodal learning and hallucination analysis. arXiv preprint arXiv:2606.22144, 2026. [46] Abu Kowshir Bitto, Md. Hasan Imam Bijoy, Kamrul Hassan Shakil, Aka Das, Khalid Been Badruzzaman Biplob, Imran Mahmud, and Syed Md. Minhaz Hossain. GastroEndoNet: Comprehensive endoscopy image dataset for GERD and polyp detection. Data in Brief, 60: 111572, 2025. doi: 10.1016/j.dib.2025.111572. [47] Konstantin Pogorelov, Kristin Ranheim Randel, Carsten Griwodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato, Duc-Tien Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt, Michael Riegler, and Pål Halvorsen. KVASIR: A multi-class image dataset for computer aided gastrointestinal disease detection. In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys’17), pages 164–169, 2017. doi: 10.1145/3083187. 3083212. [48] Xiang Li, Wenhai Wang, Lijun Wu, Shuo Chen, Xiaolin Hu, Jun Li, Jinhui Tang, and Jian Yang. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020. arXiv:2006.04388. [49] Yiting Ma, Xuejin Chen, Kai Cheng, Yang Li, and Bin Sun. LDPolypVideo benchmark: A large-scale colonoscopy video dataset of diverse polyps. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2021, volume 12905 of Lecture Notes in Computer Science, pages 387–396. Springer, 2021. doi: 10.1007/978-3-030-87240-3_37. [50] Debesh Jha, Nikhil Kumar Tomar, Vanshali Sharma, et al. PolypDB: A curated multi-center dataset for development of AI algorithms in colonoscopy, 2024. [51] Konstantin Pogorelov, Kristin Ranheim Randel, Thomas de Lange, Sigrun Losada Eskeland, Carsten Griwodz, Dag Johansen, Concetto Spampinato, Mario Taschwer, Mathias Lux, Peter Thelin Schmidt, Michael Riegler, and Pål Halvorsen. Nerthus: A bowel preparation quality video dataset. In Proceedings of the 8th ACM on Multimedia Systems Conference (MMSys), pages 170–174, 2017. doi: 10.1145/3083187.3083216. [52] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019. arXiv:1711.05101. [53] Alexey Bochkovskiy, Chien-Yao Wang, and Hong-Yuan Mark Liao. YOLOv4: Optimal speed and accuracy of object detection, 2020. [54] F. Wang, W. Liao, H. Su, L. Xu, H. Xue, L. Jiang, Y. Yang, M. Piao, J. Sun, and T. Li. A histopathologically verified dataset of magnifying narrow-band imaging endoscopy for classifying the gastric precancerous cascade. Scientific Data, 2026. doi: 10.1038/s41597-026-07567-8.

woma: real-time endoscopy foundation model

23

[55] Zeyu Li, Yingying Fang, Yiming Song, Haotian Bai, Tianhao Li, Yaqi Wang, Qingwei Zhang, and Xiaobo Li. EGID: A comprehensive multi-label endoscopic image dataset for gastritis classification. Scientific Data, 2026. doi: 10.1038/s41597-026-07666-6. Online 22 June 2026. [56] Mojgan Forootan, Mahziar Setayeshfar, Ali Darvishi, Mohammad Tashakoripour, and Hamidreza Bolhasani. GIM-ENDO: A multimodal endoscopic image and video dataset for gastric intestinal metaplasia morphology and pathology. arXiv preprint arXiv:2606.20919, 2026. [57] Sharib Ali, Mariia Dmitrieva, Noha Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Yun Bo Guo, Bogdan Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Le Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, and Jens Rittscher. Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy. Medical Image Analysis, 70:102002, 2021. doi: 10.1016/j.media.2021.102002. [58] Kenshi Yao. The endoscopic diagnosis of early gastric cancer. Annals of Gastroenterology, 26 (1):11–22, 2013. PMID 24714327.

woma: real-time endoscopy foundation model

24

Supplementary material S1. Head designs. Table S1 gives every head’s design; both models share foundation model, the neck and the head designs, and only the class lists differ. S2. How the two fine-tuned models were developed. Figures S1 and S2 give each model’s complete recipe: data and what was kept out of training, training settings, what is measured while a run trains, what is measured on held-out data for every snapshot, pass marks, and the rule that chose shipped weights. S3. Program as delivered. test clips (Section 3.4).

Table S2 gives delivered rate, latency and per-stage cost on both

Table S2: Program as delivered, on one GPU (f16 by default; medians over three runs). Delivered is the interval between finished frames, the badge’s number; the stage figures are per frame with a device sync after every GPU stage.

delivered FPS: mean delivered FPS: median delivered FPS: 5th percentile delivered FPS: minimum frames delivered at ≥24 FPS frame to results, mean ms frame to results, median ms latency, frame read to encoded: mean ms latency, frame read to encoded: median ms

gastroscopy clip 1350×1080, 15 fps

colonoscopy clip 316×256, 30 fps

100.3 99.2 90.4 82.7 100 % 6.9 5.4 19.7 15.3

118.3 117.6 114.3 109.8 100 % 6.9 5.4 18.5 14.1

Clips are 151 and 150 frames. Delivered is the interval between finished frames, the badge’s number; frame to results covers upload, foundation model, heads, decode and NMS. Stage medians on the gastroscopy clip, ms per frame: letterbox and normalise 8.9 (CPU, producer thread), upload 0.9, foundation model 2.8 (5.4 in f32), detection head and decode 0.55, outline head 1.1, station and findings 0.1, NMS and post 0.0, overlay 0.3, encode 6.5 mean (CPU, encoder thread).

S4. Station label spaces and their groupings. Table S3 lists every station class both models carry, and the coarser space each is also scored in. Grouping never changes a prediction: a model answers in its own space, and grouping is applied to both the answer and the reference before counting. S5. Lesion, polyp-type and findings label spaces. Table S4 lists what each detector and each descriptive head answers in. Both detectors carry one class: a box says where, and a separate head says what. S6. Bowel preparation scale. Table S5 gives the published segment scale the preparation head answers in [18].

woma: real-time endoscopy foundation model

25

Table S1: Head designs. Both models share the foundation model, the neck and the head designs; only the class lists differ. head

reads

design and output

detection

P3 , P4 , P5

outline station

P4 , P5 pooled P5

polyp type / findings bowel preparation

pooled P5 pooled P5

feature-pyramid neck under an anchor-free head; box edges as distributions over 16 bins; one class a small decoder to a 6402 mask two-layer classifier to 10 colon segments or 22+1 gastric stations linear classifier to 4 polyp types or 7 upper-GI findings linear classifier to the 4 Boston classes

The neck follows Lin et al. [34], the head Jocher et al. [35] and the distributional box edges Li et al. [48]. A frame that carries no label for a head adds no loss through it, so a head is never trained on a guess.

Table S3: Station labels as shipped, with the coarser spaces used for reporting. Colon merges follow the published single-frame benchmark so accuracy is comparable to it; gastroscopy regions merge the four walls photographed at one protocol level. colonoscopy — CAS-Colon, ten classes, merged to six for comparison [17] index 0 1 2 3 4 5 6 7 8 9

label

segment

merge

termIleum cecum ascend hepFlex transv splFlex descend sigmoid rectum anal

terminal ileum caecum ascending colon hepatic flexure transverse colon splenic flexure descending colon sigmoid colon rectum anal canal

0 1 2 2 2 3 3 3 4 5

gastroscopy — GastroHUN, 22 sites plus OTHERCLASS, merged to seven regions [10] region 1 2 3 4 5 6 7

codes

protocol level

view

A1 L1 P1 G1 A2 L2 P2 G2 A3 L3 P3 G3 A4 L4 P4 G4 A5 L5 P5 A6 L6 P6 OTHERCLASS

antrum distal (lower) body upper-middle body fundus and cardia body incisura none of the 22

anterograde anterograde anterograde retroflex retroflex retroflex —

A letter names the wall and a digit names the protocol level: A anterior wall, L lesser curvature, P posterior wall, G greater curvature, as the dataset’s own figure defines them, following the systematic screening protocol for the stomach [58]. Greater curvature is photographed at four levels rather than six, giving 22 sites. Each one: A1 antrum, anterior wall; L1 antrum, lesser curvature; P1 antrum, posterior wall; G1 antrum, greater curvature; A2 lower body, anterior wall; L2 lower body, lesser curvature; P2 lower body, posterior wall; G2 lower body, greater curvature; A3 middle-upper body, anterior wall; L3 middle-upper body, lesser curvature; P3 middle-upper body, posterior wall; G3 middle-upper body, greater curvature; A4 fundus and cardia, anterior wall; L4 fundus and cardia, lesser curvature; P4 fundus and cardia, posterior wall; G4 fundus and cardia, greater curvature; A5 body in retroflexion, anterior wall; L5 body in retroflexion, lesser curvature; P5 body in retroflexion, posterior wall; A6 incisura, anterior wall; L6 incisura, lesser curvature; P6 incisura, posterior wall. Levels 1–3 are photographed anterograde and 4–6 in retroflexion. Regions here merge the walls at one level, which is the grouping the landmark-region figure is read in; a seventh group holds OTHERCLASS, which the dataset assigns when the intended site is not clearly visible or a lesion is present, and which the live program never ticks. Protocol levels are named from the dataset paper’s withdrawal sequence, which describes steps rather than naming six levels; its count of 12 anterograde and 10 retroflex photographs matches digits 1–3 and 4–6 exactly. Colon merges drop the benchmark’s seventh class, “indistinguishable”, which CAS-Colon has no counterpart for. Six-class merge: 0 terminal ileum, 1 caecum, 2 ascending with hepatic flexure and transverse, 3 splenic flexure with descending and sigmoid, 4 rectum, 5 anal canal.

woma: real-time endoscopy foundation model

Polyp detection and polyp type KUMC PolypsSet (26,593 frames, type per box) · REALColon studies 1–3 (27,859 polyp frames, 99,174 normal frames) · LDPolypVideo (5,875) · PolypDB (3,934) · five public still-image sets (1,855) never used for training: PolypGen (six centres), REALColon study 4, the KUMC held-out patients

26

The other heads colon segment: CAS-Colon, 63 videos, 26,186 frames, ten segments (15 videos held out) · polyp outline: 1,305 Kvasir-SEG and CVC-ClinicDB image–mask pairs · bowel preparation: Nerthus and HyperKvasir, 5,960 frames with a Boston score

1 Build the training mixture — one fixed random seed; duplicate images removed across sources; every video kept whole on one side of the split; polyps whose histology reads “no polyp” dropped; a frame without a box is an explicit “nothing here”; 165,290 training frames, 66,116 with boxes

2 Train all heads together on three GPUs — woma foundation model (pretrained), heads from random weights · each step trains one head, drawn by weight (detection 45 %, segment 30 %, outline 8 %, polyp type 10 %, preparation 7 %) · 20 epochs of 8,500 steps, 4 images per GPU per step · AdamW, foundation-model learning rate 10−4 , heads 10−3 , cosine decay · 640-pixel letterbox, mosaic, scale, shift, colour and flip augmentation · an exponential moving average of the weights is what gets evaluated and saved

3a Every epoch, on the run’s own validation data detection mAP, segment recall, outline Dice, polyp-type and preparation recall → curves on the dashboard only; never used to choose the model

3b Every second epoch a snapshot of every weight is written and scored on the held-out data below while the run continues

4 Held-out tests, per snapshot (data the training never saw) — PolypGen, six centres, 1,347 polyp frames and 193 clean frames: share of polyps found at precision ≥0.85 · REAL-Colon study 4, 15 full videos with 19 polyps: polyps found per procedure and false alarms per procedure · 15 held-out CAS-Colon videos: ten-segment recall, with and without an 8-second memory · HyperKvasir landmarks, another hospital · polyp type per lesion and per frame on held-out patients · outline Dice on three other-centre sets · preparation score agreement

5 The bars, fixed before the run, and the rule PolypGen ≥0.90 of polyps found at precision ≥0.85 · REAL-Colon ≥0.93 of polyps found at ≤2 false alarms per procedure · other-centre outline Dice ≥0.82 · segment recall ≥0.70 · polyp type ≥0.80 and preparation adequacy F1 ≥0.90 (both reported) rule: a snapshot must pass the two detection bars, then the one that meets the most bars wins; the training curve is not consulted → kept: epoch 6 (0.960 of PolypGen polyps found; 19 of 19 REAL-Colon polyps at 1.6 false alarms per procedure)

6 Ship — one model file (foundation model, five heads, class names, thresholds, a fingerprint of every source weight) · acceptance: the file reproduces the kept snapshot’s held-out record exactly · the live program runs it at 115 frames per second on the test workstation

Figure S1: How the colon model was developed: the data and what was kept out of training, the training settings, what is measured while the run trains, what is measured on held-out data for every snapshot, the bars, and the rule that chose the shipped weights. Sources: KUMC PolypsSet [37], REAL-Colon [15], LDPolypVideo [49], PolypDB [50], the five still-image sets [11, 22–24, 43], PolypGen [14], CAS-Colon [36], Nerthus [51], HyperKvasir [9]; optimiser [52], mosaic augmentation [53].

woma: real-time endoscopy foundation model

Sources that may ship (CC BY licences) GastroHUN: 22 stations plus “other”, 6,165 training frames, 793 validation and 803 test frames from other patients · HyperKvasir: oesophagitis, Barrett’s, normal landmarks; half of the landmark frames are the lesion gate’s clean frames · EndoWLI-NBI · EGID · GIM-ENDO

27

Sources for research use only EDD2020: the only boxes and outlines (Barrett’s, suspicious, high-grade dysplasia, cancer; 386 images) · UGIAD (its published test split kept as a test) · SAGE · GastroEndoNet · Kvasir v2 · GastroVision, split in half: one half trains, the other tests

1 Build the training set — seven findings classes (oesophagitis and Barrett’s merged into “oesophageal mucosal change”) · split by patient wherever a patient key exists, by image otherwise · a frame that has a box stays a detection frame; a frame that only names a finding never counts as “no lesion” for the detector · per-patient atrophy labels are placed only on frames the station head reads as the labelled region

2 Train, phase A: all heads together — the same foundation model, neck, heads and optimiser as the colon model; heads: lesion detection (one class), outline, station (22 + 1), findings (7) · 20 epochs of 1,500 steps on three GPUs; each step trains one head, drawn by weight (detection 50 %, station 25 %, findings 15 %, outline 10 %); rare findings drawn more often than their share Phase B: the findings head alone — retrained for two epochs on the frozen epoch-10 foundation model; the shipped model is that foundation model with its detector, outline and station heads and the retrained findings head

3a Every epoch, on the run’s own validation data station recall, detection mAP on a fixed card, outline Dice, findings recall → curves only

3b Every second epoch a snapshot of every weight, scored on the held-out data below while the run continues; assembled snapshots scored the same way

4 Held-out tests, per snapshot (data the training never saw) — station: 793 GastroHUN validation frames from other patients: landmark-region accuracy (7 regions) and 23-code recall; the 803-frame test split read once, for reporting · lesion gate: 39 EDD2020 frames with a neoplasia box against 1,347 lesion-free frames from another hospital: share of lesions flagged at the loosest threshold that keeps specificity ≥0.90 · findings: validation recall over seven classes; UGIAD test accuracy; GastroVision-half accuracy · outline: EDD2020 validation Dice

5 The bars, fixed before the run, and the rule tier 1: landmark-region accuracy ≥0.92 (23-code recall ≥0.85 reported) · tier 2: lesion gate ≥0.93 of lesions flagged at specificity ≥0.90; findings recall ≥0.85; UGIAD coarse accuracy ≥0.90 · tier 3: GastroVision-half accuracy ≥0.85; Dice reported rule: a snapshot must pass the lesion gate, then the one that meets the most bars wins, ties by how close it comes on the rest → kept: the assembled epoch-10 snapshot, five of five gated bars (region 0.923, gate 0.949 at specificity 0.912, findings 0.877, UGIAD 0.932, GastroVision 0.861)

6 Ship — one model file (foundation model, four heads, station and class names, thresholds, a fingerprint of every source weight) · acceptance: the file reproduces the kept snapshot’s held-out record exactly · the live program runs it at 98 frames per second on a 1080p clip; the lesion heads are research heads because EDD2020 is licensed for noncommercial use

Figure S2: How the upper-GI model was developed: the two licence classes of its data, the training set’s routing rules, the two-phase training, what is measured while training and on held-out data, the tiered bars, and the rule that chose the shipped weights. Sources: GastroHUN [10], HyperKvasir [9], EndoWLI-NBI [54], EGID [55], GIM-ENDO [56], EDD2020 [38, 57], UGIAD [19], SAGE [45], GastroEndoNet [46], Kvasir v2 [47], GastroVision [39].

woma: real-time endoscopy foundation model

28

Table S4: Lesion and finding labels as shipped, with the coarser space each is also scored in. colonoscopy index

label

meaning

held-out reporting

—

polyp

detection, one class

—

0 1 2 3

adenoma hyperplastic serrated other

adenomatous polyp hyperplastic polyp sessile serrated lesion any other histology

reported reported trained, not in the held-out split trained, not in the held-out split

gastroscopy index

label

meaning

UGIAD coarse-five

—

lesion

detection, one class

—

0 1 2 3 4 5 6

normal eso_mucosal neoplasia gastric_polyp gastric_inflammatory gastric_metaplasia varices

no finding oesophagitis or Barrett’s suspicious, high-grade dysplasia, cancer gastric polyp ulcer, inflammation, blood intestinal metaplasia or atrophy oesophageal or gastric varices

normal oesophageal mucosal change gastric lesion gastric lesion gastric lesion metaplasia varices

Polyp type is a partial-label task: a frame whose polyp has no histology appears in no class list and contributes detection loss alone, so the head is never taught a class the record does not support. Its reported figure is macro recall over adenoma and hyperplastic, because the held-out patient split carries no serrated or other frame at all; both classes stay in the output space, and the model does predict them on other sets. Training keeps oesophagitis and Barrett’s apart as separate classes and the shipped head merges them into eso_mucosal, because neither is what the detector is asked to flag and the two are not reliably separable on a single frame. Coarse-five is the space the UGIAD comparison is read in. Both detectors carry one class deliberately: box and description are separate answers, so a wrong description cannot suppress a correct box.

Table S5: Boston Bowel Preparation Scale, scored per colonic segment. score

label

segment as the published scale defines it

0

bbps0

1

bbps1

2

bbps2

3

bbps3

Unprepared segment, mucosa not seen because of solid stool that cannot be cleared. Part of the mucosa seen, other parts not, because of staining, residual stool or opaque liquid. Minor residual staining, small stool fragments or opaque liquid, mucosa seen well. Entire mucosa seen well, no residual staining, stool fragments or opaque liquid.

Adequacy, the figure reported against a pass mark, treats 2 and 3 as adequate and 0 and 1 as inadequate. Training labels are partial where the source is: Nerthus frames carry an exact score, HyperKvasir frames carry a grouped one — either {0,1} or {2,3} — and the grouped ones train through a partial-label loss that asks only for the group to be right, rather than guessing which of the two the annotator meant.

Record · ID 919489 · SHA-256 2d4bb29767d3d9e3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.