Conceptio › Archive › arXiv CS
arXiv CSopen access

From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic–Geometric Contracts Runze Li1,2 , Yukun Zhao3 , Can Xu4 , Yucheng Shen5 , Shuaiqiang Wang1 , Jianmin Wu1 , Lingyong Yan1∗ , Dawei Yin1

arXiv:2609.17326v1 [cs.AI] 15 Sep 2026

2

1 Baidu Inc., Beijing, China Harbin University of Science and Technology, Harbin, China 3 Shandong University, Jinan, China 4 East China Normal University, Shanghai, China 5 Soochow University, Suzhou, China

Abstract Scientific poster generation distills a multimodal paper into a single-page visual artifact, forcing strict trade-offs between informational coverage and readability under a fixed spatial budget. Existing methods pass plans as transient prompts and validate individual stages in isolation. This strategy causes requirements to drift across content and layout modules, and previous checks to be silently invalidated. We introduce PosterVisor, a control framework that shifts poster generation from transient prompts to persistent control. An Orchestrator grounds rubrics in the paper and visual assets, compiling them into a Semantic–Geometric Contract (SGC) that binds claims and sources to required visuals, budgets, and spatial commitments. Only fully instantiated records become executable assertions; other usable requirements remain soft guidance. Recursive Contract Enforcement (RCE) dynamically triggers checks across stages as evidence emerges. Crucially, during repairs, RCE rechecks affected checkpoint states, preventing repair-induced regressions from propagating silently. We instantiate PosterVisor in HTML/CSS and editable PPTX generators. On the 100-paper Paper2Poster benchmark, PosterVisor-PPT improves observed mean poster-grounded QA accuracy over PosterGen (64.47% vs. 58.53%) and is preferred by human judges in 72.5% of non-tied pairwise comparisons (95% CI, 61.6–83.4%). A secondary 30-paper study also yields higher VLM Overall and PaperQuiz means. These results support rubric-compiled contracts and stage-conditioned enforcement for controllable poster synthesis.

Introduction Scientific posters provide a compact visual medium for presenting research papers, distilling a full paper into a singlepage narrative (Qiang et al. 2019; Zhong et al. 2025; Tanaka, Wang, and Ushiku 2024). Automating this process goes beyond conventional text summarization: under a fixed spatial budget, a system must jointly select scientific claims, retain supporting evidence, and coordinate text, visuals, and layout to clearly convey the paper’s core message while keeping its visual evidence legible (Sun et al. 2026; Pang et al. 2025; Zhang et al. 2026; Choi et al. 2026). Recent LLM- and VLM-based systems frame scientific poster generation as a multi-stage process. They combine ∗

Corresponding author.

content–layout planning, specialized agents, hierarchical or editable representations, and visual feedback to coordinate content, layout, and revision (Pang et al. 2025; Sun et al. 2026; Zhang et al. 2026; Choi et al. 2026). More recent work improves efficiency and auditability through token compression, targeted violation detection, provenance tracking, and geometry checks (Tang et al. 2026; Yang et al. 2026). However, these pipelines still bind plans and checks to individual stages rather than maintaining a shared, paper-specific acceptance state throughout generation. This limitation creates two control challenges. First, semantic and spatial requirements are frequently weakened as a poster evolves from a textual plan to structured content and rendered outputs. Early plans merely specify what to generate, failing to bind claims, evidence, and spatial allocations into a persistent specification. Second, since requirements become observable at different stages, later modifications can silently invalidate earlier checks. Without tracking the scope and dependencies of each repair, the system cannot determine which results remain valid. To address these challenges, we introduce PosterVisor, an Orchestrator-centered framework that turns poster generation from transient prompts to persistent contracts. It turns paper-specific planning decisions into a Semantic–Geometric Contract (SGC) shared across the content generation, layout construction, and rendering stages. Starting from a fixed catalog of poster-quality criteria, the Orchestrator grounds relevant criteria in the source paper and its visual assets, binding narrative claims and source evidence to required visuals, content budgets, and coarse spatial commitments. Requirements with explicit targets and available validators are compiled into executable assertions, while other useful requirements are retained as soft guidance. Because these requirements become observable at different stages, Recursive Contract Enforcement (RCE) evaluates each executable assertion when the evidence needed to check them becomes available. When an assertion fails, RCE routes the violation to a scoped and bounded repair and then revalidates the assertions that may have been affected by the change. SGC therefore limits the drift of semantic and spatial commitments, while RCE prevents affected validation results from being silently treated as valid after repair. Together, SGC and RCE extend stage-local planning and checking into persistent, repair-aware control

throughout poster generation. We instantiate PosterVisor in two systems: a P2P-derived HTML/CSS generator and a PosterGen-derived editablePPTX generator. On the primary benchmark of 100 papers, both implementations obtain higher observed VLM Overall and Raw PaperQuiz means than their matched baselines and tie for the highest displayed automatic VLM Overall of 3.82. PosterVisor-PPT improves Raw PaperQuiz by 5.94 points and is preferred to PosterGen in 72.5% of non-tied human comparisons (95% CI, 61.6–83.4%). Across four additional model backbones, PosterVisor-HTML improves both objectives in all four settings, while PosterVisor-PPT improves each objective in three of the four settings. Archived execution logs further show that all 159 recorded deterministic repair events are followed by rule-based rechecking, providing direct evidence of the intended enforcement behavior. A secondary study on 30 papers provides additional evidence of transferability. Our contributions are summarized as follows: • We formulate scientific poster generation as a crossrepresentation control problem in which coupled semantic, evidential, and spatial commitments become observable at different stages and may be invalidated by subsequent repairs. • We introduce a rubric-grounded Semantic–Geometric Contract that preserves paper-specific commitments as a shared acceptance state, together with Recursive Contract Enforcement, which evaluates executable assertions at appropriate checkpoints and revalidates affected assertions after scoped and bounded repairs. • We instantiate the framework as HTML/CSS and editable PPTX generators, and demonstrate its effectiveness across two benchmarks in terms of output quality, cross-model robustness, human preference, and recorded control behavior.

Related Work Scientific Poster Generation. Early scientific-poster systems select and arrange content using learned statistics or neural components (Qiang et al. 2019; Xu and Wan 2022). Subsequent resources support poster generation, layout analysis, structural parsing, and summarization (Zhong et al. 2025; Tanaka, Wang, and Ushiku 2024; Tanaka, Hashimoto, and Ushiku 2026; Saxena, Minervini, and Keller 2025). Recent LLM/VLM pipelines use structured intermediate representations to transform parsed document content into editable visual structures for posters, slides, and webpages (Pang et al. 2025; Sun et al. 2026; Zhang et al. 2026; Choi et al. 2026; Sun et al. 2021; Zheng et al. 2025; Ma et al. 2025). A typical agentflow decomposes the task into paper parsing, content and figure selection, narrative planning, content writing, layout construction, rendering, and output review (Pang et al. 2025; Sun et al. 2026; Zhang et al. 2026; Choi et al. 2026; Tang et al. 2026). PosterHarness further advances contract-oriented control through placeholder-first planning, staged QA, and deterministic source-figure composition (Yang et al. 2026). Any2Poster and ResearchStudio-

Reel broaden the task to diverse inputs and editable multiartifact outputs (Vinaykumar et al. 2026; Xiao et al. 2026). Rubric-Guided Evaluation and Recursive Enforcement. Rubrics decompose open-ended quality into interpretable criteria for evaluation, diagnosis, and reward or verifier design; recent work also automates their synthesis and extraction (Hashemi et al. 2024; Sharma et al. 2025; Gunjal et al. 2025; He et al. 2026; Shao et al. 2025; Liu et al. 2026; Xie et al. 2025; Li et al. 2026). Prior poster-generation systems likewise employ fine-grained criteria, stage-specific checkers, rendered-image critics, deterministic layout tests, and rejection rules (Pang et al. 2025; Sun et al. 2026; Zhang et al. 2026; Choi et al. 2026; Tang et al. 2026; Yang et al. 2026). These mechanisms address local defects, but their criteria usually remain attached only to the current artifact at a particular stage. For example, a plan may require a key figure to be both included and readable. A content checker may confirm its inclusion, while an overflow checker may confirm that the layout fits; yet the figure may remain too small to convey its evidence. Enlarging it may then cause text overflow unless the affected constraints are rechecked. These mechanisms therefore provide useful local feedback but offer limited cross-stage persistence and repair-aware revalidation.

Problem Formulation Given a multimodal source paper P , a fixed poster-quality rubric catalog R, and a target output format f , we formulate controlled poster generation as (P, R, f ) 7−→ (X, Y, L), where X is an editable artifact, Y is its rendered poster, and L is an audit log. Requirements instantiated from P and R under format f combine scientific content, source evidence, and spatial presentation. The control problem is to maintain and evaluate this paper-specific requirement set across generated content, editable geometry, and rendered output, recognizing that requirements become verifiable at different stages and repairs may affect previously satisfied constraints.

Method In this section, we first summarize PosterVisor’s end-toend workflow and then describe its two core components: Semantic–Geometric Contract (SGC) construction and Recursive Contract Enforcement (RCE).

Overview PosterVisor first uses an Evidence Reader to transform the source paper P into indexed text T and a provenancepreserving visual inventory A. The Orchestrator uses (T, A) to determine the thesis, priorities, and spatial allocation. It grounds applicable criteria from R to freeze the SGC K, binding claims to required visuals, budgets, and spatial commitments. Guided by K, the Content Agent writes the poster content C. The Layout Agent combines the checked content, selected visuals, and spatial requirements into the editable artifact X, which a format-specific renderer converts into Y .

Evidence Reader Paper PDF

evidence

Quality Rubrics

Target Format

fixed dimensions

format rules

Orchestrator CONTRACT-GUIDED PRODUCTION LANE

Orchestrator: ground, compile & freeze

Semantic–Geometric Contract CONTENT CONTRACT

LAYOUT CONTRACT

Thesis & panels Claims & sources Required visuals Text budgets

Columns & reading order Visual roles & emphasis Size & coverage targets

EXECUTABLE ASSERTIONS

SOFT GUIDANCE

Targets & predicates Checkpoints & dependencies Allowed repairs

Grounded preferences without supported validators

Semantic–Geometric Contract (SGC)

violation scoped repair & affected-assertion recheck

RCE scheduler & router

Content Agent

Layout Agent

Outputs

Renderer

Editable Artifact Rendered Poster Audit Log

frozen SGC K

FEEDBACK LOOP Content C

Editable Structure X

Rendered Poster Y

CHECKPOINT LANE

Checker

checks each artifact against the shared plan and target-format rules

All stages read the frozen contract; repairs update artifacts

Recursive Contract Enforcement (RCE)

Figure 1: Overview of PosterVisor. The Evidence Reader extracts traceable paper evidence, and the Orchestrator grounds rubric dimensions to construct a persistent SGC. Content and layout modules generate artifacts under the frozen contract. RCE evaluates assertions as state becomes available, routes violations to bounded repairs, and rechecks affected assertions.

At the content, editable-layout, and rendered-output checkpoints, the Checker evaluates applicable assertions against the corresponding artifact state and returns structured violations. Under RCE, the Orchestrator routes each actionable violation to an allowed repair operator, and the Checker rechecks the assertions affected by the repair. We instantiate this shared SGC/RCE abstraction in a P2P-derived HTML/CSS generator and a PosterGen-derived editablePPTX generator. Figure 1 summarizes the workflow.

Semantic–Geometric Contract Construction The fixed catalog R defines the general quality dimensions considered in our experiments: scientific-content coverage, visual-evidence sufficiency, information density, visual readability, figure–text balance, and spatial balance. For a source paper P , the Orchestrator grounds the applicable dimensions in the indexed text T and visual inventory A and represents the resulting paper-specific requirements as an SGC, K = (Kc , Kl , E, S). Kc stores the poster thesis, ordered panels, claims or content intents, source anchors, required visual evidence, and text budgets. Kl stores panel-to-column assignments, reading order, spatial emphasis, visual roles, key-visual protection, and paper-specific size or coverage targets. E contains fully instantiated, checkable assertions, whereas S retains grounded records that can guide generation but are not executable. Each panel record thereby binds scientific intent and evidence to a content budget and a coarse spatial commitment. Only assertions in E determine verified contract compliance. Format-wide measurement rules, default tolerances, checkpoints, and admissible repair operators are specified separately by the format policy Θf . Planning first produces a candidate record set G. For each applicable rubric dimension, grounding yields zero or more

records gj = ⟨rj , qj , ηbj , pj ⟩, where rj identifies the criterion, qj is a typed target query, ηbj is a required value or policy key, and pj ⊆ T ∪ A records paper provenance. The Orchestrator also adds schema-derived records for required sections, evidence bindings, and budgets. Normalization resolves qj and ηbj against the draft contract and Θf , producing either gj′ = ⟨rj , τj , ηj , pj ⟩ or an explicitly unresolved record in G ′ . Each executable assertion has the form ei = ⟨ci , τi , ηi , ϕi , si , Di , Ui ⟩, where ci is a violation code, τi is a typed target, ηi is the expected value or bound, ϕi is the checking predicate, and si is its checkpoint. Di lists read dependencies, and Ui is a repair allow-list that may be empty. The format policy supplies a finite catalog of typed predicate templates. A policy key is not executable by itself: before a record enters E, its target, expectation, and checking rule must be materialized by the run’s contract state and format-policy configuration. Accordingly, for gj′ ∈ G ′ , the compiler emits ei = Cf (gj′ ; Kc , Kl , Θf ) ∈ E only when τi and ηi are resolved and a compatible template defines (ϕi , si , Di , Ui ). Otherwise, the record enters S only if it remains usable as generation guidance; unusable records are retained as unresolved in L. To illustrate this compilation mechanism, consider a rubric requiring “visual balance.” Grounding might resolve this abstract concept to extracting Figure 3 from the paper. During compilation, if the format policy Θf supports measuring image area, this becomes an executable assertion in E (e.g., “Figure 3 must occupy ≥ 15% of the column”). Conversely, if the system cannot reliably validate a specific aesthetic style, that requirement degrades to soft guidance in S, prompting

the generator without forcing a hard check. For a compiled assertion, ϕi (Z[Di ], τi ; ηi , Θf ) returns {pass, fail, ⊥}, where ⊥ denotes an unavailable observation or checker failure and is never treated as a pass. Indexed text, captions, visual references, asset geometry, and readability or importance estimates provide the evidence used to resolve record targets and expectations. The construction process is summarized as (T, A) = E(P ), e G) = Π(T, A; R), (K, e G; T, A, Θf ), (Kc , Kl , G ′ ) = N (K, (E, S) = Cf (G ′ ; Kc , Kl , Θf ). Here, E, Π, N , and Cf denote evidence extraction, paperspecific planning and grounding, normalization, and assere is the draft contract protion compilation, respectively; K duced during planning. Normalization resolves identifier, anchor, budget, and estimated-load conflicts before the final contract K = (Kc , Kl , E, S) is frozen. Any conflict that cannot be resolved remains a plan violation. After freezing, K remains unchanged, and state-specific projections of the same contract are supplied to generation, checking, and repair. Format-specific modules may instantiate supplementary artifact-specific checks from Θf for diagnosis or auxiliary repair, but these checks neither modify K nor expand E.

Recursive Contract Enforcement Recursive Contract Enforcement (RCE) evaluates executable assertions when their required state is available at checkpoint st . Content checkpoints expose sections, claims, evidence references, sources, and budgets. Layout checkpoints expose structural feasibility and realized geometry. Within Qst , executable assertions are evaluated either by deterministic validators for computable structural and geometric requirements or by task-scoped model predicates for properties that are difficult to fully formalize, such as plan coverage, figure–text consistency, and rendered content quality. Model outputs affect Qst only when they implement the predicate of an assertion already compiled into E; other critic findings remain diagnostic or may trigger auxiliary repair without affecting contract compliance. For an artifact state Zt ∈ {C, X, Y }, the Checker and repair router compute Vt ut Zt+1 Wt It Γt

= = = = = =

Qst (Zt ; Est (K), Θf ), H(vt , Zt ; K, Θf ) ∈ U(vt ), ρ(ut , Zt ; K, Θf ), Wdecl (ut ) ∪ diff(Zt , Zt+1 ), {ei ∈ E : Di ∩ Wt ̸= ∅}, clK,Θf {e(vt )} ∪ It ,

for an actionable violation vt ∈ Vt . Here, Qst returns violations among the assertions active at checkpoint st , and U(vt ) is the failed assertion’s nonempty repair allow-list. Wt combines the repair’s declared scope with the observed artifact difference. It identifies directly affected assertions, while Γt adds the failed assertion e(vt ) and closes this set under contract- and format-specific dependencies. If a reliable structural difference is unavailable, the Checker reruns all executable assertions over the modified artifact state.

The closure Γt additionally includes transitive geometry dependencies and global structural guards. When reliable field-level dependencies are unavailable, the full-checkpoint fallback conservatively includes all applicable assertions at the modified state. This dependency-aware rechecking is our key departure from single-pass self-correction, which may silently retain cascading errors. For example, after an image is shrunk to fix spatial overflow, the recheck scope includes image-text readability and adjacent alignment, allowing any resulting regression to be detected rather than silently accepted. After repair, the Checker reevaluates Γt , and a target is considered resolved only when its predicate passes. If an assertion changes from pass on Zt to fail or ⊥ on Zt+1 , RCE records a typed regression violation. The frozen contract K provides the comparison reference. Thus, the observeddelta closure makes regressions detectable, whereas contract immutability alone neither prevents nor repairs them. Repair types depend on the violation, not the checkpoint (e.g., layout overflow may trigger geometric tweaks or content shortening). RCE uses deterministic operators for directly measurable failures and scoped model calls when regeneration is required. It continues only while the relevant verification results improve and stops when all executable assertions applicable to the final artifact return pass, no permitted repair remains, verification no longer improves, or the format-specific repair budget is exhausted. Contract compliance is assigned only when all applicable assertions pass. Otherwise, the system returns the latest available artifact and records failed or unresolved assertions, exhausted repairs, and detected regressions in L. Regression detection is limited by the declared dependencies, observed differences, and underlying checkers; soft guidance and unmodeled properties remain outside the compliance claim.

Experiments Experimental Setup Datasets. Our primary benchmark is Paper2Poster (Pang et al. 2025), containing 100 source papers, author posters, and factual QA data. We additionally use a 30-paper secondary evaluation set (Zhang et al. 2026). Lacking PaperQuiz data for this set, we reproduce the Paper2Poster protocol to generate 50 verbatim and 50 interpretive questions per paper. All generators in the main comparison use GPT-4o, whereas the secondary comparison uses GPT-4.1. Baselines. We compare against eight automatic systems: PosterAgent (Pang et al. 2025), P2P (Sun et al. 2026), PosterGen (Zhang et al. 2026), PPTAgent (Zheng et al. 2025), PosterForest (Choi et al. 2026), Any2Poster (Vinaykumar et al. 2026), EfficientPosterGen (Tang et al. 2026), and PosterHarness (Yang et al. 2026). P2P and PosterGen are the respective base generators for PosterVisor-HTML and PosterVisor-PPT, enabling direct within-generator comparisons. Superscripts in Table 1 distinguish reported reference values from our reproductions. Evaluation Metrics. Following Paper2Poster (Pang et al. 2025), VLM-as-Judge uses GPT-4o to score six criteria from 1 to 5. Element quality, layout balance, and engagement

VLM-as-Judge (↑) Method

Aesthetic

PaperQuiz (↑)

Information

Overall

Elem. Layout Engage. Avg. Clarity Content Logic Avg.

Raw Accuracy

Density-Augmented

Verb. Interp. Overall Verb. Interp. Overall

Paper† GT Poster†

4.05 4.07

3.89 3.90

2.80 2.70

3.58 4.00 3.56 4.09

4.68 3.96

3.98 4.22 3.89 3.98

3.90 3.77

84.41 79.69 82.05 92.97 87.75 90.36 68.25 74.73 71.49 127.75 140.29 134.02

PosterAgent‡ PosterGen⋆ PPTAgent⋆ P2P⋆ PosterForest⋆ Any2Poster⋆ EfficientPosterGen⋆ PosterHarness⋆

3.92 3.59 2.76 3.99 3.96 3.10 3.18 3.42

3.05 2.98 2.83 3.50 3.56 2.87 2.99 3.40

2.98 2.81 1.93 3.01 3.01 2.39 2.21 3.02

3.32 3.13 2.51 3.50 3.51 2.79 2.79 3.28

3.98 4.14 3.80 3.99 3.90 3.76 4.09 4.29

3.97 3.93 2.68 3.92 3.91 3.69 3.96 4.14

3.55 3.32 2.31 3.83 3.70 3.29 3.43 4.14

3.83 3.80 2.93 3.91 3.84 3.58 3.83 4.19

3.58 3.46 2.72 3.71 3.67 3.18 3.31 3.74

57.81 76.39 43.90 73.15 22.24 60.09 55.67 76.60 51.05 80.42 49.91 79.89 55.08 76.71 55.50 76.68

PosterVisor-PPT 3.97 PosterVisor-HTML 3.73

3.76 3.45

3.09 2.96

3.61 4.45 3.38 4.62

4.05 3.92

3.59 4.03 4.25 4.26

3.82 3.82

53.02 75.91 64.47 106.04 151.83 128.93 58.23 75.40 66.82 115.95 150.21 133.08

67.10 58.53 41.17 66.14 65.73 64.90 65.89 66.09

111.40 147.29 129.34 87.71 146.19 116.95 44.48 120.19 82.33 111.19 153.03 132.11 101.99 160.74 131.37 99.59 159.45 129.52 109.98 153.18 131.58 110.94 153.28 132.11

Table 1: Results on the 100-paper Paper2Poster benchmark. † marks values reported by Paper2Poster, ‡ marks our LaTeX PosterAgent reproduction, and ⋆ marks our other reproductions. Paper and GT Poster are reference inputs excluded from automatic-method ranking. Bold and underline indicate the best and second-best automatic results; rounded ties share a mark. form Aesthetic; clarity, content completeness, and logical flow form Information; Overall averages all six. PaperQuiz measures source-information recoverability with 50 verbatim and 50 interpretive questions answered by GPT-4o, GPT-4omini, and o3. The Paper reference provides up to 20 rendered pages, whereas each poster method provides one image. Unsupported answers count as incorrect, and reader scores are averaged per sample. Density augmentation is computed before cross-paper averaging as sA = sR (1+1/ max(1, l/w)), where l is the OCR-text token count and w = 1335.5 is the median over re-evaluated author posters. These criteria are external output-quality measures rather than the generationtime rubric catalog R or direct tests of contract compliance. Implementation Details. Both realizations share the rubric catalog R and the SGC/RCE abstraction but use format-specific artifact states, validators, and repair operators. PosterVisor-HTML checks generated content, DOM structure, and Playwright-rendered geometry. It uses a 650word budget, balanced density, at least three required visuals, up to two content and DOM repair rounds, and three layout–render–check attempts. PosterVisor-PPT checks its panel blueprint, editable PPTX geometry, and LibreOfficerendered output and permits at most three repair–render rounds. Only executable validators determine contract compliance; if a repair budget is exhausted, the latest artifact is returned with its remaining violations recorded rather than being marked compliant. Each evaluated system produces one poster per paper without candidate selection; within each benchmark, newly evaluated outputs share the same judge and PaperQuiz protocol. Each run retains the frozen SGC, structured violations, applied repair patches, geometry

reports, unresolved assertions, and final artifacts for audit.

Main Results Table 1 shows higher observed means for both matched comparisons. PosterVisor-HTML raises VLM Overall from 3.707 to 3.822 and Raw PaperQuiz Overall from 66.14 to 66.82, while attaining the highest Density-Augmented Overall of 133.08. PosterVisor-PPT raises the corresponding VLM and Raw PaperQuiz scores from 3.460 to 3.818 and from 58.53 to 64.47. Both realizations round to the highest displayed automatic VLM Overall of 3.82. Finegrained results show complementary strengths: PosterVisor-HTML leads clarity, logical flow, and aggregate Information, whereas PosterVisor-PPT leads engagement and improves most strongly over its base in Aesthetic. Cross-model experiments with Qwen3.7-Plus, GPT-5.5, Gemini-3.1-Pro-Preview, and Claude-Opus-4.8 further test whether the control abstraction depends on GPT-4o. PosterVisor-HTML improves both VLM Overall and Raw PaperQuiz over P2P in all four settings, by 0.019–0.127 VLM points and 1.56–3.90 QA points. PosterVisor-PPT improves VLM Overall in three settings and Raw PaperQuiz in three settings, revealing stronger model-dependent content– presentation trade-offs. Full results appear in the supplement. Figure 2 illustrates missing evidence, excessive whitespace, text overload, and undersized visuals in representative baseline outputs. The final outputs use space more fully and foreground key visuals; mechanism evidence comes from the ablation configurations and repair logs below. On the 30-paper secondary set, PosterVisor-PPT raises PosterGen’s VLM Overall from 3.91 to 4.01 and Raw Pa-

Figure 2: Qualitative comparison for one Paper2Poster paper across the author poster, seven automatic baselines, and the two PosterVisor realizations.

Configuration

VLM-as-Judge (↑) Aes.

Info.

Raw PaperQuiz (↑)

All Verb. Interp.

Configuration

All

VLM QA Vis. Struct. Overall Overall demot. repair

P2P (base)

3.500 3.913

3.707 55.67

76.60 66.14

Full system

3.818

64.47

27

217

SGC + C/S + C/S/G Full

3.530 3.900 3.562 3.938 3.657 3.980 3.380 4.263

3.715 40.42 3.729 45.28 3.818 49.74 3.822 58.23

77.78 76.38 69.76 75.40

w/o Evidence Grounding w/o Visual Constraint w/o Geometry Verifier w/o Repair Guard

3.537 3.622 3.573 3.496

64.60 66.40 66.50 67.30

18 33 0 29

162 179 200 6

59.10 60.83 59.75 66.82

Table 2: HTML ablations on 100 papers. C, S, and G denote content-, structure-, and geometry-level RCE; “+” indicates cumulative addition to SGC. Each configuration is generated independently; the base and full-system results match Table 1.

perQuiz Overall from 78.9 to 85.5, ranking first among the evaluated automatic methods on both. Full results appear in the supplement.

Ablation Study We conduct format-specific ablations. For HTML, we compare combinations of SGC, content/structure RCE, and rendered-geometry RCE against the P2P base. For PPTX, we remove one implementation component from the full system at a time. Crucially, Evidence Grounding and the Repair Guard jointly implement the dependency closure (Γt ), acting as local constraints that uphold the global persistent contract (SGC). Specifically, Evidence Grounding supplies required visual dependencies, the Visual Constraint prevents incompatible large assets from sharing a column, the Geometry Verifier checks pre-render spatial feasibility, and the Repair Guard prevents local edits from hiding global structural failures. Table 2 shows non-monotonic trade-offs among the HTML ablations, with only the full system exceeding P2P on both aggregate objectives. In Table 3, removing any PPTX component lowers VLM Overall from 3.818 to 3.496–3.622. Removing the Geometry Verifier eliminates visual-demotion

Table 3: Diagnostic leave-one-component-out ablations for the PPTX realization. VLM and QA are Overall scores; event columns report totals over each run.

events, while removing the Repair Guard reduces structuralrepair events from 217 to 6, consistent with their intended control roles. The QA changes reveal a content–presentation trade-off rather than uniform gains. Complete switch definitions, additional PPTX configurations, and event counts appear in the supplement.

Analysis Failure-Mode Reduction. Figures 3 and 4 show complementary reductions in output-level failure signals: HTML overflow falls from 31% to 8% and missing required figures from 27% to 5%; PPTX severe text crowding drops from 18% to 0% and illegible figures from 14% to 2%. RCE Repair Activity. The archived HTML runs contain 159 deterministic local repairs, all followed by rule-based rechecking: 41 are followed by a subsequent full check, while 118 retain only an inline residual check. Regression analysis is restricted to the 41 repair stages with a subsequent full check. Comparing the same rule codes across each adjacent full-check pair yields 75 code-level assertion transitions, of which eight are pass-to-fail transitions across eight different posters. Seven appear only after model-based regeneration following a deterministic patch, while one is already present in the deterministic residual check. Crucially, RCE’s postrepair rechecking detects these observed regressions and

HTML Outputs

PPTX Outputs 31%

Content overflow

8% 27%

Missing required figures

5%

9% 18%

Reading-order violations

6% 0

10

P2P PosterVisor-HTML

20

0% 6%

Content gaps

1% 14%

Illegible figures

22%

Illegible figures

18%

Severe text crowding

2% 4%

Reading-flow disruptions 0

30

5

Case Violation and scoped repair

Recheck

H01 Three required visuals missing → restore Content and affected placeholders global checks pass H02 Required visual lacks two-column span Violations: → local layout repair and rerender 2→1→0 U01 Four-column contract, three-column out- Unresolved and put → two bounded local retries logged

Table 4: Selected auditable HTML repair trails. These examples illustrate RCE behavior rather than aggregate repair reliability.

routes them through the bounded repair loop. Five regressed codes are eliminated before finalization; the remaining three persist after the repair budget is exhausted and are explicitly recorded as residual violations. This directly demonstrates an operational advantage over single-pass checking: subject to its repair budget, RCE prevents observed regressions from propagating silently. Generation Overhead. PosterVisor-PPT increases calls from 8.05 to 19.15 and mean latency from 160.32 to 505.49 seconds per poster. Its decomposed requests nevertheless reduce total tokens from 104,251 to 67,066 and the token-only cost estimate from $0.297 to $0.228 per poster.

10

15

20

Failure-signal rate (%)

Failure rate (%)

Figure 3: Output-level failure-signal rates across 100 matched HTML cases: PosterVisor-HTML versus P2P.

PosterGen PosterVisor-PPT

0%

Figure 4: Output-level failure-signal rates across 100 matched PPTX cases: PosterVisor-PPT versus PosterGen. (a) Ranking over seven poster sources Method Top-1 Top-2 Top-3 Score (↑) Rank (↓) Author Poster

341

198

186

1605

3.18

P2P PosterGen PosterAgent PPTAgent

164 95 59 64

225 128 98 62

209 199 131 55

1151 740 504 371

3.73 4.30 4.68 4.93

PosterVisor-PPT PosterVisor-HTML

254 184

277 169

183 176

1499 1066

3.29 3.89

(b) Pairwise preference per pipeline Comparison O/B/T

Win % (95% CI)

PPT / PosterGen HTML / P2P

72.5 (61.6–83.4) 54.7 (45.6–63.8)

898/341/57 653/540/99

Table 5: Human evaluation: seven-source ranking (top) and matched-pipeline pairwise outcomes (bottom), with clusterrobust 95% CIs.

matic methods and wins 72.5% of non-tied comparisons with PosterGen (95% CI, 61.6–83.4%). PosterVisor-HTML secures more Top-1 selections than P2P but yields a 54.7% win rate (CI includes 50%). This statistical tie is expected, as P2P is already highly optimized for aesthetics. Crucially, PosterVisor achieves visual parity while significantly improving factual recoverability (Table 1). CIs are clustered by annotator and paper (Cameron, Gelbach, and Miller 2011).

Human Evaluation We recruit 15 annotators with AI backgrounds and graduatelevel research training: two doctoral researchers and 13 current master’s students or master’s degree holders. After standardized training, all annotators complete 1,175 blind sevensource ranking trials; 13 also complete 2,588 blind withinpipeline pairwise trials. In each ranking trial, annotators may select up to three poster sources or abstain at any position. Top-1, Top-2, and Top-3 selections receive 3/2/1 points, and unselected sources share the mean of the remaining ranks. Pairwise trials compare each PosterVisor realization with its matched base generator, and win rates exclude ties. Table 5 shows PosterVisor-PPT ranks first among auto-

Conclusion We presented PosterVisor, which coordinates scientific content, visual evidence, and geometry through a persistent SGC and RCE. On the primary benchmark, both realizations obtain higher observed VLM Overall and Raw PaperQuiz Overall than their matched baselines and tie for the highest displayed automatic VLM Overall of 3.82. PosterVisor-PPT additionally gains 5.94 Raw PaperQuiz points and receives 72.5% of non-tied preferences over PosterGen. These results support persistent contracts, scoped repair, and rechecking as practical control mechanisms for the two pipelines studied here.

References Cameron, A. C.; Gelbach, J. B.; and Miller, D. L. 2011. Robust Inference with Multiway Clustering. Journal of Business & Economic Statistics, 29(2): 238–249. Choi, J.; Park, S.; Song, S.; and Shim, H. 2026. PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 379–401. San Diego, California, United States: Association for Computational Linguistics. Gunjal, A.; Wang, A.; Lau, E.; Nath, V.; He, Y.; Liu, B.; and Hendryx, S. 2025. Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains. arXiv:2507.17746. Hashemi, H.; Eisner, J.; Rosset, C.; Van Durme, B.; and Kedzie, C. 2024. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13806–13834. Bangkok, Thailand: Association for Computational Linguistics. He, Y.; Li, W.; Zhang, H.; Li, S.; Mandyam, K.; Khosla, S.; Xiong, Y.; Wang, N.; Peng, X.; Li, B.; Bi, S.; Patil, S. G.; Qi, Q.; Feng, S.; Katz-Samuels, J.; Pang, R. Y.; Gonugondla, S. K.; Lang, H.; Yu, Y.; Qian, Y.; Fazel-Zarandi, M.; Yu, L.; Benhalloum, A.; Awadalla, H. H.; and Faruqui, M. 2026. AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 18003– 18022. San Diego, California, United States: Association for Computational Linguistics. Li, S.; Zhao, J.; Ren, H.; Wei, Z.; Zhou, Y.; Yang, J.; Liu, S.; Zhang, K.; and Wei, C. 2026. RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 31320–31344. San Diego, California, United States: Association for Computational Linguistics. Liu, T.; Xu, R.; Yu, T.; Hong, I.; Yang, C.; Zhao, T.; and Wang, H. 2026. OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 17417–17437. San Diego, California, United States: Association for Computational Linguistics. Ma, Q.; Wang, S.; Chen, Y.; Tang, Y.; Yang, Y.; Guo, C.; Gao, B.; Xing, Z.; Sun, Y.; and Zhang, Z. 2025. HumanAgent Collaborative Paper-to-Page Crafting for Under $0.1. arXiv:2510.19600. Pang, W.; Lin, K. Q.; Jian, X.; He, X.; and Torr, P. 2025. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers. In Belgrave, D.; Zhang, C.; Lin, H.-T.; Pascanu, R.; Koniusz, P.; Ghassemi, M.; and Chen, N., eds., Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc.

Qiang, Y.-T.; Fu, Y.-W.; Yu, X.; Guo, Y.-W.; Zhou, Z.-H.; and Sigal, L. 2019. Learning to Generate Posters of Scientific Papers by Probabilistic Graphical Models. Journal of Computer Science and Technology, 34(1): 155–169. Saxena, R.; Minervini, P.; and Keller, F. 2025. PosterSum: A Multimodal Benchmark for Scientific Poster Summarization. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, 1828–1844. Mumbai, India: The Asian Federation of Natural Language Processing and The Association for Computational Linguistics. Shao, R.; Asai, A.; Shen, S. Z.; Ivison, H.; Kishore, V.; Zhuo, J.; Zhao, X.; Park, M.; Finlayson, S. G.; Sontag, D.; Murray, T.; Min, S.; Dasigi, P.; Soldaini, L.; Brahman, F.; Yih, W.-t.; Wu, T.; Zettlemoyer, L.; Kim, Y.; Hajishirzi, H.; and Koh, P. W. 2025. DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research. arXiv:2511.19399. Sharma, M.; Zhang, C. B. C.; Bandi, C.; Wang, C.; Aich, A.; Nghiem, H.; Rabbani, T.; Htet, Y.; Jang, B.; Basu, S.; Balwani, A.; Peskoff, D.; Ayestaran, M.; Hendryx, S. M.; Kenstler, B.; and Liu, B. 2025. ResearchRubrics: A Benchmark of Prompts and Rubrics for Evaluating Deep Research Agents. arXiv:2511.07685. Sun, E.; Hou, Y.; Wang, D.; Zhang, Y.; and Wang, N. X. R. 2021. D2S: Document-to-Slide Generation Via Query-Based Text Summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 1405–1418. Online: Association for Computational Linguistics. Sun, T.; Pan, E.; Yang, Z.; Sui, K.; Shi, J.; Cheng, X.; Li, T.; Zhang, G.; Huang, W.; Yang, J.; and Li, Z. 2026. P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark. In The Fourteenth International Conference on Learning Representations. Tanaka, S.; Hashimoto, A.; and Ushiku, Y. 2026. SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2753–2762. Tanaka, S.; Wang, H.; and Ushiku, Y. 2024. SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters. In 35th British Machine Vision Conference 2024, BMVC 2024, Glasgow, UK, November 25–28, 2024. BMVA. Tang, W.; Xiao, J.; Gong, Y.; Ran, F.; Xia, T.; Liu, J.; Lam, M. H.; Wang, W.; and Lyu, M. R. 2026. EfficientPosterGen: Semantic-Aware Efficient Poster Generation via Token Compression and Accurate Violation Detection. arXiv:2603.00155. Vinaykumar, A.; Li, A.; Huang, S.; and Liu, S. 2026. Any2Poster: Any-Source Poster Generation Across Modalities and Domains. arXiv:2606.02915. Xiao, L.; Dai, Y.; Huang, Y.; Zhao, Q.; Wu, W.; He, H.; Chen, R.; Jiang, J.; Ma, Q.; Zhang, J.; Zhang, X.; Xin, Y.; Ou, Y.; Xia, Y.; Li, S.; Huang, L.; Zhang, Z.; He, Y.; Hui,

Y. K.; and Lu, Y. 2026. ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog. arXiv:2607.04438. Xie, L.; Huang, S.; Zhang, Z.; Zou, A.; Zhai, Y.; Ren, D.; Zhang, K.; Hu, H.; Liu, B.; Chen, H.; Liu, Z.; and Ding, B. 2025. Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling. arXiv:2510.17314. Xu, S.; and Wan, X. 2022. PosterBot: A System for Generating Posters of Scientific Papers with Neural Models. Proceedings of the AAAI Conference on Artificial Intelligence, 36(11): 13233–13235. Yang, T.; Fu, D.; Wu, Y.; Kou, Z.; Chen, L.; Jiang, R.; Wang, Z.; and Li, Q. 2026. PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark. arXiv:2607.03006. Zhang, Z.; Zhang, X.; Wei, J.; Xu, Y.; and You, C. 2026. PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation Via Multi-Agent LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 9813–9823. Zheng, H.; Guan, X.; Kong, H.; Zhang, W.; Zheng, J.; Zhou, W.; Lin, H.; Lu, Y.; Han, X.; and Sun, L. 2025. PPTAgent: Generating and Evaluating Presentations Beyond Text-toSlides. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 14402–14418. Suzhou, China: Association for Computational Linguistics. Zhong, X.; Tan, Z.; Li, J.; Gao, S.; Ma, J.; Feng, S.; and Chiu, B. 2025. Scientific Poster Generation: A New Dataset and Approach. Pattern Recognition, 164: 111507.

Supplementary Material for From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic–Geometric Contracts

1

Operationalizing the SGC and RCE

The two implementations specialize the frozen contract defined in the main paper through format-specific states, validators, and bounded repair operators.

1.1

Worked Archived SGC Instance

This section provides an implementation-level example of a frozen SGC and its checkpoint evidence. The archived HTML record binds a warm-start claim to a required source visual, its planned placement, executable survival checks, and readability guidance. The assignments and validator bindings are read directly from the saved contract and repair trail, without additional model inference or post-hoc reclassification. The initial content omitted the required visual. Deterministic repair restored its reference, and the residual recheck passed. The editable-HTML checkpoint then raised COLUMN_COUNT_DRIFT; a bounded layout retry restored the frozen column target, and the final checkpointwide recheck passed. The case therefore illustrates both executable constraints in E and soft preferences in S without treating every stored plan field as a compliance test.

2

Format-Specific Instantiations

The HTML GlobalPlan and PPTX blueprint become the format-specific frozen SGC representations only after normalization and assertion compilation have instantiated Kc , Kl , E, and S. Before that point, they are candidate plans or blueprints; the resulting complete contract is then frozen and serialized as K.

2.1

PosterVisor-HTML

The PosterVisor-HTML runs on the main 100-paper Paper2Poster benchmark (Pang et al. 2025) use GPT-4o for generation and visual analysis. The frozen plan uses a 650-word poster budget, balanced density, and at least three required visual assets. HTML Contract Schema The GlobalPlan links section identifiers, sources, budgets, and positions to visual roles, spans, size constraints, readability guidance, and grounded rubric records.

The planner grounds claim-oriented sections and visual roles in extracted text, captions, and image geometry. It distributes the text budget, prioritizes the key visual, and emits schema-valid JSON; dense tables may be represented as result cards or simplified tables. Pre-Render Content Verification The content checkpoint verifies required sections and visual placeholders, the 650word poster budget, and a per-section limit of 1.5 times the planned budget. Section-count drift is blocking only when a required section is missing; readability is deferred to rendered-output inspection. Deterministic repair restores planned order, required visual references, and budget compliance. A missing section triggers a targeted retry restricted to the implicated section, followed by the same content checks. Structure and Rendered-Geometry Verification The editable-layout checkpoint checks required image references, the key visual’s intended span, and COLUMN_COUNT_ DRIFT. DOM repair can restore references or span metadata; column drift triggers a bounded layout retry and the same structural recheck. Playwright then measures rendered bounding boxes. Blocking assertions cover overlap above 2% of the smaller element, overflow beyond 8 pixels, aspect-ratio distortion above 10%, text columns below 180 pixels or 12% of poster width, within-row width ratios above 4:1, content-width coverage below 70%, section blank ratios above 35%, and figure containers below 15% of their column width. Supported figure-size predicates are enforced only when compiled into E; other readability preferences remain in S. Content and editable-layout verification each permit up to two repair rounds; the subsequent layout–render–check loop permits at most three attempts along the same trajectory. A retained artifact may be returned after budget exhaustion, but any unresolved executable assertion keeps compliance false.

2.2

PosterVisor-PPT

The PPTX implementation normalizes the candidate blueprint before freezing and then verifies content, editable geometry, typography, and rendered output. Blueprint Contract The Evidence Reader supplies indexed paper sections and a scored visual inventory. The Orchestrator uses the highest-ranked candidates to construct a

Grounded requirement

Saved contract realization

Operational class

Validator and archived evidence

Warm-start claim and Kc : claim-bearing section and required Kc /Kl contract fields; the Fields are retained in the frozen source visual source visual; Kl : assigned column and linked requirement is exe- contract and supplied to content and multi-column span cutable only with a validator layout generation binding Required visual vives writing

sur- The required visual must appear as a E: executable content assertion MISSING_FIGURE_ PLACEHOLDER; failed bevalid numeric Markdown placeholder fore repair and passed after after content generation deterministic reference restoration

Required visual vives layout

sur- Editable HTML must retain the cor- E: executable layout assertion responding image reference before source-asset embedding

Readable, logically Keep trends and annotations visible; S: soft guidance placed visual place figures logically to support the storyline

MISSING_VISUAL_REF; passed at the archived layout checkpoint Retained as generation guidance unless compilation produces a supported measurable predicate

Table S1: One archived paper-specific SGC instance. The E/S assignments and their validator bindings are read from the saved frozen-contract record; checkpoint artifacts provide the observed pass/fail evidence.

five-to-eight-panel blueprint with grounded claims, sources, budgets, and visual bindings. Before normalization and assertion compilation, the planner output is only a candidate blueprint. The finalized blueprint binds panel claims, evidence, budgets, and layout priorities to the compiled E/S records and is then frozen as the PPTX-specific SGC. During normalization, the Orchestrator designates exactly one high-importance anchor panel, places it at the top of the middle column, and binds it to the key visual. The PPTX format policy clamps per-panel budgets to 50–140 words and normalizes them toward a poster-wide target of approximately 400 body words. Before the SGC is frozen, required visual and source identifiers are resolved against the supplied inventories; unresolved identifier or anchor conflicts remain plan violations. Adaptive Global Geometry During SGC normalization, the Orchestrator selects a near-equal global column profile from panel roles, estimated load, and visual evidence. The resulting allocations instantiate Kl , and panel budgets in Kc are rescaled by the assigned column width before freezing. The Geometry Verifier estimates column coverage from text budgets, typography, padding, and visual aspect ratios, using an accepted band of 0.75–1.10. For an overfull candidate, bounded normalization may move one nonanchor panel, replace one non-protected visual binding with grounded textual evidence, or split one long text panel. Feasibility is recomputed after each action; unresolved conflicts remain plan violations. During the same normalization stage, the PPTX format policy imposes a deterministic large-visual allocation constraint: no column may contain two large figures or a large figure together with a large table. For this constraint, a figure is considered large when its estimated height exceeds 45% of the column height, and a table is considered large when its estimated height exceeds 35%. The key visual and any visual bound to the anchor panel cannot be converted into text-only evidence requirements. If the allowed candidate-

contract normalization actions cannot resolve the allocation conflict, it remains a plan violation. Plan/Content Verification and Repair After the Content Agent produces the poster content C, the content checkpoint compares the storyboard with the frozen SGC blueprint. Deterministic predicates check panel presence (SECTION_PRESENT), title identity (SECTION_TITLE), required visual evidence (REQUIRED_FIGURE), and panel and poster-wide word budgets. Two task-scoped model predicates evaluate the semantic records already compiled into E: UNCOVERED_CLAIM tests coverage of each panel’s main_claim and must_mention facts, while OUT_OF_SOURCE_EVIDENCE tests whether the stated evidence is consistent with the panel’s frozen source_sections. Both codes are blocking executable assertions. A model-call failure returns ⊥ and leaves the assertion unresolved rather than passing it. Deterministic repairs restore missing panels or titles, insert required figures, and truncate text to the applicable budget. The two semantic violation codes permit only a scoped_content_patch: the implicated panel is rewritten to cover its frozen claim and required facts using only its allowed source sections. Each applied patch is followed by the same content checks. The format policy allows up to three scoped content-repair calls and stops when verification no longer improves. Layout Verification and Typed Repair At the editablelayout checkpoint before rendering, the Checker evaluates executable assertions in E for page coverage, per-column coverage, column imbalance, bottom whitespace, key-visual and main-result size, aspect-ratio distortion, and overlap. In the reported PPTX format policy, the instantiated blocking bounds are 0.82–0.94 for page coverage, 0.78–0.96 for column coverage, at most 0.18 for column imbalance, and at most 0.12 for the bottom-blank ratio. After typography is applied, the post-typography layout checkpoint remeasures text bounds and collisions and enforces a minimum body-

font size of 36 points. Failed assertions are routed through bounded RCE and must pass before the artifact can be certified contract-compliant. At the rendered-output checkpoint, the Checker applies full-poster task-scoped VLM visual-QA together with rulefirst panel- and column-level verification. The rule-based path recomputes deterministic geometry from the rendered styled layout, while task-scoped VLM predicates evaluate already compiled assertions in E, including applicable figure– text consistency and numerical- evidence requirements. A model finding affects contract compliance only when it evaluates an assertion in E; other critic observations remain diagnostic. The configured severity threshold determines which actionable violations are routed to repair, but it does not convert a failed executable assertion into a pass. A model-call failure returns ⊥ and is never treated as verified compliance. At rendered-output RCE, the repair router selects a typed operator from Ui , such as resizing or moving a visual within its column, adjusting text, adding compact evidence, or dropping a non-key visual. The dependency closure is implemented jointly by Evidence Grounding and the Repair Guard: Evidence Grounding supplies required-visual dependencies, while the Repair Guard enforces the resulting repair scope by validating targets and parameters, preserving the canvas and frozen column assignments, and preventing removal of the key visual. Format-wide readability bounds remain enforced, including minimum scales of 0.65 for ordinary figures and 0.75 for tables and the key visual. Up to three repair–rerender–recheck iterations are allowed; exhausted budgets return the latest artifact with unresolved assertions recorded in L and compliance kept false.

3

Mechanism Validation

This section audits the archived generation-time logs and artifacts from the full PosterVisor-HTML and PosterVisor-PPT GPT-4o runs underlying their rows in the main 100paper comparison. It invokes no additional model inference and uses no evaluation-model logs.

3.1

HTML Recheck and Regression Audit

At the immediate post-repair check, all eight detected passto-fail regression transitions were recorded as fail rather than ⊥. Five were subsequently resolved before finalization, whereas three remained as residual violations after the repair budget was exhausted.

3.2

PPTX Mechanism Audit

SGC construction boundary. The audit observes 21 candidate-blueprint geometry mutations across 18 posters. They occur before the blueprint is frozen and are therefore SGC construction, not post-freeze RCE repair. Audit recheck coverage is 1.0 for content, editable-layout, and post-typography repair records. Recheck coverage for the rendered-output critic loop is also 1.0 (29/29 critic-loop repair trajectories): all 29 posters whose critic loop applied repairs completed the required post-repair recheck. Each trajectory completed violation detection, repair writing, repair

Control stage Candidate-blueprint normalization Content-checkpoint RCE Editable-layout RCE Post-typography RCE Rendered-output RCE Total

Events Mechanism role 21 Pre-freeze SGC construction 153 Content repair and recheck 121 Geometry repair and recheck 24 Styled-geometry repair and recheck 103 Guarded repair, rerendering, and recheck 422

Table S2: PPTX control events observed in the mainbenchmark GPT-4o PosterVisor-PPT runs over 100 papers. Counts are per-event repair records, not violation counts or state transitions; one record may bundle multiple violation codes. The 21 pre-freeze blueprint-geometry mutations are excluded from the 401 post-freeze records by definition. application, rerendering, and full-poster rechecking. The denominator is poster-level critic-loop repair trajectories, not the 103 rendered-output event records in Table S2; rechecking is performed over the full poster state rather than paired to individual violation codes. Thus, no repaired trajectory inherits a preceding check result without revalidation. The reported repair counts therefore use distinct scopes. The main paper’s 159 events are deterministic local repairs from the HTML pipeline; Table S2 counts 422 PPTX control records across checkpoints, including 21 pre-freeze mutations; and the 217 value in main-paper Table 3 is a PPTX categoryspecific structural-repair counter. These quantities do not share a denominator and should not be compared as alternative totals.

4

Ablation Study and Visual Evidence

For the main 100-paper GPT-4o benchmark, this section reports the complete PPTX progressive study and an expanded decomposition of the leave-one-component-out results summarized in Table 3 of the main paper. The quantitative HTML ablation table is not repeated here.

4.1

PPTX Progressive Configurations

The progressive study begins with the PosterGen baseline (Zhang et al. 2026). L2 adds only a blueprint scaffold, not the finalized blueprint used as the SGC. L3 adds sourcegrounded claims, evidence anchors, and required-visual dependencies; after normalization and assertion compilation, the resulting blueprint is frozen as the grounded SGC. L4 activates the pre-render Visual Constraint, Geometry Verifier, and Repair Guard, and L5 adds rendered-output RCE. Tables S3 and S4 report the switches and results.

4.2

Same-Paper Visual Ablations

Figure S1 shows same-paper outputs for the cumulative HTML configurations and diagnostic PPTX leave-one-out configurations.

Key figure too small Dense figure unreadable

FULL Unused space

P2P

SGC

Secondary plots oversized

Unused space

bounded repair

+C/S RCE

+C/S/G RCE

(a) Cumulative HTML configurations

FULL

Unbalanced layout

Crowded visuals

w/o Repair Guard

w/o Visual Constraint

Background retained Core evidence missing

Unbalanced columns

w/o Geometry Verifier

w/o Evidence Grounding

(b) Diagnostic PPTX leave-one-out configurations Figure S1: Same-paper visual ablations. In (a), C, S, and G denote content-, structure-, and rendered-geometry-level RCE, respectively. The +C/S/G configuration adds all three RCE levels to SGC; Full additionally enables visual analysis, readabilityaware figure-size enforcement, and bounded layout–render–check retries with layout compaction, all performed under the frozen SGC. Callouts mark undersized or unreadable figures, unused space, crowded layouts, column imbalance, and misplaced visual priority across the HTML and PPTX configurations.

5 5.1

Extended Effectiveness

Cross-Model Robustness

Table S6 reports the six systems under four additional model settings. Within each setting, the named model handles every model-based call in the generation and internal-control

pipeline, including figure scoring and visual analysis, Orchestrator and blueprint planning, Content Agent and section writing, internal semantic and rendered-poster critics, and Repair Writer. Thus, the Qwen3.7-Plus setting uses Qwen3.7-Plus throughout this pipeline, and the same rule

Layer

Added

EG VC GV RG RR

L1 PosterGen baseline – L2 blueprint scaffold scaffold × L3 grounded SGC EG ✓ L4 pre-render VC/GV/RG ✓ L5 full rendered RCE ✓

– × × ✓ ✓

– × × ✓ ✓

– × × ✓ ✓

– × × × ✓

Table S3: Switches in the progressive PosterVisor-PPT study. EG: Evidence Grounding; VC: Visual Constraint; GV: Geometry Verifier, including pre-render layout verification; RG: Repair Guard; RR: rendered-output RCE.

GPT-4.1 generates the automatic posters, GPT-4o provides all VLM-as-Judge scores, and GPT-4o, GPT-4o-mini, and o3 form the image-only PaperQuiz ensemble. All methods share the frozen questions, prompts, scoring, and aggregation; density augmentation uses the main-paper formula with the transfer benchmark’s separately computed GT-Poster reference median w = 1335.0. The 100-paper primary and crossmodel experiments instead use their benchmark-specific median w = 1335.5.

6 6.1

applies to GPT-5.5, Gemini-3.1-Pro-Preview, and ClaudeOpus-4.8. The six-system subset was fixed when these experiments were run; baselines added later to the main comparison were omitted because their public implementations were unavailable or not reproducible at that time. GPT-4o does not participate in poster generation, internal checking, or repair. It is used only during external evaluation, both as the VLM-as-Judge and as one member of the common PaperQuiz reader ensemble.

5.2

Failure-Signal Audit

For the HTML comparison in Figure 3 of the main paper, each percentage is computed over the final rendered posters. The four categories capture output-level signals of content overflow, missing required figures, illegible figures, and reading-order disruption. These signals are first identified from poster-level VLM judgments and then manually reviewed against the final rendered posters; the reported percentages use the reviewed labels. The categories are not mutually exclusive, and one poster may contribute to multiple categories. These measurements characterize observable final-output failures rather than internal checkpoint activity or contract-compliance status. For the PPTX comparison in Figure 4 of the main paper, we apply fixed rules to poster-level VLM-as-Judge outputs for paired PosterGen and PosterVisor-PPT posters. Severe text crowding requires an engagement score of at most 2 and an explanation explicitly identifying dense or excessive text. A content gap requires a content-completeness score of at most 3 and an explanation identifying missing or underdeveloped content. An illegible figure requires a relevant score of at most 3 and an explanation identifying a small or unreadable visual. A reading-flow disruption is counted only when the explanation explicitly identifies a broken order or narrative flow. Categories are not mutually exclusive. All ruleflagged signals are manually reviewed against the final rendered posters, and the reported percentages use the reviewed labels. These measurements capture VLM-observable output signals and are not direct evaluations of executable assertions in E.

5.3

PosterGen Transfer Set

We evaluate transfer on the 30-paper set provided by the PosterGen authors. Because it lacks PaperQuiz questions, we generate 50 verbatim and 50 interpretive questions per paper once and freeze them before comparing methods.

Evaluation and Human-Study Protocols Automatic Evaluation Protocol

The primary protocol uses the 100-paper Paper2Poster benchmark; cross-model and transfer exceptions are defined in Sections 5.1 and 5.3. Each system produces one poster per paper without independent candidate selection. HTML and PPTX outputs are rendered with Playwright and LibreOffice, respectively, and all automatic metrics evaluate only the final poster image. Automatic metrics. For external VLM-as-Judge evaluation, all comparisons use GPT-4o with the benchmark’s six criteria and five-level anchors. PaperQuiz uses 50 frozen verbatim and 50 frozen interpretive questions per paper, an image-only prompt, common scoring, and the same ensemble of GPT-4o, GPT-4o-mini, and o3. Scores are averaged across readers per paper and then across the benchmark; missing, invalid, or unsupported answers are incorrect. For the 100-paper primary and cross-model experiments, density augmentation follows the main-paper formula with the fixed GT-Poster median w = 1335.5; the 30-paper transfer exception is defined in Section 5.3.

6.2

Human Evaluation Protocol

Fifteen trained annotators with AI research backgrounds completed the seven-way ranking task, and 13 also completed the paired comparisons. Method identities were concealed; poster order and pairwise left–right placement were randomized. The seven-source subset was fixed when the study was run; baselines added later to the main comparison were omitted because their public implementations were unavailable or not reproducible at that time. Seven-way ranking. For mean-rank calculation in the seven-way task, unselected systems share the average of the remaining ranks; when three systems are selected, each of the other four receives rank 5.5. Across 1,175 trials, annotators provide 1,161 top-1, 1,157 top-2, and 1,139 top-3 selections, with the differences arising from permitted abstentions. Restricting the sensitivity analysis to the 11 annotators who completed all 100 papers preserves the direction of every paired mean-rank difference. Within-pipeline pairs. The study contains 1,296 PPTX and 1,292 HTML comparisons. Because annotators and papers are repeatedly observed, primary 95% confidence intervals use two-way cluster-robust standard errors (Cameron, Gelbach, and Miller 2011). A 20,000-replicate crossedcluster bootstrap provides sensitivity intervals, and Holm correction is applied across the two tests.

VLM-as-Judge (↑)

Configuration L1 PosterGen L2 blueprint scaffold L3 + grounded SGC L4 + pre-render controls L5 full PosterVisor-PPT

Raw PaperQuiz Accuracy (↑)

Aesth.

Info.

Overall

Verbatim

Interpretive

Overall

3.130 3.040 3.060 3.100 3.607

3.800 4.100 4.030 4.050 4.030

3.460 3.570 3.540 3.580 3.818

43.90 45.17 48.91 47.09 53.02

73.15 76.10 75.49 74.70 75.91

58.53 60.64 62.20 60.89 64.47

Table S4: Complete progressive PosterVisor-PPT results on the main GPT-4o benchmark. Because configurations are independently generated, adjacent rows describe configuration-level transitions rather than isolated marginal effects. (a) Leave-one-out quality and information retention Configuration

VLM-as-Judge

(b) Recorded mechanism events

Raw PaperQuiz Accuracy

Aesth. Info. Overall Verbatim Interpretive Overall full

3.607 4.030 3.818

53.02

75.91

64.47

w/o Visual Constraint w/o Geometry Verifier w/o Evidence Grounding w/o Repair Guard

3.080 3.057 3.057 3.016

60.80 61.50 58.80 62.40

72.00 71.50 70.30 72.20

66.40 66.50 64.60 67.30

4.163 4.090 4.017 3.975

3.622 3.573 3.537 3.496

Config.

Active geom. Vis. Struct. reports demot. repairs

Leave-one-out configurations full 100 w/o GV 0 w/o VC 100 w/o RG 100 w/o EG 100

27 0 33 29 18

217 200 179 6 162

Progressive-stage controls L2 blueprint scaffold 0 L3 grounded SGC 0 L4 pre-render controls 100

0 0 29

0 0 112

Table S5: Diagnostic PosterVisor-PPT leave-one-component-out configurations and recorded control events. VLM quality and Raw PaperQuiz Accuracy are aggregate results for each configuration. Whenever the geometry-gate node is instantiated, it emits a report file; when GV is disabled in the leave-one-out ablation, that file is a status="disabled" stub with no actions. Active geometry reports therefore count only status="ok" records, rather than files. In the focal full versus w/o GV comparison, the active gate produced 21 geometry mutations across 18 posters, versus zero mutations when GV was disabled. Visual demotions and structural repairs are event totals rather than poster counts or repair success rates. EG, VC, GV, and RG follow the component names defined in Table S3.

VLM-as-Judge (↑) Method

Aesthetic

PaperQuiz (↑)

Information

Overall

Elem. Layout Engage. Avg. Clarity Content Logic Avg. Qwen3.7-Plus GT Poster

Raw PaperQuiz Accuracy

Density-Augmented PaperQuiz

Verbatim Interpretive Overall Verbatim Interpretive Overall

4.07

3.90

2.70

3.56 4.09

3.96

3.89 3.98

3.77

68.25

74.73

71.49

127.75

140.29

134.02

PosterAgent 2.92 PosterGen 3.53 PPTAgent 3.75 P2P 2.90 PosterVisor-PPT 3.72 PosterVisor-HTML 3.20

3.01 3.62 3.45 3.00 3.89 3.37

1.76 2.38 2.43 1.93 2.87 2.14

2.56 3.18 3.21 2.61 3.49 2.90

2.07 4.71 4.61 4.35 4.66 4.53

3.89 3.90 4.02 4.18 4.09 3.87

4.90 3.62 4.94 4.52 4.91 4.51 4.98 4.50 4.97 4.57 4.99 4.46

3.09 3.85 3.86 3.56 4.03 3.68

60.53 53.53 63.48 66.03 55.64 71.83

72.91 72.97 76.45 76.31 74.87 77.12

66.72 63.25 69.97 71.17 65.26 74.47

114.21 107.07 126.96 131.11 110.79 143.18

137.90 145.93 152.91 151.56 149.08 153.77

126.06 126.50 139.93 141.33 129.94 148.47

GPT-5.5 GT Poster

4.07

3.90

2.70

3.56 4.09

3.96

3.89 3.98

3.77

68.25

74.73

71.49

127.75

140.29

134.02

PosterAgent 2.51 PosterGen 3.06 PPTAgent 2.94 P2P 2.63 PosterVisor-PPT 3.13 PosterVisor-HTML 2.85

2.37 2.99 3.12 2.96 2.99 3.02

2.06 2.62 2.92 2.01 2.63 2.20

2.31 2.89 2.99 2.53 2.92 2.69

2.19 3.75 3.52 3.44 3.58 3.48

3.48 3.29 2.96 3.99 3.45 3.78

3.90 3.19 3.89 3.64 3.85 3.44 4.00 3.81 3.89 3.64 3.97 3.74

2.75 3.27 3.22 3.17 3.28 3.22

67.75 56.45 48.75 73.75 60.02 75.60

75.54 74.21 75.55 77.15 75.58 79.19

71.65 65.33 62.15 75.45 67.80 77.39

125.69 112.80 97.49 124.82 119.48 149.11

140.68 148.25 151.11 130.43 150.42 156.20

133.19 130.52 124.30 127.62 134.95 152.66

Gemini-3.1-Pro-Preview GT Poster 4.07

3.90

2.70

3.56 4.09

3.96

3.89 3.98

3.77

68.25

74.73

71.49

127.75

140.29

134.02

PosterAgent 2.66 PosterGen 2.62 PPTAgent 2.85 P2P 2.16 PosterVisor-PPT 2.72 PosterVisor-HTML 2.40

2.46 2.56 2.64 2.26 2.42 2.73

2.53 2.58 2.17 2.09 2.55 2.29

2.55 2.59 2.55 2.17 2.56 2.47

2.29 3.73 2.88 3.62 3.74 3.87

3.19 3.40 2.08 3.57 3.24 3.09

4.45 3.31 4.10 3.74 4.23 3.06 4.35 3.85 4.02 3.67 4.18 3.71

2.93 3.17 2.81 3.01 3.12 3.09

65.07 55.67 32.66 55.96 71.11 63.03

76.17 74.61 73.14 75.06 78.12 75.79

70.62 65.14 52.90 65.51 74.62 69.41

123.50 111.23 65.32 110.17 142.15 125.99

144.95 149.09 146.28 147.80 156.17 151.53

134.22 130.16 105.80 128.99 149.16 138.76

Claude-Opus-4.8 GT Poster

4.07

3.90

2.70

3.56 4.09

3.96

3.89 3.98

3.77

68.25

74.73

71.49

127.75

140.29

134.02

PosterAgent 2.48 PosterGen 2.80 PPTAgent 2.91 P2P 2.50 PosterVisor-PPT 2.70 PosterVisor-HTML 2.44

2.86 2.97 3.43 3.20 3.12 3.17

2.20 2.59 3.00 2.58 2.81 2.63

2.51 2.79 3.11 2.76 2.88 2.75

3.69 3.75 3.60 3.78 3.90 3.90

3.98 3.51 3.98 3.89 4.00 3.87 3.85 3.88 3.99 3.96 3.91 3.92

3.01 3.34 3.49 3.32 3.42 3.34

70.49 65.05 58.06 72.68 63.58 74.79

77.50 76.33 75.98 77.14 76.51 78.14

74.00 70.69 67.02 74.91 70.04 76.47

130.27 130.05 116.12 128.39 126.72 146.77

143.46 152.59 151.96 136.66 152.48 153.35

136.87 141.32 134.04 132.53 139.60 150.06

2.86 3.93 4.01 4.00 4.00 3.96

Table S6: Cross-model robustness results. In each setting, the named model is used for every model-based call in the generation and internal-control pipeline. GPT-4o does not participate in poster generation, internal checking, or repair; during external evaluation, it serves as the VLM-as-Judge and as one member of the common PaperQuiz reader ensemble. The complete GT Poster row reuses the main-comparison reference values. Bold and underlined values denote the best and second-best automatic methods within each setting; rounded ties share the same mark.

VLM-as-Judge (↑) Method

Aesthetic

PaperQuiz (↑)

Information

Overall

Elem. Layout Engage. Avg. Clarity Content Logic Avg.

Raw PaperQuiz Accuracy

Density-Augmented PaperQuiz

Verbatim Interpretive Overall Verbatim Interpretive Overall

GT Poster

3.97

3.73

3.10

3.60 4.40

4.00

4.37 4.26

3.93

71.27

87.93

79.60

135.67

167.80

151.73

PosterAgent PosterGen PPTAgent P2P PosterForest Any2Poster EfficientPosterGen PosterHarness

4.03 4.03 2.07 3.73 4.00 2.00 3.90 3.07

3.70 3.60 3.40 3.37 3.33 2.33 3.57 2.77

3.13 3.07 2.53 3.07 3.03 2.33 2.93 2.27

3.62 3.57 2.67 3.39 3.46 2.22 3.47 2.70

3.90 4.53 4.33 4.50 4.57 4.03 4.37 4.07

4.00 4.00 3.07 4.03 4.00 3.43 3.80 4.33

4.23 4.23 3.67 4.23 4.13 3.53 4.10 3.30

4.04 4.26 3.69 4.26 4.23 3.67 4.09 3.90

3.83 3.91 3.18 3.82 3.84 2.94 3.78 3.30

70.07 65.60 24.80 67.60 70.67 53.53 59.33 66.67

90.00 92.27 46.87 87.73 89.20 81.40 82.33 87.73

80.03 78.93 35.83 77.67 79.93 67.47 70.83 77.20

134.12 130.44 49.60 133.07 135.56 107.07 116.39 133.33

172.36 183.45 93.73 172.81 171.00 162.80 161.75 175.47

153.24 156.94 71.67 152.94 153.28 134.93 139.07 154.40

PosterVisor-PPT 4.00 PosterVisor-HTML 3.90

3.83 3.23

3.07 3.13

3.63 4.87 3.42 4.53

4.00 4.10

4.30 4.39 4.23 4.29

4.01 3.86

82.60 71.00

88.40 91.73

85.50 81.37

164.23 137.33

175.88 177.11

170.06 157.22

Table S7: Transfer results on the 30-paper PosterGen evaluation set. Automatic posters use GPT-4.1 for generation; all VLMas-Judge scores in this table are produced by GPT-4o using the same six criteria as in the main experiment. PaperQuiz uses the shared three-reader image-only ensemble of GPT-4o, GPT-4o-mini, and o3; scores are averaged across readers per poster and then across papers. We report Raw PaperQuiz Accuracy and per-poster Density-Augmented PaperQuiz scores using the fixed reference w = 1335.0. The GT Poster is included only as a reference and is excluded from automatic-method ranking. Bold and underlined values denote the best and second-best automatic results; rounded ties share the same mark. Table S8: Sensitivity analyses for pairwise preference. Holm p adjusts both tests; bootstrap CIs cross-cluster by annotator and paper. Comparison

Holm p

Bootstrap CI

LOO range

PPT vs. PosterGen HTML vs. P2P

.0015 .278

62.5–82.4% 46.3–63.7%

70.6–74.6% 53.0–57.2%

Leave-one-annotator-out results do not attribute the PPTX preference to one annotator. The HTML intervals include 50%, so that result remains descriptive.

7

Extended Qualitative Results

Figures S2 and S3 compare HTML- and PPTX-oriented methods on the same four papers. They support direct inspection of density, evidence use, readability, and spatial organization, but not aggregate performance estimates or component-level causal attribution. Light-blue headers identify PosterVisor outputs. Figure S4 additionally shows PosterHarness outputs on the same paper set.

References Cameron, A. C.; Gelbach, J. B.; and Miller, D. L. 2011. Robust Inference with Multiway Clustering. Journal of Business & Economic Statistics, 29(2): 238–249. Pang, W.; Lin, K. Q.; Jian, X.; He, X.; and Torr, P. 2025. Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers. In Belgrave, D.; Zhang, C.; Lin, H.-T.; Pascanu, R.; Koniusz, P.; Ghassemi, M.; and Chen, N., eds., Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc. Zhang, Z.; Zhang, X.; Wei, J.; Xu, Y.; and You, C. 2026. PosterGen: Aesthetic-Aware Multi-Modal Paper-to-Poster Generation Via Multi-Agent LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 9813–9823.

Figure S2: Representative HTML-oriented outputs on four papers. Columns show Author Poster, PPTAgent, Any2Poster, P2P, and PosterVisor-HTML.

Figure S3: Representative PPTX-oriented outputs on the same four papers. Columns show PosterAgent, PosterForest, EfficientPosterGen, PosterGen, and PosterVisor-PPT.

Figure S4: Representative PosterHarness outputs on the same four papers.

Record · ID 919438 · SHA-256 70586a98e5efc174
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.