2026-9-7
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding Shenxi Wu*1,2 , Yuhong Liu*1,2 , Haosong Zhang3 , Tongjin Zou4 , Yanxun Zhang4 , Gaochang Chen5 , Liang Dun6 , Jiaqi Wang7 , Zhecan James Wang2 , Yuhang Zang2 and Dahua Lin†1,2,8 1
The Chinese University of Hong Kong, 2 Shanghai Artificial Intelligence Laboratory, 3 New York University, 4 Fudan
University, 5 Shanghai Jiao Tong University, 6 Harbin Institute of Technology, 7 JD Explore Academy, 8 Centre for
arXiv:2609.05141v1 [cs.AI] 4 Sep 2026
Perceptual and Interactive Intelligence (CPII) Limited
Scientific papers require models to integrate evidence across text, equations, figures, tables, code, and datasets while preserving its provenance. Existing benchmarks typically assess these capabilities separately, making it difficult to determine how well models support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark containing 124 expert-authored and difficulty-screened questions across seven research-assistant capability groups, 19 subtasks, and five scientific domains. Each question is instantiated in four matched settings formed by pairing its bilingual variants with the All Images First and Markdown Interleaved document representations, yielding 496 evaluation instances. The strongest evaluated model, Claude-Opus-5, scores 62.6 out of 100, with remaining gaps in evidence localization, structured information extraction, crossdocument synthesis, and robustness to document representation. To convert these diagnostics into scalable training signals, we introduce SciDocIR, a structured representation of scientific document objects, layout and crossreference relations, and provenance. Using SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement-learning samples across 14 verifiable subtasks. Together, SciDocBench, SciDocIR, and SciDocDataset form an evaluation-to-training framework that connects capability diagnosis with verifiable training-data construction for scientific-document assistants. The project is available at https://github.com/InternLM/SciDocBench. (a)Why we need SciDocBench? Current Doc Bench: What is the last sentence in the abstract section of the article?
Manual & Challenging 124 tasks
5 Scientific domains contained 7 Task categories, 19 subtasks
LLM: Paper input
(b) SciDocBench
Easy! Let me use the OCR tool… Tool call……
Got! “Experiments have proven that our method makes the trained model increase by 9.4% compared to the baseline.”
3.9k Avg. Text Tokens 14.2 Avg. images
Example Signle-Doc Science Question
Dual-Axis Evaluation Bilingual x Input Format
Doc representation 1. All Images first
2 Question language
Page images
English · Chinese
2 Document representation Ground Truth
All Images First · Markdown Interleaved
3 Scoring methods combined
2. Markdown Interleaved
Rule-Based · LLM-as-a-Judge · Execution-Based
text
SciDocBench(Ours):
Figure/table
text
Question 1. English
Now that you have identified the claim that model performance increased by 9.4%, which table in the article supports this statement?
Figure 4a shows the random-forest decision regions based on IgG3/4Bis and IgG1Gal. For a patient with(IgG3/4Bis, IgG1Gal) = (0.0175, 0.175), what is the predicted organ-involvement class: LOW or HIGH?
2. Chinese
LLM: I’m really confused……
OCR ability or SciDoc Understanding?
Capability Profiles of group leading models
Evaluation setting Profiles of group leading models
图 4a 展示了基于 IgG3/4Bis 和 IgG1Gal 的随机森林决策区 域。对于一名指标为 (IgG3/4Bis, IgG1Gal) = (0.0175, 0.175) 的患者,模型预测的器官受累类别是 LOW 还是 HIGH?
Figure 1: Overview of SciDocBench. (a) From text extraction to evidence grounding. A direct text-extraction prompt can be answered by recovering visible content, whereas the SciDocBench example additionally requires locating the document evidence supporting a scientific claim. (b) Benchmark and evaluation design. SciDocBench contains 124 expert-authored and difficulty-screened questions across five scientific domains, seven capability groups, and 19 subtasks. Its dual-axis protocol pairs bilingual question variants with the All Images First and Markdown Interleaved document representations, producing four matched settings per question. Responses are assessed using task-appropriate rule-based, LLM-as-a-judge, or execution-based evaluators. The plots summarize capability-group and evaluation-setting profiles for group-leading models, while the example illustrates the matched question and document-representation variants. ∗ Equal contribution.
† Corresponding author: Dahua Lin ([email protected]).
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
1. Introduction Scientific papers communicate claims through text, equations, figures, tables, captions, appendices, citations, code, and datasets [52, 112]. A useful scientific assistant must therefore do more than retrieve passages or summarize prose: it must locate supporting evidence, recover notation and experimental details, verify numerical and logical relations, compare findings across papers, trace code and data provenance, and convert scientific content into reusable outputs [20, 52, 93]. These capabilities underpin literature review, claim verification, result comparison, and reproducibility analysis, making scientific document understanding an important foundation for reliable research assistance. Scientific document understanding presents three major challenges. 1) Heterogeneous document structure. Relevant information is distributed across text, equations, figures, tables, and other scientific objects whose interpretation depends on layout, reading order, captions, and cross-references [66, 112]. 2) Compositional and cross-document reasoning. A single task may require locating a claim, identifying its supporting evidence, recovering reported values, and checking their consistency; the required evidence may also span multiple papers or associated repositories and datasets [20, 52, 93]. 3) Verifiable outputs and controlled evaluation. Research-oriented outputs must preserve scientific roles, numerical fidelity, structure, and provenance so that they can be verified or reused. Evaluation must also control for sensitivity to bilingual question variants and document representations. Existing benchmarks provide strong foundations for individual components of this problem. Documentunderstanding benchmarks evaluate visual question answering, layout-aware representation learning, and long-context document reading [10, 66, 68, 119]. Scientific question-answering datasets ground questions in research papers or biomedical abstracts [20, 41], while chart and plot benchmarks assess visual, logical, and numerical reasoning over graphical data [67, 70]. Large-scale scientific-document resources further support structural analysis and question answering across papers [52, 62, 112]. However, coverage of full scientific documents, scientific objects, cross-document reasoning, and code or data provenance remains distributed across separate benchmarks, as summarized in Table 1. This fragmentation makes it difficult to evaluate complete scientific research workflows or identify where such workflows fail. Figure 1 shows the distinction between text extraction and evidence-grounded scientific reasoning. A direct extraction prompt asks a model to recover a visible sentence, whereas the scientific-document task asks which table supports the extracted claim. The latter requires the model to locate the relevant evidence, verify their relationship to the claim, and identify the supporting source. Motivated by this distinction, we introduce SciDocBench, a workflow-centered benchmark in which each question represents a concrete research task. SciDocBench contains 124 expert-authored and difficulty-screened questions across five scientific domains, seven capability groups, and 19 subtasks. It covers both detailed single-document analysis and multi-document tasks that integrate notation, datasets, evidence, or experimental results across papers. To distinguish task capability from sensitivity to input presentation, SciDocBench uses a matched dual-axis evaluation protocol. Each question has bilingual variants and is presented using either All Images First, which supplies the paper as page images, or Markdown Interleaved, which places extracted text and visual elements in document order. Pairing these axes produces four matched settings and 496 evaluation instances derived from the 124 underlying questions. The task semantics, reference-answer semantics, and evaluation criteria remain fixed across settings, enabling controlled comparisons of question-language and document-representation sensitivity. Questions require task-appropriate structured or executable outputs, which are assessed using rule-based matching, LLM-as-a-judge, or execution-based evaluation. Evaluation can identify capability gaps, but it does not by itself provide scalable training supervision. Scientific papers also vary substantially in template, source representation, and page layout, making direct task generation difficult to control and verify. We therefore introduce SciDocIR, a normalized representation that aligns visual and semantic document objects while preserving layout, cross-reference relations, and provenance. Based on SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement learning samples across 14 verifiable subtasks. These training subtasks are selected because their targets can be checked through structured records, controlled perturbations, exact operations, numerical tolerances, or execution. Benchmark questions are not converted into training examples. We evaluate a broad collection of proprietary and open multimodal models on SciDocBench. ClaudeOpus-5 achieves the highest overall score of 62.6 out of 100, and no model leads more than two of the seven 2
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
capability groups. The results reveal persistent gaps in scientific information extraction and cross-document synthesis, despite stronger performance on some evidence-verification tasks. Model rankings also change across evaluation settings: for 11 of the 14 evaluated models, the document-representation gap is larger than the question-language gap. These findings indicate that scientific-document performance depends both on the capabilities required by a task and on how equivalent questions and evidence are presented. Our contributions are as follows: 1) Workflow-centered benchmark. We introduce SciDocBench, a workflow-centered benchmark that evaluates concrete scientific research tasks across full documents, multiple papers, and associated code or data, using a matched bilingual and dual-representation protocol. 2) Systematic model evaluation. We systematically evaluate proprietary and open multimodal models, revealing complementary capability profiles, persistent weaknesses in structured extraction and cross-document synthesis, and substantial sensitivity to document representation. 3) Evaluation-to-training resources. We introduce SciDocIR and construct SciDocDataset, connecting benchmark diagnostics with scalable supervised fine-tuning and reinforcement learning data across verifiable scientific-document tasks.
2. Related Work Document understanding benchmarks. Document understanding encompasses layout parsing, page segmentation, table extraction, field extraction, visual question answering, and long-context reading. PubLayNet provides large-scale layout annotations for scientific articles [124], while LayoutLM models interactions between text and document layout [119]. DocVQA and DUE evaluate question answering over document images [10, 68], and MMLongBench-Doc extends multimodal evaluation to lengthy documents containing rich layouts and visual elements [66]. ArXivDoc directly compares image-based, text-based, and interleaved representations for retrieving evidence from scientific documents [43]. Building on these foundations, SciDocBench evaluates how document perception and long-context understanding operate within scientific workflows involving evidence localization, verification, cross-document synthesis, and provenance-aware reconstruction. Scientific question answering and research assistance. QASPER evaluates information-seeking questions over full NLP papers and provides supporting-evidence annotations [20], whereas PubMedQA focuses on biomedical research questions grounded primarily in abstracts [41]. M3SciQA extends scientific question answering to multimodal and multi-document settings [52], while PaperScope evaluates agentic deep research over multimodal evidence distributed across collections of scientific papers [116]. PaperBench evaluates agents on end-to-end replication of AI research, including code development and experiment execution [93]. Complementing these resources, SciDocBench evaluates the document operations that underlie broader research workflows, including reading-flow recovery, claim grounding, notation and experiment extraction, consistency checking, and structured result integration. Chart, table, and scientific-object reasoning. PlotQA [70] and ChartQA [67] evaluate visual, logical, and numerical reasoning over plots and charts. SCITAB focuses on compositional reasoning and claim verification over scientific tables [63]. Scientific-document assistance, however, may require evidence distributed across a figure, caption, table, surrounding text, appendix, or another paper. SciDocBench therefore embeds figure, table, and equation tasks within full scientific documents and evaluates whether models can recover values, verify cross-object consistency, and preserve provenance when integrating results across documents. Structured scientific-document resources and multimodal training data. S2ORC provides scholarly metadata, resolved references, structured full text, and linked mentions of citations, figures, and tables [62]. DocGenome further structures multimodal scientific-document objects, layout attributes, source code, and relations for document-oriented training and evaluation [112]. Beyond scientific documents, visual instructiontuning methods use generated image–instruction pairs to support general multimodal interaction [57], while MM-IFEngine constructs multimodal training data with compositional output constraints and corresponding rule- and model-based evaluation [24]. DocSeeker combines supervised fine-tuning with evidence-aware reinforcement learning to train structured evidence localization and reasoning over long visual documents [120]. Our framework combines these directions through task-specific verifiability: SciDocIR aligns rendered PDF content with available LaTeX semantics and records block-level layout, relations, and provenance, while SciDocDataset uses these records to construct supervised fine-tuning and reinforcement learning data across 14 verifiable scientific-document subtasks. Positioning and scope. Table 1 summarizes representative benchmarks along four dimensions central to 3
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 1: Focused comparison with representative scientific-document benchmarks. ✓denotes primary coverage, –partial or adjacent coverage, and ✗that the dimension is not an evaluation target. Cross-doc. requires integrating multiple documents; code/data prov. denotes explicit reasoning over code, datasets, or provenance. Full paper Scientific Cross- Code/data or PDF objects doc. prov.
Benchmark
Primary focus
DocVQA [68] QASPER [20] ChartQA [67] SCITAB [63] MMLongBench-Doc [66] DocGenome [112] M3SciQA [52] PaperBench [93] ArXivDoc [43] PaperScope [116]
Document-image question answering Evidence-grounded full-paper QA Chart question answering Scientific-table claim verification Long-context document multimodal QA Scientific-document structure and QA Multimodal multi-document QA End-to-end research replication Scientific retrieval across representations Agentic multi-document deep research
✗ ✓ ✗ ✗ ✓ ✓ ✓ ✓ ✓ ✓
✗ – – ✓ – ✓ ✓ – ✓ ✓
✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✓
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗
SciDocBench
Workflow-centered scientific-document understanding
✓
✓
✓
✓
SciDocBench: full-document input, reasoning over figures, tables, and equations, cross-document understanding, and code or data provenance. A “partial” designation indicates that a benchmark contains relevant signals without making that dimension a central evaluation target. For example, QASPER includes evidence annotations but does not center multimodal reasoning over scientific objects [20]; ChartQA evaluates chart reasoning outside full-paper workflows [67]; and MMLongBench-Doc studies long-context multimodal documents without centering paper–code alignment or cross-document result integration [66]. Accordingly, our contribution is not that these individual capabilities are new, but that SciDocBench jointly evaluates their composition within concrete scientific research workflows.
3. The SciDocBench Benchmark 3.1. Benchmark Overview and Design Principles SciDocBench contains 124 expert-authored questions organized into seven capability groups and 19 subtasks. Each question is instantiated under four matched settings formed by combining two question languages with two document representations, producing 496 evaluation instances. Together, these instances evaluate whether multimodal models can locate evidence, interpret scientific objects, verify reported information, integrate findings across documents, and connect papers with associated code and data. The benchmark follows three design principles. Workflow Realism: each question is derived from a concrete research activity, such as reviewing a claim, extracting experimental details, comparing results across papers, reproducing an analysis, tracing a dataset, or converting a scientific figure into reusable data. Evidence Fidelity and Controlled Variation: the supplied materials retain the scientific content required by the task, while matched evaluation settings isolate sensitivity to question language and document representation. Verifiable Structure: questions require structured or executable outputs when the corresponding research activity demands them, enabling task-appropriate evaluation rather than relying exclusively on free-form responses. Positioning against existing benchmarks. Table 1 compares SciDocBench with representative document, scientific QA, chart/table, and long-document benchmarks. The comparison focuses on whether a benchmark provides complete papers or PDFs, evaluates reasoning over scientific objects, requires evidence integration across documents, and covers code or data provenance. These dimensions capture operations that arise when researchers use multimodal models as paper-reading assistants but are only partially represented in conventional document QA. 3.2. Task Taxonomy SciDocBench organizes its questions into seven capability groups. Document Perception and Structure, Scientific Information Extraction, and Evidence Alignment and Verification evaluate whether models can locate, interpret, extract, and check information within scientific documents. Cross-Document Understanding 4
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
evaluates the integration of notation, datasets, and results across papers. Reconstruction and Execution, Paper–Code Alignment, and Dataset Understanding examine whether models can transform scientific content into reusable artifacts and connect papers with associated computational resources. The complete inventory of 19 subtasks, together with their definitions and evaluation focus, is provided in Appendix B. 3.3. Evaluation Instances and Scoring An evaluation instance specifies the materials provided to the model, the task instruction, the expected form of the response, and the method used to score that response. Model inputs. Each instance provides a question in either English or Chinese together with all source materials required to answer it. Most tasks use one or more complete scientific papers. In the All Images First setting, every page of the source papers is supplied as an image before the question. In the Markdown Interleaved setting, extracted text and visual elements are arranged in their original document order, followed by the question. Multi-document tasks provide all required papers under the same representation. For paper–code alignment and dataset understanding tasks, the context may additionally contain repository file trees, code or pseudocode excerpts, dataset descriptions, or structured metadata. The model is not given the reference answer, capability label, subtask label, evaluator type, or correctness metadata. Output contract. The question specifies what the model must return and, when necessary, the exact response format. Depending on the task, the required output may be a label, numerical value, text span, ordered list, JSON object, table, graph structure, reconstructed data, LaTeX expression, or executable program. Structured formats are used when the corresponding research task requires an output that can be parsed, verified, or reused directly. Evaluation and score aggregation. We use three evaluator families according to the required output. Rulebased matching is used for deterministic answers with well-defined structures or values. LLM-as-a-judge is used when correctness requires semantic assessment that cannot be reduced to exact matching. Execution-based evaluation is used for runnable code or other executable artifacts. Every evaluation instance receives a score 𝑠𝑖 ∈ [0, 1], and failed, empty, or otherwise unusable responses receive zero. Let 𝑁 = 496 denote the total number of evaluation instances. The standard overall score is 𝑁
Overall =
100 ∑︁ 𝑠𝑖 . 𝑁 𝑖=1
(1)
Capability-group scores and evaluation-setting scores are computed as means over their corresponding instances and reported on a 0–100 scale. The four matched instances derived from each benchmark question preserve the same source content, task semantics, reference-answer semantics, and evaluation criteria. Their score differences therefore support paired analysis of sensitivity to question language and document representation. Full task definitions and benchmark statistics are provided in Appendix B; parsing rules and task-specific scorers are included in the evaluation code. 3.4. Benchmark Construction As summarized in Figure 2, benchmark construction consists of four stages: paper selection, question authoring and review, difficulty screening, and controlled augmentation. The objective is to obtain questions that are scientifically meaningful, answerable from the supplied materials, unambiguous in their expected outputs, and challenging for current multimodal models. Representative benchmark samples are provided in Appendix F.
5
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Benchmark pipeline
Solved by all
MinerU
Candidate question 1 Filtered ✓
Candidate question 2 Paper PDF Candidate question 3
Paper PDF
Data Collection
···
Expert Design
Ambiguity check Evaluability check Answer verification
Output (a) Interleaved markdown
Pdf2Img
(b) Page images (c) English Instruction
Verified Seed
Difficulty Check
Translation
(d) Chinese Instruction
Augmentation
Figure 2: SciDocBench construction pipeline. 1) Data collection: information-rich papers are selected from arXiv, bioRxiv, OpenReview, and PubMed. 2) Expert design: researchers author candidate questions from realistic scientific-reading activities. 3) Difficulty control: candidates solved by all screening models are removed, while the remaining questions undergo ambiguity, evaluability, and answer-verification checks. 4) Controlled augmentation: each verified question is paired with Markdown Interleaved or All Images First document inputs and English or Chinese instructions, yielding four matched evaluation instances.
Paper selection. We collect 116 publicly accessible papers spanning five scientific domains from arXiv, bioRxiv, OpenReview, PubMed, and corresponding journal pages. We retain information-rich papers containing several types of scientific objects, including figures, tables, equations, captions, appendices, and experimental descriptions. We summarize the document characteristics using the profile (︀ )︀ 𝑟(𝑝) = 𝑛fig , 𝑛tab , 𝑛eq , 𝑛cap , 𝑛app , 𝑛exp .
(2)
Papers that support only generic summarization questions or lack sufficiently stable information for reliable evaluation are filtered out. We also consider the types of questions naturally supported by different papers. Theoretical papers commonly support questions about notation, equations, and dependency relations, whereas empirical papers support questions about experimental settings, table consistency, figure interpretation, and result integration. Construction statistics are provided in Appendix B. Question authoring and review. Questions are designed from realistic scientific reading and analysis scenarios. Each candidate includes one or more source documents, a user-facing question, a reference answer, capabilitygroup and subtask labels, and a task-appropriate evaluation method. When a structured response is required, the expected format is specified directly in the question. The review process verifies that the question is scientifically meaningful, answerable from the supplied materials, associated with a stable reference answer, and unambiguous in both its instruction and expected output. Field-level authoring requirements, the independentverification procedure, and the annotation interfaces are documented in Appendix C. 3.5. Quality Control and Evaluation Reliability Difficulty screening. To remove questions that are already reliably solved by strong systems, we evaluate each candidate once using Claude Opus 4.8, Gemini 3.1 Pro, ChatGPT 5.5, and Kimi K2.5. We denote this screening set by ℳbench . Let 𝜑𝑡 (ˆ 𝑦 , 𝑦) ∈ [0, 1] denote the task-specific scorer for subtask 𝑡. A candidate is considered overly easy when every model in ℳbench answers it correctly: [︀ ]︀ easy(𝑞) = I ∀𝑚 ∈ ℳbench , 𝜑𝑡(𝑞) (ˆ 𝑦𝑚 , 𝑦𝑞 ) = 1 .
(3)
A valid candidate is retained only if at least one screened model fails. Questions that all screened models fail are not accepted automatically; they are retained only when independent review confirms that the supplied materials are sufficient, the reference answer is stable, and the expected output is unambiguous. The screening stage therefore removes both trivial questions and questions whose apparent difficulty results from ambiguity or insufficient information. Controlled augmentation. Each benchmark question is augmented along two controlled dimensions: question language and document representation. We construct semantically equivalent English and Chinese questions and pair each language version with the two document representations defined in Section 3.3. The source 6
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Dataset pipeline
Example Question Figure
Paper PDF
Visual Parsing
Block Bbox
Doc Metadata
Paragraph
Page Info.
Equation Latex Source
Semantic Parsing
SciDocIR
Table ···
Block Details
Build SciDocIR
External Material (if needed)
Re-render ( if needed )
The surrounding text reports the correct values, but one bar in Figure 4 is drawn at the wrong height. Identify the affected country and year, and report the textual value and the visually represented value. Your answer should follow the format ……
Source Paper Modified Text
PlotQA Img
Corrupted Info. ( if needed )
With Seed Prompt
Seed Instruction
LLM Generate Tasks
Pass Rate Check
GT Paper Images
<think> … </think> { "country": "Kazakhstan", "year": "2013", "textual_value": "15.0%", "visual_value": "approximately 7.0%" }
Figure 3: SciDocDataset construction pipeline. SciDocIR construction: visual parsing of PDF pages and semantic parsing of available LaTeX sources are aligned into document-, page-, and block-level records. Seeded task generation: task-specific prompts combine selected SciDocIR records with external or synthesized material when required, while deterministic source fields remain fixed. Evidence rendering: controlled figures or text are inserted into recorded block regions, while tasks requiring no modification retain the original pages. Validation: task-specific checks produce training samples containing an instruction, verifiable ground truth, and corresponding paper images.
content, task semantics, reference-answer semantics, required output format, and evaluation criteria are preserved across the resulting four matched instances. Each instance is checked for semantic equivalence and answerability. This procedure increases evaluation coverage without treating translated or reformatted instances as independently authored questions. Reliability of semantic evaluation. We audit GPT-5.4-mini judgments on 100 response-level instances sampled from the LLM-as-a-judge portion of the evaluation. The sample is balanced across the four evaluation settings, with 25 instances per setting, and spans 11 evaluated models and all seven capability groups. A human expert reviewed the question, reference answer, and anonymized model response without access to the model identity or GPT judgment. Under a strict three-level mapping, where scores of 0 and 1 denote incorrect and correct responses and every intermediate score denotes a partially correct response, the expert and GPT-5.4-mini achieve 92.0% exact agreement and a quadratic-weighted Cohen’s 𝜅 of 0.866. A binary threshold at 0.5 yields 70.0% agreement and 𝜅 = 0.355, indicating that distinctions involving partial credit are better represented by the ordinal analysis than by a hard threshold. An independent Codex blind review provides a third set of ratings; across the expert, GPT-5.4-mini, and Codex ratings, ordinal Krippendorff ’s 𝛼 is 0.854, with a 95% bootstrap confidence interval of [0.757, 0.924]. Appendix C.1 describes the protocol, interface, and pairwise results.
4. SciDocIR and SciDocDataset Scientific papers use heterogeneous publication templates, source-code conventions, and page layouts. The same semantic object can appear as a LaTeX environment, a rasterized PDF region, or several visually separated blocks. Direct task generation from these heterogeneous inputs makes evidence retrieval, controlled editing, and answer verification difficult. We therefore normalize each source paper into SciDocIR and construct SciDocDataset through verifiable operations over this shared representation. Figure 3 summarizes the complete pipeline. 4.1. SciDocIR Construction Each source paper provides a PDF and, when available, its LaTeX source. The PDF supplies the rendered document view. We convert each page into an image and perform visual parsing to identify figures, paragraphs, equations, tables, captions, and other layout blocks. For every detected block, we record its page index, bounding box, category, visual crop, and local spatial relations. The LaTeX source supplies the corresponding semantic view. We parse the section structure and document environments to recover source representations for text, equations, tables, figures, labels, and references. Visual and semantic records are aligned using textual content, labels, caption associations, and page geometry. This alignment preserves both the appearance of the published page and the source representation of its scientific content. SciDocIR stores the aligned result at three levels. Document metadata records identity, source, domain,
7
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
page count, and provenance. Page information records page indices, dimensions, rendered images, and page-level layout. Block details record the block type, bounding box, plain text, LaTeX or source code, visual crop, caption target, reading order, cross-references, and task-specific fields. Relations connect blocks that form continuations, parent–child structures, caption–target pairs, or local evidence chains. The complete field definitions and a simplified record schema are provided in Appendix D. 4.2. Verifiable Training-Task Design We distinguish benchmark subtasks from training subtasks. The 19 benchmark subtasks provide diagnostic coverage of broad and difficult scientific workflows, whereas training subtasks are selected according to whether their answers can be verified from SciDocIR fields, controlled perturbations, exact set operations, numeric tolerances, structured parsing, or execution. SciDocDataset instantiates 14 verifiable subtasks in eight construction directions, listed in Appendix E. Layout and reading-flow recovery includes block-role classification, parent–child linking, reading-order prediction, and next-hop prediction. Evidence verification covers table logic and chart visual consistency. Cross-document understanding includes dataset intersection and real or simulated table integration, while reconstruction includes single-point and multi-point chart readout. The remaining subtasks address notation extraction, symbol disambiguation, and citation-role classification. This separation keeps SciDocBench broad and diagnostic while allowing SciDocDataset to provide scalable, task-specific supervision. Benchmark questions are not converted into training examples; training samples are generated through task-aligned operations rather than derived from held-out benchmark items. 4.3. Seeded Generation and Source Authority For each training subtask, we define a seed instruction specifying the scientific operation, available records, output schema, and rejection conditions. For tasks grounded in real papers, a retrieval program selects the relevant SciDocIR pages and blocks before GPT naturalizes the question or performs a bounded semantic check. For synthetic or externally sourced tasks, the generator receives locked values, labels, or formulas together with the rendered page context. Fields obtained from deterministic rules, source data, or construction logs remain immutable. GPT is therefore used for linguistic naturalization, bounded candidate generation, semantic classification, or solvability checks rather than as the authority for ground truth. Appendix E.2 provides the task-specific templates and identifies the source records, model roles, and ground-truth authority for each subtask. Different task families use different evidence sources. Real cross-document table integration retrieves compatible tables from the SciDocIR paper collection. Its simulated counterpart splits one real table into two partially overlapping tables with permuted columns and embeds them in separate paper contexts. Chart readout uses PlotQA charts and their underlying series data, while chart visual consistency uses the same values to render controlled radar, line, scatter, or bar plots with Matplotlib. Dataset-intersection tasks draw names from an internal lexicon of approximately 50 common scientific datasets and resources. Notation tasks use domain-specific formula and layout templates rendered with XeLaTeX. These sources provide exact values and construction records, while SciDocIR supplies realistic scientific-document context when applicable. 4.4. Controlled Block-Level Re-rendering Synthetic evidence is embedded into realistic paper context through block-level re-rendering. We select a figure or table block from SciDocIR, resize the synthesized visual to its recorded bounding box, and replace precisely that region in the page image. When required by the task, the corresponding caption or nearby textual statement is updated in its own block while unrelated page content remains unchanged. This procedure preserves the page template, typography, column structure, and surrounding scientific context while providing exact control over the modified evidence. Samples that require no modification retain the original page images. Each rendering record stores the source page, target block identifier, bounding box, inserted content, textual modification, and expected answer. These fields connect the rendered evidence to the operation that produced it and support subsequent verification and auditing.
8
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 2: SciDocBench leaderboard. Overall is computed over all 496 evaluation instances derived from 124 underlying questions, with failed or unusable responses assigned a score of zero. Evaluation-setting scores pair English or Chinese questions with the All Images First (Images) or Markdown Interleaved (Interleaved) document representation. Capability scores correspond to the seven groups defined in Section 3. Model families are ordered by their best overall score, and models within each family are sorted by overall score. In each score column the best, second, and third values are shaded from dark to light; ties receive the same shading.
Model
English
Overall
Chinese
Capability scores
Images
InterInter- Document Information Evidence CrossReconstr. Paper– Dataset Images leaved leaved Perception Extraction Verification Document & Exec. Code Underst.
Claude-Opus-5 [6] Claude-Opus-4.8 [5] Claude-Sonnet-4.6 [7] Claude-Haiku-4.5 [4]
62.60 53.41 49.57 40.35
63.19 53.63 52.25 38.18
61.34 52.43 51.18 42.72
65.21 54.13 53.62 38.33
60.68 53.46 41.21 42.16
61.37 46.59 44.80 37.76
55.52 57.12 51.51 42.83
74.47 58.63 54.25 40.33
62.81 56.65 47.43 37.64
59.65 57.09 52.64 41.62
51.75 44.00 52.17 51.58
43.75 53.12 59.38 59.38
GPT-5.6-Sol [76] GPT-5.6-Terra [77] GPT-5.6-Luna [75]
61.00 54.19 50.58
61.07 51.88 47.55
60.32 55.69 53.99
61.32 52.83 46.65
61.31 56.37 54.13
54.08 44.01 38.68
56.78 51.74 47.52
78.16 66.28 65.15
54.62 54.01 56.07
69.70 65.75 53.86
46.69 61.88 54.25
59.38 56.25 62.50
Gemini-3.6-Flash [33]
59.87
59.19
62.98
59.68
57.61
54.00
55.19
73.28
62.42
55.42
58.31
56.25
Qwen3.8-Max [84] Qwen3.7-Plus [82] Qwen3.8-27B [83]
57.15 55.88 52.89
59.55 57.17 58.65
55.33 53.34 47.25
58.13 57.57 54.21
55.57 55.42 51.46
48.17 49.00 47.43
50.79 53.84 50.50
73.84 63.72 62.87
63.02 59.04 54.80
53.60 55.31 54.45
55.75 65.06 54.06
59.38 71.88 31.25
Kimi-K2.5 [45]
50.38
54.74
45.42
55.04
46.33
45.69
50.62
53.05
53.61
50.41
70.25
40.62
GLM-4.6V [121]
42.44
43.30
46.70
38.21
41.56
37.15
46.58
41.63
44.57
43.90
57.27
59.38
MiMo-V2.5 [113]
40.16
42.26
39.26
44.66
34.46
34.11
37.66
49.60
36.40
44.84
55.30
43.75
4.5. Validation and Output Records Each generated candidate is checked against the output schema and rejection conditions specified by its training subtask. Depending on the task, validation uses stored layout relations, locked labels or values, exact set operations, numeric tolerances, structured parsing, visual-solvability checks, or execution-based tests. Candidates are rejected when required fields are missing, generated content conflicts with locked source records, the rendered evidence is insufficient to recover the answer, or a task-specific consistency check fails. Each accepted sample contains the instruction, associated document input, and verifiable target answer. The accompanying construction metadata retains source identifiers, locked fields, rendering records, and rejection outcomes. This separation preserves an auditable connection between model-assisted language generation and the deterministic or traceable source of each answer.
5. Experiments 5.1. Experimental Setup We evaluate 14 proprietary and open multimodal models on SciDocBench. Each model is evaluated on the same 124 underlying questions under the four matched language and document-representation settings defined in Section 3.3, producing 496 evaluation instances per model. We report the overall score, four evaluation-setting scores, and seven capability-group scores. 5.2. Overall Leaderboard Table 2 reports the SciDocBench leaderboard. Claude-Opus-5 ranks first with an overall score of 62.60, followed by GPT-5.6-Sol with 61.00 and Gemini-3.6-Flash with 59.87. Qwen3.8-Max and Qwen3.7-Plus obtain 57.15 and 55.88, respectively. No evaluated model reaches 63, and overall scores span 22.44 points across the evaluated systems.
9
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
5.3. Capability-Level Analysis Capability leadership is distributed across model families. Claude-Opus-5 leads Group A with 61.37, ClaudeOpus-4.8 leads Group B with 57.12, GPT-5.6-Sol leads Groups C and E with 78.16 and 69.70, Qwen3.8-Max leads Group D with 63.02, Kimi-K2.5 leads Group F with 70.25, and Qwen3.7-Plus leads Group G with 71.88. No model leads more than two capability groups. Claude-Opus-5 obtains the highest overall score by remaining comparatively strong across several groups rather than dominating the complete taxonomy. Document perception and scientific extraction form common bottlenecks. Averaged equally across the 14 evaluated models, Group A has the lowest score at 45.92, followed by Group B at 50.59. Group C has the highest mean at 61.09. The capability ceilings exhibit a related pattern: the best Group B score is 57.12, while Groups A and D peak at 61.37 and 63.02. These results indicate that locating document elements, recovering structured scientific information, and integrating evidence across papers remain difficult even when models perform comparatively well on evidence verification. High overall performance does not imply a balanced capability profile. The difference between each model’s strongest and weakest capability-group scores ranges from 14.58 to 31.62 points. This variation remains large for the two strongest models: Claude-Opus-5 spans 30.72 points across capability groups, while GPT-5.6-Sol spans 31.47 points. Their overall scores therefore conceal distinct weaknesses, including dataset understanding for Claude-Opus-5 and paper–code alignment for GPT-5.6-Sol. Evaluating only an aggregate score would obscure these capability-specific limitations. 5.4. Language and Representation Robustness Aggregate language differences conceal model-specific effects. Averaged equally over the 14 models and two document representations, English questions score 52.52 and Chinese questions score 51.83, a difference of 0.69 points. Individual model effects are considerably larger and differ in direction. Qwen3.7-Plus improves by 1.24 points with Chinese questions, whereas GLM-4.6V and Claude-Sonnet-4.6 decrease by 5.12 and 4.30 points, respectively. Language sensitivity is therefore model dependent rather than a uniform translation penalty. Document representation changes both scores and rankings. Across all models and languages, All Images First exceeds Markdown Interleaved by 1.52 points on average. However, nine models favor All Images First and five favor Markdown Interleaved. Kimi-K2.5 and Qwen3.8-27B favor All Images First by 9.02 and 7.07 points, whereas GPT-5.6-Luna and Claude-Haiku-4.5 favor Markdown Interleaved by 6.96 and 4.19 points. Claude-Opus-5 leads both All Images First settings, Gemini-3.6-Flash leads EN-IL, and GPT-5.6-Sol leads ZH-IL. Neither representation is therefore uniformly preferable across systems. Language and representation effects can interact. For Gemini-3.6-Flash, Markdown Interleaved exceeds All Images First by 3.79 points with English questions, but All Images First exceeds Markdown Interleaved by 2.07 points with Chinese questions. Qwen3.8-27B exhibits a different reversal: Chinese questions reduce its All Images First score by 4.44 points relative to English but improve its Markdown Interleaved score by 4.21 points. These reversals show that the two controlled axes should be analyzed jointly rather than interpreted as independent sources of performance variation. Cross-setting stability is distinct from peak accuracy. The average range across the four evaluation settings is 6.41 points. GPT-5.6-Sol varies by only 1.00 point and Claude-Opus-4.8 by 1.70 points, whereas ClaudeSonnet-4.6, Qwen3.8-27B, and MiMo-V2.5 exhibit ranges above 10 points. Reporting only the best setting would therefore conflate scientific-document capability with compatibility with a particular language and document representation. 5.5. Main Findings The experiments support three conclusions. Capability composition: no evaluated model performs uniformly well across perception, extraction, verification, integration, reconstruction, and provenance-oriented tasks. Interface sensitivity: equivalent questions can produce materially different outcomes when language and document representation change, and these effects sometimes interact. Remaining headroom: the strongest overall score of 62.60, together with the low average scores for document perception and scientific information extraction, leaves substantial room for improving workflow-level scientific-document understanding. 10
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
6. Conclusion We introduced SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored questions across seven capability groups and nineteen subtasks. Each question is evaluated in two languages and with two document representations, producing 496 matched instances. The strongest evaluated model scores 62.60, and the capability winners are distributed across several model families. Performance is also sensitive to input format: for most models, the document-representation gap is larger than the question-language gap. We further introduced SciDocIR as a normalized representation of document metadata, page information, and block details, and used it to construct supervised fine-tuning and reinforcement learning data with verifiable targets. Scientific document understanding is not OCR, long-document question answering, or figure reading in isolation. It requires layout recovery, evidence grounding, structured extraction, verification, cross-document synthesis, reconstruction, and provenance tracking to operate together.
References [1] Emmanuel Abbe, Afonso S Bandeira, and Georgina Hall. Exact recovery in the stochastic block model. IEEE Transactions on information theory, 62(1):471–487, 2015. 9 [2] Dong An and Lin Lin. Quantum linear system solver based on time-optimal adiabatic quantum computing and quantum approximate optimization algorithm. ACM Transactions on Quantum Computing, 3(2): 1–28, 2022. 8, 10 [3] David Anderson, Greta Panova, and Leonid Petrov. Computation and sampling for schubert specializations. arXiv preprint arXiv:2603.20104, 2026. 10 [4] Anthropic.
Introducing Claude Haiku 4.5. https://www.anthropic.com/news/ claude-haiku-4-5, October 2025. Accessed September 4, 2026. 2
[5] Anthropic. Introducing Claude Opus 4.8. https://www.anthropic.com/news/claude-opus-4-8, May 2026. Accessed September 4, 2026. 2 [6] Anthropic. Introducing Claude Opus 5. https://www.anthropic.com/news/claude-opus-5, July 2026. Accessed September 4, 2026. 2 [7] Anthropic.
Introducing Claude Sonnet 4.6. https://www.anthropic.com/news/ claude-sonnet-4-6, February 2026. Accessed September 4, 2026. 2
[8] Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler. Benign overfitting in linear regression. Proceedings of the National Academy of Sciences, 117(48):30063–30070, 2020. 9 [9] D Basilico, G Bellini, J Benziger, R Biondi, B Caccianiga, A Caminata, A Chepurnov, DD Angelo, A Derbin, A Di Giacinto, et al. New limits on the pauli forbidden transitions in 12c nuclei obtained with the complete borexino dataset. arXiv preprint arXiv:2604.08950, 2026. 7 [10] Lukasz Borchmann, Michal Pietruszka, Tomasz Stanislawek, Dawid Jurkiewicz, Michal Turski, Karol Szyndler, and Filip Gralinski. Due: End-to-end document understanding benchmark. In NeurIPS Datasets and Benchmarks Track, 2021. 1, 2 [11] Zi Cai and Jinguo Liu. Approximating quantum many-body wave functions using artificial neural networks. Physical Review B, 97(3):035116, 2018. 10 [12] Meng Cao, Xingyu Li, Xue Liu, Ian Reid, and Xiaodan Liang. Spatialdreamer: Incentivizing spatial reasoning via active mental imagery. arXiv preprint arXiv:2512.07733, 2025. 12 [13] Yuan Cao, Valla Fatemi, Shiang Fang, Kenji Watanabe, Takashi Taniguchi, Efthimios Kaxiras, and Pablo Jarillo-Herrero. Magic-angle graphene superlattices: a new platform for unconventional superconductivity. arXiv preprint arXiv:1803.02342, 2018. 9 [14] Giuseppe Carleo and Matthias Troyer. Solving the quantum many-body problem with artificial neural networks. Science, 355(6325):602–606, 2017. 10 11
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[15] Valérie Castin, Pierre Ablin, José Antonio Carrillo, and Gabriel Peyré. A unified perspective on the dynamics of deep transformers. arXiv preprint arXiv:2501.18322, 2025. 9 [16] Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, et al. Minimax-m1: Scaling test-time compute efficiently with lightning attention. arXiv preprint arXiv:2506.13585, 2025. 10 [17] Bo Chen, Zhenmei Shi, Zhao Song, and Jiahao Zhang. Provable failure of language models in learning majority boolean logic via gradient descent. arXiv preprint arXiv:2504.04702, 2025. 9 [18] Guo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang, Yicheng Liu, Lidong Lu, De-An Huang, Wonmin Byeon, Matthieu Le, Tuomas Rintamaki, et al. Eagle 2.5: Boosting long-context post-training for frontier vision-language models. arXiv preprint arXiv:2504.15271, 2025. 11 [19] Sully F Chen, Anton Alyakin, Andreas Seas, Eunice Yang, Joanne J Choi, Jin Vivian Lee, Amelia L Chen, Pranav I Warman, Rochelle T Bitolas, Robert J Steele, et al. Llm-assisted systematic review of large language models in clinical medicine. Nature medicine, pages 1–8, 2026. 8 [20] Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A. Smith, and Matt Gardner. A dataset of information-seeking questions and answers anchored in research papers. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, 2021. 1, 2, 1, 2 [21] Isaac Davis, Julian Jara-Ettinger, and Yarrow Dunham. Inferring the internal structure of groups through the integration of statistical learning and causal reasoning. Nature Communications, 2026. 5 [22] Selvan Demir, Miguel I Gonzalez, Lucy E Darago, William J Evans, and Jeffrey R Long. Giant coercivity and high magnetic blocking temperatures for n23- radical-bridged dilanthanide complexes upon ligand dissociation. Nature Communications, 8(1):2144, 2017. 10 [23] Dong-Ling Deng, Xiaopeng Li, and S Das Sarma. Quantum entanglement in neural network states. Physical Review X, 7(2):021021, 2017. 10 [24] Shengyuan Ding, Shenxi Wu, Xiangyu Zhao, Yuhang Zang, Haodong Duan, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Dahua Lin, and Jiaqi Wang. Mm-ifengine: Towards multimodal instruction following. arXiv preprint arXiv:2504.07957, 2025. 2, 5, 6, 13 [25] Tuan Dinh, Seon-Kyeong Jang, Noah Zaitlen, and Vasilis Ntranos. Compressing the collective knowledge of esm into a single protein language model. Nature Methods, pages 1–13, 2026. 6, 11 [26] Jinshuo Dong, Aaron Roth, and Weijie J Su. Gaussian differential privacy. Journal of the Royal Statistical Society Series B: Statistical Methodology, 84(1):3–37, 2022. 9 [27] Jusaina Eyyathiyil, Subhajit Ghosh, Arunima Cheran, Silvano Geremia, Jatish Kumar, Neal Hickey, and Pakkirisamy Thilagar. Axial chirality-induced rigidification in aminoboranes enhances persistent room-temperature phosphorescence and circularly polarized luminescence. Communications chemistry, 8(1):126, 2025. 8 [28] Cheng Fang, Bei Li, and Ping Han. Semi-supervised microblog text sentiment classification based on global feature graph. Journal of Signal Processing, 37(6):1066–1074, 2021. 10 [29] Xiaomin Fang, Lihang Liu, Jieqiong Lei, Donglong He, Shanzhuo Zhang, Jingbo Zhou, Fan Wang, Hua Wu, and Haifeng Wang. Geometry-enhanced molecular representation learning for property prediction. Nature Machine Intelligence, 4(2):127–134, 2022. 11 [30] Borjan Geshkovski, Hugo Koubbi, Yury Polyanskiy, and Philippe Rigollet. Dynamic metastability in the self-attention model. arXiv preprint arXiv:2410.06833, 2024. 9 [31] Sukhpal Singh Gill, Adarsh Kumar, Harvinder Singh, Manmeet Singh, Kamalpreet Kaur, Muhammad Usman, and Rajkumar Buyya. Quantum computing: A taxonomy, systematic review and future directions. Software: Practice and Experience, 52(1):66–114, 2022. 6
12
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[32] Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, Weiyun Wang, Xiaofeng Wang, Jianbo Liu, et al. Ace-brain-0: Spatial intelligence as a shared scaffold for universal embodiments. arXiv preprint arXiv:2603.03198, 2026. 6 [33] Google DeepMind.
Gemini 3.6 Flash model card. https://deepmind.google/models/ model-cards/gemini-3-6-flash/, July 2026. Accessed September 4, 2026. 2
[34] Jiawei Guo, Tianyu Zheng, Yizhi Li, Yuelin Bai, Bo Li, Yubo Wang, King Zhu, Graham Neubig, Wenhu Chen, and Xiang Yue. Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13869–13920, 2025. 11 [35] Jing Han, Zhuochao Zhou, Rongrong Zhang, Yijun You, Zizhen Guo, Jinyan Huang, Fan Wang, Yue Sun, Honglei Liu, Xiaobing Cheng, Yutong Su, Hui Shi, Qiongyi Hu, Jialin Teng, Chengde Yang, Shifang Ren, and Junna Ye. Fucosylation of anti-dsdna igg1 correlates with disease activity of treatment-naïve systemic lupus erythematosus patients. eBioMedicine, 77:103883, 2022. doi: 10.1016/j.ebiom.2022.103883. 12 [36] Serkan Hoşten, Vadym Kurylenko, Elke Neuhaus, and Nikolas Rieke. The euler stratification for P1 × P1 × P𝑛 . arXiv preprint arXiv:2603.19184, 2026. 10 [37] Weihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. Strategies for pre-training graph neural networks. arXiv preprint arXiv:1905.12265, 2019. 11 [38] Shoham Jacoby, Alex Retzker, and Fernando Pastawski. Stairway codes: Floquetifying bivariate bicycle codes and beyond. arXiv preprint arXiv:2603.00228, 2026. 13 [39] Bin Jiang, Adrien Bouhon, Zhi-Kang Lin, Xiaoxi Zhou, Bo Hou, Feng Li, Robert-Jan Slager, and JianHua Jiang. Experimental observation of non-abelian topological acoustic semimetals and their phase transitions. Nature Physics, 17(11):1239–1246, 2021. 9 [40] Haoyi Jiang, Liu Liu, Xinjie Wang, Yonghao He, Wei Sui, Zhizhong Su, Wenyu Liu, and Xinggang Wang. Spa3r: Predictive spatial field modeling for 3d visual reasoning. arXiv preprint arXiv:2602.21186, 2026. 6, 9 [41] Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019. 1, 2 [42] William R Johnson III, Xiaonan Huang, Shiyang Lu, Kun Wang, Joran W Booth, Kostas Bekris, and Rebecca Kramer-Bottiglio. Impact-resistant, autonomous robots inspired by tensegrity architecture. arXiv preprint arXiv:2501.15078, 2025. 8 [43] Ghazal Khalighinejad, Raghuveer Thirukovalluru, Alexander H Oh, and Bhuwan Dhingra. Documentas-image representations fall short for scientific retrieval. arXiv preprint arXiv:2604.18508, 2026. 2, 1 [44] Hori Kim, Moon-Ki Jeong, Hyuk-Joon Kim, Youngsin Kim, Kisuk Kang, and Joon Hak Oh. Enhanced cycling stability of ncm811 cathodes at high c-rates and voltages via limtfsi-based polymer coating. Small, 21(30):2502816, 2025. 12 [45] Kimi Team. Kimi K2.5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026. URL https://arxiv.org/abs/2602.02276. 2 [46] Zitai Kong, Yiheng Zhu, Yinlong Xu, Mingze Yin, Tingjun Hou, Jian Wu, Hongxia Xu, and Chang-Yu Hsieh. Protflow: Flow matching-based protein sequence design with comprehensive protein semantic distribution learning and high-quality generation. bioRxiv, pages 2026–02, 2026. 6, 11 [47] Inhoe Koo, Hyunho Cha, and Jungwoo Lee. Qflownet: Fast, diverse, and efficient unitary synthesis with generative flow networks. arXiv preprint arXiv:2603.03045, 2026. 13
13
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[48] Talya S Kramer, Flossie K Wan, Sarah M Pugliese, Adam A Atanas, Sreeparna Pradhan, Alex W Hiser, Lillie M Godinez, Jinyue Luo, Eric Bueno, Thomas Felt, et al. Neural sequences underlying directed turning in caenorhabditis elegans. Nature Neuroscience, pages 1–17, 2026. 9, 10, 13 [49] Barbara Maria Latacz, Stefan R Erlewein, Markus Fleck, Julia I Jäger, Fatma Abbass, Béla P Arndt, P Geissler, Tomoka Imamura, Marcel Leonhardt, Peter Micke, et al. Coherent spectroscopy with a single antiproton spin. Nature, 644(8075):64–68, 2025. 7 [50] Olivier Laurent, Adrien Lafage, Enzo Tartaglione, Geoffrey Daniel, Jean-Marc Martinez, Andrei Bursuc, and Gianni Franchi. Packed-ensembles for efficient uncertainty estimation. arXiv preprint arXiv:2210.09184, 2022. 7 [51] Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024. 11 [52] Chuhan Li, Ziyao Shangguan, Yilun Zhao, Deyuan Li, Yixin Liu, and Arman Cohan. M3sciqa: A multimodal multi-document scientific qa benchmark for evaluating foundation models. In Findings of the Association for Computational Linguistics: EMNLP, 2024. 1, 2, 1 [53] Hongxing Li, Dingming Li, Zixuan Wang, Yuchen Yan, Hang Wu, Wenqi Zhang, Yongliang Shen, Weiming Lu, Jun Xiao, and Yueting Zhuang. Spatialladder: Progressive training for spatial reasoning in visionlanguage models. arXiv preprint arXiv:2510.08531, 2025. 12 [54] Wangda Li, Xiaoming Liu, Qiang Xie, Ya You, Miaofang Chi, and Arumugam Manthiram. Long-term cyclability of ncm-811 at high voltages in lithium-ion batteries: an in-depth diagnostic study. Chemistry of Materials, 32(18):7796–7804, 2020. 12 [55] Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Yang Zhao, Sida Peng, Hengkai Guo, Xiaowei Zhou, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views. In International Conference on Learning Representations, pages 141261–141285, 2026. URL https://proceedings.iclr.cc/paper_files/paper/2026/ file/e4cd50120b6d7e8daff1749d6bbaa889-Paper-Conference.pdf. 5 [56] Zhichao Lin, Zhichao Liang, Gaoqiang Liu, Meng Xu, Baoyu Xiang, Jian Xu, and Guanjun Jiang. Quarkmedsearch: A long-horizon deep search agent for exploring medical intelligence. arXiv preprint arXiv:2604.12867, 2026. 5 [57] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Advances in Neural Information Processing Systems, 2023. 2 [58] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26296–26306, 2024. 9 [59] Junhua Liu, Zhangcheng Wang, Zhike Han, Ningli Wang, Guotao Liang, and Kun Kuang. Twiff (think with future frames): A large-scale dataset for dynamic visual reasoning. arXiv preprint arXiv:2602.10675, 2026. 13 [60] Yuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao, Long Xing, Xiaoyi Dong, Haodong Duan, Dahua Lin, and Jiaqi Wang. Spatial-ssrl: Enhancing spatial understanding via self-supervised reinforcement learning. arXiv preprint arXiv:2510.27606, 2025. 5, 8, 9, 13 [61] Princess Stephanie Llanos, Zahra Ahaliabadeh, Ville Miikkulainen, Jouko Lahtinen, Lide Yao, Hua Jiang, Timo Kankaanpaa, and Tanja M Kallio. High voltage cycling stability of lif-coated nmc811 electrode. ACS Applied Materials & Interfaces, 16(2):2216–2230, 2024. 12 [62] Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel S. Weld. S2orc: The semantic scholar open research corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020. 1, 2
14
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[63] Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. SCITAB: A challenging benchmark for compositional reasoning and claim verification on scientific tables. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023. 2, 1 [64] Changze Lv, Yifei Wang, Yanxun Zhang, Yiyang Lu, Jingwen Xu, Xiaohua Wang, Di Yu, Xin Du, Xuanjing Huang, and Xiaoqing Zheng. Biologically plausible learning via bidirectional spike-based distillation. arXiv preprint arXiv:2509.20284, 2025. 7 [65] Shuo Ma, Muhao Chen, Hongying Zhang, and Robert E Skelton. Statics of integrated origami and tensegrity systems. International Journal of Solids and Structures, 279:112361, 2023. 7 [66] Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, et al. Mmlongbench-doc: Benchmarking long-context document understanding with visualizations. In NeurIPS Datasets and Benchmarks Track, 2024. 1, 2, 1, 2 [67] Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL, 2022. 1, 2, 1, 2 [68] Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021. 1, 2, 1 [69] Fanqing Meng, Wenqi Shao, Quanfeng Lu, Peng Gao, Kaipeng Zhang, Yu Qiao, and Ping Luo. Chartassistant: A universal chart multimodal language model via chart-to-table pre-training and multitask instruction tuning. In Findings of the Association for Computational Linguistics: ACL 2024, pages 7775–7803, 2024. 6 [70] Nitesh Methani, Pritha Ganguly, Mitesh M. Khapra, and Pratyush Kumar. Plotqa: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020. 1, 2 [71] Jonathan Mi, Wenzhe Tong, Yilin Ma, and Xiaonan Huang. Design of a variable stiffness quasi-direct drive cable-actuated tensegrity robot. IEEE Robotics and Automation Letters, 2025. 7 [72] Erik Nijkamp, Jeffrey A Ruffolo, Eli N Weinstein, Nikhil Naik, and Ali Madani. Progen2: exploring the boundaries of protein language models. Cell systems, 14(11):968–978, 2023. 13 [73] William A Nyberg, Pierre-Louis Bernard, Wayne Ngo, Charlotte H Wang, Jonathan Ark, Allison Rothrock, Gina M Borgo, Gabriella R Kimmerly, Jae Hyung Jung, Vincent Allain, et al. In vivo site-specific engineering to reprogram t cells. Nature, pages 1–10, 2026. 8 [74] Tom JN Obey, Mukesh K Singh, Angelos B Canaj, Gary S Nichol, Euan K Brechin, and Jason B Love. A delocalized mixed-valence dinuclear ytterbium complex that displays intervalence charge transfer. Journal of the American Chemical Society, 146(42):28658–28662, 2024. 7, 8 [75] OpenAI. GPT-5.6 Luna model. https://developers.openai.com/api/docs/models/gpt-5. 6-luna, 2026. Accessed September 4, 2026. 2 [76] OpenAI. GPT-5.6 Sol model. https://developers.openai.com/api/docs/models/gpt-5. 6-sol, 2026. Accessed September 4, 2026. 2 [77] OpenAI. GPT-5.6 Terra model. https://developers.openai.com/api/docs/models/gpt-5. 6-terra, 2026. Accessed September 4, 2026. 2 [78] Mineto Ota, Jeffrey P Spence, Tony Zeng, Emma Dann, Nikhil Milind, Alexander Marson, and Jonathan K Pritchard. Causal modelling of gene effects from regulators to programs to traits. Nature, 650(8101): 399–408, 2026. 13 [79] R Palas, NK Mondal, S Bhattacharya, B Das, and K Das. Removal of arsenic (iii) and arsenic (v) on chemically modified low-cost adsorbent: Batch and column operations. appl. Water Sci, 3:293–309, 2013. 5 15
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[80] Zhangzhi Peng. Ptm-mamba: a ptm-aware protein language model with bidirectional gated mamba blocks. In Proceedings of the 33rd ACM international conference on information and knowledge management, pages 5475–5478, 2024. 6, 11 [81] Tingyu Qu, Mingxiao Li, Tinne Tuytelaars, and Marie-Francine Moens. Ts-llava: Constructing visual tokens through thumbnail-and-sampling for training-free video large language models. arXiv preprint arXiv:2411.11066, 2024. 12 [82] Qwen Team. Qwen3.7-Plus. https://qwen.ai/blog?id=qwen3.7-plus, June 2026. Accessed September 4, 2026. 2 [83] Qwen Team. Qwen3.8-27B model card. https://huggingface.co/Qwen/Qwen3.8-27B, August 2026. Accessed September 4, 2026. 2 [84] Qwen Team. Qwen3.8-Max: A new bar for coding and cowork. https://qwen.ai/blog?id=qwen3. 8, August 2026. Accessed September 4, 2026. 2 [85] Guangbin Ren and Yuchen Zhang. Weighted 𝑙2 theory for the euclidean dirac operator in higher dimensions. arXiv preprint arXiv:2604.04504, 2026. 9 [86] Abdulai Salifu, Branislav Petrusevski, Emmanuel S Mwampashi, Iddi A Pazi, Kebreab Ghebremichael, Richard Buamah, Cyril Aubry, Gary L Amy, and Maria D Kenedy. Defluoridation of groundwater using aluminum-coated bauxite: optimization of synthesis process conditions and equilibrium study. Journal of Environmental Management, 181:108–117, 2016. 9 [87] Clayton Sanford, Daniel Hsu, and Matus Telgarsky. Transformers, parallel computation, and logarithmic depth. arXiv preprint arXiv:2402.09268, 2024. 9 [88] Tom Schlegel, Dennis Breu, and Michael Fleischhauer. Imaginary-time evolution of interacting spin systems in the truncated wigner approximation. arXiv preprint arXiv:2603.03950, 2026. 13 [89] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017. 6, 9, 10 [90] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024. 10 [91] Yanbin Shen, Qinkai Tang, and Yaozhi Luo. A tensegrity-inspired bidirectional quasi-zero stiffness metamaterial for buffering and energy absorption. In Structures, volume 79, page 109438. Elsevier, 2025. 10 [92] Sofia Shkunnikova, Anika Mijakovac, Lucija Sironic, Maja Hanic, Gordan Lauc, and Marina Martinic Kavur. Igg glycans in health and disease: Prediction, intervention, prognosis, and therapy. Biotechnology advances, 67:108169, 2023. 12 [93] Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. Paperbench: Evaluating ai’s ability to replicate ai research. arXiv preprint arXiv:2504.01848, 2025. 1, 2, 1 [94] Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. Saprot: Protein language modeling with structure-aware vocabulary. BioRxiv, pages 2023–10, 2023. 6 [95] Yunfei Su, Guochun Wu, Lei Yao, and Yinghui Zhang. Large time behavior of weak solutions to the inhomogeneous incompressible navier-stokes-vlasov equations in r3. Journal of Differential Equations, 402:361–399, 2024. 7 [96] Zhaochen Su, Jincheng Gao, Hangyu Guo, Zhenhua Liu, Lueyang Zhang, Xinyu Geng, Shijue Huang, Peng Xia, Guanyu Jiang, Cheng Wang, et al. Agentvista: Evaluating multimodal agents in ultra-challenging realistic visual scenarios. arXiv preprint arXiv:2602.23166, 2026. 13
16
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[97] Bambang Suryobroto. Estimation of the biological affinities of seven species of sulawesi macaques based on multivariate analysis of dermatoglyphic pattern types. Primates, 33(4):429–449, 1992. 7 [98] Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai C Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, et al. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. Advances in Neural Information Processing Systems, 37:87310–87356, 2024. 11 [99] Timothy Truong Jr and Tristan Bepler. Poet: A generative model of protein families as sequences-ofsequences. Advances in Neural Information Processing Systems, 36:77379–77415, 2023. 9 [100] Timothy Fei Truong Jr and Tristan Bepler. Understanding protein function with a multimodal retrievalaugmented foundation model. arXiv preprint arXiv:2508.04724, 2025. 13 [101] Antonella Vallenari, Anthony GA Brown, Timo Prusti, Jos HJ De Bruijne, F Arenou, Carine Babusiaux, Michael Biermann, Orlagh L Creevey, Christine Ducourant, Dafydd Wyn Evans, et al. Gaia data release 3-summary of the content and survey properties. Astronomy & Astrophysics, 674:A1, 2023. 9 [102] Hong-Yi Wang, Yu-Qi Li, Qian Wu, and Zhu-Fang Cui. Revisiting p-11 b fusion: Updated cross-sections, reactivity, and energy balance. arXiv preprint arXiv:2601.00241, 2026. 5 [103] Qian Wang, Joo Han Lee, Gregory Nachtrab, Yuan Yuan, Lei Yuan, Wei Qi, Manuel A Mohr, Jing Xiong, Mark A Horowitz, and Xiaoke Chen. Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain. Nature, pages 1–10, 2026. 7, 10 [104] Tianyuan Wang and Mark A Post. A symmetric three degree of freedom tensegrity mechanism with dual operation modes for robot actuation. Biomimetics, 6(2):30, 2021. 8 [105] Ting Wang, Xiaofei Zhu, and Gu Tang. Knowledge-enhanced graph convolutional neural networks for text classification. Journal of Zhejiang University (Engineering Science), 56(2), 2022. 10 [106] Yuyang Wang, Jianren Wang, Zhonglin Cao, and Amir Barati Farimani. Molecular contrastive learning of representations via graph neural networks. Nature Machine Intelligence, 4(3):279–287, 2022. 11 [107] Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian, Daniel Hsu, Jason D Lee, and Denny Wu. Learning compositional functions with transformers from easy-to-hard data. arXiv preprint arXiv:2505.23683, 2025. 9, 10 [108] S Wardani, H Husni, MI Sulaiman, H Desvita, and A Hadi. Simultaneous removal of heavy metals from aqueous solutions by pineapple crown and avocado peel hydrogel composites. 10 [109] Dillon Wong, Kevin P Nuckolls, Myungchul Oh, Biao Lian, Yonglong Xie, Sangjun Jeon, Kenji Watanabe, Takashi Taniguchi, B Andrei Bernevig, and Ali Yazdani. Cascade of electronic transitions in magic-angle twisted bilayer graphene. Nature, 582(7811):198–202, 2020. 5 [110] Shenxi Wu, Haosong Zhang, Xingjian Ma, Shirui Bian, Yichi Zhang, Xi Chen, and Wei Lin. Hyperparameter transfer laws for non-recurrent multi-path neural networks. arXiv preprint arXiv:2602.07494, 2026. 5, 7 [111] Bo Xia, Weimin Zhang, Guisheng Zhao, Xinru Zhang, Jiangshan Bai, Ran Brosh, Aleksandra Wudzinska, Emily Huang, Hannah Ashe, Gwen Ellis, et al. On the genetic basis of tail-loss evolution in humans and apes. Nature, 626(8001):1042–1048, 2024. 13 [112] Renqiu Xia, Song Mao, Xiangchao Yan, Hongbin Zhou, Bo Zhang, Haoyang Peng, Jiahao Pi, Daocheng Fu, Wenjie Wu, Hancheng Ye, et al. Docgenome: An open large-scale scientific document benchmark for training and testing multi-modal large language models. arXiv preprint arXiv:2406.11633, 2024. 1, 2, 1 [113] Xiaomi MiMo Team.
Xiaomi MiMo-V2.5 series. https://mimo.mi.com/docs/en-US/news/ latest/v2.5-open-sourced, 2026. Accessed September 4, 2026. 2
[114] Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, and Dahua Lin. Caprl: Stimulating dense image caption capabilities via reinforcement learning. arXiv preprint arXiv:2509.22647, 2025. 12 17
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
[115] Long Xing, Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jinsong Li, Shuangrui Ding, Weiming Zhang, Nenghai Yu, et al. Scalecap: Inference-time scalable image captioning via dual-modality debiasing. arXiv preprint arXiv:2506.19848, 2025. 9, 12 [116] Lei Xiong, Huaying Yuan, Zheng Liu, Zhao Cao, and Zhicheng Dou. Paperscope: A multi-modal multidocument benchmark for agentic deep research across massive scientific papers. In Findings of the Association for Computational Linguistics: ACL 2026, pages 8015–8040, 2026. 2, 1 [117] Guowei Xu, Peng Jin, Ziang Wu, Hao Li, Yibing Song, Lichao Sun, and Li Yuan. Llava-cot: Let vision language models reason step-by-step. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2087–2098, 2025. 5, 13 [118] Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. Pllava: Parameter-free llava extension from images to videos for video dense captioning. arXiv preprint arXiv:2404.16994, 2024. 12 [119] Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2020. 1, 2 [120] Hao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang, Minghui Liao, Jihao Wu, Wei Chen, and Xiang Bai. Docseeker: Structured visual reasoning with evidence grounding for long document understanding. arXiv preprint arXiv:2604.12812, 2026. 2, 5 [121] Z.ai. GLM-4.6V. https://docs.z.ai/guides/vlm/glm-4.6v, December 2025. Accessed September 4, 2026. 2 [122] Beichen Zhang, Yuhong Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Haodong Duan, Yuhang Cao, Dahua Lin, and Jiaqi Wang. Booststep: Boosting mathematical capability of large language models via improved single-step reasoning. arXiv preprint arXiv:2501.03226, 2025. 5 [123] Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, et al. Group sequence policy optimization. arXiv preprint arXiv:2507.18071, 2025. 7, 10, 13 [124] Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. Publaynet: Largest dataset ever for document layout analysis. In International Conference on Document Analysis and Recognition, 2019. 2 [125] Zhuochao Zhou, Yuhong Liu, Xiaotong Gu, Haowen Zhang, Panan Zhang, Yue Sun, Honglei Liu, Xiaobing Cheng, Yutong Su, Hui Shi, et al. Glycosylation of anti-dsdna igg correlates with organ involvement in treatment-naïve patients with systemic lupus erythematosus. Lupus Science & Medicine, 12(2), 2025. 8, 10, 12
A. Limitations Several limitations should be noted. First, the current benchmark contains 124 underlying questions. Although they are expert-authored and difficulty-screened, a larger question pool would strengthen statistical reliability and permit finer-grained capability analysis. The 496 evaluation instances are four matched variants of these questions and should not be interpreted as 496 independent scientific problems. Second, domain coverage emphasizes computer science, mathematics, and selected natural and biomedical sciences; social sciences, humanities, and several engineering fields remain underrepresented. Third, cross-document tasks currently involve two to four papers, while real literature reviews may span dozens of documents. Fourth, synthetic evidence enables precise control and verification but cannot reproduce every visual convention or failure mode found in naturally occurring papers. Finally, SciDocIR depends on layout parsing, source alignment, and OCR components. Errors in these upstream stages can propagate into both task generation and training-data quality.
18
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Figure 4: Overview of SciDocDataset statistics and the SciDocBench taxonomy. Panel A: benchmark task taxonomy and category proportions. Panel B: total-token distribution per SFT/RL sample. Panel C: image-count distribution per SFT/RL sample.
B. Full SciDocBench Task Taxonomy and Data Statistics Table 3 summarizes the key statistics of SciDocBench. The source documents are drawn from three channels: arXiv (primary), academic journals (including Nature, Journal of the American Chemical Society, Chemistry of Materials, Organic Letters, Biomimetics, ACS family journals, Global Journal of Environmental Science and Management and so on), and web-crawled sources (university repositories and Google Scholar). Table 3: Benchmark construction statistics of SciDocBench.
Statistic
Value
Data Source Source channels Source papers
arXiv, bioRxiv, OpenReview, PubMed 116
Benchmark Scale Underlying questions Evaluation instances Scientific domains Task categories Subtasks
124 496 5 7 19
Document & Question Characteristics Avg. images per paper 14.2 Avg. text tokens per instance 3.9K Evaluation Evaluation methods
3
Figure 4 summarizes the SciDocDataset statistics and the SciDocBench task taxonomy. Panel A shows the benchmark task taxonomy and category proportions. Panel B shows the total-token distribution per SFT/RL sample, and Panel C shows the image-count distribution per SFT/RL sample. The distributions confirm that the data is not a collection of short single-image QA pairs: it includes variable-length contexts and multi-image examples reflective of realistic scientific-document workflows.
19
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 4: Full SciDocBench task taxonomy. The benchmark contains seven capability groups and nineteen subtasks. ID
Task
A. Document Perception & Structure A1 Layout Parsing & Reading Order Recovery A2 A3
Element Association & Evidence Localization Citation Role Analysis
A4 A5
Dataset Lineage Analysis Figure Analysis
B. Scientific Information Extraction B1 Notation & Definition Extraction B2 Method & Experiment Extraction B3
Equation Dependency Graph Construction
C. Evidence Alignment & Verification C1 Numerical & Logical Consistency Checking C2 C3
Figure–Table Consistency Verification Evidence-Grounded Reasoning
D. Cross-Document Understanding D1 Cross-Document Formula Comparison D2 Cross-Document Dataset Comparison D3 Cross-Document Result Integration E. Reconstruction & Execution E1 Figure/Table Data Reconstruction E2 Figure-to-Code Translation F. Paper–Code Alignment F1 Paper-to-Code Localization F2 Pseudocode-to-Code Grounding G. Dataset Understanding G1 Dataset Source Tracing & Classification
Evaluation focus Recover local reading flow in two-column PDFs with floating figures, footnotes, and cross-page continuation. Given a claim, locate its supporting figure, table, equation, appendix, page, or block. Identify whether cited work acts as background, method basis, data resource, baseline, result support, critique, or extension. Recover the source, composition, filtering, derivation, and transformation chain of datasets. Interpret scientific figures, subfigures, legends, axes, annotations, and visual evidence supporting figure-level claims. Build notation indices containing symbols, types, definitions, contexts, and first appearances. Extract datasets, metrics, baselines, instruments, configurations, hyperparameters, and experiment stages. Convert mathematical derivations into equation nodes and dependency edges. Verify whether claims match table values, figure values, rankings, percentages, and derived quantities. Detect whether plotted positions, labels, captions, and tables express the same scientific fact. Answer questions using document evidence while avoiding unsupported free-form reasoning. Compare formula variants, assumptions, and generalization paths across papers. Identify dataset intersections, source overlap, or near-duplicate data usage across papers. Merge results from multiple papers into a canonical schema with provenance. Recover data points, structured tables, or plotting code from figures and tables. Convert scientific figures, diagrams, architectures, or flowcharts into runnable or structured code. Locate modules, scripts, or functions in a repository corresponding to paper descriptions. Complete implementation details by aligning pseudocode, derivations, and source code. Infer dataset sources, families, and provenance classes from fields, structures, samples, and source patterns.
20
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 5: The source documents of all benchmark questions in Categories A1 to A2. Subtask
A1
A2
ID
Subject
Source Document(s)
Image Cnt
1
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
6
Revisiting p-11 B Fusion: Updated Cross-sections,
7
Nuclear Theory
13
Computer Science
Depth Anything 3: Recovering the Visual Space from Any Views [55]
32
51
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
4
52
Computer Science
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
5
53
Medicine
QuarkMedSearch: A Long-Horizon Deep Search Agent for Exploring Medical Intelligence [56]
5
54
Computer Science
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks [110]
5
55
Computer Science
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding [120]
5
115
Computer Science
Depth Anything 3: Recovering the Visual Space from Any Views [55]
32
2
Computer Science
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step [117]
17
8
Environment
Removal of arsenic (III) and arsenic (V) on chemically modified low-cost adsorbent: batch and column operations [79]
17
14
Computer Science
Depth Anything 3: Recovering the Visual Space from Any Views [55]
32
56
Computer Science
Booststep: Boosting mathematical capability of large language models via improved single-step reasoning [122]
13
76
Physics
Cascade of electronic transitions in magic-angle twisted bilayer graphene [109]
14
86
Biology
Inferring the internal structure of groups through the integration of statistical learning and causal reasoning [21]
16
106
Psychology
Inferring the internal structure of groups through the integration of statistical learning and causal reasoning [21]
16
Reactivity, and Energy Balance [102]
9
21
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 6: The source documents of all benchmark questions in Categories A3 to A4. Subtask
ID
Subject
Source Document(s)
Image Cnt
3
Computer Science
Proximal Policy Optimization Algorithms [89]
12
9
Biology
ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [46]
15
15
Computer Science
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning [40]
10
77
Physics
Quantum computing: A taxonomy, systematic review and future directions [31]
24
103
Computer Science
Quantum computing: A taxonomy, systematic review and future directions [31]
24
114
Biology
ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [46]
15
119
Computer Science
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning [40]
10
121
Computer Science
Proximal Policy Optimization Algorithms [89]
12
10
Biology
SaProt: Protein Language Modeling with Structure-Aware Vocabulary [94]
22
16
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
6
17
Computer Science
ChartAssistant: A Universal Chart Multimodal Language Model via Chart-to-Table Pre-training and Multitask Instruction Tuning [69]
29
57
Computer Science
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments [32]
1
Biology
Compressing the collective knowledge of ESM into a single protein language model [25], ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [46], PTM-Mamba: a PTM-aware protein language model with bidirectional gated Mamba blocks [80]
14
107
Biology
Compressing the collective knowledge of ESM into a single protein language model [25], ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [46]
14
117
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
14
A3
A4
87
22
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 7: The source documents of all benchmark questions in Categories A5 to B1. Subtask
A5
ID
Subject
Source Document(s)
Image Cnt
58
Biology
Estimation of the biological affinities of seven species of Sulawesi macaques based on multivariate analysis of dermatoglyphic pattern types [97]
21
78
Chemistry
A Delocalized Mixed-Valence Dinuclear Ytterbium Complex That Displays Intervalence Charge Transfer [74]
7
79
Computer Science
Biologically Plausible Learning via Bidirectional Spike-Based Distillation [64]
3
80
Physics
Design of a variable stiffness quasi-direct drive cable-actuated tensegrity robot [71]
8
88
Biology
Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain [103]
1
89
Physics
New limits on the Pauli forbidden transitions in 12C nuclei obtained with the complete Borexino dataset [9]
13
90
Physics
Statics of integrated origami and tensegrity systems [65]
13
104
Computer Science
Biologically Plausible Learning via Bidirectional Spike-Based Distillation [64]
3
108
Physics
12 C nuclei obtained with the complete Borexino dataset
4
Computer Science
Group Sequence Policy Optimization [123]
7
11
Physics
Coherent spectroscopy with a single antiproton spin [49]
11
12
Maths
Large time behavior of weak solutions to the inhomogeneous incompressible Navier-Stokes-Vlasov equations in R3 [95]
23
18
Computer Science
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks [110]
23
81
Computer Science
Packed-ensembles for efficient uncertainty estimation [50]
10
116
Computer Science
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks [110]
23
118
Computer Science
Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural Networks [110]
23
New limits on the Pauli forbidden transitions in
B1
[9]
13
23
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 8: The source documents of all benchmark questions in Category B2. Subtask
B2
ID
Subject
Source Document(s)
Image Cnt
5
Computer Science
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
20
19
Medicine
Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [125]
9
20
Biology
LLM-assisted systematic review of large language models in clinical medicine [19]
17
21
Biology
In vivo site-specific engineering to reprogram T cells [73]
31
59
Maths
Quantum linear system solver based on time-optimal adiabatic quantum computing and quantum approximate optimization algorithm [2]
28
69
Chemistry
Axial chirality-induced rigidification in aminoboranes enhances persistent room-temperature phosphorescence and circularly polarized luminescence [27]
11
82
Chemistry
A Delocalized Mixed-Valence Dinuclear Ytterbium Complex That Displays Intervalence Charge Transfer [74]
5
91
Physics
A symmetric three degree of freedom tensegrity mechanism with dual operation modes for robot actuation [104]
10
92
Physics
Impact-resistant, autonomous robots inspired by tensegrity architecture [42]
9
113
Computer Science
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
20
24
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 9: The source documents of all benchmark questions in Categories B3 to C2. Subtask
B3
C1
C2
ID
Subject
Source Document(s)
Image Cnt
6
Computer Science
Proximal Policy Optimization Algorithms [89]
12
22
Physics
Observation of non-Abelian topological acoustic semimetals and their phase transitions [39]
41
23
Biology
Poet: A generative model of protein families as sequences-of-sequences [99]
37
70
Maths
Weighted 𝐿2 theory for the Euclidean Dirac operator in higher dimensions [85]
17
24
Computer Science
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
17
27
Maths
Learning compositional functions with transformers from easy-to-hard data [107]
7
28
Maths
Dynamic metastability in the self-attention model [30]
12
60
Environment
Defluoridation of groundwater using aluminum-coated bauxite: optimization of synthesis process conditions and equilibrium study [86]
3
61
Computer Science
Exact recovery in the stochastic block model [1]
12
62
Maths
Gaussian differential privacy [26]
13
71
Computer Science
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning [40]
10
72
Computer Science
A unified perspective on the dynamics of deep transformers [15]
17
73
Computer Science
Transformers, parallel computation, and logarithmic depth [87]
10
74
Maths
Benign overfitting in linear regression [8]
6
83
Physics
Magic-angle graphene superlattices: a new platform for unconventional superconductivity [13]
18
84
Computer Science
Provable failure of language models in learning majority boolean logic via gradient descent [17]
10
25
Computer Science
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
1
29
Physics
Gaia data release 3-summary of the content and survey properties [101]
23
63
Computer Science
Scalecap: Inference-time scalable image captioning via dual-modality debiasing [115]
1
64
Biology
Neural sequences underlying directed turning in Caenorhabditis elegans [48]
6
93
Computer Science
Improved baselines with visual instruction tuning [58]
15
25
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 10: The source documents of all benchmark questions in Categories C3 to D1. Subtask
ID
Subject
Source Document(s)
Image Cnt
26
Computer Science
Semi-supervised microblog text sentiment classification based on global feature graph [28], Knowledge-enhanced graph convolutional neural networks for text classification [105]
16
30
Medicine
Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [125]
9
31
Environment
Simultaneous removal of heavy metals from aqueous solutions by pineapple crown and avocado peel hydrogel composites [108]
24
32
Maths
The Euler Stratification for P1 × P1 × P𝑛 [36]
31
33
Maths
Computation and sampling for Schubert specializations [3]
33
34
Maths
Learning compositional functions with transformers from easy-to-hard data [107]
11
65
Maths
Quantum linear system solver based on time-optimal adiabatic quantum computing and quantum approximate optimization algorithm [2]
28
66
Biology
Neural sequences underlying directed turning in Caenorhabditis elegans [48]
1
75
Biology
Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain [103]
1
85
Chemistry
Giant coercivity and high magnetic blocking temperatures for N23− radical-bridged dilanthanide complexes upon ligand dissociation [22]
9
94
Physics
A tensegrity-inspired bidirectional quasi-zero stiffness metamaterial for buffering and energy absorption [91]
23
105
Chemistry
Giant coercivity and high magnetic blocking temperatures for N3− 2 radical-bridged dilanthanide complexes upon ligand dissociation [22]
9
Computer Science
Proximal Policy Optimization Algorithms [89], Deepseekmath: Pushing the limits of mathematical reasoning in open language models [90], Group Sequence Policy Optimization [123], Minimax-m1: Scaling test-time compute efficiently with lightning attention [16]
9
Physics
Approximating quantum many-body wave functions using artificial neural networks [11], Quantum entanglement in neural network states [23], Solving the quantum many-body problem with artificial neural networks [14]
8
C3
D1
35
67
26
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 11: The source documents of all benchmark questions in Category D2. Subtask
ID
Subject
Source Document(s)
Image Cnt
36
Computer Science
Eagle 2.5: Boosting long-context post-training for frontier vision-language models [18], Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale [34]
2
Computer Science
Cambrian-1: A fully open, vision-centric exploration of multimodal llms [98], Llava-onevision: Easy visual task transfer [51], Mammoth-vl: Eliciting multimodal reasoning with instruction tuning at scale [34]
11
Chemistry
Strategies for pre-training graph neural networks [37], Geometry-enhanced molecular representation learning for property prediction [29], Molecular contrastive learning of representations via graph neural networks [106]
5
Biology
ProtFlow: Flow Matching-based Protein Sequence Design with Comprehensive Protein Semantic Distribution Learning and High-quality Generation [46], PTM-Mamba: a PTM-aware protein language model with bidirectional gated Mamba blocks [80], Compressing the collective knowledge of ESM into a single protein language model [25]
10
Chemistry
Molecular Contrastive Learning of Representations via Graph Neural Networks [106], ChemRL-GEM: Geometry Enhanced Molecular Representation Learning for Property Prediction [29], Strategies for Pre-Training Graph Neural Networks [37]
42
Computer Science
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs [98], LLaVA-OneVision: Easy Visual Task Transfer [51], MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale [34]
11
D2 37
95
96
109
120
27
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 12: The source documents of all benchmark questions in Category D3. Subtask
D3
ID
Subject
Source Document(s)
Image Cnt
38
Computer Science
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery [12], Spatialladder: Progressive training for spatial reasoning in vision-language models [53]
39
39
Computer Science
Ts-llava: Constructing visual tokens through thumbnail-andsampling for training-free video large language models [81], Pllava: Parameter-free llava extension from images to videos for video dense captioning [118]
32
47
Medicine
Glycosylation of anti-dsDNA IgG correlates with organ involvement in treatment-naive patients with systemic lupus erythematosus [125], IgG glycans in health and disease: Prediction, intervention, prognosis, and therapy [92]
44
97
Computer Science
Caprl: Stimulating dense image caption capabilities via reinforcement learning [114], Scalecap: Inference-time scalable image captioning via dual-modality debiasing [115]
2
Chemistry
High voltage cycling stability of LiF-coated NMC811 electrode [61], Long-term cyclability of NCM-811 at high voltages in lithium-ion batteries: an in-depth diagnostic study [54], Enhanced Cycling Stability of NCM811 Cathodes at High C-Rates and Voltages via LiMTFSI-Based Polymer Coating [44]
6
Computer Science
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing [115], CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning [114]
15
Chemistry
Long-Term Cyclability of NCM-811 at High Voltages in Lithium-Ion Batteries: an In-Depth Diagnostic Study [54], Enhanced Cycling Stability of NCM811 Cathodes at High C-Rates and Voltages via LiMTFSI-Based Polymer Coating [44], High Voltage Cycling Stability of LiF-Coated NMC811 Electrode [61]
29
Medicine
IgG glycans in health and disease: Prediction, intervention, prognosis, and therapy [92], Fucosylation of anti-dsDNA IgG1 correlates with disease activity of treatment-naïve systemic lupus erythematosus patients [35]
13
98
110
111
123
28
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 13: The source documents of all benchmark questions in Categories E, F, and G. Subtask
E1
E2
F1
F2
G1
ID
Subject
Source Document(s)
Image Cnt
40
Computer Science
Group Sequence Policy Optimization [123]
7
41
Computer Science
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step [117]
17
48
Physics
Stairway Codes: Floquetifying Bivariate Bicycle Codes and Beyond [38]
18
49
Physics
Imaginary-time evolution of interacting spin systems in the truncated Wigner approximation [88]
7
68
Biology
Neural sequences underlying directed turning in Caenorhabditis elegans [48]
1
99
Biology
Causal modelling of gene effects from regulators to programs to traits [78]
1
101
Biology
Causal modelling of gene effects from regulators to programs to traits [78]
1
112
Biology
Causal modelling of gene effects from regulators to programs to traits [78]
3
124
Physics
Imaginary-time evolution of interacting spin systems in the truncated Wigner approximation [88]
7
42
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
2
50
Physics
QFlowNet: Fast, Diverse, and Efficient Unitary Synthesis with Generative Flow Networks [47]
7
43
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24]
6
122
Computer Science
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning [59], AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios [96], MM-IFEngine: Towards Multimodal Instruction Following [24], Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
34
44
Computer Science
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step [117]
17
100
Biology
On the genetic basis of tail-loss evolution in humans and apes [111]
25
102
Biology
On the genetic basis of tail-loss evolution in humans and apes [111]
25
45
Computer Science
MM-IFEngine: Towards Multimodal Instruction Following [24], Agentvista: Evaluating multimodal agents in ultra-challenging realistic visual scenarios [96], TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning [59], Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning [60]
12
46
Biology
Understanding protein function with a multimodal retrieval-augmented foundation model [100], Progen2: exploring the boundaries of protein language models [72]
3
29
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
C. Benchmark Annotation and Verification Protocol Benchmark construction uses two separate interfaces and two annotator roles. The authoring interface records the source PDF, expert identifier, scientific domain, capability group, subtask, finalized question, ground-truth answer, and evidence locator. The evidence locator identifies the page, section, figure, table, equation, appendix, or other document object from which the answer can be verified. Authors then run the finalized item once on each of the four designated screening models, paste the responses without editing, and assign a binary correctness label to every response. An item can be submitted only when all required fields and four screening records are present and at least one screened model is incorrect. If the question, ground truth, or evidence locator changes, the four responses are collected again. Figure 5 shows the authoring interface. Instructions for Question Authors. Read the supplied paper and create one scientifically grounded question that is answerable from the paper but challenging for current multimodal models. 1. Upload the source PDF; provide the finalized prompt and ground-truth answer; and identify the page, section, figure, table, equation, appendix, or other scientific object that supports the answer. 2. Write a clear and self-contained prompt with an unambiguous output format and a stable answer that can be verified directly from the supplied document. Do not rely on unstated external knowledge. 3. Run the finalized question once on each designated screening model. Paste every response verbatim and label it Correct or Incorrect according to the ground truth. Do not edit, merge, or selectively rerun individual responses. 4. Submit the item only when every required field and all four screening records are complete and at least one valid screening response is incorrect. If the prompt, ground truth, or evidence locator changes, collect all four responses again. Every submitted item is subsequently inspected by a reviewer who did not author the question. The reviewer loads the immutable submission and verifies four conditions: the question is answerable from the supplied PDF, the wording and requested output format are unambiguous, the ground truth is correct, and at least one screening model is incorrect. The reviewer also audits every response label and can reverse an author label when it does not agree with the ground truth. An item is accepted only when all four checks pass. Items with correctable defects are returned for revision, while items with unstable evidence or irreparable ambiguity are rejected. Figure 6 shows the independent-verification interface. The annotation system stores each submission as an atomic record containing a schema version, item identifier, author and reviewer identifiers, timestamps, task metadata, the question and ground truth, the evidence locator, the source-PDF filename and SHA-256 digest, the four verbatim responses, the author labels, reviewer label audits, and the final decision. The PDF digest prevents accidental substitution of the source document after screening. Screening responses are used only to remove items already solved by all designated models; they are not used to construct the ground truth. This separation preserves expert evidence as the basis of correctness while retaining an auditable record of question difficulty and reviewer agreement.
30
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Figure 5: Expert question-authoring interface. The interface presents the source PDF and rendered pages alongside the task metadata, question, ground truth, evidence locator, and verbatim screening-model responses with author-assigned correctness labels. The populated fields constitute a representative interface record used to illustrate the annotation procedure and are not additional leaderboard results.
31
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Figure 6: Independent-verification interface. Reviewers inspect the same PDF, task metadata, evidence locator, question, ground truth, and four screening records; audit each correctness label; complete the acceptance checklist; and record an Accept, Revise, or Reject decision.
32
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
C.1. LLM-as-a-Judge Reliability Audit We assess the reliability of semantic scoring on 100 response-level instances drawn from accepted LLM-as-ajudge evaluation runs. The audit set contains 25 instances from each of the four matched evaluation settings, covers 11 evaluated models, and includes questions from all seven capability groups. Sampling is approximately balanced across the judge’s initial incorrect, partial, and correct score bands. Deterministic rule-based and execution-based evaluations are excluded because their correctness is established by programmatic scorers rather than semantic judgment. The human expert was shown only the question, reference answer, and anonymized model response. Model identity, the GPT-5.4-mini score, and the judge rationale were hidden until all 100 ratings had been submitted. The expert assigned one of three scores: 0 for an incorrect response, 0.5 for a partially correct response, and 1 for a fully correct response. For the ordinal comparison, continuous GPT-5.4-mini and Codex scores are mapped using the same strict endpoint rule: 0 is incorrect, 1 is correct, and every value strictly between 0 and 1 is partially correct. The audit therefore measures consistency in applying the reference-based scoring rubric; it does not constitute an independent re-annotation of each source document. Table 14: Pairwise agreement on the 100-instance semantic-judge audit. Binary results threshold the original scores at 0.5. QWK denotes quadratic-weighted Cohen’s 𝜅, and MAE denotes mean absolute score difference. Rater pair Human expert vs. GPT-5.4-mini Codex vs. GPT-5.4-mini Human expert vs. Codex
Three-level exact
QWK
Binary exact
Binary 𝜅
MAE
92.0% 89.0% 93.0%
0.866 0.818 0.872
70.0% 80.0% 74.0%
0.355 0.593 0.405
0.169 0.128 0.145
The average pairwise three-level exact agreement is 91.3%. To summarize all three rating sources jointly, we compute ordinal Krippendorff ’s 𝛼 over the expert, GPT-5.4-mini, and Codex labels. The resulting 𝛼 is 0.854, with a 95% confidence interval of [0.757, 0.924] obtained from 10,000 bootstrap resamples of the audit instances. The stronger ordinal agreement than binary agreement reflects the intended use of partial credit: thresholding at 0.5 converts small differences within the intermediate score range into categorical disagreements.
33
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Figure 7: Complete blind expert-review interface for the English audit item HJ067. The interface presents the item metadata, question, reference answer, anonymized model response, three-level rubric, rating controls, navigation controls, and the agreement summary revealed after all 100 ratings were completed. Model identity and the GPT-5.4-mini judgment remain hidden during rating. Browser chrome and unrelated tabs are omitted from the capture.
D. SciDocIR Schema Figure 8 presents the simplified schema of SciDocIR. The representation stores document-level metadata, page-level metadata, block-level records, relations, and task-specific metadata. Optional fields are populated only when they are relevant to task generation or reward verification.
34
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Record SciDocIR document : {doc_id, source_dir, num_pages, provenance} categories: List<{id, name, supercategory}> pages : List<Page> blocks : List<Block> relations : List<Relation> stats : {num_blocks, num_relations, category_counts} End Record Record Page page_index : Integer width, height : Float image_path : String End Record Record Block block_id, page_index : Integer bbox : {x1, y1, x2, y2, width, height} category : {id, name, supercategory} content : {source_code, format, plain_text, latex} structure : {previous_block, parent_block, next_block} references : {labels, outgoing_ref_labels, incoming_ref_from} metadata : Metadata End Record Record Metadata title_level caption_text caption_target figure_info table_structure equation_latex nearby_block_ids task_fields End Record
: Integer? : String? : {type, block_ids}? : {crop_path, visual_ocr, chart_type, subfigures}? : TableStructure? : String? : List<BlockId> : Dict<String, Any>
Record Relation type : "adj" | "sub" | "peer" | "identical" | "continuation" from : BlockId to : BlockId End Record
Figure 8: A simplified schema of SciDocIR. Optional fields are filled only when useful for generation or verification.
35
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
E. Training-Data Directions and Verifiable Subtasks Table 15 lists the eight construction directions and fourteen verifiable subtasks instantiated in SciDocDataset. The list reflects the final data pipeline rather than the broader benchmark taxonomy. Table 15: Training-data construction directions and their instantiated subtasks. Direction
Training subtasks
Layout and reading-flow recovery
Block role classification; parent–child linking; reading-order prediction; next-hop reading-target prediction. Table logic consistency Table logic consistency check. Cross-document table integration Real cross-document table merge; single-document simulated table merge. Chart readout and recovery Single-point chart readout; multi-point or series readout. Chart visual consistency Chart visual consistency. Cross-document dataset comparison Cross-document dataset intersection. Notation understanding Notation extraction; symbol ambiguity disambiguation. Citation-role analysis A3-lite citation-role classification.
E.1. Construction Sources and Model Roles The seed instruction is not an assertion that every sample is generated end to end by an LLM. It is a task-specific contract governing the part assigned to GPT-5.4. Deterministic labels, values, alignments, and perturbation records are locked before model invocation whenever possible. GPT-5.4 is used for linguistic naturalization, bounded candidate generation, semantic classification, or solvability checks. Tables 16 and 17 identify the records supplied to each template and the authority used for ground truth. Table 16: Construction sources and GPT-5.4 roles for subtasks 1–7. “None” denotes no external dataset beyond the SciDocIR paper collection. Subtask
SciDocIR or document records
External or synthetic source GPT-5.4 role and GT authority
Light question rewriting. The layout label remains Page image; target page index, bounding None box, type, and layout annotation the locked GT. Parent–child linking Block text, page and bounding box; head- None Question rewriting and optional plausibility check. ing level; caption pairs; footnote and conThe verified relation remains GT. tinuation relations Order and reading annotations; block None Question rewriting only. GT follows the stored order Reading-order prediction pages and boxes; section and discourse and layout rules. roles Current and successor blocks; caption pairs; None Question rewriting and optional next-hop Next-hop reading target figure/table references; cross-page and applausibility check. The stored successor remains GT. pendix links Table logic consistency Table LaTeX, normalized JSON, caption, None Generates a natural corrupted statement, repair, and distractors, then self-checks them. Rules lock headers, cells, citing context, page, and box the perturbed field and answer. Real cross-document table Table schemas, rows, columns, values, cap- None Polishes context and question. Rules align merge tions, and nearby text from different papers compatible schemas and compute the merged GT. Simulated table merge One real table and its context, split into two Programmatic table splitting Adjusts table LaTeX, local prose, and question partially overlapping tables with permuted and a simulated second pa- without changing locked cells. Rules compute the columns per page complete merged GT. Block role classification
36
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
Table 17: Construction sources and GPT-5.4 roles for subtasks 8–14. Subtask
SciDocIR or document records
External or synthetic source GPT-5.4 role and GT authority
Single-point chart readout
Figure and caption blocks; page, box, and PlotQA chart, axes, series la- Naturalizes the question and performs repeated nearby text used for page replacement bels, and source values visual solvability checks. PlotQA values remain GT. Multi-point or series readout Same figure placement and context records Complete PlotQA series data Rewrites the question and checks the returned list. The source series remains GT. as the single-point task Chart visual consistency Figure position, caption, and page context PlotQA values; Matplotlib Rewrites the question and checks solvability. Rules used to reinsert a controlled chart radar, line, scatter, and bar move one plotted value and retain the original renderers annotation; the corruption log defines GT. Dataset intersection Dataset-like entities mined from text, ta- Internal lexicon of approxi- Checks name plausibility and polishes page prose. bles, and captions when available; simu- mately 50 CS, AI, mathemat- Exact set intersection defines GT. lated two-paper text ics, and natural-science resources; alias and composition pools Notation extraction Final batch does not require real SciDocIR Mathematics, machine- May polish surrounding prose while formulas and records learning, and physics definitions stay locked. The saved symbol table formula templates; XeLaTeX defines GT. page rendering Symbol ambiguity disam- Final batch does not require real SciDocIR Formula and context tem- May naturalize local contexts without changing biguation records plates; one-column, two- definitions. The rule log for each symbol occurrence column, and spanning lay- defines GT. outs A3-lite citation role Block LaTeX and text; real citation keys; No external bibliography Rules propose a role from lexical cues and section section, page, block identifier, and local database context; GPT-5.4 classifies the observed use. context Ambiguous or disagreeing cases are rejected.
E.2. Seed Instruction Templates The templates below show the shared generation contract and the task-specific seed appended to it. Anglebracketed expressions are populated by the retrieval, synthesis, or rendering program. A locked field cannot be altered by the model. The model-facing training instance contains the generated question and its associated document input; construction metadata and source identifiers are retained only for generation, verification, and auditing. You construct one scientific-document training instance from supplied records. Inputs may include: - IR_RECORDS: selected SciDocIR pages, blocks, relations, and metadata. - DOCUMENT_VIEW: the page image, crop, or rendered document shown to the learner. - STRUCTURED_SOURCE: source table, chart data, formula template, or entity lists. - CONSTRUCTION_RECORD: locked labels, values, mappings, perturbations, and tolerances. Use only the supplied material. Do not invent a scientific fact, citation, identifier, value, unit, relation, or source. Preserve every field marked LOCKED. Write one self-contained question that is answerable from DOCUMENT_VIEW and that states the required answer format. Do not reveal the answer or hidden construction metadata in the question. Return JSON only: { "question": "<user-facing instruction>", "answer": <task-specific schema>, "source_ids": ["<records used for audit>"], "quality_check": {"answerable": true, "single_interpretation": true} } If a locked gold answer is supplied, copy it exactly. If the task asks for a bounded semantic label, select only from the supplied ontology. Return {"reject": "<reason>"} when the evidence is incomplete, ambiguous, illegible, or inconsistent with the construction record.
37
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
E.2.1. Layout and Reading-Flow Recovery 1. Block role classification. Inputs: PAGE_IMAGE, TARGET_BLOCK {page_index, bbox, type}, LAYOUT_ANNOTATION, ROLE_LABELS. Mark TARGET_BLOCK visibly in the page view. Rewrite the base instruction as a concise classification question asking for exactly one label from ROLE_LABELS. Use nearby layout only as context and do not mention the stored type. Copy LAYOUT_ANNOTATION[TARGET_BLOCK] as the answer. Reject a crop containing several inseparable semantic blocks or a label absent from ROLE_LABELS. Answer schema: "<one label from ROLE_LABELS>"
2. Parent–child linking. Inputs: BLOCK_A, BLOCK_B, local page context, heading levels, caption pairs, footnote links, and continuation relations; LOCKED_RELATION. Present A and B with visible labels and enough context to determine their structural relation. Ask whether B is the content targeted by A as a caption, a continuation of A, or unrelated. Use exactly caption_target, continuation, and none. Naturalize the wording without changing block contents. Copy LOCKED_RELATION as the answer and reject pairs that require unseen context or admit more than one relation. Answer schema: "caption_target" | "continuation" | "none"
3. Reading-order prediction. Inputs: PAGE_IMAGE, CANDIDATE_BLOCKS with page_index and bbox, ORDER_ANNOTATION, READING_ANNOTATION, and discourse roles; LOCKED_ORDER. Assign visible labels A, B, C, ... in the supplied shuffled order. Ask for the order in which the selected blocks should be read. Ensure that the view retains columns, headings, captions, and any cross-page cue needed to solve the task. Copy LOCKED_ORDER, expressed as a label sequence, as the answer. Reject cyclic, disconnected, or visually ambiguous selections. Answer schema: "<sequence such as A>B>D>C>"
4. Next-hop reading-target prediction. Inputs: CURRENT_BLOCK, CANDIDATE_BLOCKS, verified successor, caption pairs, textual figure/table references, cross-page relations, and appendix links; LOCKED_NEXT_LABEL. Label two to six candidates A through F. Ask which block should be read immediately after CURRENT_BLOCK in order to continue the document or follow the explicit reference. Keep plausible distractors from the same local context. Copy LOCKED_NEXT_LABEL as the answer. Reject the sample if two candidates are reasonable next hops or the required link is not visible. Answer schema: "<one candidate letter>"
E.2.2. Table Logic and Integration 5. Table logic consistency. Inputs: TABLE_LATEX, TABLE_JSON, caption, headers, cells, nearby citing text, and PERTURBATION_RECORD with one locked change. Using the locked perturbation, write one natural but false statement about a value, entity, rank, comparison direction, or metric direction. Also write its minimal correction and three plausible statements that are exactly supported by the table. Shuffle the four statements, ask which one is inconsistent, and report the corresponding letter. Independently check every option against TABLE_JSON. Do not create another inconsistency. Answer schema: {"choice": "A|B|C|D", "repair": "<corrected statement>"}
38
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
6. Real cross-document table merge. Inputs: TABLE_A and TABLE_B from different papers, captions, nearby text, SCHEMA_MAP, and LOCKED_MERGE. Write a concise cross-paper question asking the learner to merge the compatible rows and columns under the canonical schema in SCHEMA_MAP. Preserve method names, dataset splits, metric direction, units, and paper provenance exactly. Do not request a comparison that the mapping does not support. Copy LOCKED_MERGE as the answer; rules, not the language model, determine all aligned keys and values. Answer schema: {"columns": ["..."], "rows": [{"row_key": "...", "values": {...}, "sources": {...}}]}
7. Single-document simulated table merge. Inputs: one real SOURCE_TABLE, SPLIT_RECORD, two simulated paper contexts, and LOCKED_MERGE. The construction program has split SOURCE_TABLE into TABLE_A and TABLE_B with different column orders and partially overlapping rows. Improve the two LaTeX tables and their surrounding prose so they read as independent paper excerpts, but preserve every locked header and cell. Ask for a complete merge with duplicate rows reconciled by SPLIT_RECORD. Copy LOCKED_MERGE as the answer and reject any rewrite that changes a value, unit, or row identity. Answer schema: {"columns": ["..."], "rows": [{"row_key": "...", "values": {...}, "sources": {...}}]}
E.2.3. Chart Readout and Visual Consistency 8. Single-point chart readout. Inputs: a PlotQA chart and source data, TARGET_POINT {series, x, value, unit}, target IR figure bbox, caption, and nearby paper text. After the chart is placed in the paper page, write a natural question identifying one unambiguous point by series and x-axis condition. Ask for its value and unit and state the allowed precision when needed. Copy TARGET_POINT.value and unit as the answer. Perform repeated visual checks that the point, axes, and legend are legible; reject overlapping, clipped, or uncertain points. Answer schema: {"series": "<name>", "x": "<condition>", "value": <number>, "unit": "<unit>"}
9. Multi-point or series readout. Inputs: a PlotQA chart, FULL_SERIES_DATA, target IR figure bbox, caption, and nearby paper text; LOCKED_TARGETS. Write a question asking for either several specified points or one complete short series in displayed x-axis order. State the unit and numeric precision. Preserve the locked series and x labels exactly. Copy the requested values from FULL_SERIES_DATA and visually verify the complete list after rendering. Reject stacked, occluded, discontinuous, or dual-axis cases without an explicit unambiguous mapping. Prefer single-point conversion if the full series is not reliably readable. Answer schema: {"series": "<name>", "x": ["..."], "values": [<numbers>], "unit": "<unit>"}
10. Chart visual consistency. Inputs: PlotQA source values, Matplotlib chart rendering, IR figure bbox, caption and page context, and CORRUPTION_RECORD for exactly one moved plotted value. The rendered chart contains one visual error while its label, caption, or surrounding text preserves the correct value. Write a question asking for the affected series or entity, x-condition, correct value, and visually represented value. Do not reveal which mark was moved. Copy all fields from CORRUPTION_RECORD. Check that every other mark is consistent and that both values are visually recoverable within TOLERANCE. Answer schema: {"entity": "<series or entity>", "condition": "<x-label>", "correct_value": <number>, "visual_value": <number>, "unit": "<unit>"}
39
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
E.2.4. Cross-Document Dataset Intersection 11. Cross-document dataset intersection. Inputs: DATASET_LEXICON, ALIAS_MAP, two simulated paper records or IR-mined dataset mentions, and LOCKED_DATASET_SETS. Check that each selected name plausibly denotes a dataset or scientific resource. Write or polish two short paper-like passages that mention the locked dataset sets without adding another dataset name. Ask for the resources common to both documents and require canonical names. Compute the answer by exact set intersection after ALIAS_MAP normalization. Reject unresolved family/version distinctions or prose that leaks the intersection explicitly. Answer schema: {"intersection": ["<canonical dataset names>"]}
E.2.5. Notation Understanding 12. Notation extraction. Inputs: GENERATED_FORMULAS, DEFINITION_SENTENCES, SAVED_SYMBOL_TABLE, and a XeLaTeX page template. Real IR input is not required. Polish only the unlocked surrounding prose while preserving every formula, symbol, definition, and first-occurrence position. Ask for all symbols explicitly defined in the displayed scope, their definitions, and their first visible locations. Do not include symbols that are merely used. Copy SAVED_SYMBOL_TABLE as the answer and reject a rendering in which a formula or definition is clipped or separated from its scope. Answer schema: {"notations": [{"symbol": "<LaTeX>", "definition": "<text>", "first_location": "<visible location>"}]}
13. Symbol ambiguity disambiguation. Inputs: repeated TARGET_SYMBOL, OCCURRENCE_RECORDS with distinct definitions, formula and context templates, and a one-column, two-column, or spanning page layout. Naturalize only unlocked context sentences. Preserve each occurrence of TARGET_SYMBOL and its locally defined meaning. Ask the learner to map every labeled occurrence to the correct definition using its paragraph, formula, or section context. Do not imply that the symbol has one global meaning. Copy the occurrence-to-definition mapping from OCCURRENCE_RECORDS and reject contexts that do not uniquely disambiguate every occurrence. Answer schema: {"symbol": "<LaTeX>", "occurrences": [{"label": "A", "definition": "..."}]}
E.2.6. Citation-Role Analysis 14. A3-lite citation-role classification. Inputs: real block LaTeX and text, CITATION_KEY, section name, page, block id, local citing context, ROLE_ONTOLOGY, and HEURISTIC_ROLE. Classify how CITATION_KEY is used in the supplied text, not what the cited paper is generally about. Select exactly one role from Background, MethodBasis, DataResource, Baseline, ResultSupport, Critique, or Extension. Do not consult or invent an external bibliography entry. Return a concise question that includes the observed citing context and the selected role. Reject the sample if several citations share the same grammatical role span, if two roles are equally plausible, or if the classification conflicts with the rule-based prior after review. Answer schema: {"role": "<one role from ROLE_ONTOLOGY>"}
The released construction metadata retains source identifiers, locked fields, rendering records, and rejection outcomes. This separates model-assisted linguistic generation from the deterministic or auditable source of each answer.
40
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
F. Prompt and Data Examples
A1. Layout Parsing & Reading Order Recover 1. Question Authoring Prompt: You are constructing adversarial multimodal evaluation items for scientific document parsing. Use the page images plus the IR-derived block summaries to write a hard but unambiguous task. Do not invent document facts. Keep the final stem concise. Return JSON only with keys: question_stem, answer_format, difficulty_rationale.
User payload: {
"instruction": "Design a very hard ... task", "facts_for_question_design_only": { "task": "<block_role / parent_child / reading_order / next_hop>", "page_indices": [...], "candidate_blocks": [...], "gold_answer": "..." }, "important_constraints": [ "The final question must be answerable from the page image(s) and the provided candidate block metadata.", "Do not mention hidden gold labels or hidden relation names that are not part of the answer space.", "Maximize challenge by emphasizing visually confusing alternatives or reading-flow traps.", "The final answer must remain exactly the provided gold answer after validation." ] }
2. Light Question-Rewriting Prompt: You : are verifying a scientific-document multimodal task. Solve the question using the provided text and images. If hidden validation-only context is present, use it only to judge whether the gold answer is reasonable and unambiguous. Return JSON only with keys: predicted_answer, confidence, ambiguity, short_reason.
3. Answerability Verification: You are rewriting a single benchmark question sentence. Make only a light paraphrase that preserves the exact meaning. Do not add any hints, examples, rationale, block ids, page numbers, spatial descriptions, quoted text, or any new information. Keep the rewritten sentence similar in length to the original, or shorter. Return JSON only with keys: rewritten_question_stem, rewrite_notes.
Figure 9: Prompt templates for A1 layout parsing and reading-order recovery.
41
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
C1. Numerical & Logical Consistency Checking 1. Claim-Rewriting Prompt: You write diverse but faithful scientific-paper result statements. Preserve answer correctness and all grounding facts. Vary sentence structure and phrasing when you can do so safely. Return JSON only.
User payload: { "task": "table_text_logic_rewrite", "instructions": [ "Rewrite the corrupted statem ent and answer options so they sound like realistic scientific-paper prose.", "You may clean up minor OCR or parsing artifacts ...", "Do not change the underlying facts, target row/column, entity names, numbers, ordering words, or the correct answer.", "Keep the corrupted statement to exactly one sentence.", "Keep each answer option to exactly one sentence.", "Produce m ultiple stylistically distinct but semantically equivalent candidates when possible." ], "caption_text": "<table caption>", "style_reference_text": "<local paper text>", "table_source_latex": "<truncated LaTeX>", "fact_summary": "<normalized fact>", "canonical_corrupted_statement": "<rule-built wrong claim>", "canonical_correct_option": "<gold fix>", "canonical_distractor_options": ["...", "...", "..."], "must_preserve_tokens": { "corrupted_statement": ["<wrong token>"], "correct_option": ["<correct token>"] } }
2. Self-Check Prompt: You are solving a scientific-paper table consistency task. Use only the provided images and prompt text. Return JSON only with keys: predicted_answer, confidence, short_reason.
Figure 10: Prompt templates for C1 numerical and logical consistency checking.
42
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
D3. Cross-Document Result Integration 1. Page-Polishing Prompt:
You help create cross-document scientific table-merging tasks grounded in real tables. Preserve all metrics, entities, and row semantics. Return JSON only.
User payload: { "task": "cross_docum ent_real_table_merge_page_polish", "table_a": { "caption_text": "...", "headers": [...], "nearby_titles": [...], "nearby_text_snippets": [...], "sample_rows": [...] }, "table_b": { "caption_text": "...", "headers": [...], "nearby_titles": [...], "nearby_text_snippets": [...], "sample_rows": [...] }, "heuristic_alignment": { "b_to_a_column_mapping": {...}, "row_overlap": ..., "matched_column_coverage": ... }, "proposed_canonical_schema": { "canonical_headers": [...], "unmatched_columns_from_a": [...], "unmatched_columns_from_b": [...] }, "instructions": [ "Treat the two tables as coming from different papers but describing the sam e benchmark.", "Do not change the table semantics or invent new columns.", "Ground the page titles, running headers, section headings, and contextual paragraphs in the provided real document context whenever possible.", "Write the task stem as a realistic user request ..." ] }
2. Merge Self-Check Prompt: Solve the user's table integration task carefully. Return only valid JSON that follows the requested schema. Preserve exact numeric values and keep source provenance.
Figure 11: Prompt templates for D3 cross-document result integration.
43
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
E1. Figure/Table Data Reconstruction
1. Caption-and-Question Rewriting Prompt : You rewrite captions and questions while preserving exact answer-critical semantics.
User payload: { "task": "figure_numeric_readout_rewrite", "instructions": [ "Rewrite the caption and question so they sound like realistic scientific-document text.", "Do not change the chart identity, target series, x position/category, or numeric answer.", "Keep the caption concise and plausible for a paper figure caption.", "Keep the question to exactly one sentence.", "Do not mention synthetic data, replacement, or PlotQA." ], "chart_type": "...", "series_name": "...", "target_x": "...", "canonical_caption": "...", "canonical_question": "...", "answer_format": "..." }
Figure 12: Prompt template for E1 figure and table data readout.
44
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
A3. Citation Role Analysis 1. Caption-and-Question Rewriting Prompt : You are a strict citation-role classifier for scientific papers. The citation keys are provided from the original LaTeX source and must not be changed or expanded. Choose one coarse role from the allowed list using only the local context.
User prompt: Allowed roles: Background, MethodBasis, DataResource, ComparisonBaseline, ResultSupport, CritiqueGap, Extension Role definitions: - Background: ... - MethodBasis: ... - DataResource: ... - ComparisonBaseline: ... - ResultSupport: ... - CritiqueGap: ... - Extension: ... Exact citation keys: ["..."] Citation command: \cite{...} Page: ... Block id: ... Section: ... Rule guess: ... Local context: ... Return JSON only: {"citation_keys": ["..."], "role": "...", "evidence_quote": "...", "rationale": "...", "confidence": 0.0}
Figure 13: Prompt template for A3 citation role analysis.
45
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
A4. Dataset Lineage Analysis
1. Concept-Graph Extraction Prompt: You are an expert scientific-document analyst. Your job is to construct an evidence-grounded concept composition graph from real paper pages and extracted IR snippets. Do not simply list all named datasets. Identify a meaningful target concept/resource/process, then recover the components, sources, evaluation resources, assumptions, or intermediate objects explicitly linked to it.
User prompt: Choose ONE central target concept that is well-supported by the evidence. The target may be a dataset, benchmark, evaluation protocol, methodological data object, simulation result, or other scientific concept whose relationships can be recovered from the paper. Return JSON only with this schema: { "task_title": "...", "question": "...", "center": {...}, "nodes": [...], "edges": [...] } Allowed relation labels: constructed_from, input_to, feature_source, evaluated_with, ... Rules: - Every node and every edge must be supported by a quoted evidence sentence from the IR snippets. - Prefer 5 to 10 edges. Do not force weak edges. - Exclude mere citations unless the citation is itself the target evidence. - Use supports_analysis only when no specific label fits. IR evidence snippets: ...
Figure 14: Prompt template for A4 dataset lineage and concept-graph extraction.
46
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
A. Document Perception & Structure Ground Truth Our MM-IFEval comprises \textbf{400 humanannotated questions}: 300 \textit{compose-level} open-ended questions and 100 \textit{perceptionlevel} questions with ground truth. With 32 distinct constraint categories and an average of 5.1 constraints per question, MM-IFEval is substantially more challenging than prior benchmarks. 2504.07957 (Rendered)
Question You are given a scientific PDF document. Your task is to extract information about the composition and data sources of the MM-IFInstruct-23k dataset described in the paper, and reproduce it as a TikZ radial diagram in LaTeX. Follow these rules strictly: Extraction rules: - Identify all source datasets that contribute to MM-IFInstruct-23k. - For each source dataset, extract the exact data quantity (e.g., number of samples or instances) if explicitly stated in the paper. If a quantity is not stated, omit the label but still include the node. - Do not infer or fabricate any information not explicitly present in the document. Diagram rules: - Place MM-IFInstruct-23k as a single center node. - Place each source dataset as a surrounding leaf node, evenly distributed in a radial (spoke) layout around the center. - Draw arrows pointing from each leaf node toward the center node. - If a data quantity is available for a source, annotate it on the corresponding arrow as a label. - Use distinct visual styles for the center node (e.g., filled circle) and the leaf nodes (e.g., rounded rectangle). Output rules: - Output a complete, self-contained LaTeX file using the standalone document class. - Do not include any commentary, explanation, or text outside the LaTeX code. - Output only the complete LaTeX source, ready to compile.
Figure 15: Representative benchmark example for Category A: Document Perception and Structure.
47
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
B. Scientific Information Extraction Ground Truth {
2510.27606
"Model": [ { "name": "Spatial-SSRL-3B", "base_model": "Qwen2.5-VL-3B" }, { "name": "Spatial-SSRL-7B", "base_model": "Qwen2.5-VL-7B" }, { "name": "Spatial-SSRL-4B", "base_model": "Qwen3-VL-4B" } ], ······
Question Please read the attached paper thoroughly and extract the following information in JSON format: 1. Model: List all models proposed in this paper. For each model, provide its name and the base model it is built upon. 2. Data: Provide the name of the training dataset constructed in this paper, and list all source datasets used to build it. 3. Training: List all training stages/methods used in the paper's training pipeline, in order. For each stage, provide its name and all hyperparameters explicitly mentioned in the paper. Output the result as a single JSON object with exactly three keys: "Model", "Data", and "Training". Follow the schema below strictly: { "Model": [ { "name": "<model name>", "base_model": "<base model name>" } ], "Data": { "name": "<dataset name>", "source": ["<source dataset 1>", "<source dataset 2>"] }, "Training": [ { "name": "<training stage name>", "parameters": { "<param_name>": "<value>" } } ] } Notes: - For numerical hyperparameters, use string representation (e.g., "1e-5", "128"). - Only include parameters that are explicitly stated in the paper. Do not infer or compute values. - List training stages in the order they are applied in the pipeline. - Output only the JSON, nothing else.
Figure 16: Representative benchmark example for Category B: Scientific Information Extraction.
48
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
C. Evidence Alignment & Verification Ground Truth
{"BITS":"10110101111001010 0"}
2410.06833
Question Read the paper end-to-end. Use only this paper. Do not explain. Return only a JSON object of the form {"BITS":"<BITSTRING>"} where <BITSTRING> is exactly a 18character bitstring. Use 1 iff the claim is fully supported by the paper; otherwise use 0. A. Definition 1.1 requires β > 1 and ε ∈ (0, 1/16). B. In Definition 1.1, α is defined using points from S_i(ε) × S_j(ε) for i ≠ j. C. Theorem 1.2 applies to both (SA) and (USA). D. In Theorem 1.2, if x_i(0) ∈ S_q(ε), then x_i(t) stays in S_q(2ε) for all t ∈ [0, T2]. E. Theorem 1.2 asserts that by time T2 all particles have merged into a single cluster. F. Remark 1.3 says one upper bound on λ is used to ensure T2 > T1, while the other is used in a propagation-ofsmallness argument. G. The proof of Theorem 1.2 in §2 relies crucially on the gradient-flow interpretation of E_β. H. Section 3.2 carries out the Hessian/slow-manifold verification on the circle S^1 and primarily for (USA), with the extension to (SA) deferred to Remark 3.7. I. Figure 5 depicts the slow manifold as an almost-flat zone surrounded by regions where E_β satisfies a PL inequality. J. Proposition 4.2 proves that Gaussian-mixture samples projected onto the sphere are (β, ε)-separated with probability at least 1 − 2e^(−d). K. Corollary 4.4 says that for d ≫ n and suitable β, uniformly random points on S^{d−1} are (β, ε)-separated with ε = 4 log d / d and probability at least 1 − 2 n^2 d^(−1/64). L. Remark 4.5 interprets Corollary 4.4 as saying particles initialized uniformly at random in d ≫ n move rapidly and merge on polynomial time scales. M. Definition 5.1 allows arbitrary mixture weights in μ0 = Σ_q p_q ν_q. N. Remark 5.3 says Theorem 5.2 is only a partial generalization: T1 and T2 are of the same order, and only the variance inside a cap is exponentially small. O. Remark 5.4 says Gaussian mixture laws supported mostly, but not exactly, inside the caps are handled by the proof after lower bounding Z_{β,μ(t)}(x). P. Definition 6.3 modifies (USA) by enforcing a merge once the first pair reaches distance 1/sqrt(β log β), provided exactly two indices are involved. Q. Theorem 6.4 is proved for the original flow without modifying the dynamics and yields a piecewise-linear limiting energy profile. R. Theorem 6.4 proves uniform convergence of the rescaled energy on the whole half-line, including the jump times T_i. Return only the JSON object.
Figure 17: Representative benchmark example for Category C: Evidence Alignment and Verification.
49
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
D. Cross-Document Understanding Ground Truth
2504.15271
Question
2412.05237
\documentclass[border=8pt] {standalone} \usepackage{enumitem} \begin{document} \begin{minipage}{10cm} \textbf{Intersection of Eagle-2.5 and MAmmoTHVL-Instruct (12M) source datasets} \begin{itemize} \item DocVQA \item LVideo-NeXT-Q \end{itemize} \end{minipage} \end{document}
You are given two scientific PDF documents. The first paper introduces Video, multipage document, and long text dataset used in Eagle-2.5 and the second paper introduces the MAmmoTH-VL-Instruct (12M) training dataset. Your task is to identify the datasets that appear in BOTH dataset compositions and output the intersection as a sorted list in LaTeX. Follow these rules strictly: Extraction rules: - Compute the intersection: retain only those dataset names whose spelling is exactly the same (case-insensitive) in both documents. Do not match datasets based on abbreviations, alternate spellings, or inferred equivalences, even if they appear to refer to the same dataset. - Do not infer or fabricate any dataset name not explicitly present in the documents. Sorting rules: - Sort the intersecting dataset names in ascending alphabetical order by their first character (case-insensitive). - If two names start with the same letter, apply standard lexicographic ordering for the subsequent characters. Output rules: - Output a complete, self-contained LaTeX file using the standalone document class. - Render the sorted intersection as a single-column itemized list (\begin{itemize}) with one dataset name per \item. - Above the list, include a short header line stating: "Intersection of Eagle-2.5 and MAmmoTH-VL-Instruct (12M) source datasets" - Do not include any commentary, explanation, or text outside the LaTeX code. - Output only the complete LaTeX source, ready to compile.
Figure 18: Representative benchmark example for Category D: Cross-Document Understanding.
50
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
E. Reconstruction & Execution Ground Truth
2507.18071
Question You are an expert Data Visualization Engineer and Python Developer (Matplotlib/Seaborn Specialist). Objective: Your goal is to Reverse Engineer the provided scientific chart image into an executable Python script. The Gold Standard: 1.Data Fidelity: The data points in your code must be as close as possible to the pixels in the image. 2. Visual Fidelity: The generated plot must look identical to the original image in terms of: Axis labels (including LaTeX math equations). Legend position and text. Line styles (dashed, solid, dotted), markers (circles, triangles), and colors. Axis limits and scales (log vs. linear). Grid lines and background style. Output: A single, self-contained Python code block.
Figure 19: Representative benchmark example for Category E: Reconstruction and Execution.
51
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
F. Paper–Code Alignment Ground Truth { "BLANK_1": "reward_mean + Z * reward_std", "BLANK_2": "sorted_scores[1]", "BLANK_3": "backtrack_cutoff" } 2411.10440
Question
Code Repository of 2411.10440
You are given a research paper (PDF) and its code repository. Your task is to fill in the blanks in the code by carefully reading the paper. --- YOUR TASK --1. Locate and read Algorithm 1 in Appendix D of the paper. 2. Use what you learn from the paper to fill in the three blanks from the code. --- OUTPUT FORMAT --Respond only in the following JSON format with no additional text: { "BLANK_1": "<expression>", "BLANK_2": "<expression>", "BLANK_3": "<expression>" }
Figure 20: Representative benchmark example for Category F: Paper-Code Alignment.
52
SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding
G. Dataset Understanding Question
2508.04724
Ground Truth
2206.13517
Below are 4 dataset entries extracted from a mixed data pool. For each entry, identify: - The Source Paper (ProGen2 or PoET-2). - The Primary Database the data was drawn from. - The Key Preprocessing Step (e.g., clustering threshold or filtering method) mentioned in the text. [Dataset Entries] Entry A: A collection of 1.5B antibody sequences from 80 studies, later clustered at 85% sequence identity to reduce redundancy. Entry B: A mixture of UniRef90 and a dataset approximately 1/3 its size, primarily from metagenomic sources, often containing non-fulllength proteins. Entry C: A dataset derived from UniprotKB, Metaclust, SRC, and MERC, clustered at 90% identity, containing sequences with at least 3 cluster members. Entry D: A dataset used for zero-shot variant effect prediction, specifically focusing on indels and multiple mutations using a hierarchical transformer. [Output Format] Return the result to a JSON array of objects.
[
{ "entry": "Entry A", "source_paper": "ProGen2: Exploring the Boundaries of Protein Language Models", "primary_database": "Observed Antibody Space (OAS)", "key_preprocessing": "Clustered at 85% sequence identity using Linclust", "verification_snippet": "OAS is a curated collection of 1.5B antibody sequences... clustered the OAS sequences at 85% sequence identity [cite: 104, 106]" }, { "entry": "Entry B", "source_paper": "ProGen2: Exploring the Boundaries of Protein Language Models", "primary_database": "Uniref90 + BFD30", "key_preprocessing": "Mixed Uniref90 with BFD30 (1/3 size of Uniref90, metagenomic sources)", "verification_snippet": "The standard PROGEN2 models are pretrained on a mixture of Uniref90 and BFD30... BFD30 dataset is approximately 1/3 the size of Uniref90 [cite: 88, 90]" }, { "entry": "Entry C", "source_paper": "ProGen2: Exploring the Boundaries of Protein Language Models", "primary_database": "BFD90", "key_preprocessing": "Mixed Uniref90 with sequences having at least 3 cluster members from multiple databases at 90% identity", "verification_snippet": "For the PROGEN2-BFD90 model, Uniref90 is mixed with representative sequences with at least 3 cluster members... at 90% sequence identity " }, { "entry": "Entry D", "source_paper": "Understanding protein function with a multimodal retrieval-augmented foundation model", "primary_database": "Evolutionary constraints / family-specific MSA", "key_preprocessing": "In-context learning of family-specific evolutionary constraints with optional structure conditioning", "verification_snippet": "PoET-2... incorporates in-context learning of family-specific evolutionary constraints... excelling at scoring variants with multiple mutations and challenging indel mutations " } ]
Figure 21: Representative benchmark example for Category G: Dataset Understanding.
53