AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot Joydeep Biswas1 , Sheila Schoepp2 , Gautham Vasan2 , Anthony Opipari3 , Arthur Zhang1 , Zichao Hu1 , Sebastian Joseph1 , Matthew Lease4 , Junyi Jessy Li5 , Peter Stone1,7 , Kiri L. Wagstaff6 , Matthew E. Taylor2,8 , Odest Chadwicke Jenkins3
arXiv:2604.13940v1 [cs.AI] 15 Apr 2026
1
Department of Computer Science, The University of Texas at Austin; 2 Department of Computing Science, University of Alberta; 3 Department of Electrical Engineering and Computer Science, University of Michigan; 4 School of Information, The University of Texas at Austin; 5 Department of Linguistics, The University of Texas at Austin; 6 The Valley Library, Oregon State University; 7 Sony AI; 8 Alberta Machine Intelligence Institute. [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected] Abstract Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer review, yet a key unresolved question is whether AI can generate technically sound reviews at real-world conference scale. Here we report the first large-scale field deployment of AIassisted peer review: every main-track submission at AAAI26 received one clearly identified AI review from a state-ofthe-art system. The system combined frontier models, tool use, and safeguards in a multi-stage process to generate reviews for all 22,977 full-review papers in less than a day. A large-scale survey of AAAI-26 authors and program committee members showed that participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions. We also introduce a novel benchmark and find that our system substantially outperforms a simple LLMgenerated review baseline at detecting a variety of scientific weaknesses. Together, these results show that state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research.
1
Introduction
The scientific peer review process is under significant strain. The AAAI Conference on Artificial Intelligence, a major artificial intelligence (AI) research conference, received more than 30,000 initial submissions for 20261 , up from approximately 15,000 for 2025. This dramatic growth is not unique to AAAI; submissions have grown rapidly for other venues too, such as Nature [1] and NeurIPS [2]. Unfortunately, despite this rapid growth, the peer review process has remained largely static, with a large cohort of human reviewers providing detailed reviews and ratings for papers, and a smaller group of senior researchers comparCopyright © 2026, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved. 1 The AAAI-26 review process ran from Aug.–Nov. 2025.
ing those reviews and ratings to make final paper acceptance recommendations. While this established review process has long endured, the rising scale of submissions means we face increasingly overburdened reviewers with more papers assigned, the need to recruit an ever-wider pool of potentially less experienced reviewers, and increasingly compressed timelines. Maintaining the quality, consistency, and timeliness of peer review is thus increasingly challenging. For example, the scale of AAAI-26 submissions required the recruitment and oversight of over 28,000 Program Committee members, Senior Program Committee members, and Area Chairs, nearly three times the size of the committee in AAAI-25 [3]. At the same time, there have been rapid advances in stateof-the-art AI systems, particularly in mathematical, coding [4, 5], and other technical domains [6]. Autonomous AI scientists [7, 8, 9, 10, 2] now perform iterative feedback loops of automated paper writing, generating critiques, and revising the writing in response. Most relevant, a growing body of work is now investigating whether and how AI systems can be used to assist with scientific peer review and help alleviate the growing strain on the review process [11, 12, 13, 14, 15, 16, 17, 18, 19, 20]. Against the backdrop of increasing strain on human peer review, the reviewer population has started using AI reviewing against explicit guidance [21], while conference venues grapple with the question of how to effectively and meaningfully integrate AI into the review process in a way that is beneficial to the community. There have been synthetic, benchmark-based, and posthoc studies of AI-generated reviews and AI-assisted peer review on existing papers and reviews, including retrospective analyses of reviewers’ use of AI during peer review [12, 14, 16, 18, 17, 13, 22]. Evaluation remains challenging, however, because existing benchmarks and evaluation datasets measure only limited aspects of reviewing. These aspects include specific error types, evaluation of structured outputs as opposed to unstructured review text, or similarity of scores to human reviews, rather than end-toend review quality [23, 24, 25, 26].
Based on these encouraging findings, there have been a small number of live studies of AI systems providing limited assistance within conference workflows, notably author checklist assistance in NeurIPS 2024 [27] and feedback to reviewers in ICLR 2025 [28]. However, neither of these studies deployed official AI-generated reviews on live submissions. Prior to AAAI-26, there had been no conferencewide live study of AI-generated reviews deployed on real submissions at a major conference. Thus, despite substantial recent progress, a key question remained: could stateof-the-art AI systems generate technically meaningful and practically useful reviews in a live peer-review process at conference scale? The AAAI-26 AI Review Pilot Program was the first fullscale live study of AI-generated reviews on real submissions at a major conference. Every paper that entered the full review phase (22,977 in total) in the main track at AAAI-26 received one clearly labeled AI review, generated by a stateof-the-art custom-developed AI review system. Consistent with prior results, we found that simply asking off-theshelf LLMs to review papers does not lead to high-quality reviews [14, 16, 29]. Recent work has therefore explored more structured review systems based on deeper multi-stage reasoning, hierarchical question decomposition, and multimodal workflows [30, 31, 32]. In response, we developed a novel, multi-stage, multi-tool, LLM-based review pipeline that does lead to very high-quality reviews. AAAI-26 used a double-blind review process, so reviewers and authors were anonymized to one another during evaluation, but both reviewers and authors could identify the AI review. The AI review was added during Phase 1 of the two-phase review process, alongside at least two human reviews. The AI review system did not include any scores or recommendations, and no human reviewers were replaced in the process. Instead, the AI reviews were intended to provide additional input to the peer-review process [33]. Senior Program Committee members (SPCs) and Area Chairs (ACs) (who are responsible for making paper recommendations and normalizing within their batch), were able to view the AI reviews along with the human reviews and use them to help make their decisions about whether to promote papers to Phase 2 of the review process. Papers that were promoted to Phase 2 received additional human reviews, and the authors had a chance to respond to all reviews, including the AI review, before the reviewers, SPCs, and ACs discussed the papers and made their final decisions in light of all reviews, including the AI review. An optional survey was sent to authors, reviewers, SPCs, and ACs to assess both human and AI reviews on a variety of criteria. The design, development, and deployment of the AI Review Pilot Program was overseen through extensive and careful consideration by the AAAI Executive Council, the AAAI Conference Committee, the AAAI Ethics Committee, and the AAAI-26 Program Committee. The survey design was reviewed by the Institutional Review Boards (IRB) of the University of Texas at Austin and the University of Michigan, as well as the Research Ethics Board (REB) of the University of Alberta. We found that AI-generated peer reviews are opera-
tionally feasible at conference scale, with a modest cost of less than $1 per paper (covered through an in-kind donation of API credits from OpenAI as a AAAI-26 sponsor)2 . All reviews were generated in less than 24 hours using a state-ofthe-art frontier LLM in a multi-stage workflow with coding and web search tools. By creating a new benchmark for scientific review based on synthetic perturbations to published papers, we further showed that the AAAI-26 AI Review System significantly improves upon the ability of a base LLM to catch scientific errors in the story, presentation, evaluations, correctness, and significance of scientific papers. Finally, analysis of the survey of authors, reviewers, SPCs, and ACs (5,834 responses) indicates that the community not only found the AAAI-26 AI reviews helpful, but that AI reviews were actually preferred to human reviews on several important criteria spanning technical accuracy, review focus, and research suggestions. Moreover, respondents believed that AI reviews would be useful in future peer review processes. Qualitative responses further highlighted both the success and the current limitations of the system. Strengths mentioned by respondents included providing an impartial review as a safeguard against human variability and some forms of malicious activity; and structured and in-depth technical feedback. Weaknesses included errors in reading some equations and tables, difficulty in prioritizing the significance of issues (an area of ongoing research), and producing reviews that were longer than readers preferred (a limitation that is straightforward to mitigate via tighter output-length controls). Overall, we found that the technical capabilities of this AI review system are already sufficient to usefully assist in scientific peer review in ways that the community finds helpful. Further study is needed to ascertain how best to integrate the complementary strengths of AI systems and human reviewers to improve the process of evaluating and advancing scientific research.
2
The AAAI-26 AI Review System
The AAAI-26 AI Review System integrates learnings from prior studies of AI-generated reviews [12, 16, 29, 7]. A key design goal for the system was to ensure that the reviews considered scientific accuracy of all forms — including mathematical and algorithmic correctness, sufficiency of the evaluation methods, and positioning of the work in the context of the previous state-of-the-art. Previous studies have shown that prompting LLMs with different ‘personas’ [24, 18, 17] or criterion-specific prompts [14], rather than asking them to directly produce full reviews, can improve identification of specific types of scientific errors. Recent systems have also explored hierarchical question decomposition, deeper staged reasoning, and multimodal agent designs with shared memory for paper review [31, 30, 32]. The AAAI-26 AI Review System thus consists of five core scientific review stages intended to identify errors in: (1) story, (2) presentation, (3) evaluations, (4) 2 Even with 30K submissions, this cost is a small fraction of the conference budget.
Paper PDF
Resampling Resampled Paper PDF
Logging
Story
2
Presentation
3
Evaluations
4
Correctness
5
Significance
Paper Submission
Checkpointing Reporting
Phase 1 Review Phase 1 Decision
Human Oversight
olmOCR
Initial Review
Critical Judgments
Paper Markdown
Self-Critique
Quality-Checking Critic
Refinement
Final AI Review
AAAI-26 AI Review System
Phase 2 Review Author Response Discussion
Cumulative reviews completed
1
20000 15000
Story Presentation Evaluations Correctness Significance
Initial review Critique Revision Final Review Human check
10000 5000 0 00:00 03:20 06:40 10:00 13:20 16:40 20:00 23:20 26:40 Elapsed time from start (HH:MM)
Phase 2 Decision AAAI-26 Review Process
(b) Review generation timeline
(a) Review system
Figure 1: The AAAI-26 AI review system (a) and review generation timeline (b). For every submission to the AAAI-26 main track that entered the full review phase, the AI review system took as input the submitted paper in PDF form and generated an AI review that was then included in the first phase of the two-phase AAAI-26 review process. The AI review system first pre-processed the paper to resample all images to a consistent resolution of 250 DPI, and then used olmOCR [34] to convert the PDF version of the paper to markdown. Both PDF and markdown versions of the paper were provided to the the subsequent stages. Five core review stages assessed the story, presentation, evaluations, correctness, and significance of the paper. An initial review was generated, compiling findings from all stages, into a consistent format. A self-critique stage checked the review for unsubstantiated claims, missing details, and inconsistencies with the paper. The final review stage revised the review to address the self-critique and compiled a final review. Logs, checkpoints, and review reports generated at all stages were saved for auditing and human oversight. The final review was assessed by a quality-checking critic, and a human inspected the critical judgments to identify potential concerns such as ethical concerns, revealing author identities, or missing structural elements in the review. The review generation timeline (b) shows the progress of all staged generations for all 22,977 papers. An initial batch of 30% of the papers were run through the core review stages and the outputs inspected manually, before proceeding to complete all remaining stages and process all remaining papers. After the final manual inspection stage, a few papers were flagged for re-processing after manual PDF conversions to handle exceptional graphics formats. correctness, and (5) significance. The review system takes each PDF paper as input and generates a textual review with markdown notation [35] for math and tables, such that it can be rendered on the review interface. Fig. 1 shows the multi-stage AAAI-26 AI Review System that we built based on the desiderata and insights above. After preprocessing, subsequent stages of the review system include both PDF and markdown versions of the paper, along with a system prompt to provide context for when to pay particular attention to one version over the other. Each review stage includes both a stage-specific prompt as well as the prompts and results from all previous stages. The evaluations and correctness stages include a Python code interpreter made available to the LLM to allow it to check for errors by testing out math and code snippets. The significance stage includes a web search tool to assist with literature search — with specific instructions to restrict references to published work at relevant venues. After all targeted stages, the system generates an initial review and then revises it through a self-critique stage. The initial review generation and revision prompts include specific instructions to ensure that each review contains the following structural elements: (1) the title of the paper, (2) a brief synopsis of the paper, (3) a summary of the review, (4) a detailed list of strengths,
(5) a detailed list of weaknesses, and (6) a list of references cited in the review, in APA citation format. Prompt details for each stage are provided in Section A. We implemented a quality-checking workflow to identify potential issues in the generated reviews, similar to previous work on “peer reviews of peer reviews” [36]. Details on the quality-checking workflow and additional checks for citation hallucinations are provided in Section A.
3
Review Survey and Findings
To assess the value and impact of the AI reviews on the AAAI-26 review process, we conducted a survey of all the participants: authors, PC, SPC, and ACs. Participation in the survey was voluntary, and the study design was reviewed by the University of Texas at Austin (IRB Protocol STUDY00007931), the University of Michigan (IRB Protocol HUM00280758), and the University of Alberta (REB Protocol Pro00159777).
3.1
Survey Description
The survey had a questionnaire associated with every review (both human and AI) that was made available to authors, PC, SPC, and ACs. Authors and reviewers received the survey when the reviews were made available to them. For papers
This Review Was A Thorough Review For This Conference (+) This Review Raised Points That I Had Not Previously Considered (+) This Review Accurately Conveyed The Significance And Impact (+) This Review Overemphasized Minor Issues (-) This Review Accurately Identified Technical Errors (+) This Review Made Technical Errors In The Review (-) This Review Provided Useful Suggestions To Improve The Paper Presentation (+) This Review Provided Useful Suggestions To Improve Research Design (+) This Review Provided Suggestions That Were Wrong Or Unhelpful (-)
N
p-value
+0.48
5639 (3032, 2607)
1.73e-35
+0.61
5542 (2966, 2576)
2.77e-77
+0.33
5643 (3038, 2605)
2.12e-15
-0.38
5437 (2909, 2528)
1.83e-30
This review demonstrated capabilities beyond what I n=1939 expected from AI
+0.67
5130 (2666, 2464)
6.01e-97
-0.22
5180 (2780, 2400)
This review found concerns that a human reviewer would n=1329 have difficulty catching (R)
1.09e-16
+0.54
5461 (2898, 2563)
3.97e-59
+0.49
5519 (2951, 2568)
1.26e-45
Overall 5327 (2896, 2431) 5.03e-06 Authors PC+SPC+ACs
-0.11
0.6 0.4 0.2 0.0 0.2 0.4 0.6 Difference in mean response ( Human preferred, AI preferred )
Overall AAAI26 AI reviews were n=1935 useful Overall AI reviews would be useful in future peer review n=1940 processes Overall went into review thinking AI reviews would be n=1948 useful
This review overlooked points that a human reviewer would n=1337 likely have caught (R) This review changed my evaluation or interpretation n=1335 of the paper (R) This review raised important points that the human written n=561 reviews omitted (A)
Strongly disagree
(a) AI vs. human review comparisons
80 40 0 40 80 Percent of responses (left = disagreement, right = agreement) Disagree
Neutral
Agree
Strongly agree
(b) Responses to AI review questions
Figure 2: Survey responses: AI vs. human review comparisons (a) and AI review questions (b). The left figure shows the differences in the mean response score between AI and human reviews for each of the nine review-quality criteria. In six out of nine criteria, AI reviews were rated higher than human reviews. The preference towards AI reviews was stronger for authors than for PC, SPC, and ACs. All p-values show strong statistical significance at the α = 0.01 level. The right figure shows the responses to the questions posed for just the AI reviews. Overall, there is strong support among the respondents that the AI reviews were useful in the AAAI-26 pilot and that they would be useful in future peer review processes. Respondents also expressed that the AI reviews demonstrated capabilities that were beyond what they expected from AI, despite indicating that they went into the process expecting that the AI reviews would be useful. Responses to questions posed specifically to the reviewers (labelled ‘(R)’) show that the AI reviews both raised points that a human reviewer would have difficulty catching, while simultaneously overlooking points that a human reviewer would likely have caught — indicating the complementary strengths of AI vs. human reviews. The majority of reviewers indicated that the AI reviews did not change their evaluation or interpretation of the paper, though a minority did indicate that they did. Authors (labelled ‘(A)’) were mixed in their views about the AI reviews raising points that the human reviews omitted. Table 1: Number of responses by the respondent roles and review types, including the reviewer program committee (PC), senior program committee (SPC), and area chairs (AC). Response counts are further broken down by review type being surveyed (AI or human reviews). There were 5,834 responses in total. Review Type
Authors
PC
SPC
AC
AI Reviews Human Reviews
575 2500
1305 879
117 433
12 13
Total
3075
2184
550
25
rejected in phase 1, this happened with the release of the phase 1 decisions, while for papers that proceeded to phase 2, this happened at the start of the author rebuttal period. Each questionnaire consisted of a set of statements about the review, with a five-point Likert scale to assess the degree of agreement with the statement (-2: strongly disagree, -1: disagree, 0: neutral, 1: agree, 2: strongly agree). The statements were grouped into four sections: (1) overall impressions, (2) review emphasis, (3) technical accuracy, and (4) suggestions for research and writing. An open-ended field
allowed free-form comments to capture details about the review, the AAAI-26 AI Review Pilot, or any other comments related to the future of AI in peer review. The survey received a total of 5,834 responses from authors, PC, SPC, and ACs, enabling strong statistical significance in the comparisons of AI and human reviews among respondents, though self-selection bias must also be considered. Table 1 shows the number of responses by role, for AI and human reviews.
3.2
Quantitative Analysis
We analyzed the survey responses to assess the overall attitudes toward the AAAI-26 AI Review Pilot, including comparing the distributions of responses for AI and human reviews. Fig. 2 shows the primary findings from the survey. The full survey text is included in Section A. We compared AI and human reviews on the nine reviewquality criteria shown in Fig. 2a by comparing the means of the empirical response distributions for each question. Overall, AI reviews were preferred to human reviews on six of the nine criteria, and all nine AI-human differences were statistically significant under a Mann-Whitney U test at α = 0.01. The largest advantages for AI reviews were in
Top Positive Themes
5.3% 5.2% 5.0%
Actionable Revision Guidance Breadth and Thoroughness Technical Error Detection Relative Objectivity and Consistency Presentation and Writing Polish
4.3% 4.2%
Top Negative Themes Weak Big-Picture Judgment on Novelty, Significance, and Impact Nitpicking and Overemphasis on Minor Issues Excessive Verbosity and Cognitive Overload Other Factual Errors and Misreadings Shallow Contextual and Domain Understanding
9.1% 8.5% 8.3% 7.7% 7.6%
Figure 3: Top five most frequent positive and negative themes specific to the AAAI-26 AI Review Pilot found in written feedback from authors and program committee members. The percentages refer to the percentage of mentions belonging to that particular theme over all classified mentions specific to this pilot. identifying technical errors (+0.67), raising previously unconsidered points (+0.61), suggesting improvements to presentation (+0.54) and research design (+0.49), and overall thoroughness (+0.48). At the same time, respondents also judged AI reviews as more likely to overemphasize minor issues (−0.38), somewhat more likely to contain technical errors themselves (−0.22), and slightly more likely to include wrong or unhelpful suggestions (−0.11). The effect sizes were consistently larger for authors than for PC, SPC, and AC respondents. We next consider the AI-only questions in Fig. 2b. Overall, 53.9% of respondents judged the AI reviews useful, compared with 20.2% who did not, and 61.5% expected AI reviews to be useful in future peer review, compared with 14.5% who did not. Reported expectations before the pilot were already positive, with 57.8% expecting AI reviews to be useful and 16.8% not expecting them to be useful. Even so, 55.6% reported that the AI reviews demonstrated capabilities beyond what they had expected from AI, whereas 17.3% did not. The final set of questions addressed the AI reviews’ complementary role alongside human reviews. Among PC, SPC, and AC respondents, 46.6% agreed that the AI review found concerns that a human reviewer would have had difficulty catching, whereas 21.3% disagreed. At the same time, 49.4% agreed that the AI review overlooked points that a human reviewer would likely have caught, whereas 14.0% disagreed, indicating that respondents saw AI reviews as complementary rather than interchangeable with human review. Consistent with this interpretation, only 13.8% reported that the AI review changed their evaluation or interpretation of the paper, whereas 55.6% disagreed. Authors were similarly mixed on whether the AI review raised important points omitted by the human reviews, with 34.4% agreeing and 37.6% disagreeing. The full distributions of responses for the overall thoroughness of human vs. AI reviews, as re-
ported by authors and separated by whether the paper was accepted, are provided in Section A. We found that among both accepted and rejected papers, AI reviews were rated as more thorough than human reviews, and the distributions for human reviews were wider than the distributions for AI reviews.
3.3
Qualitative Analysis
There were 320 usable free-form responses to the openended written feedback section in the survey. Details on the process for aggregating a collective taxonomy of themes, which we then used to classify the points made in each freeform response, are provided in Section A. Fig. 3 describes the five most frequent positive and the five most frequent negative themes specific to the AAAI-26 AI Review Pilot found in all the written feedback. Among the positive themes, the most frequently discussed was the aspect of Actionable Revision Guidance, which describes how AI reviews tend to turn criticisms into actionable revision suggestions. Other aspects which were thought to be useful were the AI reviews’ thorough and exhaustive analysis (Breadth and Thoroughness), the flagging of technical mistakes overlooked by human reviewers (Technical Error Detection), the consistent and objective nature of the reviews compared to humans (Relative Objectivity and Consistency), and the reviews’ detection of typos and other presentation errors (Presentation and Writing Polish). In general, AI reviews appeared to be valued for their actionable and thorough feedback. The most frequently discussed negative point was that reviews often failed to accurately assess the novelty, significance of contributions, and overall scientific impact of a paper (Weak Big-Picture Judgment of Novelty, Significance, and Impact). Other commonly discussed negative points concerned the AI review focus on minor issues (Nitpicking and Overemphasis on Minor Issues), while overly lengthy
Cross-Stage Detection Rates by Family Source Latex
Candidate Sampling
Perturbed Latex
Perturbation Description
Judging Criteria
Curated Set
Paper Curation
Source Validation + Compilation
Rejected Perturbation
Human Checking Perturbed PDF
Perturbation Generation
0.67
0.47
0.16
0.15
Targetedness margin 0.33
0.7
+0.20
0.6
Review
Matched Set Latex Compilation Test
Story
Perturbed PDF
Review Generation
Perturbation LLM
Candidate Pool
Paper Matching
Perturbation Prompt
Perturbation Family
Conference Proceedings
Judging LLM
Presentation
0.36
0.52
0.20
0.27
0.01
Evaluations
0.58
0.56
0.67
0.48
0.04
+0.16
0.5 0.4
+0.09 0.3
Correctness
0.56
0.62
0.06
0.69
0.05
Significance
0.25
0.11
0.09
0.03
0.53
+0.08
0.2 0.1
Judging Notes
Reviewing + Judging
(a) The SPECS benchmark
+0.28 0.0
Result
ry
Sto
on
ati ent Pres
s nce ions nes ifica luat rect Eva Cor Sign Review Stage
0.0
0.1 0.2 0.3 expected - best off-target
(b) Stage-by-criterion detection rates
Figure 4: The SPECS Review Benchmark curation and analysis workflow (a) and stage-by-criterion detection rates (b) from evaluating each stage in the AAAI-26 AI Review System for the SPECS criteria. Starting from the accepted papers listed in a venue proceedings (AAAI-25 in our case), the SPECS curation system samples across proceedings categories, matches each paper to an arXiv source release, and retains papers whose LATEX sources compile successfully. For each curated paper, an LLM generates controlled source-level perturbations targeting the five SPECS criteria: Story, Presentation, Evaluations, Correctness, and Significance. Each perturbation is accepted only if the cited source span matches the original file and the edited paper recompiles without errors, yielding executable synthetic errors for evaluation. The full AAAI-26 AI Review System and the criterion-targeted intermediate stages are then run on these perturbed papers, and a judge model determines whether each review explicitly identifies the injected error with supporting evidence from the review text. The stage-by-criterion detection-rate matrix in (b) summarizes criterion-wise detection behavior: each row corresponds to a perturbation type, while each column corresponds to a core stage in the AAAI-26 AI Review System. Each cell is computed independently as the fraction of perturbations of the row criterion detected by the column stage; rows do not sum to 1 because a perturbation can be detected by multiple stages (or none). The diagonal entries indicate criterion-aligned detection of the target core stage and the corresponding criterion, while off-diagonal entries reveal cross-criterion detection patterns. The targetedness margin is the difference between a row’s diagonal value and its largest off-diagonal value, indicating how much more strongly each stage detects its intended criterion than the other stages. AI reviews created additional cognitive work for authors and other reviewers (Excessive Verbosity and Cognitive Overload). Such verbosity and nitpicking highlights the downside of the thoroughness found in AI reviews. Finally, the feedback also discussed Factual Errors and Misreadings found in AI reviews, indicative of a continuing gap in LLMs understanding scientific content compared to an academic (human) established within their field. This ties into the the fifth most discussed theme (Shallow Contextual and Domain Understanding), which specifically points out how the AI reviews struggled to provide feedback that was appropriate to a particular research domain. 9 out of the 33 themes we found within the written feedback were general opinions about the use of AI in academic review process that were not specific to the AAAI-26 AI Review Pilot. Representative free-form survey responses illustrating these patterns are provided in Section A. Respondents also spoke positively about using AI as a presubmission assistance tool and a means of generating metalevel summaries [37]. They highlighted its potential to scale peer review and support overburdened reviewers, while acknowledging that these systems are poised for rapid future improvement. Not all of these opinionated themes were pos-
itive. Respondents also emphasized that AI reviews had the potential to mislead reviewers and other decision-makers in the review process. There were also concerns that authors might optimize papers for AI preferences rather than scientific quality, and that reliance on these tools could lead to a long-term decline in reviewing skill. Adding to this, many respondents voiced principled objections, arguing that the use of AI undermines the trust, human effort, and essential value of the peer review process.
4
The SPECS Review Benchmark
A full peer review of a paper involves evaluating multiple criteria, including the (1) problem definition and formulation; (2) technical accuracy of the proposed approach; (3) sufficiency of the evaluation methods; (4) accurate positioning of the work in the context of the related work; and (5) readability and clarity of the paper. While there are several existing benchmarks for evaluating AI- and human-written reviews, they either focus on a specific criterion [23], require specific structured outputs rather than full reviews [25, 23], or assess consistency of review scores between human and AI reviews [7, 24, 16, 29]. We introduce the SPECS review benchmark to evaluate
Table 2: SPECS benchmark results by criterion, comparing the AAAI-26 AI Review System (Final) to the baseline (Baseline) and criterion-targeted stages (Targeted). Improvements over baseline are listed in the ∆ columns, and corresponding p-values are reported for each comparison. Criterion
n
Baseline
Targeted
Final
∆ T–B
p-value
∆ F–B
p-value
−12
Story Presentation Evaluations Correctness Significance
153 173 159 144 154
0.3529 0.4162 0.5157 0.6111 0.2597
0.6667 0.5202 0.6730 0.6944 0.5325
0.6732 0.5665 0.7547 0.7639 0.4481
+0.3137 +0.1040 +0.1572 +0.0833 +0.2727
2.9 × 10 0.0051 0.0006 0.0576 6.5 × 10−7
+0.3203 +0.1503 +0.2390 +0.1528 +0.1883
1.5 × 10−12 4.2 × 10−5 1.6 × 10−9 2.7 × 10−5 0.0003
All SPECS criteria
783
0.4291
0.6143
0.6386
+0.1852
4.8 × 10−21
+0.2095
6.1 × 10−30
the effectiveness of AI review generation algorithms over multiple criteria, and to provide a comprehensive assessment of the quality of AI reviews. Specifically, it assesses the ability to catch errors in the Story, Presentation, Evaluations, Correctness, and Significance of the paper. The SPECS review benchmark follows a data curation and evaluation process similar to the FLAWS review benchmark [23]. While FLAWS focuses on just technical errors and evaluates LLM responses in a structured output format, SPECS evaluates multiple criteria in a full free-form review format. Fig. 4 shows the process of curating the SPECS review benchmark dataset. The SPECS benchmark includes both a process and a dataset resulting from the process. The process can be repeated to generate a new dataset, e.g., for a different venue, or to generate a new set of perturbations. Details on the curation process and the dataset that we generated by running the process on the AAAI-25 proceedings are provided in the Appendix. We focus here on the primary results from using the SPECS benchmark to evaluate the AAAI-26 AI Review System. Given the set of PDFs compiled from the synthetic perturbations in the SPECS benchmark, the full AAAI-26 AI Review System was used to generate a review for each synthetically pertubed paper. Results from each review stage were additionally logged for analysis and comparison. As a baseline, we generated reviews for all papers using a single review prompt with no intermediate stages. This yielded a total of 5,481 reviews: 783 reviews from the baseline singleprompt approach, 3,915 reviews from the five intermediate review stages (one review per perturbation, per review stage), and 783 reviews from the final review stage. The reviews were judged (details in Section A) to identify whether they caught the specific error introduced in the perturbation stage for the corresponding paper. Table 2 reports the recall rates of the AAAI-26 AI Review System (column ‘Final’), the baseline (column ‘Baseline’), and the criterion-targeted stages (column ‘Targeted’) for each SPECS criterion. The ‘∆’ columns show the absolute gains over baseline, and the ‘p-value’ columns show the statistical significance of the gain compared to baseline, computed using a two-sided exact McNemar test. For every criterion, the AAAI-26 AI Review System significantly improves upon the baseline, with statistical significance at the α = 0.01 level and an average gain of +0.21 across all criteria. Of the five targeted stages, all but the correctness stage
demonstrate statistically significant improvements over the baseline, with an average gain of +0.19 across all criteria. Note that for one criterion, Significance, some errors are correctly identified in the targeted stage of the pipeline, but the final generated review does not explicitly identify them. Fig. 4b shows the stage-by-criterion detection rates for the AAAI-26 AI Review System, revealing that each stage is indeed most effective (compared to other stages) at catching errors for its intended criterion.
5
Conclusions
The AAAI-26 AI Review Pilot Program demonstrated that AI-generated peer reviews are operationally feasible at conference scale, and are capable of generating reviews that are helpful to the authors and reviewers. The large-scale survey of AAAI-26 authors, reviewers, senior program committee members, and area chairs found that participants broadly found AI reviews useful and preferred them to human reviews on key dimensions such as technical accuracy and research suggestions, but also identified some limitations and areas for improvement including technical errors in reading some equations and tables, difficulty in prioritizing the significance of issues, and producing reviews that were longer than readers preferred. The quantitative and qualitative analyses combined indicate the complementary strengths of AI systems and human reviewers, and suggest that future work should explore how best to integrate the two to leverage their strengths.
6
Acknowledgments
We thank the AAAI Executive Council, Ethics Committee, and Conference Committee for their support and guidance in shaping the AAAI-26 AI Review Pilot Program. Zico Kolter initiated the idea of the AI review pilot, and served as liaison to OpenAI for securing a sponsorship to support the initiative. We thank OpenAI for providing API credits as a AAAI-26 sponsor to support the development, deployment, and evaluation of the AI review system. Meredith Ellison (AAAI Executive Director) and Marc Pujol (AAAI Director of Program Operations & Systems) assisted in the operations and logistics of deploying the AI review pilot for AAAI-26. We are grateful to the OpenReview staff for their assistance in integrating the AAAI-26 AI Review System reviews and survey forms into the OpenReview platform.
References [1] Nature Index. Why did the nature index grow by 16% in 2024? Nature Index, 2025. URL https://www.nature.com/nature-index/news/why-didthe-nature-index-grow-by-sixteen-percent-in-twentytwenty-four. Accessed: 2026-03-28. [2] Chenhao Tan and Haokun Liu. The mirage of autonomous ai scientists. https://chenhaot.com/papers/ mirage ai scientist.pdf, 2026. February 2. [3] Association for the Advancement of Artificial Intelligence. AAAI-26 Review Process Update: Scale, Integrity Measures, and Pathways to Sustainability. https://aaai.org/conference/aaai/aaai-26/reviewprocess-update/, 2025. Accessed: 2026-03-14. [4] Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, 2024. URL https: //arxiv.org/abs/2402.03300. [5] Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. DeepSeek-Coder: When the Large Language Model Meets Programming – The Rise of Code Intelligence, 2024. URL https://arxiv.org/abs/2401.14196. [6] David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. In Proceedings of the Conference on Language Modeling, 2024. URL https://openreview.net/forum?id= Ti67584b98. [7] Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. The AI scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292, 2024. [8] Peter Jansen, Oyvind Tafjord, Marissa Radensky, Pao Siangliulue, Tom Hope, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Daniel S Weld, and Peter Clark. CodeScientist: End-to-end semi-automated scientific discovery with code-based experimentation. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 13370–13467, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 9798-89176-256-5. doi: 10.18653/v1/2025.findings-acl. 692. URL https://aclanthology.org/2025.findings-acl. 692/. [9] Chenglei Si, Zitong Yang, Yejin Choi, Emmanuel Candès, Diyi Yang, and Tatsunori Hashimoto. Towards execution-grounded automated AI research. arXiv preprint arXiv:2601.14525, 2026. [10] Aniketh Garikaparthi, Manasi Patwardhan, and Arman Cohan. Researchgym: Evaluating language model
agents on real-world AI research. arXiv preprint arXiv:2602.15112, 2026. [11] Jiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, and Wenpeng Yin. LLMs assist NLP researchers: Critique paper (meta-)reviewing. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 5081–5099, Miami, Florida, USA, November 2024. Association for Computational Linguistics. URL https://aclanthology.org/2024.emnlp-main.292/. [12] Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Smith, Yian Yin, Daniel McFarland, and James Zou. Can Large Language Models Provide Useful Feedback on Research Papers? A LargeScale Empirical Analysis. NEJM AI, 1(8), 2024. doi: 10.1056/AIoa2400196. URL https://ai.nejm.org/doi/ abs/10.1056/AIoa2400196. [13] Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, Daniel Mcfarland, and James Y. Zou. Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 29575–29620, Vienna, Austria, 2024. PMLR. URL https://proceedings.mlr.press/ v235/liang24b.html. [14] Ryan Liu and Nihar B. Shah. ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing, 2023. URL https://arxiv.org/abs/ 2306.00622. [15] Jianxiang Yu, Zichen Ding, Jiaqi Tan, Kangyang Luo, Zhenmin Weng, Chenghua Gong, Long Zeng, RenJing Cui, Chengcheng Han, Qiushi Sun, Zhiyong Wu, Yunshi Lan, and Xiang Li. Automated peer reviewing in paper SEA: Standardization, evaluation, and analysis. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 10164– 10184, Miami, Florida, USA, November 2024. Association for Computational Linguistics. URL https: //aclanthology.org/2024.findings-emnlp.595/. [16] Ruiyang Zhou, Lu Chen, and Kai Yu. Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks. In Proceedings of the 2024 Joint International Conference on
Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 9340–9351, Torino, Italia, 2024. ELRA and ICCL. URL https: //aclanthology.org/2024.lrec-main.816/. [17] Nicolas Bougie and Narimawa Watanabe. Generative Reviewer Agents: Scalable Simulacra of Peer Review. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 98–116, Suzhou, China, 2025. Association for Computational Linguistics. URL https: //aclanthology.org/2025.emnlp-industry.8/. [18] Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Ting Liu, and Yuzhuo Fu. ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews, 2025. URL https://arxiv.org/abs/2503.08506. [19] Hyungyu Shin, Jingyu Tang, Yoonjoo Lee, Nayoung Kim, Hyunseung Lim, Ji Yong Cho, Hwajung Hong, Moontae Lee, and Juho Kim. Mind the blind spots: A focus-level evaluation framework for LLM reviews. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 35630–35656, Suzhou, China, November 2025. Association for Computational Linguistics. URL https://aclanthology.org/2025. emnlp-main.1805/. [20] Zhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig, and Arman Cohan. Can LLMs identify critical limitations within scientific research? a systematic evaluation on AI research papers. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20652– 20706, Vienna, Austria, July 2025. Association for Computational Linguistics. URL https://aclanthology. org/2025.acl-long.1009/. [21] Miryam Naddaf. More than half of researchers now use AI for peer review—often against guidance. Nature, 649(8096):273–274, 2026. [22] Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R. Davidson, Veniamin Veselovsky, and Robert West. The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates, 2024. URL https://arxiv.org/abs/2405.02150. [23] Sarina Xi, Vishisht Rao, Justin Payan, and Nihar B. Shah. FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers, 2025. URL https: //arxiv.org/abs/2511.21843. [24] Gaurav Sahu, Hugo Larochelle, Laurent Charlin, and Christopher Pal. ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review. arXiv preprint arXiv:2510.08867, 2025. [25] Renze Lou, Hanzi Xu, Sijia Wang, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Jian Xie, Yuxuan Sun, Yusen Zhang, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Wenchao Ma, Xi Li, Kai Zhang, Congying Xia,
Lifu Huang, and Wenpeng Yin. AAAR-1.0: Assessing AI’s Potential to Assist Research. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 40361–40383, Vancouver, Cananda, 2025. PMLR. URL https://proceedings.mlr. press/v267/lou25c.html. [26] Madhav Krishan Garg, Tejash Prasad, Tanmay Singhal, Chhavi Kirtani, Murari Mandal, and Dhruv Kumar. ReviewEval: An evaluation framework for AIgenerated reviews. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Findings of the Association for Computational Linguistics: EMNLP 2025, pages 20542– 20564, Suzhou, China, 2025. Association for Computational Linguistics. URL https://aclanthology.org/ 2025.findings-emnlp.1120/. [27] Alexander Goldberg, Ihsan Ullah, Thanh Gia Hieu Khuong, Benedictus Kent Rachmat, Zhen Xu, Isabelle Guyon, and Nihar B. Shah. Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS’24 Experiment, 2024. URL https://arxiv.org/ abs/2411.03417. [28] Nitya Thakkar, Mert Yuksekgonul, Jake Silberg, Animesh Garg, Nanyun Peng, Fei Sha, Rose Yu, Carl Vondrick, and James Zou. A large-scale randomized study of large language model feedback in peer review. Nature Machine Intelligence, pages 1–11, 2026. doi: 10.1038/s42256-026-01188-x. URL https://doi. org/10.1038/s42256-026-01188-x. [29] Tianmai M. Zhang and Neil F. Abernethy. Reviewing Scientific Papers for Critical Problems With Reasoning LLMs: Baseline Approaches and Automatic Evaluation, 2025. URL https://arxiv.org/abs/2505.23824. NeurIPS 2025 Workshop on AI for Science: The Reach and Limits of AI for Scientific Discovery. [30] M. Zhu, Y. Weng, L. Yang, and Y. Zhang. DeepReview: Improving LLM-Based Paper Review with Human-Like Deep Thinking Process. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 29330–29355, Vienna, Austria, 2025. Association for Computational Linguistics. doi: 10.18653/ v1/2025.acl-long.1420. URL https://aclanthology.org/ 2025.acl-long.1420/. [31] Y. Chang, Z. Li, H. Zhang, Y. Kong, Y. Wu, H. K.H. So, Z. Guo, L. Zhu, and N. Wong. TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-Based Scientific Peer Review. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15651– 15682, Suzhou, China, 2025. Association for Computational Linguistics. URL https://aclanthology.org/ 2025.emnlp-main.790/. [32] K. Lu, S. Xu, J. Li, K. Ding, and G. Meng. Agent Reviewers: Domain-Specific Multimodal Agents with Shared Memory for Paper Review. In Proceedings
of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 40803–40830, Vancouver, Canada, 2025. PMLR. URL https://proceedings.mlr. press/v267/lu25p.html. [33] Vianney Renata and John Lee. AI Reviewers: Are Human Reviewers Still Necessary? In Proceedings of the Human Factors and Ergonomics Society Annual Meeting, volume 69, pages 338–342. SAGE Publications Sage CA: Los Angeles, CA, 2025. [34] Jake Poznanski, Aman Rangapur, Jon Borchardt, Jason Dunkelberger, Regan Huff, Daniel Lin, Christopher Wilhelm, Kyle Lo, and Luca Soldaini. olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models, 2025. URL https://arxiv.org/abs/ 2502.18443. [35] John Gruber. Markdown: Syntax. http://daringfireball. net/projects/markdown/syntax. Retrieved on June, 2012. [36] Alexander Goldberg, Ivan Stelmakh, Kyunghyun Cho, Alice Oh, Alekh Agarwal, Danielle Belgrave, and Nihar B. Shah. Peer reviews of peer reviews: A randomized controlled trial and other experiments. PLOS ONE, 20(4):e0320444, 2025. doi: 10.1371/journal. pone.0320444. URL https://journals.plos.org/plosone/ article?id=10.1371/journal.pone.0320444. [37] Eftekhar Hossain, Sanjeev Kumar Sinha, Naman Bansal, R. Alexander Knipper, Souvika Sarkar, John Salvador, Yash Mahajan, Sri Ram Pavan Kumar Guttikonda, Mousumi Akter, Md. Mahadi Hassan, Matthew Freestone, Matthew C. Williams Jr., Dongji Feng, and Santu Karmaker. LLMs as meta-reviewers’ assistants: A case study. In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7763–7803, Albuquerque, New Mexico, April 2025. Association for Computational Linguistics. URL https://aclanthology.org/2025.naacl-long.395/. [38] Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. OpenAI GPT-5 System Card. arXiv preprint arXiv:2601.03267, 2025. [39] OpenAI. File inputs. https://developers.openai.com/ api/docs/guides/file-inputs/, 2026. OpenAI API documentation. Accessed: 2026-03-16. [40] GPTZero. GPTZero Developer Docs. https://gptzero. me/docs, 2026. API documentation. Accessed: 202603-16. [41] Michael J. Parker, Caitlin Anderson, Claire Stone, and YeaRim Oh. A large language model approach to educational survey feedback analysis. International Journal of Artificial Intelligence in Education, 35(2):444– 481, Jun 2025. ISSN 1560-4306. doi: 10.1007/s40593-
024-00414-0. URL https://doi.org/10.1007/s40593024-00414-0. [42] Alex Liu and Min Sun. From voices to validity: Leveraging large language models (llms) for textual analysis of policy stakeholder interviews. AERA Open, 11:23328584251374595, 2025. doi: 10.1177/ 23328584251374595. [43] Paul Rozin and Edward B. Royzman. Negativity Bias, Negativity Dominance, and Contagion. Personality and Social Psychology Review, 5(4):296–320, 2001. doi: 10.1207/S15327957PSPR0504 2. [44] Reanna M. Poncheri, Jennifer T. Lindberg, Lori Foster Thompson, and Eric A. Surface. A Comment on Employee Surveys: Negativity Bias in Open-Ended Responses. Organizational Research Methods, 11(3): 614–630, 2008. doi: 10.1177/1094428106295504.
A A.1
Appendix
Prompt design
We report the prompt structure and design for the AAAI26 AI Review System. The exact prompts are withheld to reduce the risk of prompt-targeted optimization or prompt injection attacks in future deployments. The review pipeline uses a layered prompt design rather than a single monolithic instruction. This sequence of decomposition, synthesis, and validation is intended to improve coverage across all review criteria and minimize errors in the review process. A persistent base instruction establishes the overall reviewing objective, the expectation of factual and objective analysis, and guidance on how to use the PDF and OCR-derived markdown versions of the paper together. This base instruction is reused across all later stages so that each stage operates with the same review context and source-handling policy. On top of this shared base layer, the system uses a fixed stack of specialized prompt components, listed in Table 3. These components are invoked in a defined order and play distinct roles in the pipeline: a base review instruction, the story, presentation, evaluations, correctness, and significance prompts, an initial review prompt, a self-critique prompt, and a final review prompt. The main purpose of this decomposition is to reduce overload, encourage deeper checking of each review criterion, and preserve intermediate findings for later synthesis. Upon completion of all core stages, the system uses a sequence of prompts to compile the findings into a single review, generate a self-critique, and revise the review to address the self-critique to produce the final review.
A.2
Quantitative Analysis Details
For each question that was included in both the AI and human review questionnaires, we compared the responses for human AI AI and human reviews. Let Rauthors and Rauthors be the collections of responses for the human and AI review questionhuman naires from the authors, respectively, and let Rreviewers and AI Rreviewers be the collections of responses for the human and AI reviewers questionnaires, respectively. Let M (R) be the mean of the response collection R: |R|
M (R) =
1 X Ri , |R| i=1
where |R| is the number of responses in the collection R. The difference in the distribution means for the human and AI reviews for the authors and reviewers separately are given by AI human M (Rauthors ) − M (Rauthors ), and AI human M (Rreviewers ) − M (Rreviewers ), AI respectively. The overall mean M (Roverall ) for the AI reviews from all respondents (both authors and reviewers) is given by AI AI AI AI |Rauthors |M (Rauthors ) + |Rreviewers |M (Rreviewers ) , AI AI |Rauthors | + |Rreviewers |
human and the overall mean M (Roverall ) for the human reviews from all respondents (both authors and reviewers) is given by human human human human |Rauthors |M (Rauthors ) + |Rreviewers |M (Rreviewers ) . human human |Rauthors | + |Rreviewers |
The overall difference in the distribution means for the human and AI reviews is thus given by AI human M (Roverall ) − M (Roverall ),
We used a Mann-Whitney U test to assess the statistical significance of the differences in the response distributions for AI and human reviews, testing against the null hypothesis that the samples are drawn from the same distribution. For example, for the author responses, we tested the null hypothAI human esis that Rauthors and Rauthors are drawn from the same distribution, against the alternative hypothesis that the two response collections are drawn from different distributions. This test was performed for each AI vs. human review comparison for overall, authors, and reviewers separately.
A.3
Survey Questionnaire
Table 4 summarizes the questionnaire items shown to respondents. There were four questionnaire variants, corresponding to whether the respondent was an author or a reviewer (PC, SPC, or AC), and whether the questionnaire was attached to an AI review or a human review. All closed-form items used the same five-point Likert scale described in Section 3, and each questionnaire included an open-ended field for free-form comments.
A.4
Preprocessing PDFs
To ensure that the multimodal tokenization of the paper PDF does not exceed the context window limit of the LLM due to the inclusion of high-resolution images, all paper PDFs are resampled to a consistent resolution of 250 DPI. We found through initial testing that the PDF version of the paper alone is insufficient for accurately reading equations and tables — initial testing demonstrated errors in the generated reviews caused by the LLM misinterpreting mathematical notation and table structures. To overcome this issue, a specialized optical character recognition (OCR) model, olmOCR [34], is used to convert the PDF version of the paper to markdown. The markdown version includes LATEX notation for the equations and inline math symbols, and a structured layout for the tables in the paper.
A.5
LLM Details for AAAI-26 AI Review System
We used the OpenAI gpt-5 LLM [38] with ‘high’ reasoning effort for all stages of the AAAI-26 AI Review System. This model was the most capable model among the OpenAI models available at the time of deployment, and had a September 30, 2024 knowledge cutoff. All OpenAI API calls were made using a special Zero Data Retention (ZDR) agreement in place to ensure that confidentiality was upheld to the highest degree — model inputs and outputs were not logged, and only existed in ephemeral memory on the OpenAI servers. The context window for this model
Table 3: Descriptions for prompts used in the AAAI-26 AI review system. The table lists the prompt components, where each is used, and a high-level synopsis of what each prompt asks the model to do. Prompt component
Where it is used
Synopsis of Prompt
Base instruction
Prepended to all stages
Story prompt
Story stage
Presentation prompt
Presentation stage
Evaluations prompt
Evaluations stage
Correctness prompt
Correctness stage
Significance prompt
Significance stage
Initial review prompt
Initial review stage
Self-critique prompt
Self-critique stage
Final review prompt
Final review stage
Explains the task to the model: to generate a factual, objective review without scores or recommendations through a sequence of subsequent steps. Also explains how to use the PDF and OCR-derived markdown as complementary sources. Asks the model to analyze the paper’s problem formulation, claimed gap in prior work, core contribution, and whether the evidence supports the main story. Asks the model to assess clarity, organization, readability, and whether the technical narrative is easy for a researcher to follow. Asks the model to inspect baselines, datasets, metrics, statistical evidence, empirical support for claims, and reproducibility-related weaknesses, using the available code interpreter tool if needed. Asks the model to verify equations, proofs, algorithms, figures, and tables, using the available code interpreter tool if needed. Asks the model to identify closely related published work at toptier venues and use it to contextualize novelty, competitiveness, and missing comparisons, with specific instructions to restrict results to published work only, and to disregard preprints or non-peer-reviewed work. Asks the model to synthesize the accumulated findings into a standardized draft review with a fixed structure. Includes details about the specific markdown variant used by the frontend, and formatting instructions to ensure rich text formatting including math notation and tables, if any, are rendered correctly. Asks the model to reread the draft review against the paper and flag unsupported claims, factual problems, and unclear or incomplete citations. Asks the model to revise the draft to address the self-critique and compile a final review.
was 400,000 tokens with 128,000 max output tokens, which included the reasoning tokens in addition to the final output tokens. For PDF inputs, OpenAI’s file-input documentation [39] states that vision-capable multimodal models process both extracted text and rendered page images, so PDF context usage reflects both textual and visual content rather than text alone. The evaluations and correctness stages included the built-in code interpreter tool with ‘auto’ container type. The significance stage included the built-in web search preview tool with ‘approximate’ user location set to ‘US’ and ‘medium’ search context size. All OpenAI LLM calls were made with the flex tier, with errorhandling and retry logic implemented to ensure robustness. During deployment, since review generation is executed in parallel for a large batch of papers, occasional service errors due to rate limits, server load, or other issues are inevitable. Upon error, the LLM calls were retried up to 5 times with exponential backoff.
A.6
Quality-Checking and Human Oversight of AI Reviews
After all reviews were generated, a second critic LLM (OpenAI o4-mini) was prompted to identify specific potential issues in the reviews as well as editorial concerns about the papers. The critic LLM was not provided with the paper or any of the original prompts, and was not told that the reviews
were generated by an AI system. The list of review issues assessed included (1) revealing author identities; (2) potentially offensive language or content in the review; (3) judgments that may have been biased based on gender, geography, or other factors; and (4) missing structural elements in the review. Editorial concerns included (1) ethical concerns, (2) inclusion of information in the paper that may reveal author identities, and (3) paper violations of conference policies such as formatting requirements. We additionally asked the critic LLM to judge auxiliary criteria including (1) whether the review appeared to be written by an LLM, (2) whether the review appeared to be written by an unqualified reviewer, (3) the apparent effort that went into the review, and (4) an overall rating of the quality of the review. The critic LLM’s judgments were compiled into a spreadsheet, after which the authors manually inspected all reviews that were identified as potentially containing one or more of the above issues. Through this process, we identified several papers that violated author anonymity through one of several means including the inclusion of specific acknowledgments or links to public non-anonymous websites. There were a few papers flagged for ethical concerns including potential software license violations and potential identity leakage in the dataset released with a paper.
Table 4: Questionnaire items by respondent group and review type. ‘A’ denotes author questionnaires and ‘R’ denotes the shared reviewer-form structure used for PC, SPC, and AC respondents. ‘AI’ and ‘H’ denote questionnaires attached to AI reviews and human reviews, respectively. Questionnaire item
A/AI A/H R/AI R/H
Overall impressions This review was a thorough review for this conference This review demonstrated capabilities beyond what I expected from AI This review changed my evaluation or interpretation of the paper This review changed my evaluation of the paper Overall, AAAI-26 AI reviews were useful Overall, AI reviews were harmful in AAAI-26 Overall, AI reviews would be useful in future peer review processes Overall, AI reviews would be harmful in future peer review processes Overall, I went into review thinking AI reviews would be useful
✓ ✓
Suggestions for research and writing This review provided useful suggestions to improve research design This review provided useful suggestions to improve the paper presentation This review provided suggestions that were wrong or unhelpful This review correctly suggested related work This review incorrectly suggested nonexistent work This review raised points to be addressed in camera-ready or resubmission This review raised points to be addressed in a revision Free response General feedback on the review and the AAAI-26 AI Review Pilot
A.6.1 Citation Hallucination Checking We randomly sampled 100 reviews generated by the AAAI-26 AI review system and checked for hallucinated citations using the GPTZero [40] API. There were 1356 citations in the sampled reviews. GPTZero identified 1346 of the citations as valid, and matched them to published work at the cited venues, with matching authors and titles. It labelled 8 citations as ‘unsure,’ and 2 citations as ‘fake.’ We further manually inspected the 8 ‘unsure’ citations and the 2 ‘fake’ citations and found that the ones marked as ‘unsure’ were indeed correctly matched to published work, and one of the ‘fake’ citations actually existed, but was a (correct) citation to a technical reference manual rather than a published work, which the tool is designed to detect. The second ‘fake’ citation had an incorrect venue — the paper existed by the cited authors and paper title, but at a different venue than the one cited.
✓ ✓ ✓
✓
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓ ✓ ✓ ✓
✓ ✓
✓ ✓
✓ ✓
✓ ✓ ✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓ ✓ ✓ ✓
✓ ✓ ✓ ✓
✓ ✓
✓
A.7
✓ ✓ ✓
✓ ✓ ✓ ✓ ✓ ✓
Review emphasis This review accurately conveyed the significance and impact This review overemphasized minor issues This review raised points that I had not previously considered This review raised important points that others might have missed This review raised important points that the human-written reviews omitted This review found concerns that a human reviewer would have difficulty catching This review overlooked points that a human reviewer would likely have caught This review discovered important points that AI would have missed This review overlooked points that AI would have caught Technical accuracy This review accurately identified technical errors This review made technical errors in the review
✓
✓
✓
✓
Survey Taxonomy Creation and Feedback Classification
We used the OpenAI gpt-5.4 LLM with ‘high’ reasoning effort to summarize the free-form survey responses into positive and negative categories, similar to previous LLMassisted survey coding approaches [41, 42]. The LLM was prompted to detail the name of the category, a one sentence summary of the category, and a more detailed description and rationale for why this category was created. We initially extracted 13 positive themes and 20 negative themes through this process; the larger number of negative themes is consistent with the well-documented negativity bias in openended survey feedback, where dissatisfied respondents are more likely to comment and comments tend to be disproportionately negative in tone [43, 44]. Of these categories, via manual inspection, we identified 24 out of the 33 categories as being specific to the AAAI-26 AI Review Pilot, vs. 9 as being general opinions about the use of AI in the academic review process, both present and future.
After creating this taxonomy, we applied it to extract individual instances of such categories mentioned in every piece of written feedback. As in the taxonomy creation process, we used gpt-5.4 to analyze the written feedback and identify if any of the categories were mentioned in the feedback, and if so, to provide attribution for which excerpts in the feedback support which categories of feedback. Each feedback item could support multiple categories, and the LLM was prompted to provide a complete list of identified categories, with attribution. The LLM was also prompted to provide a rationale for every category identified.
A.8
Sample Free-Form Survey Responses
Here, we show a couple of examples of the free-form survey responses we received that cover many of the commonly found themes, both positive and negative, among all responses. A.8.1
Anonymous Author Response 1
Overall, the AI review pilot demonstrates promising potential in enhancing the peer review process. Its most notable positive aspect is the unexpected thoroughness in technical scrutiny unlike some human reviews that might overlook niche details, the AI effectively identified overlooked technical gaps and provided contextually relevant, authentic references, which significantly aids authors in refining their work. Objectivity is another strength, as the AI’s evaluations appear less influenced by subjective biases, focusing more on alignment with established methodologies and literature. However, a potential area for improvement lies in nuanced contextual understanding; while technically precise, the AI sometimes lacks the depth of insight that comes from a human expert s long-term immersion in a specific subfield, such as recognizing the broader, unstated implications of a study’s limitations or the novelty of work that deviates slightly from mainstream approaches. Looking ahead, AI could best serve peer review as a complementary tool handling the initial screening of technical accuracy, reference verification, and structural coherence to free up human reviewers to focus on higher-level assessments of innovation, significance, and real-world impact, creating a more efficient and balanced process. A.8.2
Anonymous Author Response 2
I appreciate being able to participate in the AI review pilot. This particular review felt impressively thorough, constructive, and fair—it engaged deeply with the technical and experimental aspects of the paper and made thoughtful and highly actionable suggestions regarding methodology, reproducibility, related work, and presentation. The level of detail and breadth of recent references exceeded my initial expectations for automated reviewing. A notable strength is the AI’s ability to check consistency, pinpoint technical ambiguities (such as tokenretention accounting and KV-cache handling), and
call for more rigorous baseline and efficiency comparisons. These are aspects that even experienced human reviewers sometimes miss or only mention superficially. The AI provided a rich set of pointers for both technical and presentational improvement. On the other hand, while the AI review raised almost all the necessary concerns, it sometimes leaned toward exhaustive checklists rather than contextual prioritization or nuanced judgment about what issues most affect the overall scientific contribution. I also noticed that some points (e.g., writing/formatting inconsistency) might be less helpful in earlier-stage submissions where major revisions are expected. Overall, I see significant potential for AI to support or supplement peer review, ensuring greater thoroughness and standardization. However, human insight, especially in evaluating novelty, potential real-world impact, and the balance between theoretical elegance and practical feasibility, remains indispensable. Combining AI’s systematic coverage with human expertise could improve review quality and fairness. Thank you for running this pilot—I found it thoughtprovoking and valuable.
A.9
SPECS Benchmark Curation Process
We describe the SPECS benchmark curation process in detail. The LLM used for this process was OpenAI gpt-5.4 with ‘high’ reasoning effort. A.9.1 Paper Selection The benchmark dataset curation starts with a collection of recently accepted papers from a relevant venue — AAAI-25 in our case. It then samples papers to cover a representative set of topics across the proceedings category organization to yield a candidate pool of papers. For each paper, it searches for a matching paper on arXiv, and downloads the LATEX source code of the paper. A paper is included in the benchmark if (1) the bibliographic information in the proceedings matches that on arXiv, and (2) the LATEX source code compiles successfully without errors. The curation process yielded 120 papers from the AAAI-25 proceedings for the benchmark, with a track distribution of Game Theory and Economic Paradigms (39, 32.5%), Machine Learning (30, 25.0%), Knowledge Representation and Reasoning (14, 11.7%), Computer Vision (12, 10.0%), Constraint Satisfaction and Optimization (11, 9.2%), Data Mining & Knowledge Management (8, 6.7%), Intelligent Robots (4, 3.3%), and Humans and AI (2, 1.7%). The specific distribution was a result of the distribution of accepted papers in the AAAI-25 proceedings, random sampling from the distribution, the failure rate to find arXiv source, and the failure rate to compile the LATEX source. A.9.2 Perturbation Generation For each paper, an LLM is prompted to generate synthetic edits to the LATEX source of the paper to introduce one of five perturbation types of scientific errors: story, presentation, evaluation, correctness, and significance. The prompt used for each perturbation type describes the perturbation type and includes specific examples. Perturbation types evaluation, correctness,
and presentation include several subtypes to further classify the type of error — for example, the evaluation subtypes include missing baseline, missing metric, and data misinterpretation. The LLM is prompted with (1) the perturbation type and subtype with their associated descriptions and examples and (2) the LATEX source code of the paper; and instructed to output (1) a description of the perturbation being introduced and why it is significant to catch in a peer review, (2) the original source LATEX being edited and the associated file name and line numbers, and (3) the modified LATEX source code. A generated perturbation is accepted if the identified original source being edited matches the paper source at the correct file name and line numbers, and if the modified source compiles successfully without errors. The perturbation process yielded 783 synthetic perturbations, with a distribution of 153 (19.5%) for story, 173 (22.1%) for presentation, 159 (20.3%) for evaluation, 144 (18.4%) for correctness, and 154 (19.7%) for significance. A.9.3 Judging Each review was judged by an LLM prompted to identify if the review caught the specific error introduced in the perturbation stage for the corresponding paper. A review is considered to have caught the error only if it (1) explicitly identified the specific error, and (2) substantiated the identification with a text excerpt from the review that matches the synthetically perturbed text. To allow human audit of the judging results, the judge LLM was also prompted to output a justification for its judgment, including which excerpt from the review supported the judgment, or why the review did not catch the error. A.9.4 Human Oversight To ensure that the perturbations generated by the LLM were indeed significant scientific errors that should be identified in a peer review, a randomly sampled set of 35 perturbations were independently inspected by two human reviewers — 6 story, 7 presentation, 8 evaluation, 7 correctness, and 7 significance perturbations. Table 5: SPECS benchmark human oversight. Reviewers R1 and R2 independently reviewed each perturbation. The responses here are the verdicts of the perturbation that they declared valid. Criterion
n
R1
R2
Consensus
Story Presentation Evaluations Correctness Significance
6 7 8 7 7
5 3 5 7 3
5 4 6 6 4
5 3 5 6 3
All SPECS criteria
35
23
25
22
From the manual inspection, both reviewers agreed that 22 perturbations were valid scientific errors that should be identified in a peer review, and 9 were minor errors that did not warrant being noted in a peer review. The reviewers were split on 4 perturbations. The presentations and significance perturbations were particularly difficult — the reviewers unanimously assessed that only 3/7 for both of these types were significant enough to make a significant
impact on paper quality. The story, correctness, and evaluations perturbations, on the other hand, were relatively successful, with unanimous agreement on 5/6, 6/7, and 5/8 respectively. The Appendix summarizes the human-oversight results by criterion (Table 5). Of the valid perturbations, two additional reviewers reviewed 40 judging results from the LLM, randomly drawing from judging results from the baseline, targeted stages, and final system. All but one of the 40 judgments were unanimously agreed upon as being correct.