ConceptioArchivearXiv CS
arXiv CSopen access

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment Marı́a Eugenia Curi1 , Germán Capdehourat1*, Isabel Amigo1 , Magdalena Romano2 , Rosana Serra2 , Adrián Silveira2 , Andrés Peri2

arXiv:2609.05143v1 [cs.CL] 4 Sep 2026

1

2

Ceibal, Avda. Italia 6201, Montevideo, 11500, Uruguay. ANEP, Av. Libertador 1409, Montevideo, 11100, Uruguay.

*Corresponding author(s). E-mail(s): [email protected]; Abstract The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows. Keywords: automated essay scoring, large language models, rubric-based assessment

1

1 Introduction The recent advance of large language models (LLMs) has renewed interest in the use of artificial intelligence (AI) to support educational assessment processes, particularly in tasks involving written production [1]. Automated essay scoring has been studied for decades; however, recent advances in generative AI have significantly expanded the range of feasible approaches, especially for short argumentative texts that require holistic and rubric-based evaluation [2]. Despite these advances, evidence from real large-scale operational contexts remains limited, and important questions persist regarding reliability, validity, and the appropriate role of human oversight [1]. This study presents the integration of LLM-based scoring into the evaluation process of Spanish written productions within a national large-scale assessment. The target task involves short argumentative texts of approximately 150-200 words, evaluated using a detailed analytic rubric by trained human raters. The assessment constitutes a high-stakes certification pathway for over-age students seeking to complete lower secondary education, which places strong requirements on scoring quality, consistency, and fairness. A key strength of this work lies in the availability of operational data at scale. The study draws on two recent editions of the national test (2024 and 2025), each comprising approximately 5,000–6,000 participants. This setting provides a rare opportunity to examine AI-assisted scoring under realistic conditions, including established human scoring workflows, expert-designed rubrics, and longitudinal comparability across test forms. The objective of this research is to design and validate an AI-assisted assessment process that can be integrated into the existing evaluation workflow. Beyond measuring raw agreement between AI and human raters, the study investigates how AI-based scoring behaves within the full decision pipeline, including proficiency-level classification and pass–fail outcomes. In addition, the work explores the consistency of the model, its generalization across test editions, and its potential to reduce human scoring workload while preserving decision quality. Based on the empirical findings, we propose a Human-in-the-Loop (HITL) evaluation framework that strategically combines automated scoring with expert human review. The proposed approach aims to accelerate result reporting, optimize the use of human expertise, and maintain the quality standards required in high-stakes assessment contexts. To guide the empirical evaluation, the study addresses the following research questions: RQ1: To what extent can a large language model, guided by detailed prompting, achieve agreement with expert human raters in the assessment of argumentative writing in Spanish within a large-scale national examination? RQ2: Can a Human-in-the-Loop (HITL) evaluation framework be designed so that potentially critical AI scoring errors are systematically identified and reviewed by expert human raters, thereby preserving the fairness and reliability required for high-stakes pass/fail decisions? RQ3: To what extent can AI-assisted scoring reduce human scoring workload while maintaining the quality of assessment decisions? The main contributions of this paper are the following: 2

• An empirical evaluation of LLM-based scoring in a real national large-scale writing assessment. • A cross-year analysis (2024–2025) examining robustness and generalization of prompt-based scoring. • A detailed comparison between AI scoring and human inter-rater agreement at the rubric-item level. • The design and simulation of a Human-in-the-Loop operational framework for AIassisted scoring. • An estimation of the potential reduction in human scoring workload under realistic deployment assumptions. Overall, this work contributes new evidence on the practical integration of LLMs into high-stakes educational assessment and outlines a feasible pathway toward responsible, scalable AI-assisted evaluation.

2 Literature Review and Research Positioning Automatic evaluation of student responses has been the subject of research for many years. Several solutions based on NLP techniques, prior GPT technologies, were explored and proposed [3]. In recent years, the subject has regained attention, since LLMs provide not only the possibility to increase accuracy in those cases where previous results exist, but also promise to expand the domains of application, as well as allowing for personalized and timely feedback. This ultimately results in a significant reduction in teacher workload, while fostering the improvement of learning outcomes. All of this without requiring labeled data for training. Automated grading has been proposed for various educational settings. Indeed, a large set of use cases focus on the linguistic domain. Evaluations are in the vast majority text-based, and might include short answers, open-ended questions and essays. Most studies focus on English language, see for instance [4] for a recent result. Further text-based use-cases involve computer science domain [5] and science [6], with databases that are usually in English. Fewer studies focus on the chemistry (see e.g. [7]) and mathematics domains (see e.g. [8]), where handwritten responses are usually given, and evaluation involves combining visual and textual information, thus being more complex to handle. A whole other range of use cases involve video or audio automatic evaluation, we shall leave out of the scope of this review such cases. We also focus on recent studies, as the rapid advances of technology capture the more relevant solutions. Research specifically addressing AI-assisted grading of Spanish writing remains limited. Existing studies have primarily focused on proof-of-concept evaluations using relatively small datasets or educational settings with limited operational constraints (e.g. [9]), and few have investigated rubric-based assessment of argumentative writing in large-scale examinations (e.g. [10]). As a result, evidence regarding the applicability of LLM-assisted grading to high-stakes Spanish-language assessments is still scarce, leaving important questions about reliability, operational deployment, and human oversight largely unexplored.

3

In this context, we are particularly interested in reviewing related work to gain insights on to what extent have previous results succeeded in these automated tasks, what metrics have they used in order to establish validation of the method and which techniques have been used to optimize such metrics. In addition, how human oversight is addressed by the proposed methods is of paramount importance to our real use case, as well as what human perceptions about introducing AI the evaluation flow are. It is with these different lenses that we analyze and summarize related work in the following lines.

2.1 Techniques for LLM-assisted grading In [1], an exhaustive and recent review on automated grading is presented. The authors reviewed 42 papers, out of which the most frequently represented disciplines are computer science and foreign languages at the top, followed by mathematics and medicine and fewer cases of engineering and finally social sciences. This plethora of use cases makes the related literature rich, although the results are not easy to generalize. Studies on Spanish writing assessment remain comparatively scarce and are mostly limited to specific educational settings, such as Spanish as a second language [10]. To the best of our knowledge, no previous work has investigated LLM-assisted grading of Spanish argumentative writing within a large-scale, highstakes assessment. Not surprisingly, [1] highlights that LLMs work best in English, although they are evolving and increasing their capabilities in other languages. Also, according to the same review, results are better when rubrics are used, or examples are given, and when tasks require short and structured answers, rather than long ones or where opinion is involved. For use cases more similar to ours, [1] concludes that expert oversight is necessary, along with detailed rubrics and carefully defined prompts. This is also the approach of the framework proposed in [11]. [8], [12] and [13] are also examples where the need of well designed and detailed rubric is emphasized. The need of human oversight is indeed pointed out by several studies, not only for legal and ethical reasons but also for performance ones (see e.g. [7] and [1]). How to measure performance and when to trigger human intervention is addressed in the following.

2.2 Metrics for validating LLM-assisted grading systems Accuracy is one of the most used metrics in the literature and is defined as the level of agreement between AI and human expert evaluators, with some variations in order to capture non random agreements such as Cohen’s Kappa and Quadratic Weighted Kappa (QWK) to adjust for chance agreement, as well as Fleiss’ Kappa and Krippendorff’s Alpha for multi-rater scenarios. In addition to accuracy or precision, some articles use standard metrics such as F1 score, recall, or even normed F1, to handle highly granulated grades [7]. [14] highlights that beyond accuracy, for ethical and legal reasons, it is of paramount importance to take into account metrics based on false positives or false negatives, which translates into measuring the performance of the method in terms of over-grading and under-grading. Indeed, under-grading poses equity concerns as it unfairly penalizes students, for instance with conceptually correct

4

answers but with atypical phrasing [5], while over-grading are errors that are difficult to detect, as it is less likely that an audit is requested by students in this case [7]. Other metrics aim to capture trust in the evaluation model, or in other words, the reliability of the prediction. This concept is frequently referred as uncertainty. In [15] a comprehensive benchmark of different uncertainty metrics is performed for different settings. The study provides valuable insights and concludes that no metric is universally optimal, and selecting the proper one should be carefully studied in each setting. Among the insights, they highlight that categoric metrics have shown to be more effective in capturing uncertainty than semantic ones. In the following, we illustrate different definitions of uncertainty metrics found in the literature. Typically, uncertainty metrics are obtained by calculating some index based on the output of querying the model repeatedly with the same input. Semantic metrics group the output according to their semantic meaning. In particular, semantic entropy (see e.g. [16]) is based on querying the model several times to generate explanations to the same response. Then explanations are clustered according to their meaning. Finally, the probability of each cluster is calculated, and the Shannon formula for entropy is applied. High values of entropy are interpreted as a signal of either poor rubric specification, difficult to understand response, or high uncertainty of the model. Low values of entropy, in turn, are obtained when all explanations are fairly similar semantically, and interpreted as a signal of confidence on the model output. The method is elegant and as shown in [16] it can capture uncertainty quite well, though it also fails when the LLM is confident but incorrect. Also, results do not necessarily generalize to different settings. On the other hand, categoric approaches rely on the labels or grades given by the model, rather than semantic explanations. Categoric entropy [15] is calculated in a way similar to Semantic entropy but frequency of grades are considered (rather than concepts). A simpler version is Max-Agree-Rate (MAR) [15], that measures the proportion of the most frequent grade. A large MAR value is obtained when the model is consistent and is interpreted as a signal of low uncertainty. A low value signals high uncertainty. Yet another approach is the one based on Item Response Theory (IRT). Rather than measuring uncertainty related to the model’s internal consistency, this approach measures how “surprising” is one output based on the student demonstrated abilities and the item difficulty [7]. The uncertainty is then measured as the difference of the expected grade (obtained from the IRT modelling) and the grade given by the LLM. Regardless of the way of calculating uncertainty, its effectiveness must be carefully study for each particular setting. Effectiveness of an uncertainty metric can be measured using for instance methods like AUC or C-Index. AUC (Area Under the Receiver Operating Characteristic curve) [17] measures the probability that a randomly chosen incorrect response is assigned an uncertainty value higher than a randomly chosen correct response. The higher the AUC value, the better the metric is at discriminating reliable from unreliable model outputs. The C-Index (Concordance Index introduced in [18]) is a rank-based metric, it assesses whether responses that have larger true errors also receive higher uncertainty scores. The higher the C-Index the more effective the uncertainty metric is.

5

Beyond accuracy and uncertainty, other metrics are found in the literature that are worth citing. For instance, robustness, defined as sensitivity to synonyms in prompts and to prompt injection, is measured in [6]. Beyond technical metrics, some works focus on measuring explainability and transparency of outputs of the model, for instance using SHAP (see [19] and [11]), these studies also highlight that human in the loop approaches increase transparency. Finally, the time saved by automated grading is measured in [14].

3 Research Gap and Contribution Taken together, the literature provides a rich and nuanced set of metrics for evaluating LLM-assisted grading systems, ranging from agreement-based measures to more sophisticated approaches for estimating uncertainty, robustness, and explainability. However, comparatively little attention has been devoted to how these models can be safely integrated into operational assessment workflows, particularly in high-stakes contexts where grading outcomes directly determine certification decisions. Likewise, evidence on large-scale Spanish-language assessments remains scarce. Consequently, an important research gap lies not only in achieving high agreement with expert raters, but also in designing deployment strategies that preserve fairness and reliability despite the inevitable errors of AI models. In large-scale, high-stakes evaluation settings such as the one considered in this study, the practical relevance of a given metric depends not only on its theoretical properties but also on how it informs operational decisions within the overall assessment process. In this work, while we draw on standard metrics such as accuracy, inter-rater agreement, and consistency to validate the alignment between AI and human raters, our primary focus lies elsewhere. Specifically, the evaluation perspective is shifted from model-centric performance to decision-oriented impact, analyzing how AI-generated scores affect final outcomes when integrated with other components of a standardized test, such as multiple-choice sections. This approach prioritizes the identification of cases in which AI predictions may alter pass/fail decisions, thereby requiring human review, rather than relying solely on uncertainty estimates to determine model trustworthiness. In doing so, the work emphasizes the design of an operational workflow that balances efficiency and reliability in real-world conditions, aligning evaluation practices with the ultimate goal of fair and accurate decision-making. All in all, automated evaluation based on LLMs and well-designed rubrics and prompts has shown promising results for English written assessments. However, as aforementioned, relatively few studies focus on Spanish texts. Given that LLMs are not equally trained across different languages, findings from English-based studies are not necessarily transferable to other contexts, such as Spanish [1]. Addressing this gap, a key contribution of this study is the use of real-world Spanish data, combined with expert validation, within a human-in-the-loop evaluation framework. Building on state-of-the-art models, this study designs and implements an AIassisted evaluation system for Spanish argumentative writing in a high-stakes national exam. A prompt-engineering-based approach is developed and validated on largescale datasets, and is integrated into a human-in-the-loop workflow that explicitly

6

balances efficiency with expert supervision. The proposed framework is guided by a central principle: ensuring that no final decision affecting test outcomes is made without appropriate human oversight. Beyond evaluating model performance, this work therefore contributes an operational framework for integrating LLM-based scoring into a high-stakes assessment process, demonstrating how Human-in-the-Loop decision making can mitigate fairness risks while substantially reducing expert grading workload.

4 Context and Case Study The study focuses on a knowledge accreditation test called Prueba Nacional de Acreditación de la Educación Media Básica (Acredita EB)1 , organized by the National Public Education Administration (ANEP), the public authority responsible for primary and secondary education in Uruguay. This is a large-scale assessment that has been administered annually since 2020, with 3,604 participants in its first edition and reaching its highest number of participants in 2024 (6,204). The target population consists of individuals over 21 years of age who completed primary school but subsequently dropped out of the educational system. The test provides an opportunity to obtain lower secondary education certification, equivalent to completing the third grade of basic secondary school. The continuity of its administration over the years has enabled the compilation of a substantial volume of human-scored data, contributing to the validity and reliability of the results. Because the test tasks vary each year and the rubric has evolved over time, only data from the two most recent editions were considered for this study. The 2024 scoring data were first used to develop an AI-based model to support the grading process, and the 2025 data were subsequently used for validation.

4.1 Structure of the Exam The test consists of three sections: Reading Comprehension, Problem Solving, and Writing. It is administered individually on a computer, without access to study materials or internet resources. All sections are completed within a single test session lasting 2 hours and 50 minutes, including a 10-minute break. The sections are graded independently, with separate technical teams and examiners responsible for each of them. The Reading Comprehension section evaluates the ability to understand, interpret, and critically analyze written texts. It includes locating explicit information, inferring global themes and implicit meanings, identifying the author’s intentions, and critically reflecting on texts in relation to social and communicative contexts. In addition, it assesses linguistic awareness, such as recognizing and interpreting the use of punctuation, connectors, graphic resources, and the meanings of words and expressions—both literal and figurative—within a text. The Problem Solving section assesses individuals’ ability to understand and characterize problems, design and implement appropriate solution strategies, and evaluate 1

Prueba Acredita EB - ANEP: https://acredita.anep.edu.uy/.

7

Calibration phase

Double-evaluation phase

Single-evaluation phase

Random sample of written texts

New random sample of written texts

Remaining sample of texts

Human evaluation of the sample

Two evaluators per text using rubric

One evaluator per text using rubric

Improvements to evaluation rubric

Human supervision of the sample

Human supervision of the sample

Fig. 1: Evaluation workflow for the Writing section.

and communicate results. This includes identifying and organizing relevant information, relating variables and making inferences, applying knowledge and basic procedures, evaluating the quality of information sources and solutions, and clearly communicating strategies, reasoning, and evidence-based conclusions. In the Writing section, candidates are asked to produce an argumentative text on a specific topic. Each year, a different theme is defined, addressing different types of issues, such as environmental–ecological, ethical–social, and other related topics. Participants are expected to develop a clear position, provide supporting arguments, and demonstrate control of text organization, cohesion, and language conventions. The first two sections consist primarily of multiple-choice questions and are therefore scored automatically. In contrast, the written production is evaluated manually using a detailed rubric and a procedure established by a specialized language team. As a result, the evaluation of the Writing section is the most time- and resource-intensive component of the overall scoring process, and applicants typically receive their final results no sooner than three months after the examination.

4.2 Evaluation Procedure Since the multiple-choice sections are scored automatically, most of the human effort in the test evaluation process is devoted to the Writing section. Two different groups of professionals participate in this process: the language assessment expert team and the raters. The language experts define the evaluation rubric, which is described in detail in the next subsection. They also train and supervise the raters’ application of the rubric. Figure 1 illustrates the current evaluation workflow for the Writing section. The first stage, known as the calibration phase, involves the expert team rating a random set of written texts (typically around 50), with each response evaluated by multiple raters. This stage concludes with discussion meetings in which the coding of each response is reviewed in order to reach a consensus score. The outcome of this stage may lead to adjustments to the rubric. In Figure 1, this phase corresponds to the blocks grouped within the red frame.

8

In the second stage, based on a new and larger random sample of written productions, a double-scoring procedure is conducted. Each text is scored by two raters and supervised by a member of the expert team. The goal of this stage is to measure interrater agreement in order to verify that the rubric and its application are sufficiently clear and reliable. This phase is highlighted within an orange box in Figure 1. In the third stage, shown within a blue box in Figure 1, the remaining written responses are evaluated by a single rater. At this stage, a random sample of responses is selected for expert team review. By the end of this stage, all written responses have been assigned scores. After the evaluation process is complete, each of the three test sections receives a numerical score for every participant. Next, the same methodology is applied independently to each section in order to determine the sufficiency levels. An Item Response Theory (IRT) model [20–22] is used to estimate the latent variable representing each student’s ability based on their responses. The bookmark method [23, 24] is then applied, combining the statistical calibration from the IRT model with expert judgment to define the cut scores that establish three possible performance levels for each test section: Proficient, Close to Proficiency, and Insufficient. Finally, based on the resulting proficiency levels across the test sections, a pass/fail decision is made for each student. In order to pass the test, participants must obtain at least two sections rated as Proficient, and the remaining section must be rated as either Proficient or Close to Proficiency.

4.3 Writing Evaluation Rubric As previously mentioned, the evaluation of the Writing section each year is based on a scoring rubric defined by a team of language assessment experts. This rubric constitutes one of the most important technical components of the assessment. It is organized into three complementary domains: discursive, textual, and orthographic. The first focuses on content and communicative intention, the second on organization and textual flow, and the third on technical correctness. Each domain groups specific indicators that assess key aspects of written production, assigning scores according to predefined criteria. The evidence associated with each indicator follows differentiated coding schemes, which may take the form of either dichotomous criteria or progressive scales depending on the type of performance being evaluated. Due to space limitations and confidentiality restrictions associated with the operational assessment, the full rubric and its application manual are not included in this article; however, the domains and their respective items are described below.

Discursive Domain The discursive domain refers to the ability of the text to fulfill its communicative purpose, to be appropriately structured, and to employ language suitable for an academic context. It evaluates how the student addresses the task, organizes ideas, and uses language to communicate a clear and coherent message. This domain includes three areas, each of which contains two rubric items. Communicative adequacy is evaluated by verifying whether the text expresses an opinion and whether it presents a defined position supported by arguments. The structure of the text is assessed through two 9

items that verify whether it is properly organized, including an appropriate introduction and conclusion. Finally, language use is evaluated through the items register and vocabulary. Register is assessed according to whether the text is appropriate for the communicative situation, without markers of informality or orality. Vocabulary evaluation examines whether the words used are appropriate for the context and sentence structure, without usage errors or unnecessary repetitions that could affect fluency. All items in this domain use binary scoring values, except for argumentation, which has three levels of proficiency.

Textual Domain The textual domain focuses on the use of various grammatical mechanisms at both the text and sentence levels. These mechanisms contribute to the overall unity of the text, either by maintaining referential continuity or by preserving the relationships between different ideas. This domain is organized into four areas comprising seven scoring items. Coherence is evaluated through two items: thematic progression and paragraph structure, highlighting the importance of a clear organization that facilitates reading. Two additional items correspond to the evaluation of textual cohesion: the use of connectors or linking expressions, and the use of pronominal references. The rubric also evaluates grammatical agreement, both nominal—between nouns and their determiners, or between nouns and their attributes or predicatives—and verbal, that is, between subject and verb. Finally, the last item in this domain corresponds to syntax, which evaluates the structural completeness and coherence of sentences. All items in this domain use binary scoring values. Orthographic Domain The orthographic domain evaluates the correct application of the conventions governing written language, including punctuation, accentuation, and spelling, as well as the overall formal presentation of the written production. This domain focuses on identifying errors that compromise legibility and affect the effective communication of the message. The evaluation considers both the absence and the incorrect use of orthographic signs and rules. This domain is coded using two items: punctuation and spelling. Punctuation is evaluated on a scale from 0 to 2, where 0 corresponds to three or more errors, 1 corresponds to one or two errors, and 2 corresponds to the absence of punctuation errors. For spelling, the number of errors is counted (up to a maximum of 10) and divided into quartiles to assign levels of sufficiency, resulting in four possible performance levels. Rubric Summary In summary, the 15 rubric items are: Opinion, Argumentation, Introduction, Conclusion, Register, Vocabulary, Thematic progression, Paragraph structure, Connector use, Pronominal references, Nominal agreement, Verbal agreement, Syntax, Punctuation, and Spelling. The preceding description of each domain and its respective items illustrates the technical complexity involved in interpreting the rubric for scoring purposes. This complexity is an important aspect to consider when attempting to replicate the task using a language model and when designing the corresponding 10

prompts. Finally, it is important to note that the writing evaluation rubric does not assign predefined weights to its 15 items. Instead, proficiency levels are determined as previously mentioned, by applying Item Response Theory (IRT) to estimate the psychometric properties of the items, followed by the Bookmark method to establish the cut scores associated with each proficiency level.

5 Methodology In this section, we describe the methodology used to develop the AI-based model for evaluating written responses. We begin with an overview of the dataset and then present an illustrative example of the general prompt structure.

5.1 Dataset Description The dataset considered in this study corresponds to the last two editions of the test (2024 and 2025). For each edition, we have the proficiency levels achieved by each participant in each test section (Reading Comprehension, Problem Solving, and Writing). Furthermore, for the Writing section, the dataset includes the written responses produced by each participant and the item-level scores assigned by the raters for each rubric criterion. For the 50 texts used in the calibration phase, we also have the item-level scores assigned independently by each member of the language expert team, as well as the final consensus scores reached after discussion. For the remaining texts, only a single score assigned by a rater is available; this score represents the final result, validated through random sample supervision by the team of language experts. For the purposes of this study, the available data allow us to analyze agreement among human evaluators during the calibration stage. In 2024, ten evaluators participated in this process, where they all separately corrected each of the 50 text productions in the calibration sample. This enables us to establish an empirical upper bound on the performance that could reasonably be expected from an AI-based model. Figure 2 presents the results of this analysis for the 2024 calibration phase. For each rubric item, the figure reports the minimum, average, and maximum agreement between individual raters and the consensus score reached after expert discussion. It can be observed that there is no item for which all evaluators agree on the final assigned score, and that some items show greater discrepancies than others. The results show that the minimum value, that is, the rater exhibiting the greatest discrepancy with the consensus score, generally achieves agreement below 80%. In contrast, the maximum agreement exceeds 90% for several rubric items, although for others it remains closer to 80%. It is usually considered that the level of agreement between raters for this type of test should be at least 70% to be acceptable. To complement the percentage agreement analysis, we also computed Cohen’s Kappa coefficient for each rubric item, providing a chance-corrected measure of interrater agreement. According to the interpretation scale proposed by [25], more than half of the rubric items achieved either substantial agreement (Introduction, Register, Paragraph structure, and Nominal agreement) or moderate agreement (Conclusion, Pronoun usage, Subject–verb agreement, Syntax, and Punctuation). Two additional 11

100

Max Mean Min

Inter-rater Agreement (%)

80

60

40

20

0

Re g

Pa O I T A N Co P Sp Sy V Pr Su C o nta oca rag pini ntrod hem rgum omi ell bje onn nc unc in bu nou tu n lu ati uc x rap on er n u ct ve ecto lar c p enta al ag sion atio g tio hs rs y s n rb n tio rog ree ag tru ag e res n me ctu ree sio nt re me n

ist

nt

Fig. 2: Inter-rater agreement analysis in the 2024 Calibration Phase.

items, Vocabulary and Thematic progression, showed slight agreement. Finally, Opinion, Argumentation, and Connectors exhibited fair agreement. Inspection of these latter cases revealed that the relatively low Kappa values are largely explained by the skewed distribution of scores, as the vast majority of responses satisfy these rubric criteria, reducing the coefficient despite the high observed agreement among raters. Finally, considering the operational human scores as the reference labels for this study, we are able to evaluate the performance of the AI-based model. It is important to note that multiple independent human ratings are only available for the calibration sample discussed in the previous section. For the remaining responses, a single operational score assigned by an expert rater is available, which is the score used in the official assessment process and is therefore adopted as the reference label throughout this study. As the inter-rater agreement analysis showed, some level of disagreement between equally qualified human raters is expected, implying that these reference labels may contain a limited number of inconsistencies. Nevertheless, they represent the only labels available for the complete dataset and therefore constitute the operational ground truth against which the AI model is evaluated. For this purpose, a random sample of 1,000 written responses was selected. This procedure was applied to both the 2024 and 2025 datasets. In both cases, the data used for performance evaluation were not used to refine the corresponding prompts. The selected sample

12

size ensures the statistical validity of the performance estimates while avoiding the increased costs associated with running the AI model on the entire dataset.

5.2 Prompt Engineering The strategy used to develop the AI-based model consisted of operationalizing each rubric item through dedicated prompts that assign a score (generally binary) to each written response, identify the most relevant text excerpt, and provide a justification for the assigned score. The original dataset does not include evidence supporting the score assigned to each rubric item. However, generating such evidence is considered useful for supporting and explaining the scores produced through AI-based evaluation. To choose the model we only had two possible options: open models running locally but with reduced computing capacity, or the use of an institutional licensed account to run proprietary models. After preliminary tests, the selected LLM was the OpenAI GPT-5 model, which was the most advanced reasoning model from OpenAI at the time of conducting the study. The default values were used for the parameter configuration, with the reasoning effort level set to Medium. Additionally, the structured output functionality was used to ensure the JSON format of the model’s output. A prompt-based strategy was implemented: for each rubric item, a specific prompt was designed in Spanish, based on the rubric criteria and the detailed description of how the item should be evaluated. Each prompt was initially tested on a sample of 10 texts. This iterative process made it possible to identify cases not explicitly covered in the rubric description but implicitly considered by human raters during manual scoring. After refinement, the system was applied to a sample of 1,000 written responses (approximately one fifth of the total dataset) to estimate the agreement between AIbased scoring and the final human-assigned scores for each response. Human evaluation was treated as the ground truth, as the objective was to approximate as closely as possible the assessments performed by expert human raters. The final prompts follow the structure outlined below:

− Introduction. E.g.: “Correct errors of ITEM in an argumentative text, strictly according to the rubric detailed below. It is not allowed to assume rules that are not explicitly stated here. Definitions related to item.” − Item description. E.g.: “What is evaluated? Agreement between the noun and its determiners, or the noun and its predicative.” − Evaluation Rules. Rules applied by the evaluator to penalized or not the text using the item description. − Explanation of Codes. “Assign Code 1 when ...” and ”Assign Code 0 when ...” − List of examples. Correct Example (Code 1) and Incorrect Example (Code 0). − Output Format. E.g.: “Output Format (maximum 8 tokens in fragment and explanation, maximum 3 errors; if Code 1 → “errors”: [])”: { "code": "0", "errors": [ {"fragment": ’este crianza puede verse’,

13

"explanation": ’Determinante ("este") no concuerda con crianza’}] }

5.3 Evaluation Metrics The model outputs were compared with human scores at both the item level and the aggregate level. Accuracy was computed as the proportion of exact matches between the model predictions and the final human evaluations. In addition, we conducted a cross-year validation to assess prompt generalization. Specifically, models were tested on the 2025 responses using rubric prompts and contextual patterns derived from the 2024 data, without any retraining. This procedure provides insight into the robustness of the approach across exam editions.

6 Results This section presents the results of the experiments conducted with the proposed AIbased model. First, we analyze its performance during the calibration phase of the 2024 test and evaluate the consistency of its outputs (i.e., whether the AI model assigns the same score consistently). Next, we examine how this performance generalizes to a large sample from the 2024 test. We then assess whether the model generalizes to data from the following year’s test (2025). Finally, we analyze how the AI model’s results affect the decision-making process regarding writing proficiency levels and, ultimately, the pass/fail outcomes.

6.1 Calibration AI-model Accuracy and Consistency Analysis The first analysis focuses on the 2024 calibration phase. Based on these data, minor adjustments were made to the prompts derived from the evaluation rubric. Therefore, these results can be interpreted as reflecting the model’s performance during a development stage, since the same data used for evaluation also informed prompt refinement. Furthermore, each of the 50 texts in the calibration phase was evaluated 10 times by the AI model to assess output consistency. Figure 3 reports the average AI model accuracy across the 10 runs. For reference, the average inter-rater agreement is also shown for comparison. As expected, the model’s results fall slightly below the agreement observed between human raters. However, for more than half of the rubric items, this difference is approximately 5%, and for the remaining items, it never exceeds 15%. In absolute terms, AI model accuracy ranges between 60% and 80%. Regarding the AI-model output consistency, the results are presented in Figure 4. For the vast majority of rubric items, consistency approaches or exceeds 90%. This indicates that, on average, the same score is assigned in 9 out of 10 repeated evaluations, reflecting a high degree of stability. It is important to note that the AI model’s output is inherently non-deterministic, which makes consistency analysis particularly relevant.

14

Mean inter-rater

90

Mean AI model

Agreement/Accuracy (%)

80 70 60 50 40 30 20 10 0

Re g

Pa O I T A N C P S S V P S C ra pin ntro hem rgu om onc unc pel ynt oca rono ubje onn bu er grap ion duct atic men inal lusi tuat ling ax lar un u ct ve ector ion ion hs pro tati agr on y s ag rb a s ee tru gre on e gre me ctu ss em nt re ion e

ist

nt

Fig. 3: AI model accuracy vs. average inter-rater agreement in the 2024 Calibration Phase.

The only item that shows substantially lower consistency corresponds to spelling, which aligns with the typical behavior of large language models. For this reason, the use of the AI model for this rubric item was replaced with a deterministic NLPbased approach. The selected method was LanguageTool [26], an open-source grammar checker that enables the detection of grammatical and spelling errors. A filter was also applied to the LanguageTool output to remove corrections that the rubric explicitly assigns to other items rather than to spelling (e.g. nominal or subject-ver agreement). This approach was validated with the calibration sample, where a comparable performance was achieved with that obtained with LLMs (68%), but with the advantage that this approach, being deterministic, all executions give the same result (its consistency is 100%). In the following results, this method is used to score the spelling-related rubric item, while the remaining items are evaluated using the developed AI model.

6.2 Large-scale AI model performance for 2024 and 2025 tests After the evaluation using the calibration-phase data, the first challenge was to determine whether the model maintained its performance in a large-scale evaluation using data that had not been used during the prompt adjustment stage. To this end, a random sample of 1,000 texts from the 2024 test writing section was selected, and the model’s accuracy was analyzed using the scores assigned by human raters as ground 15

100

Mean Consistency (%)

80

60

40

20

0

Op in

ion

Su

bje

Int N Pu Re P C Ar Pa Th Sy C V Sp g rod omi e nta onn oca nc r ell gis rono onc ing b tu n lus ume agra mat ec t uc x ct un ic tor ular nta ph ion tio al ag atio er ve u p y s s n rb n tio rog str ree ag ag n u e res me ctu ree sio nt re me n nt

Fig. 4: AI model consistency analysis in the 2024 Calibration Phase. Error bars indicate 95% confidence intervals.

truth. Figure 5 shows the results. The most notable finding is that the model’s performance remained similar to that observed in the calibration phase, with accuracy ranging between 60% and 80% for most rubric items. The items with the lowest scores are those related to vocabulary, syntax, and spelling. This can be explained by the relative complexity of the rubric in specifying what is penalized and what is not—distinctions that are difficult for an LLM operating at the token level to capture. The next step was to extend the evaluation to data from a subsequent year. For this purpose, a new sample of 1,000 texts was considered, this time from the writing section of the 2025 test. At the prompt level, only minor adjustments were made to those rubric items that explicitly refer to the writing task instructions, as the specific topic assigned to candidates changes every year. Figure 5 presents the results, showing that the observed performance variations with respect to 2024 were very small. This is a relevant finding, as it indicates that prompts developed for one year’s assessment can generalize to another year’s test without major effort, assuming a similar level of agreement with human raters. To complement the agreement analysis, Cohen’s Kappa coefficients were also computed for the 2025 evaluation sample, providing a chance-corrected measure of agreement between the AI model and the operational human scores. Using the same

16

100

2024 2025

Accuracy Human-AI (%)

80

60

40

20

0

Op in

Ar Re Th Co Int No Co Pa Su Pr Pu Vo Sy Sp g o r ion ume giste ema nclu odu mina nnec ragra bject nou nctu cabu ntax elling tic sio cti ati nu lar r la tor nta ph ve o o p n y g s s n r n tio rog str ba ree ag n uc res me tur gree e sio nt e me n nt

Fig. 5: Human–AI agreement by rubric item for random samples from the 2024 and 2025 assessments. Error bars indicate 95% confidence intervals.

interpretation scale proposed by [25] as in the human inter-rater analysis, the agreement levels were, as expected, lower than those observed between human evaluators during the calibration phase. Most rubric items achieved either moderate agreement (Introduction, Conclusion, Register, Nominal agreement, and Subject–verb agreement) or fair agreement (Opinion, Paragraph structure, Pronoun usage, Syntax, and Punctuation). The remaining items (Argumentation, Vocabulary, Thematic progression, and Connectors) showed slight agreement. Similar to the human inter-rater analysis, some of these lower Kappa values are explained by the skewed distribution of scores for rubric items that are satisfied by the vast majority of responses, while others correspond to the rubric dimensions exhibiting the lowest agreement throughout the study. Overall, the results are consistent with the expected reduction in agreement when comparing AI-based evaluations against operational human scores instead of comparisons among expert human raters.

6.3 Accuracy on Proficiency Levels and Pass/Fail Decision As described in Section 4.2, the final score assigned to each written production does not depend directly on individual item scores, but rather on cut scores established after applying Item Response Theory and the Bookmark method. Accordingly, to

17

0.3 Insufficient

12.3%

2.1%

0.5

0.1%

Insufficient

7.1%

0.7%

0.2%

0.2 Close to proficiency

4.0%

8.8%

0.2%

0.15

0.1

Proficient

11.3%

29.2%

32.0%

Insufficient

Close to proficiency

Proficient

0.4

Human correction

Human correction

0.25

0.05

0.3 Close to proficiency

3.5%

0.4% 0.2

Proficient

AI correction

3.2%

17.5%

13.4%

53.9%

Insufficient

Close to proficiency

Proficient

0.1

AI correction

(a) Level confusion matrix 2024

(b) Level confusion matrix 2025

Fig. 6: Agreement between AI-based and human scoring across performance levels

approximate the final evaluation of the texts, this process was replicated in an automated manner by first applying IRT and then defining cut scores based on the number of correctly answered items. Specifically, the automatically derived cut scores on the IRT scale correspond to a 67% probability of correctly answering at least seven rubric items for the Close to proficiency level and ten rubric items for the Proficient level. These thresholds are very close to the operational cut scores established by expert raters through the Bookmark standard-setting procedure, indicating that the simplified automated approach provides a good approximation of the official proficiency-level classification while enabling a consistent comparison between AI- and human-generated rubric scores. The same automated procedure was applied to both the AI-generated rubric scores and the human scoring results, ensuring that proficiency levels were assigned using identical cut-score criteria. This provides a fair and consistent basis for comparing the two evaluation approaches, independently of the original operational scoring workflow. Figure 6 presents the confusion matrices for 2024 and 2025. In both cases, the AI model tends to penalize responses more strictly, indicating a consistent and predictable bias rather than random error. This systematic behavior suggests that AI-based scoring could effectively support the overall evaluation process when integrated into an appropriate Human-in-the-Loop workflow, as discussed in the following section. Finally, the proficiency-level data from the other sections of the test (Reading Comprehension and Problem Solving) were combined with the AI-model results for the Writing section in order to evaluate agreement in pass–fail decisions for the overall test (see Figure 7). The results again highlight the effect of the AI model’s more stringent scoring of written production, which leads to a substantial number of cases in which tests classified as passing by human raters would be classified as failing under AI-based scoring. These cases (15.3% in 2024 and 16.5% in 2025) are of particular interest, as they provide evidence supporting the need for human review of texts for which the AI predicts a failing outcome.

18

0.4 0.4 Fail

41.2%

0.2%

Fail

0.2

Pass

15.3%

43.2%

Fail

0.6%

Human correction

Human correction

0.3

36.7%

0.3

0.2

Pass

0.1

Pass

16.5%

46.2%

Fail

0.1

Pass

AI correction

AI correction

(a) Pass-fail confusion matrix 2024

(b) Pass-fail confusion matrix 2025

Fig. 7: Agreement in pass-fail decisions for the overall assessment

Importantly, this bias is systematic, with very few instances observed in the opposite direction—namely, cases in which the AI model assigns a passing decision contrary to the judgment of human evaluators. A more detailed analysis by proficiency level showed that this conservative behavior is not uniform across all categories. The AI model achieved higher agreement for responses classified as Proficient than for those classified as Insufficient, a pattern consistent with the confusion matrices presented in Figure 6. Consequently, the observed bias is primarily associated with under-grading rather than over-grading. While such behavior would be problematic if AI scores were used as final decisions, it provides valuable information for the design of the proposed Human-in-the-Loop workflow. In particular, responses classified as passing by the AI model can be accepted with a high degree of confidence, whereas responses classified as failing are systematically routed to expert human raters for verification before any final certification decision is made. Under this operational framework, the identified bias does not propagate to the final assessment outcome, since the cases with the greatest equity risk are precisely those subject to mandatory human review.

7 Human-in-the-Loop AI-Assisted Evaluation The results obtained are encouraging and support the feasibility of incorporating AI into the scoring of written productions while maintaining quality standards and consistency with human evaluation. Although variability is observed in the LLM’s accuracy when scoring individual rubric items, often not reaching high levels of agreement with human raters, the overall tendency of the model is conservative, systematically identifying errors that are not penalized by human evaluators. While this behavior might initially appear unfavorable, when performance is examined at the proficiency-level

19

Calibration phase (50100 texts) AI ratings for all Writing responses

Reading Comprehension & Problem Solving proficiency levels

Set proficiencylevels cut scores Fails due to other test sections?

Yes

No Is sufficient in all test sections?

Yes

Update proficiencylevels cut scores

Final Writing evaluation results

No Human review of AI ratings

Fig. 8: Proposed Human-in-the-Loop AI-assisted evaluation framework.

stage and subsequently combined with the results from the other test sections, the model demonstrates acceptable overall performance. Based on these findings, a new scoring process is proposed that integrates AI-assisted Writing assessment to improve the efficiency of the overall evaluation workflow. The other sections of the test—namely Reading Comprehension and Problem Solving—are proposed to continue being scored using the current operational procedures. Figure 8 presents a diagram of the proposed Human-in-the-Loop evaluation framework. The first stage of the process corresponds to the calibration phase, which would be conducted following procedures similar to those used in previous editions, with the possibility of increasing the number of written productions considered (between 50 and 100). This stage would serve to validate the AI model, adapt the prompts to the specific topic of the current test, and identify potential issues arising from the calibration texts. The second stage consists of the automated scoring of all written productions using the AI model. This step produces an initial set of item-level scores for each rubric criterion. Subsequently, IRT and the Bookmark method are applied to establish cut

20

scores and assign the three proficiency levels defined in the scale. Candidates whose pass–fail outcome depends on the Writing section are then reviewed by human raters, supported by the AI-based scoring. It is important to recall that, to pass the test, candidates must achieve a Proficient level in at least two sections, and the remaining section must be rated at least Close to Proficient. This implies that candidates who fail to reach at least Proficient in one section and Close to Proficient in another (considering Reading Comprehension and Problem Solving) cannot pass the test, regardless of their performance in the Writing section. Therefore, reviewing the AI-based Writing score in such cases would not affect the overall test outcome and is consequently unnecessary. The next decision point in the workflow occurs when all three test sections are classified as Proficient. This situation arises when Reading Comprehension and Problem Solving are rated Proficient and, following AI-based scoring, the Writing section is also classified as Proficient. In this case, the candidate passes the test. The only scenario that could reverse this outcome would be if the Writing section were actually Insufficient, since even a Close to Proficient rating would still meet the minimum requirement for passing. Therefore, the likelihood of a candidate passing due to an AI scoring error in this context is 0.2% in 2024 and 0.6% in 2025, given that it would require a two-level misclassification in the Writing proficiency scale. In our analyses, such cases were not observed among candidates who had achieved Proficient in the other test sections. Finally, a recalibration of the proficiency-levels cut scores is proposed using again IRT+Bookmark. This aims to account for potential adjustments resulting from human review, which may lead to slight shifts in performance levels and, consequently, in candidates’ final outcomes. Based on data from the most recent test editions and the AI model performance observed in this study, the proposed process would reduce by at least 50% the number of written productions requiring full human scoring. The exact magnitude of this reduction would vary across editions, as it also depends on candidate performance in the other test sections.

8 Conclusion and Future Directions This study explored the feasibility of integrating large language models into the scoring process of a large-scale standardized writing assessment. Using real operational data from the 2024 and 2025 test editions, we developed and evaluated a prompt-based AI scoring approach aligned with an existing analytic rubric and human evaluation workflow. The results show that, although item-level agreement between the AI model and human raters varies across rubric criteria, the overall behavior of the model is stable and systematically conservative. When performance is examined at the proficiencylevel stage and within the full pass–fail decision process, the AI-assisted approach achieves acceptable alignment with human-based evaluation. Cross-year validation further suggests that the proposed prompting strategy generalizes across test editions with minimal adjustments.

21

A key contribution of this work is the proposal of a Human-in-the-Loop (HITL) evaluation framework that strategically integrates AI-based scoring into the operational workflow. Simulation results indicate that this approach could reduce by at least 50% the volume of written responses requiring full human scoring, while preserving decision quality in high-stakes outcomes. Beyond efficiency gains, the proposed framework maintains expert oversight in critical cases, supporting a balanced approach between automation and human judgment. Despite these encouraging findings, several limitations and open questions remain. First, the analysis relies on agreement with human raters as ground truth, which itself exhibits non-negligible variability. Second, the study was conducted under offline experimental conditions, and operational deployment may surface additional challenges related to monitoring, drift, and edge cases. Future work will therefore focus on multiple directions. A primary next step is the controlled pilot implementation of the proposed HITL workflow in upcoming test editions, in order to measure its impact under real operational conditions. Continued longitudinal monitoring of AI–human alignment will also be necessary to detect potential performance drift over time. Another important research line concerns human–AI interaction effects. In particular, it will be important to study whether prior exposure to AI-generated scores influences human raters’ judgments, potentially introducing anchoring or automation bias. Experimental studies with blinded and non-blinded rating conditions could help quantify this effect. Additional work is also needed to refine rubric-sensitive prompting strategies, especially for linguistic micro-level features such as vocabulary, syntax, and punctuation, where agreement gaps remain larger. Exploring hybrid approaches that combine LLM-based reasoning with deterministic NLP tools represents a promising direction. Overall, the findings provide empirical evidence that carefully designed Humanin-the-Loop AI scoring systems can support large-scale writing assessment processes. With continued validation and responsible deployment, such approaches have the potential to improve scalability, timeliness, and consistency in educational assessment contexts.

Acknowledgements We gratefully acknowledge the work of the expert human raters for their participation.

Declarations Funding: This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Data privacy management: The study was conducted using secondary assessment data provided by the National Public Education Administration (ANEP) of Uruguay. The dataset made available to the authors was fully anonymized and contained only the written responses together with their corresponding human-assigned scores. No personally identifiable information was included in the data used for this research.

22

LLM-based evaluations were performed through the OpenAI Enterprise API. According to the enterprise service terms, data submitted through the API are not retained for model training or used to improve future OpenAI models, thereby preserving the confidentiality of the assessment data. Data availability: The dataset analyzed in this study consists of responses from an operational nationwide assessment administered by the National Public Education Administration (ANEP) of Uruguay. Due to confidentiality restrictions, the dataset is not publicly available.

References [1] Jukiewicz, M. & Wyrwa, M. Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback. Applied Sciences 16, 680 (2026). URL https://www.mdpi. com/2076-3417/16/2/680. [2] Mizumoto, A. & Eguchi, M. Exploring the potential of using an AI language model for automated essay scoring. Research Methods in Applied Linguistics 2, 100050 (2023). URL https://www.sciencedirect.com/science/article/pii/ S2772766123000101. [3] Burrows, S., Gurevych, I. & Stein, B. The Eras and Trends of Automatic Short Answer Grading. International Journal of Artificial Intelligence in Education 25 (2015). [4] Li, J., Huang, J., Wu, W. & Whipple, P. B. Evaluating the role of ChatGPT in enhancing EFL writing assessments in classroom settings: A preliminary investigation. Humanities and Social Sciences Communications 11, 1268 (2024). URL https://www.nature.com/articles/s41599-024-03755-2. [5] Pecuchova, J., Benko, Ľ. & Drlik, M. Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models. International Journal of Artificial Intelligence in Education 35, 3813–3846 (2025). URL https://link. springer.com/10.1007/s40593-025-00517-2. [6] Deng, H., Farber, C., Lee, J. & Tang, D. Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness. arXiv preprint arXiv:2601.08843 (2026). URL https://arxiv.org/abs/2601.08843. Version Number: 1. [7] Cvengros, J. & Kortemeyer, G. Assisting the grading of a handwritten general chemistry exam with artificial intelligence. arXiv preprint arXiv:2509.10591 (2025). URL https://arxiv.org/abs/2509.10591. [8] Caraeni, A., Scarlatos, A. & Lan, A. Evaluating gpt-4 at grading handwritten solutions in math exams. arXiv preprint arXiv:2411.05231 (2024). URL https: //arxiv.org/abs/2411.05231. 23

[9] Capdehourat, G., Amigo, I., Lorenzo, B. & Trigo, J. On the effectiveness of llms for automatic grading of open-ended questions in spanish (2025). URL https://arxiv.org/abs/2503.18072. arXiv:2503.18072. [10] Zimotti, G. Local Language Models for Spanish L2 Writing Assessment: Human Baselines, RAG, and Fine-Tuning. Journal of Computers in Education (2026). URL https://doi.org/10.1007/s40692-026-00399-w. [11] Selvam, M. & González Vallejo, R. Human-in-the-loop models for ethical ai grading: Combining ai speed with human ethical oversight. EthAIca 4, 413 (2025). URL https://ai.ageditor.ar/index.php/ai/article/view/413. [12] Chu, Y., Li, H., Yang, K., Copur-Gencturk, Y. & Tang, J. LLM-Based Automated Grading with Human-in-the-Loop. 2025 IEEE International Conference on Teaching, Assessment, and Learning for Engineering (TALE) 1–8 (2025). URL https://ieeexplore.ieee.org/document/11346771/. Conference Name: 2025 IEEE International Conference on Teaching, Assessment, and Learning for Engineering (TALE) ISBN: 9798331598419. [13] Sonkar, S. et al. Olney, A., Chounta, I.-A., Liu, Z., Santos, O. & Bittencourt, I. (eds) Automated long answer grading with ricechem dataset. (eds Olney, A., Chounta, I.-A., Liu, Z., Santos, O. & Bittencourt, I.) Artificial Intelligence in Education - 25th International Conference, AIED 2024, Proceedings, Lecture Notes in Computer Science, 163–176 (Springer Science and Business Media Deutschland GmbH, Germany, 2024). [14] Craig, P. et al. AI-Marking Assistant: A Web-Based Application for Humanin-the-loop GAI Assisted Assessment Marking and Feedback. Asean Journal of Engineering Education 9, 162–170 (2025). URL https://ajee.utm.my/index.php/ ajee/article/view/213. [15] Li, H. et al. How uncertain is the grade? a benchmark of uncertainty metrics for llm-based automatic assessment (2026). URL https://arxiv.org/abs/2602.16039. arXiv:2602.16039. [16] Iyer, K., Ravikiran, M., Pendse, P. & Mohanty, S. Towards Transparent AI Grading: Semantic Entropy as a Signal for Human-AI Disagreement. arXiv preprint arXiv:2508.04105 (2025). URL https://arxiv.org/abs/2508.04105. Version Number: 1. [17] Xia, Z., Xu, J., Zhang, Y. & Liu, H. Che, W., Nabende, J., Shutova, E. & Pilehvar, M. T. (eds) A Survey of Uncertainty Estimation Methods on Large Language Models. (eds Che, W., Nabende, J., Shutova, E. & Pilehvar, M. T.) Findings of the Association for Computational Linguistics: ACL 2025, 21381– 21396 (Association for Computational Linguistics, Vienna, Austria, 2025). URL https://aclanthology.org/2025.findings-acl.1101/.

24

[18] Raykar, V. C., Steck, H., Krishnapuram, B., Dehing-Oberije, C. & Lambin, P. Platt, J. C., Koller, D., Singer, Y. & Roweis, S. T. (eds) On ranking in survival analysis: bounds on the concordance index. (eds Platt, J. C., Koller, D., Singer, Y. & Roweis, S. T.) Proceedings of the 21st International Conference on Neural Information Processing Systems, Vol. 20 of NIPS’07, 1209–1216 (Curran Associates Inc., Red Hook, NY, USA, 2007). [19] Pasupuleti, M. K. Human-in-the-Loop AI: Enhancing Transparency and Accountability. International Journal of Academic and Industrial Research Innovations(IJAIRI) 05, 574–585 (2025). URL https://www.nationaleducationservices. org/humanintheloop-ai-enhancing-transparency-and-accountability/ pid-2230865795. [20] Hambleton, R. K., Swaminathan, H. & Rogers, H. J. Fundamentals of item response theory Fundamentals of item response theory (Sage Publications, Inc, Thousand Oaks, CA, US, 1991). Pages: x, 174. [21] van der Linden, W. J. (ed.) Handbook of Item Response Theory: Three Volume Set 1 edn (Chapman and Hall/CRC, 2016). [22] Chen, Y., Li, X., Liu, J. & Ying, Z. Item response theory—a statistical framework for educational and psychological measurement. Statistical Science 40, 167–194 (2025). [23] Karantonis, A. & Sireci, S. G. The bookmark standard-setting method: A literature review. Educational Measurement: Issues and Practice 25, 4–12 (2006). URL https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1745-3992.2006.00047.x. [24] Mitzel, H., Lewis, D., Patz, R. & Green, D. The bookmark procedure: Psychological perspectives, 249–281 (2001). [25] Landis, J. R. & Koch, G. G. The measurement of observer agreement for categorical data. Biometrics 33, 159–174 (1977). [26] Naber, D. A rule-based style and grammar checker. GRIN Verlag Munich, Germany (2003).

25

Record · ID 660872 · SHA-256 8e8f1ae4c0f1bd46
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.