Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs Harry Rogers1 , Sally Shiels2 , Ashley Tomlinson2 , James Thomas2 , James Aylward3 , Nathan Gauge2 , Helen Higham2 , Alison Noble1 1 Department of Engineering Science, University of Oxford, Oxford, UK Nuffield Department of Clinical Neurosciences, University of Oxford, Oxford, UK 3 Department for Continuing Education, University of Oxford, Oxford, UK {harry.rogers, alison.noble}@eng.ox.ac.uk {sally.shiels, ashley.tomlinson, james.thomas, nathan.gauge, helen.higham}@ndcn.ox.ac.uk [email protected] 2
arXiv:2607.19063v1 [cs.AI] 21 Jul 2026
Abstract Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable to examiner subjectivity, fatigue, and cognitive bias. Standard examiner validation via inter-rater statistics lacks explanatory power regarding the source of errors, as it neither analyzes examiner reasoning nor verifies examiner claims against actual events. Thus, we introduce Quality Action Assurance (QAA), a multimodal framework that verifies examiner claims in Virtual Reality (VR) pediatric OSCEs by comparing actions claimed by examiners against the true sequence of events, constructed from video, VR logs, and actor data. QAA combines a constrained temporal action alignment model, which performs action localization and actor source attribution, with a large language model that extracts examiner claims and checks them against the record. Across a 5fold cross-validation, QAA achieves 99.2% ± 0.7% Actor F1 and 93.4% ± 1.9% W@16 for temporal alignment. Overall, QAA detects examiner errors with 70.0% precision and 76.7% recall, improving factual correctness from 39.2% to 79.2%, enabling fairer OSCE assessment.
Introduction Objective Structured Clinical Examinations (OSCEs) are the gold standard for certifying clinical competence (Khan et al. 2013). Virtual Reality (VR) OSCEs offer scalable standardization through consistent scenario playback (Neher et al. 2025; Mühling et al. 2025); yet, assessment still depends on human examiners observing, interpreting, and judging under time pressure. Marking is susceptible to cognitive noise (Haviari et al. 2024; Touma, Paco, and MacIntyre 2024): examiners may reconstruct events using schemas rather than faithfully recording them (Tversky and Kahneman 1974; Reason 1990), producing systematic factual errors undetectable by inter-rater reliability. Most published work on automated assessment in Surgical Data Science and OSCE Artificial Intelligence (AI) targets student performance from video, kinematics, or logs (Maier-Hein et al. 2017; Gazis et al. 2025; Tekin et al. 2025; Bentegeac et al. 2025), implicitly treating expert judgments as a ground truth despite known label variability (Maier-Hein et al. 2021). Psychometric studies monitor extreme examiner
effects (Fuller et al. 2016), while cognition-focused work uses think-aloud protocols to characterize scoring assumptions (Chahine, Holmes, and Kowalewski 2015; Roduta Roberts, Cook, and Chao 2020; Scully 2023). These approaches document that examiner cognition contributes to variance but do not verify whether specific claimed reasoning is factually correct. Automating the assessment itself is not the answer, because OSCE assessment is fundamentally a complex observation task that extends beyond a simple checklist of actions. Examiners must track the sequence of events, the specific actors addressed, and qualitative aspects of performance: a student’s confidence, decisiveness, and timeliness are clinically meaningful proxies for safety, and a technically correct action performed with hesitation may still indicate poor practice. System logs, such as ones within VR, record that an action occurred, not how it was performed, so they cannot replace examiner judgment. Examiners must therefore remain the assessors, yet every qualitative judgment rests on factual premises about what was done, when, and by whom, and it is precisely these premises that can be checked against an objective record. To understand how often those premises fail and how failures affect grades, we combined psychometrics with grounded verification: we examined inter-rater reliability alongside examiner verbal reasoning, and checked verbalized claims against an actor-attributed event record. Using concurrent and retrospective verbal protocols (Ericsson and Simon 1993), we identified two recurring cognitive errors: (i) Inferential, where an examiner asserts an action occurred based on suggestive cues rather than definitive completion (Schacter 1999), and (ii) Source Misattribution, where the examiner identifies the correct action but assigns it to the wrong actor (Johnson 1997). We introduce Quality Action Assurance (QAA), a multimodal framework that treats examiner verbalizations as verifiable claims and grounds them against an actor-attributed record derived from egocentric video aligned to VR action traces. QAA is designed as a decision-support system: it flags and explains potentially incorrect claims for human review rather than replacing examiner judgment. Our contributions are: 1. A cognitive-error analysis of examiner reasoning in VR OSCEs, based on two-phase verbal protocols from four
(a)
(b)
(c)
Figure 1: Egocentric view: (a) actors, (b) hover options, (c) action menu. clinical examiners, revealing a reliability paradox: examiners reach substantial-to-near-perfect agreement while making factual errors in over 60% of assessments, errors large enough to shift the borderline-regression pass mark. 2. A constrained temporal action alignment model that combines dual encoders, a monotonic alignment objective with a minimum temporal gap, and scenario knowledge constraints to localize each action and attribute it to the correct actor. 3. An LLM verifier that extracts examiner claims from transcripts, checks them against the aligned actor-attributed record, and classifies mismatches as Inferential or Source Misattribution errors with an explicit rationale. 4. A grading-impact analysis showing that correcting detected errors significantly changes per-student grades (paired Wilcoxon p < 10−4 ), lowers the cohort cut-score, and moves a borderline student from an automatic fail into the mandatory-adjudication zone.
Dataset and Cognitive Error Analysis We compiled egocentric videos and VR logs from 91 final year medical students completing a 12-minute pediatric emergency scenario on the OMS VR platform (Oxford Medical Simulation Ltd.). The study began before OMS introduced AI-first voice interaction; students therefore interacted with objects and three actors (Child, Carer, Nurse) via menus (Figure 1). This modality was methodologically useful because it generated visible, binary clicks, enabling examiner claims to be verified against a ground-truth record. VR logs record action names but not target actors which can be challenging for understanding what exactly was completed. For example, asking about allergies is logged identically whether directed to the child or carer, although only the carer reports the child’s penicillin allergy. Thus, we manually relabeled all videos with actor labels (Child, Carer, Nurse, Item) and synchronized timestamps. The Joint Research Office study classification group classified this as evaluation of educational provision, not requiring further ethical approval. Students and examiners consented; data are anonymized and simulation actors are automated avatars with pre-scripted behaviors.
Verbal Protocols and Error Annotation Four clinical examiners each independently graded a subset of 30 anonymized students using a two-phase verbal proto-
Examiner Inferential Source Misattribution Total A B C D
21 4 14 13
8 19 13 12
29 23 27 25
Total
52
52
104
Table 1: Examiner-level distribution of cognitive errors across assessments.
col: (1) Concurrent observation, where examiners verbalized what they believed the student was doing and how it related to performance, and (2) Retrospective rubric completion, where examiners completed a multi-dimensional rubric covering eight assessment domains (History taking, Communication, Escalation, Physical Exam, Prescribing, Investigations, Management, Patient safety) and a Global outcome to retrospectively justify grades. Verbalizations are transcribed from both phases of the verbal protocol. Audio utterances are segmented, transcribed with WhisperX (Bain et al. 2023), and human-validated to ensure the transcription matches what the examiner states. Taxonomy and Prevalence of Cognitive Errors. We grounded examiner claims against the actor-attributed log to identify cognitive errors, defined as explicit claims unsupported by the ground-truth record. We categorized each error as: (i) Inferential, where the examiner asserts an action occurred although it is absent from the record; or (ii) Source Misattribution, where the examiner correctly identifies the action but assigns it to the wrong actor, such as attributing a carer’s statement to the child. Across the 120 transcripts, 73 (60.8%) contained at least one error, totaling 104 errors shown in Table 1; multiple errors often occurred within the same performance or recurred across different examiners assessing the same student. We further labeled errors across three dimensions to understand their impact: (1) Domain Mapping, identifying the clinical domain the examiner was discussing; (2) Propagation, labeling an error as propagated when the same claim appeared in both concurrent observation and retrospective rubric completion, indicating that an initial perceptual error persisted to directly influence the final assessment; and (3) Direction of Impact, analyzing the examiner’s evaluative language to determine if the
(a)
(b)
(c)
Figure 2: (a) Propagated and non-propagated errors with grading effect, (b) Domain error rate vs. inter-rater reliability, (c) Domain error counts with grading effect. error inflated or deflated the student’s score. For example, a propagated error resulting in grade deflation occurred where an examiner stated during observation, “They requested a lumbar puncture,” and later penalized the student during rubric completion “...getting a lumbar puncture [was not] indicated... putting all of those at poor” despite the student never requesting the procedure in the ground-truth record. The Reliability Paradox. We measured inter-rater reliability using Gwet’s AC2 with quadratic weights (Gwet 2014) to accommodate the non-linear ordinal scales: domain-specific grades are Poor, Satisfactory, Excellent, and the Global outcome is Clear Fail, Borderline Fail, Borderline Pass, Clear Pass, Above Expected. Despite substantial-to-near-perfect agreement (AC2 ≈ 0.65–0.86), cognitive errors persist, creating a reliability paradox where examiners are statistically consistent yet factually incorrect. These errors extend beyond commentary: of 104 total errors, 53 were grade-affecting across all domains except Escalation and the Global outcome. However, since domain scores serve as anchors for standard setting using borderline regression for OSCEs and surgical assessment (Kramer et al. 2003; de Montbrun, Satterthwaite, and Grantcharov 2015), these factual errors directly compromise the calculated pass mark. Applying an oracle correction, resolving all 53 grade-affecting errors, shifted the composite cut-score from 46.3% to 45.2% (R2 = 0.69) and moved one student from fail to pass. Under a factually correct record this student would have passed; their failure was therefore an artifact of examiner error rather than a lack of clinical ability. Crucially, even non-propagated errors influenced grades, suggesting examiners introduce factual errors from memory without prior verbalization. Figure 2 visualizes the relationship between error rates, reliability, and grading effects.
Methodology: Quality Action Assurance (QAA) The QAA framework facilitates the verification of examiner reasoning via a three-stage pipeline: (1) Fine-tuned feature extraction, (2) Temporal Action Alignment, and (3) LLMbased verification shown in Figure 3.
Feature Extraction We fine-tuned three video backbones: X3D (Feichtenhofer 2020), SlowFast (Feichtenhofer et al. 2019), and VideoMAE (Tong et al. 2022), initialized from Kinetics-400 pre-training (Kay et al. 2017) on 32-frame windows of egocentric video. A 32-frame window is the temporal unit at which we pose the binary question of whether a logged action begins within it. During training we slide this window at a stride of 8 frames so that consecutive windows overlap. Backbone feature maps are spatially pooled, projected to 768 dimensions, and temporally pooled to a single embedding per window. A lightweight Transformer head (Vaswani et al. 2023) (4 layers, 8 heads) processes sequences of consecutive window embeddings after each backbone, trained on two objectives: (1) Action Presence, whether a logged action starts within the window, using focal cross-entropy (Lin et al. 2017) (γ=2); (2) Actor Classification of that action over {Child, Carer, Nurse, Item}, also with focal cross-entropy. At evaluation, a window is attributed to an actor only when the presence head fires and to a background class otherwise. For downstream alignment we discard the heads and export the per-window 768-dimensional backbone embeddings V, so the alignment stage receives purely visual features. Because the backbone is fine-tuned end-to-end through these heads, the exported features already encode actor identity as well as action presence.
Temporal Action Alignment Model We frame alignment as a structured prediction problem that maps the monotonic sequence of N scripted actions A = (a1 , . . . , aN ) onto the T overlapping stride-8 video windows summarized by the extracted features V = (v1 , . . . , vT ). The scenario prescribes the set and order of clinically valid actions but not their timing; the model must therefore recover, for every action, (i) the window in which it occurs, (ii) the actor it is directed at, and (iii) a sub-window start offset, while guaranteeing that the recovered timeline is temporally monotonic and physically plausible. Dual Encoders. Two Transformer encoders project the modalities into a shared d-dimensional space. An action encoder Ea embeds each scripted action name as a learned
Figure 3: QAA Pipeline: Feature extraction feeds a temporal action alignment model, which establishes the ground truth for LLM verification. vocabulary token and contextualizes the ordered action se(i) quence, yielding action embeddings ha ∈ Rd ; a video encoder Ev projects and contextualizes the window features, (T ) (1) yielding window embeddings Hv = (hv , . . . , hv ). We score every action–window pair with scaled dot-product attention, (i)
si,t =
(t)
⟨ha , hv ⟩ √ , d
(1)
yielding the similarity matrix S ∈ RN ×T that provides the local evidence for alignment. Constrained Monotonic Alignment. An alignment is a path P = (t1 , . . . , tN ) assigning action i to window ti . We restrict the hypothesis space to paths that are strictly ordered and separated by at least a minimum gap g, i.e. ti ≥ ti−1 + g, preventing two distinct actions from collapsing onto the same window which is the key departure from standard Connectionist Temporal Classification (CTC) (Graves et al. 2006), which permits repeats and blanks. We set this gap to the minimum temporal separation between distinct actions in our data, 32 frames; at a stride of 8 frames this is g=4 windows. The same 32-frame spacing sets the 32-frame featureextraction window: since no two actions are closer than 32 frames, a 32-frame window contains at most one action, making each window’s action-presence label unambiguous. The gap must be matched to the data: set too large, the constraint can make the ground-truth alignment infeasible, no monotonic path exists once a clip has fewer than N g windows, and set too small, it no longer keeps distinct actions apart, so in either case alignment breaks down. We score a path PN additively, score(P) = i=1 si,ti , and place a Gibbs distribution p(P | S) ∝ exp(score(P)) over the constrained set of paths. The log-partition function is computed in O(N T )
by the forward recurrence αi,t = si,t + log
t−g X
exp(αi−1,t′ ),
(2)
t′ =0
initialized with α1,t = s1,t , where αi,t accumulates the logsum of scores over all valid partial paths that align the first i actions with P action i placed at window t. The normalizer is log Z = log t exp(αN,t ). Given the ground-truth path P ⋆ , the sequence of windows containing each action’s manually annotated onset, we train by minimizing the negative loglikelihood Lalign = log Z − score(P ⋆ ), (3) which concentrates probability mass on the annotated timeline while marginalizing all competing monotonic hypotheses. Decoding. At inference the ground-truth path is unavailable, so we recover the most probable alignment with a Viterbi pass under the same gap constraint: we replace the log-sum-exp in the forward recurrence with a max, store the arg-max predecessor as a back-pointer at each step, and backtrack from arg maxt αN,t . This returns the window index ti for every action in O(N T ) time and yields a monotonic, gap-respecting timeline by construction. Actor Attribution with Knowledge Constraints. For each aligned action we attribute the target actor. A crossattention head queries the global video context Hv with (i) the action embedding ha and passes the result through a multilayer perceptron to produce actor logits zi ∈ R4 over {Child, Carer, Nurse, Item}. Because the scenario makes many action–actor pairings impossible (e.g., a history cannot be sourced from an Item), we inject scenario knowledge as a hard mask M: entries of clinically impossible pairs are set to −∞ before the softmax, p̂i = softmax(zi + mi ), confining all probability mass to admissible actors. This mask is
what distinguishes a genuine Source Misattribution from unconstrained ambiguity. Attribution is supervised with crossentropy Lactor against the relabeled actor targets. Sub-window Localization and Objective. A parallel head localizes the action onset within the aligned 32-frame window by classifying it into one of the 32 candidate frame positions δi ∈ {0, . . . , 31}, decoded by arg max at inference; it is trained with a cross-entropy loss Loff against the ground-truth offset. The three heads are optimized jointly under the weighted objective L = Lalign + λactor Lactor + λoff Loff ,
(4)
so that timing, attribution, and onset are learned end-to-end over the shared dual encoders.
LLM Examiner Verification The final stage treats each examiner transcript as a set of claims to be checked against the reconstructed record. For every assessment, a large language model (LLM) receives (1) the scenario’s action vocabulary, (2) the actor-attributed action record predicted by the alignment model for that student, rendered as actor–action pairs, and (3) the examiner’s transcribed verbalizations from both protocol phases. A single structured prompt instructs the LLM to extract every clinical action the examiner asserts, map it onto the vocabulary, and compare it against the record; it includes worked examples of a correct match, a Source Misattribution, and an Inferential error. Negative constraints filter language that does not constitute a verifiable claim: vague bundles (e.g., “doing an A2E approach”), negated or missed-opportunity critique are ignored. The LLM returns structured JSON tuples (actor, action, error type, rationale, transcript quote) under JSONconstrained decoding with default sampling. Providing the actor-attributed record is essential: actor mismatches surface as Source Misattribution, which is otherwise undetectable, while actions absent from the record surface as Inferential errors. Detections are scored against the annotated errors by greedy one-to-one matching per transcript, where normalized texts match by substring containment or character-level similarity (≥ 0.8); a transcript is fully corrected when every annotated error is detected.
Experiments and Results
Leakage is prevented fold-by-fold: features for each video are extracted by the backbone that held it out as test data, and the alignment predictions used downstream come from the fold in which that student was held out, so no stage is evaluated on students it trained on. Experiments were run with PyTorch 2.7.1 (CUDA 12.6, cuDNN 9.0.5) on a single NVIDIA RTX 6000 Ada Generation GPU (48 GB) under Ubuntu 24.04, with an Intel Xeon w5-2565X CPU and 64 GB RAM. Baselines compared each fine-tuned backbone against Kinetics-400 pre-trained weights (Kay et al. 2017). We report Raw F1 (without mask), Masked F1 (with mask M), and W@16, which counts a prediction as correct if the absolute difference between predicted and ground-truth action start frames is ≤ 16 frames (≈0.5 s). For transcript verification, we evaluated GPT-5.2 (XHigh Reasoning) (OpenAI 2025), Kimi-k2-thinking (Moonshot AI 2025), and DeepSeek-v3.2 (DeepSeek-AI et al. 2025), alongside a keyword-matching baseline. For each action in the scenario vocabulary, this baseline builds a set of lexical triggers of the action name and common clinical abbreviations to flag the action whenever a trigger occurs as a substring of the transcript, and infers the target actor from words in the surrounding context (“child”/“patient” → Child, “mum”/“carer” → Carer, otherwise the vocabulary default). Because it treats every lexical mention as a claimed action without consulting the record or filtering vague language, it isolates how much of the verifier’s performance requires grounded reasoning rather than surface lexical overlap. All verifiers are scored against the annotated errors under the identical matching protocol. To assess grading impact, we apply each method’s verified corrections to the examiner marks, recompute the borderline-regression cut-score, and test the resulting grade changes with a paired Wilcoxon signed-rank test over the 120 examiner–student checklist scores and an exact sign test over the 30 per-student composites; 95% confidence intervals for the cut-score shift come from a student-level cluster bootstrap.
Fine-tuned Feature Extraction Table 2 reports clip-level Actor F1 and Action Recall. All backbones achieve strong actor recognition, with SlowFast performing best at sequence length 5 with an Actor F1 of 98.6% ± 0.1 and an action Recall of 97.6% ± 0.3. Gains beyond this length are limited, so we use sequence length 5 in subsequent experiments.
Experimental Setup
Temporal Action Alignment
All stages use student-level 5-fold cross-validation with a fixed seed of 42 across Python, NumPy, and PyTorch (including CUDA), giving replicable splits. Feature backbones were fine-tuned with AdamW (Kingma and Ba 2014) (learning rate 5 × 10−5 , weight decay 10−4 ) for up to 50 epochs with early stopping (patience 3) on validation Actor F1, using stratified video-level folds; sequence lengths {5, 10, 15} used batch sizes 8, 4, 2. The alignment model (d=256; a 2-layer video encoder and a 1-layer action encoder, 4 heads each) was trained with AdamW (learning rate 3 × 10−4 ) for up to 30 epochs with early stopping (patience 3) on the product of W@16 and actor macro-F1, enforcing the 32-frame action gap with objective weights λactor =1.0 and λoff =0.5.
Table 3 summarizes temporal action alignment. The actor constraint mask increases Masked F1 by removing clinically impossible action actor pairs, but does not by itself resolve temporal ambiguity. Fine-tuning substantially improves both W@16 and Raw F1 relative to the Kinetics baseline, indicating that domain adaptation is necessary. SlowFast again achieves the strongest performance (W@16 93.4% ± 1.9, Raw F1 95.1% ± 4.1, Masked F1 99.2% ± 0.7), and we therefore use its predictions to ground examiner verification.
LLM Examiner Verification and Impact Using SlowFast-aligned predictions, Table 4 reports error detection performance. As QAA is decision-support, the LLM
Model SlowFast VideoMAE X3D
Seq 5 98.6 ± 0.1 95.4 ± 1.4 97.3 ± 0.2
Actor F1 (%) Seq 10 Seq 15 98.6 ± 0.2 97.9 ± 0.3 94.7 ± 1.4 94.2 ± 1.3 96.9 ± 0.3 94.6 ± 3.1
Action Recall (%) Seq 5 Seq 10 Seq 15 97.6 ± 0.3 97.6 ± 0.7 97.0 ± 0.4 93.0 ± 1.1 93.0 ± 1.4 93.2 ± 1.5 95.2 ± 0.5 95.1 ± 0.8 94.1 ± 2.1
Table 2: Cross-validation results for feature extractors. Best results in bold.
Model SlowFast VideoMAE X3D
W@16 (%) Base Seq 5 79.9 ± 1.8 93.4 ± 1.9 69.8 ± 2.7 69.1 ± 7.0 64.9 ± 2.5 93.1 ± 1.8
Raw F1 (%) Base Seq 5 67.3 ± 2.9 95.1 ± 4.1 51.0 ± 6.9 69.1 ± 17.9 67.6 ± 3.9 92.0 ± 3.0
Masked F1 (%) Base Seq 5 97.5 ± 1.0 99.2 ± 0.7 96.1 ± 1.0 95.4 ± 3.7 95.4 ± 1.3 98.4 ± 0.8
Table 3: Temporal Action Alignment Performance. Best results in bold. verifier must detect incorrect claims while avoiding excessive flags that increase reviewer burden; we therefore prioritize Inferential Recall (catching unsupported actions) and Source Misattribution Precision (avoiding incorrect actor accusations). Under this framing, GPT-5.2 offers the best balance, achieving 70.0% precision and 76.7% recall for overall detection, with an Inferential recall of 58.5% and a Source Misattribution precision of 62.0%, while producing the fewest false positives (0.44 per transcript). It fully corrects 48 of 73 error-containing transcripts, improving factual correctness from 39.2% to 79.2%. The gap to the alternatives is wide: the keyword baseline reaches only 3.0% precision at 6.29 false positives per transcript, and the other LLMs either miss more errors (Kimi-k2, 39.8% recall) or add more noise (DeepSeek, 20.5% precision). Reliable detection therefore depends on grounded reasoning over the actor-attributed record, not surface lexical cues. To quantify downstream impact, we recomputed borderline regression under each correction method (Table 5, Figure 4). The effect is broad: a full correction of all 53 errors changes the composite score of 23 of 30 students (20 fall, 3 rise; mean −1.4pp, up to −4.7pp), and the errors GPT-5.2 detects alone move 20 of 30. Examiner factual errors thus perturb the grade of the large majority of the cohort, not a few outliers, and these shifts are statistically significant, both per transcript (paired Wilcoxon over the 120 examiner– student scores, p < 10−4 ) and per student (sign test on the 30 composites, p < 10−3 ), whereas the keyword baseline moves only 4 students and is indistinguishable from no correction (p = 0.10). Significance, however, is not the substantive concern: a one- or two-point shift is immaterial for a clearly passing or failing candidate. The harm concentrates at the borderline, where the cut-score decides the outcome. Because the errors are predominantly inflations, they raise the cut-score and push marginal students down: without QAA the cut-score is 46.3% and a borderline student sits −4.1pp below it, a clear fail well outside the ±2pp review zone, produced by inflations in other students’ assessments together with two of their own actions being wrongly downgraded. Correcting the errors GPT-5.2 verifies lowers the cut-score to 45.1% and lifts the student to −1.3pp, out of an automatic fail and into the review zone where a human adjudicates; the
Figure 4: Borderline regression before (top) and after (bottom) QAA correction. Each point is a student’s mean checklist score, colored by pass/fail against the cut-score (blue line); the shaded band is the ±2 pp borderline-review zone and insets magnify it.
full (oracle) correction moves them across the line entirely (+0.1pp). A pass/fail decision was therefore determined by examiner error rather than the student’s own performance: exactly the unfairness QAA exists to surface.
Approach Keyword Baseline GPT-5.2 (XHigh) Kimi-k2-thinking DeepSeek-v3.2
Overall Detection P (%) R (%) 3.0 21.4 70.0 76.7 27.2 39.8 20.5 61.2
Inferential P (%) 100.0 53.3 21.5 24.3
R (%) 7.1 58.5 26.8 41.5
Source Misattribution P (%) R (%) 2.2 41.5 62.0 75.0 19.5 60.7 10.8 82.1
Fully Corrected
Average FPs
12 48 22 34
6.29 0.44 1.21 1.21
Table 4: LLM Verification Results. Best results in bold. Method Det. ∆cut p Without QAA 0 — — Keyword 7 +0.3 0.10 Kimi-k2 27 −1.0 <10−4 35 −1.1 <10−4 DeepSeek-v3.2 GPT-5.2 (XHigh) 42 −1.2 <10−4 53 −1.1 <10−4 Oracle
Student −4.1 (fail) −4.4 (fail) −3.1 (fail) −3.0 (fail) −1.3 (rev.) +0.1 (pass)
Table 5: Borderline-regression impact and significance of the grade shift. ∆cut is the cohort cut-score change (pp); p is a paired Wilcoxon signed-rank test over the 120 examiner– student checklist scores (corrected vs. original). “rev.” = borderline-review zone.
Discussion and Conclusion We introduced Quality Action Assurance (QAA), a multimodal framework that shifts the analytical target in medical assessment from students to the verification of examiners, grounding examiner verbalizations against an actorattributed event record. Our cognitive analysis reveals a reliability paradox: examiners exhibit substantial-to-near-perfect agreement (AC2 ≈ 0.65–0.86) yet make factual errors in over 60% of examinations, including grade-affecting errors across most domains. Standard safeguards miss this entirely, because inter-rater agreement measures whether examiners are consistent, not whether they are correct. The consequences are both broad and sharp. Broad, because these errors are not rare slips affecting a few students: correcting them moves the grades of more than three-quarters of the cohort. Sharp, because the harm concentrates at the decision boundary, where a one- or two-point shift is immaterial for a clear pass or fail but decides a borderline outcome, as our borderline case illustrates. The significance tests confirm these shifts are real rather than noise, yet the substantive point is fairness at the margin, not effect size. This vulnerability is unlikely to be specific to our scenario. Borderline regression from domain scores to a global outcome is the standard-setting method for OSCEs, and the same approach is used in surgical and other observational examinations; any assessment that aggregates human judgments into a cut-score inherits the same failure mode, in which systematic examiner error silently shifts the boundary. Because QAA depends only on synchronized video and action logs rather than on the menu-driven interface of our study, it is interface-agnostic and should transfer wherever such a record can be reconstructed. We position QAA as decision-support, not automation. It
does not grade students or overrule examiners; it converts verbalized reasoning into time-stamped, actor-attributed claims and surfaces the specific discrepancies a human should review. This makes retrospective quality assurance tractable, narrowing attention to concrete, checkable items rather than an exhaustive audit, while guarding against automation bias: the verifier can itself be incorrect, and the record it checks against is a model prediction, so its outputs must remain advisory, calibrated to a low false-positive rate, and validated per scenario, with student and examiner data subject to appropriate privacy safeguards. Experimentally, the constrained alignment attains 99.2% ± 0.7 Masked F1 for actor attribution, and the LLM verifier raises factual correctness from 39.2% to 79.2% at a low false-positive rate, moving a student who would otherwise fail outright into the adjudication range. Our study has limitations. The cognitive-error analysis draws on 30 students, four examiners, and 120 transcripts from a single menudriven pediatric scenario on one platform (alignment models trained on 91 students), so error-prevalence rates and the single observed pass/fail flip are indicative rather than definitive; establishing how common such errors are requires larger multi-scenario, multi-site studies. The menu-driven modality, chosen because it yields verifiable clicks, is less representative of emerging voice-based VR OSCEs, which would require speech-to-action grounding. Error annotations came from a single researcher; because the core determination, whether a claimed action or actor appears in the objective VR log, is a factual check rather than a quality judgment, exposure to annotator bias is limited, and multi-rater validation of the more interpretive grade-impact labels is planned. Finally, the verifier recovers 42 of 53 grade-affecting errors and checks against a predicted rather than certified record, so its outputs remain advisory and full correction requires human adjudication. Even so, QAA demonstrates that multimodal grounding makes examiner reasoning auditable, converting claims into verifiable actor–action tuples anchored to a time-aligned record; by improving factual correctness and reducing manual checking, it supports fairer standard setting across observational examinations vulnerable to cognitive bias, while preserving examiner authority through human adjudication.
Acknowledgments H.R. and A.N. acknowledge the EPSRC Turing AI Fellowship “Ultrasound Multi-Modal Video-based Human– Machine Collaboration” [EP/X040186/1]. We thank the examiners and medical students who participated in the VR
OSCE marking studies. The authors declare no competing interests.
References Bain, M.; Huh, J.; Han, T.; and Zisserman, A. 2023. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. In Interspeech 2023, 4489–4493. Bentegeac, R.; Florens, N.; Maanaoui, M.; Maisons, V.; Lanot, A.; Bobot, M.; Brilland, B.; Glowacki, F.; Gérard, E.; Hazzan, M.; Amouyel, P.; Le Guellec, B.; and Hamroun, A. 2025. ECOSBot: a multicenter validation pilot study of a generative AI tool for OSCE-based nephrology training. Clin Kidney J, 18(10): sfaf308. Chahine, S.; Holmes, B.; and Kowalewski, Z. 2015. In the minds of OSCE examiners: uncovering hidden assumptions. Adv Health Sci Educ Theory Pract, 21(3): 609–625. de Montbrun, S.; Satterthwaite, L.; and Grantcharov, T. P. 2015. Setting pass scores for assessment of technical performance by surgical trainees. Br J Surg, 103(3): 300–306. DeepSeek-AI; Liu, A.; Mei, A.; Lin, B.; Xue, B.; Wang, B.; Xu, B.; Wu, B.; Zhang, B.; Lin, C.; Dong, C.; Lu, C.; Zhao, C.; Deng, C.; Xu, C.; Ruan, C.; Dai, D.; Guo, D.; Yang, D.; Chen, D.; Li, E.; Zhou, F.; Lin, F.; Dai, F.; Hao, G.; Chen, G.; Li, G.; Zhang, H.; Xu, H.; Li, H.; Liang, H.; Wei, H.; Zhang, H.; Luo, H.; Ji, H.; Ding, H.; Tang, H.; Cao, H.; Gao, H.; Qu, H.; Zeng, H.; Huang, J.; Li, J.; Xu, J.; Hu, J.; Chen, J.; Xiang, J.; Yuan, J.; Cheng, J.; Zhu, J.; Ran, J.; Jiang, J.; Qiu, J.; Li, J.; Song, J.; Dong, K.; Gao, K.; Guan, K.; Huang, K.; Zhou, K.; Huang, K.; Yu, K.; Wang, L.; Zhang, L.; Wang, L.; Zhao, L.; Yin, L.; Guo, L.; Luo, L.; Ma, L.; Wang, L.; Zhang, L.; Di, M. S.; Xu, M. Y.; Zhang, M.; Zhang, M.; Tang, M.; Zhou, M.; Huang, P.; Cong, P.; Wang, P.; Wang, Q.; Zhu, Q.; Li, Q.; Chen, Q.; Du, Q.; Xu, R.; Ge, R.; Zhang, R.; Pan, R.; Wang, R.; Yin, R.; Xu, R.; Shen, R.; Zhang, R.; Liu, S. H.; Lu, S.; Zhou, S.; Chen, S.; Cai, S.; Chen, S.; Hu, S.; Liu, S.; Hu, S.; Ma, S.; Wang, S.; Yu, S.; Zhou, S.; Pan, S.; Zhou, S.; Ni, T.; Yun, T.; Pei, T.; Ye, T.; Yue, T.; Zeng, W.; Liu, W.; Liang, W.; Pang, W.; Luo, W.; Gao, W.; Zhang, W.; Gao, X.; Wang, X.; Bi, X.; Liu, X.; Wang, X.; Chen, X.; Zhang, X.; Nie, X.; Cheng, X.; Liu, X.; Xie, X.; Liu, X.; Yu, X.; Li, X.; Yang, X.; Li, X.; Chen, X.; Su, X.; Pan, X.; Lin, X.; Fu, X.; Wang, Y. Q.; Zhang, Y.; Xu, Y.; Ma, Y.; Li, Y.; Li, Y.; Zhao, Y.; Sun, Y.; Wang, Y.; Qian, Y.; Yu, Y.; Zhang, Y.; Ding, Y.; Shi, Y.; Xiong, Y.; He, Y.; Zhou, Y.; Zhong, Y.; Piao, Y.; Wang, Y.; Chen, Y.; Tan, Y.; Wei, Y.; Ma, Y.; Liu, Y.; Yang, Y.; Guo, Y.; Wu, Y.; Wu, Y.; Cheng, Y.; Ou, Y.; Xu, Y.; Wang, Y.; Gong, Y.; Wu, Y.; Zou, Y.; Li, Y.; Xiong, Y.; Luo, Y.; You, Y.; Liu, Y.; Zhou, Y.; Wu, Z. F.; Ren, Z. Z.; Zhao, Z.; Ren, Z.; Sha, Z.; Fu, Z.; Xu, Z.; Xie, Z.; Zhang, Z.; Hao, Z.; Gou, Z.; Ma, Z.; Yan, Z.; Shao, Z.; Huang, Z.; Wu, Z.; Li, Z.; Zhang, Z.; Xu, Z.; Wang, Z.; Gu, Z.; Zhu, Z.; Li, Z.; Zhang, Z.; Xie, Z.; Gao, Z.; Pan, Z.; Yao, Z.; Feng, B.; Li, H.; Cai, J. L.; Ni, J.; Xu, L.; Li, M.; Tian, N.; Chen, R. J.; Jin, R. L.; Li, S. S.; Zhou, S.; Sun, T.; Li, X. Q.; Jin, X.; Shen, X.; Chen, X.; Song, X.; Zhou, X.; Zhu, Y. X.; Huang, Y.; Li, Y.; Zheng, Y.; Zhu, Y.; Ma, Y.; Huang, Z.; Xu, Z.; Zhang, Z.; Ji, D.; Liang, J.; Guo, J.; Chen, J.; Xia, L.; Wang, M.; Li, M.; Zhang, P.; Chen, R.; Sun, S.; Wu, S.;
Ye, S.; Wang, T.; Xiao, W. L.; An, W.; Wang, X.; Sun, X.; Wang, X.; Tang, Y.; Zha, Y.; Zhang, Z.; Ju, Z.; Zhang, Z.; and Qu, Z. 2025. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv:2512.02556. Ericsson, K. A.; and Simon, H. A. 1993. Protocol Analysis: Verbal Reports as Data. Cambridge, MA: MIT Press. ISBN 9780262550239. Feichtenhofer, C. 2020. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 203– 213. Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 6202–6211. Fuller, R.; Homer, M.; Pell, G.; and Hallam, J. 2016. Managing extremes of assessor judgment within the OSCE. Med Teach, 39(1): 58–66. Gazis, A.; Schizas, D.; Kykalos, S.; Karaiskos, P.; and Loukas, C. 2025. Egocentric video analysis for automated assessment of open surgical skills via deep learning. Int J Comput Assist Radiol Surg, 21(2): 297–306. Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006. Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning, ICML ’06, 369–376. New York, NY, USA: Association for Computing Machinery. ISBN 1595933832. Gwet, K. L. 2014. Handbook of Inter-Rater Reliability: The Definitive Guide to Measuring the Extent of Agreement Among Raters. Gaithersburg, MD: Advanced Analytics, LLC, 4th edition. ISBN 9780970806284. Haviari, S.; de Tymowski, C.; Burnichon, N.; Lemogne, C.; Flamant, M.; Ruszniewski, P.; Bensaadi, S.; Mercier, G.; Hamaoui, H.; Université Paris Cité OSCE study group; Mirault, T.; Faye, A.; and Bouzid, D. 2024. Measuring and correcting staff variability in large-scale OSCEs. BMC Med Educ, 24(1): 817. Johnson, M. K. 1997. Source monitoring and memory distortion. Philos Trans R Soc Lond B Biol Sci, 352(1362): 1733–1745. Kay, W.; Carreira, J.; Simonyan, K.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; Suleyman, M.; and Zisserman, A. 2017. The Kinetics Human Action Video Dataset. arXiv:1705.06950. Khan, K. Z.; Ramachandran, S.; Gaunt, K.; and Pushkar, P. 2013. The Objective Structured Clinical Examination (OSCE): AMEE Guide No. 81. Part I: an historical and theoretical perspective. Med Teach, 35(9): e1437–46. Kingma, D. P.; and Ba, J. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. Kramer, A.; Muijtjens, A.; Jansen, K.; Düsman, H.; Tan, L.; and van der Vleuten, C. 2003. Comparison of a rational and an empirical standard setting procedure for an OSCE. Objective structured clinical examinations. Med Educ, 37(2): 132–139.
Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; and Dollár, P. 2017. Focal Loss for Dense Object Detection. In 2017 IEEE International Conference on Computer Vision (ICCV), 2999–3007. Maier-Hein, L.; Eisenmann, M.; Sarikaya, D.; März, K.; Collins, T.; Malpani, A.; Fallert, J.; Feussner, H.; Giannarou, S.; Mascagni, P.; Nakawala, H.; Park, A.; Pugh, C.; Stoyanov, D.; Vedula, S. S.; Cleary, K.; Fichtinger, G.; Forestier, G.; Gibaud, B.; Grantcharov, T.; Hashizume, M.; HeckmannNötzel, D.; Kenngott, H. G.; Kikinis, R.; Mündermann, L.; Navab, N.; Onogur, S.; Roß, T.; Sznitman, R.; Taylor, R. H.; Tizabi, M. D.; Wagner, M.; Hager, G. D.; Neumuth, T.; Padoy, N.; Collins, J.; Gockel, I.; Goedeke, J.; Hashimoto, D. A.; Joyeux, L.; Lam, K.; Leff, D. R.; Madani, A.; Marcus, H. J.; Meireles, O.; Seitel, A.; Teber, D.; Ückert, F.; Müller-Stich, B. P.; Jannin, P.; and Speidel, S. 2021. Surgical data science - from concepts toward clinical translation. Med Image Anal, 76: 102306. Maier-Hein, L.; Vedula, S. S.; Speidel, S.; Navab, N.; Kikinis, R.; Park, A.; Eisenmann, M.; Feussner, H.; Forestier, G.; Giannarou, S.; Hashizume, M.; Katic, D.; Kenngott, H.; Kranzfelder, M.; Malpani, A.; März, K.; Neumuth, T.; Padoy, N.; Pugh, C.; Schoch, N.; Stoyanov, D.; Taylor, R.; Wagner, M.; Hager, G. D.; and Jannin, P. 2017. Surgical data science for next-generation interventions. Nat Biomed Eng, 1(9): 691–696. Moonshot AI. 2025. moonshotai/Kimi-K2-Thinking. Hugging Face model card. Modified MIT license. Mühling, T.; Schreiner, V.; Appel, M.; Leutritz, T.; and König, S. 2025. Comparing Virtual Reality-Based and Traditional Physical Objective Structured Clinical Examination (OSCE) Stations for Clinical Competency Assessments: Randomized Controlled Trial. J Med Internet Res, 27: e55066. Neher, A. N.; Bühlmann, F.; Müller, M.; Berendonk, C.; Sauter, T. C.; and Birrenbach, T. 2025. Virtual reality for assessment in undergraduate nursing and medical education - a systematic review. BMC Med Educ, 25(1): 292. OpenAI. 2025. Introducing GPT-5.2. https://openai.com/ index/introducing-gpt-5-2/. Reason, J. 1990. Human Error. Cambridge University Press. Roduta Roberts, M.; Cook, M.; and Chao, I. C. I. 2020. Exploring assessor cognition as a source of score variability in a performance assessment of practice-based competencies. BMC Med Educ, 20(1): 168. Schacter, D. L. 1999. The seven sins of memory. Insights from psychology and cognitive neuroscience. Am Psychol, 54(3): 182–203. Scully, C. 2023. Assessor cognition and inter-rater reliability in nursing objective structured clinical examinations. Ph.D. thesis. Tekin, M.; Yurdal, M. O.; Toraman, Ç.; Korkmaz, G.; and Uysal, İ. 2025. Is AI the future of evaluation in medical education?? AI vs. human evaluation in objective structured clinical examination. BMC Med Educ, 25(1): 641. Tong, Z.; Song, Y.; Wang, J.; and Wang, L. 2022. Videomae: Masked autoencoders are data-efficient learners for
self-supervised video pre-training. Advances in neural information processing systems, 35: 10078–10093. Touma, N. J.; Paco, C. A.; and MacIntyre, I. 2024. Interobserver variance of examiner scoring in urology Objective Structured Clinical Examinations. Can Urol Assoc J, 18(4): 116–119. Tversky, A.; and Kahneman, D. 1974. Judgment under Uncertainty: Heuristics and Biases. Science, 185(4157): 1124– 1131. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2023. Attention Is All You Need. arXiv:1706.03762.