Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
arXiv:2604.17820v1 [cs.SE] 20 Apr 2026
DONGLIN LI, Anhui University of Science and Technology, China DAMING LI, Independent Researcher, USA HANYUAN SHI, Independent Researcher, China JIALU ZHANG∗ , University of Waterloo, Canada Block-based programming environments such as Scratch are widely used in introductory computing education, yet scalable and reliable automated assessment remains elusive. Scratch programs are highly heterogeneous, event-driven, and visually grounded, which makes traditional assertion-based or test-based grading brittle and difficult to scale. As a result, assessment in real Scratch classrooms still relies heavily on manual inspection and delayed feedback, introducing inconsistency across instructors and limiting scalability. We present Raven, an automated assessment framework for Scratch that replaces program-specific state assertions with instructor-specified, task-level video generation rules shared across all student submissions. Raven integrates large language models with video analysis to evaluate whether a program’s observed visual and interactive behaviors satisfy grading criteria, without requiring explicit test cases or predefined outputs. This design enables consistent evaluation despite substantial diversity in implementation strategies and interaction sequences. We evaluate Raven on 13 real Scratch assignments comprising over 140 student submissions with groundtruth labels from human graders. The results show that Raven significantly outperforms prior automated assessment tools in both grading accuracy and robustness across diverse programming styles. A classroom study with 30 students and 10 instructors further demonstrates strong user acceptance and practical applicability. Together, these findings highlight the effectiveness of task-level behavioral abstractions for scalable assessment of open-ended, event-driven programs. ACM Reference Format: Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang. 2026. Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation. In Proceedings of . ACM, New York, NY, USA, 22 pages. https://doi.org/10.1145/nnnnnnn.nnnnnnn
1
Introduction
Scratch [32] is a block-based programming language that has become a foundational platform for introductory programming education, with over 100 million learners worldwide [53]. Its eventdriven, visual programming model enables novices to create interactive animations and games, but also poses unique challenges for automated assessment. Manual grading provides rich feedback ∗ Corresponding Author
Authors’ Contact Information: Donglin Li, Anhui University of Science and Technology, Huainan, China, [email protected]; Daming Li, Independent Researcher, USA, [email protected]; Hanyuan Shi, Independent Researcher, China, [email protected]; Jialu Zhang, University of Waterloo, Waterloo, Canada, [email protected]. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. , © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM https://doi.org/10.1145/nnnnnnn.nnnnnnn , Vol. 1, No. 1, Article . Publication date: April 2026.
2
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
but is labor-intensive and, as Moreno-León et al. observe, becomes “almost impossible in large-scale settings” [34]. Educators must either set fewer assessment tasks or resign themselves to a greatly increased marking load [9], making human-centric evaluation increasingly unsustainable [2]. Beyond scalability, a more fundamental challenge lies in assessment validity. As the Scratch creator Mitchel Resnick argues, current automated methods often fall into the trap of measuring “what is easy to measure, rather than what is important” [45]. Static analysis fails to capture the dynamic, interactive nature of creative coding, obscuring the actual learning process. While the program code is visible, the visual and interactive execution behavior—the primary learning artifact in Scratch—often remains unexamined. Thus, there remains a critical dearth of valid tools [33] capable of bridging the gap between static code structure and dynamic visual execution. In recent years, several automated assessment approaches have been proposed for Scratch programs, among which the Whisker framework [58] represents one of the most advanced efforts to date. Whisker validates program behavior by simulating user interactions and checking state assertions. However, its reliance on precise, manually specified state assertions makes it difficult to accommodate the diverse implementation strategies and inherently visual evaluation requirements commonly observed in Scratch projects. As a result, existing tools have not fundamentally resolved the assessment bottleneck faced by large-scale, real-world programming education. We argue that the core limitation of prior approaches is the lack of a task-level abstraction for specifying and evaluating interactive behavior. In practice, students implement the same task using widely varying control flows, event structures, and interaction patterns, yet assessment criteria are defined at the level of the task, not individual programs. Program-specific assertions and hard-coded execution assumptions therefore fail to generalize across submissions. To address this gap, we propose Raven, a video-grounded framework that rethinks automated assessment for Scratch programs via video-based evaluation of program execution. Raven evaluates submissions at the task level using shared video-generation rules that specify how programs are exercised. Instead of relying on program-specific state assertions, it grades by analyzing execution videos. This design decouples grading criteria from individual implementations, enabling consistent assessment across diverse student solutions. The key insight of Raven is to integrate LLMs with video analysis to assess whether a program’s observed visual and interactive behavior satisfies grading criteria. LLMs enable semantic interpretation of high-level visual outcomes, such as geometric correctness or interaction responses, that are difficult to express as symbolic assertions or predefined outputs. For example, in an assignment requiring students to draw a five-pointed star, assertion-based tools struggle to verify geometric correctness or detect runtime errors such as partial drawings outside the stage. In contrast, these properties are readily observable in execution videos. By analyzing such videos, Raven can correctly assess visually valid stars regardless of drawing strategy or code path. Moreover, Raven incorporates a lightweight video understanding module that enables automated responses to interactive ask blocks, allowing it to evaluate programs with complex human–computer interaction behaviors that defeat prior tools. The main contributions of this paper are as follows: (1) We present Raven, the first automated assessment framework for Scratch that integrates LLMs with video-based execution analysis, enabling robust grading of visually grounded, interactive programs beyond traditional state-assertion-based approaches. (2) We introduce an assignment-level video generation abstraction that allows instructors to specify assessment criteria via shared interaction sequences, ensuring consistent and scalable evaluation across diverse student implementations. , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
3
Fig. 1. An example Scratch project in which a cat sprite turns around to say hello.
(3) We design a lightweight video understanding module that supports human–computer interaction patterns in Scratch, including programs that use ask blocks. (4) We evaluate Raven on 13 real-world assignments with over 140 student submissions, and a live classroom study with 30 students and 10 instructors, demonstrating substantial improvements in grading accuracy and instructor acceptance over prior tools. 2
Background
Scratch is a visual programming language designed for young and novice programmers [46]. It enables users to create interactive stories, games, and animations by dragging and snapping together colored programming blocks. As illustrated in Fig. 1, the Scratch programming environment consists of four main regions: (1) the block palette, located on the left, which provides available programming blocks; (2) the code area, located in the center, where blocks are composed into scripts; (3) the stage, located in the upper-right corner, which displays the program’s runtime behavior; and (4) the sprite list, located in the lower-right corner, which manages all sprites and backgrounds in a project. Unlike traditional textual programs, a Scratch project is composed of multiple scripts that execute concurrently. Each script is triggered by an event block (e.g., when green flag clicked) and controls sprite behavior through motion, control, and appearance blocks. As a result, Scratch programs naturally exhibit event-driven and concurrent execution behavior. As computational thinking education has expanded globally, Scratch has become the most widely adopted block-based programming platform, reaching over 100 million learners worldwide [53]. This widespread adoption has motivated research on automated techniques for analyzing and assessing Scratch programs. Automated Testing: State-of-the-art Automatic Assessment for Scratch Programs. Scratch projects differ fundamentally from conventional programs in that they lack a single entry point. Instead, execution is driven by multiple event handlers (e.g., when green flag clicked, when I receive message), which may be triggered concurrently. Automated assessment of such event-driven programs requires explicitly modeling user interactions and their effects on program execution. , Vol. 1, No. 1, Article . Publication date: April 2026.
4
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
test(" player catches apple ", () => { // 1. Initialization 3 setVariable (" score ", 0); 4 setPosition (" apple ", {x: 50, y: 100}) ; 5 setPosition (" basket ", {x: 50, y: -150}); 1 2
6 7 8 9
// 2. Interaction sequence broadcastMessage (" start_game "); wait (200) ;
10
// 3. Assertions assert ( getVariable (" score ") === 1); 13 assert ( isHidden (" apple ") === true); 14 }); 11 12
Listing 1. Example Whisker test case for a fruit-catching game.
Whisker [58] is the state-of-the-art automated testing framework for Scratch. It evaluates program behavior by simulating user interactions—such as key presses, mouse clicks, and message broadcasts—to trigger events, and by checking whether the resulting system state satisfies expected conditions under a specified interaction sequence. In Whisker, a test case is specified as a JSON-style script consisting of three core components: (1) an initial state configuration, which defines background settings, sprite positions, and variable initial values; (2) an interaction sequence, which describes how the user interacts with the program (e.g., key presses or mouse clicks); and (3) assertions, which verify whether the system state satisfies expected properties (e.g., assert(sprite.x == 100), assert(variable.score == 10)). Listing 1 shows a simplified Whisker test case for a fruit-catching game. Such test cases require precise specification of timing, spatial relationships, and program state assertions, imposing a substantial technical burden on instructors. In practice, many Scratch educators struggle to manually design test cases that are both correct and effective. To mitigate this burden, prior work [15, 16, 22] has explored the automatic generation of Whisker test cases. In this work, we adopt Whisker’s default MIO-based test generation algorithm [15], reflecting how automated testing tools are commonly used in practice. Given a generated test suite, Whisker evaluates a student program by initializing the execution state, replaying the interaction sequences, and checking assertions against observed runtime states, ultimately producing a pass/fail verdict. 3
Motivation and Technical Challenges
Despite substantial progress in automated assessment for Scratch, existing tools remain fundamentally limited in real-world educational settings. Scratch programs are open-ended, visually oriented, and highly heterogeneous, making assessment approaches based on rigid execution assumptions difficult to generalize across authentic student submissions. As the Scratch creator Mitchel Resnick emphasizes, “languages need ‘wide walls’ (supporting many different types of projects so people with many different interests and learning styles can all become engaged)” [46]. However, when deployed in classrooms, current automated assessment systems often impose restrictive assumptions that conflict with instructional practice, leading to brittle evaluations and unreliable grading outcomes. We identify four recurring technical challenges that expose this mismatch. Challenge 1: Inadequate Coverage of Open-ended Task Requirements. In real classrooms, a single Scratch assignment often admits multiple correct implementation strategies. Even within , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
Modified code
Standard code
5
Whisker Test Results
Fig. 2. Reducing the number of steps taken from 10 to 5 is a reasonable change in real-world evaluations. However, Whisker yields incorrect evaluation results.
Video-Detectable-Only Features
Image: Colorful Fireworks Key features for inspection: • Firework shape • Color transitions • Visual explosion effects
Image: Five-Pointed Stars Key features for inspection: • Star shape correctness • Spatial distribution
Image: Underwater World Key features for inspection: • Incorrect sprite types (no land or aerial sprites)
Fig. 3. Visual features – such as firework shapes, color transitions, and spatial distribution – can only be verified via video, whereas assertion-based tools like Whisker struggle to capture or evaluate them.
the same strategy, students may legitimately vary parameters such as movement speed, timing, or intermediate states while still satisfying instructional goals. For example, consider an assignment requiring a sprite to move left and right in response to arrow key presses (Fig. 2). A reference solution may update the sprite’s 𝑥-coordinate by 10 units per key press, whereas a student implementation may move the sprite by 5 units per press. Although the latter produces slower motion, it exhibits the same intended visual behavior and meets the task requirements from an instructional perspective. Nevertheless, assertion-based testing in Whisker flags such submissions as incorrect. This failure arises because Whisker relies on precise state assertions derived from a reference implementation. Small deviations in parameter choices alter subsequent execution states, leading to cascading assertion violations and incorrect grading outcomes, despite functionally acceptable observable behavior. Challenge 2: Inability to Assess Visually Defined Correctness. Many Scratch assignments define correctness primarily in terms of visual outcomes rather than program state. As illustrated , Vol. 1, No. 1, Article . Publication date: April 2026.
6
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
How to activate the “correct answer” branch
Fig. 4. In a “Math Pea Shooter” game, Whisker fails to provide the correct inputs required to trigger the specific logic for shooting zombies with shells.
in Fig. 3, examples include requiring that an underwater scene contains no land or aerial sprites, that fireworks visually explode in the sky, or that a program draws multiple five-pointed stars with specific visual properties. Such requirements are difficult to formalize using state assertions alone. In practice, evaluating these assignments requires executing the Scratch program and judging its rendered animation. Assertion-based tools such as Whisker, which reason over predefined variables and sprite states, are therefore ill-suited to assess visually grounded correctness criteria. Challenge 3: Limited Support for Human–computer Interaction Via ask Blocks. Human– computer interaction is a core component of Scratch programming. The ask block enables sprites or backgrounds to prompt users for input and conditionally alter behavior based on the response, and is often the only mechanism for implementing interactive logic in Scratch projects. However, Whisker provides limited strategies for handling ask blocks, typically relying on random inputs or simple heuristics when the input is guarded by specific conditional statements. For example, as illustrated in Fig. 4, in a “Math Pea Shooter” game, the player must correctly answer an arithmetic question to trigger a projectile attack. Despite repeated attempts using automatically generated Whisker test cases, the tool consistently fails to activate the “correct answer” branch, rendering automated assessment ineffective. Challenge 4: Fragility Under Flexible Sprite and Background Configurations. In real classroom settings, many Scratch assignments allow flexibility in the number, names, and visual appearances of sprites and backgrounds. In our dataset, a substantial fraction of assignments permit such variations. However, Whisker assumes fixed sprite identities and configurations derived from a reference program. As a result, when students choose different sprite names, costumes, or quantities, assertion-based automated tests fail to locate the expected entities. As illustrated in Fig. 5, in a “Magic Show” game (right), students use various costumes to represent magic props, leading to mismatches in sprite names and appearances. Similarly, in a “Pea Shooter (Pests)” game (left), students use varying numbers and types of pest sprites, causing test failures due to unrecognized sprite identities, even when the programs meet the task requirements. Together, these challenges highlight a fundamental mismatch between assertion-based automated testing and the realities of open-ended, visually driven, and interaction-rich Scratch programming. , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
7
Different Sprites Different Costumes Different Quantities
Fig. 5. In real classroom settings, many Scratch assignments allow flexibility in the choices, costumes, and quantities of sprites.
Assignment
LDE
VDOE
Examples of Video-Detectable Errors
Underwater World Colorful Fireworks Magic Show Draw Five-Pointed Star Ball Goes Up Big Fish Eats Small Fish Pea Shooter Thinking Graphics Catch the Eggs Hit the Bats Mental Math Race Math Pea Shooter Rock–Paper–Scissors
50% 80% 0% 71% 50% 50% 0% 100% 78% 100% 44% 0% 20%
50% 20% 100% 29% 50% 50% 100% 0% 22% 0% 56% 100% 80%
Incorrect sprite types (land/air animals) Firework shape, color transitions, explosion effect Object switching timing, hat open/close state Star shape correctness, spatial distribution Premature motion, incorrect movement direction Facing direction, smoothness, respawn animation Enemy animation, hit feedback, score visualization Dice face shape and number consistency Falling animation, catch feedback, scoring effect Random bat movement, hit animation feedback Answer feedback, finish-time visualization Projectile direction, hit animation, defeat effects Hover/click animations, sprite-choice alignment
Table 1. Comparison of Error Detectability: Logic-Detectable vs. Video-Detectable-Only. We categorize errors into two classes. (1) Logic-detectable errors (LDE), which can be identified through static or logic-based analysis; and (2) Video-detectable-only errors (VDOE), which cannot be detected by logic analysis alone and require execution videos for reliable identification. If an error is detectable by both logic-based and video-based approaches, we classify it as logic-detectable, reserving the “video-only” category strictly for cases where video evidence is essential. We intentionally reserve the term video-detectable-only for errors whose detection fundamentally depends on visual execution evidence.
Addressing these limitations requires rethinking the abstraction underlying automated assessment, rather than merely improving test generation techniques. Why Video Input Is Crucial for Scratch Program Assessment? Traditional Scratch assessment tools determine correctness by reasoning over program logic and state assertions. This paradigm assumes that correct behaviors can be specified by a fixed interaction sequence and precise system states—an assumption that rarely holds in real Scratch classrooms. In assertion-based frameworks , Vol. 1, No. 1, Article . Publication date: April 2026.
8
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
Raven Student Submission and Evaluation Stage
Task Configuration Stage
Logic evaluation Task Requirements
Logic_Checker
Logic Grading Results
Preprocessed Outputs
Check, Repair & Merge
Video_Checker
Video Grading Results
Logic Grading Rules Scratch Project
Unconfigured Assignment
Video Grading Rules Video Generation Rules
Preprocessing
Video evaluation
Video Generation Module
Grading Results
Videos
Fig. 6. Architecture of the Raven Framework. The system evaluates projects via a dual-track pipeline (Logic and Video) organized into two stages: (1) Task Configuration Stage: The instructor initializes a Unconfigured Assignment (orange-yellow box) to define specific Task Requirements, Logic Grading Rules and Video Grading Rules (light yellow boxes). (2) Student Submission and Evaluation Stage: A student’s Scratch Project (orange box) is processed by Functional Operators (white boxes, e.g., Logic_Checker, Video Generation Module, and Video_Checker). These operators generate Intermediate Results (light blue boxes), which are then synthesized into the Final Grading Results (green box).
such as Whisker, test cases must strictly follow predefined interaction sequences, and all asserted states must exactly match reference values. Even minor deviations in interaction order, timing, movement granularity, or sprite configuration can alter execution paths and trigger cascading assertion failures, causing functionally correct programs to be incorrectly labeled as wrong. Raven adopts a fundamentally different approach by evaluating correctness from execution videos. Given instructor-defined video generation rules, the system infers whether the observed visual behavior satisfies grading criteria. As a result, variation in interaction sequences, implementation strategies, sprite choices, and intermediate states does not affect grading outcomes, as long as the rendered animation meets task requirements. This shift is crucial because many Scratch assignments define correctness in terms of visual outcomes rather than internal logical states (e.g., whether fireworks visually explode or a star is drawn with the correct shape)—properties that are difficult or impossible to capture using static analysis or state assertions alone. An analysis of 13 representative Scratch assignments confirms that a substantial fraction of incorrect submissions exhibit errors detectable only through visual inspection of execution videos (Table 1). These findings show that effective automated assessment for Scratch must treat runtime visual behavior as a first-class artifact, as embodied by Raven. 4
System Overview
Raven operates in two stages: a task configuration stage and a student submission and evaluation stage. This separation enables instructors to flexibly customize assignments and grading criteria, while allowing students’ submissions to be evaluated automatically and consistently. Below, we outline the system architecture (Fig. 6) and the algorithmic workflow (Algorithm 1) and define the key terminology. , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
9
Algorithm 1 Raven: Hybrid Logic-Video Evaluation Require: Student project file S, Task description Q, Grading rules T = {T𝑙𝑜𝑔 , T𝑣𝑖𝑑 }, Video generation rules G Ensure: Final comprehensive grading report R 𝑓 𝑖𝑛𝑎𝑙 // Phase 1: Preprocessing 1: D ← Preprocess(S) ⊲ Extract project JSON and initial states // Phase 2: Logic Evaluation 2: R𝑙𝑜𝑔 ← Logic_Checker(Q, T𝑙𝑜𝑔 , D) ⊲ Static analysis based on task description // Phase 3: Video Evaluation 3: V ← GenerateVideos(S, G) ⊲ Record videos based on distinct video generation rules 4: for each 𝑣 ∈ V do ⊲ Iterate through each generated test-case video 5: for 𝑟𝑢𝑛 ← 1 to 3 do ⊲ Mitigating VLM stochasticity for the current video 6: F ← ExtractFrames(𝑣) ⊲ Sampling and encoding key frames from 𝑣 7: R𝑡𝑚𝑝 [𝑟𝑢𝑛] ← Video_Checker(F , T𝑣𝑖𝑑 , D) ⊲ Visual inspection via VLM 8:
R 𝑣𝑖𝑑 ← MergeRuns(R 𝑣𝑖𝑑 , {R𝑡𝑚𝑝 })
⊲ Update R 𝑣𝑖𝑑 using a lower-bound aggregation
// Phase 4: Result Checking and Merging 9: R 𝑓 𝑖𝑛𝑎𝑙 ← Check_Repair_Merge(R𝑙𝑜𝑔 , R 𝑣𝑖𝑑 ) 10: return R 𝑓 𝑖𝑛𝑎𝑙
4.1
⊲ Reconcile logic and video results
Task Configuration Stage
In the task configuration stage, instructors define the assignment by providing four components: (1) a task description, (2) logic grading rules (3) video grading rules, and (4) video generation rules. Hereafter, (2) and (3) are collectively referred to as grading rules. This stage offers substantial flexibility, allowing instructors to tailor assignments to different instructional goals, student skill levels, and assessment scenarios. Task Description. The task description specifies the creative requirements presented to students and describes the intended behavior of the Scratch project. Grading Rules. The grading rules define how student submissions should be evaluated in a fair and uniform manner. Each rule consists of a natural-language description and an associated score. A submission receives the score only if it fully satisfies the rule; otherwise, it receives zero. Grading rules are divided into two types. Logic grading rules specify code-level requirements (e.g., “If the project implements keyboard-controlled movement for the dog sprite, award 10 points”). Raven evaluates whether the functionality is implemented and whether the underlying logic is correct. Video grading rules specify expected runtime behaviors observable in the rendered animation (e.g., “If the ball disappears and the score increases when it collides with the dog, award 10 points”). These rules focus exclusively on visual outcomes rather than internal program states. Video Generation Rules. The video generation rules describe how Raven should interact with the program to produce execution videos for evaluation. Instructors specify one or more event sequences, which the system follows to automatically operate the Scratch program and record videos. Supported events include keyboard and mouse actions, mouse movements, and responses to ask blocks. 4.2
Student Submission and Evaluation Stage
After task configuration, students submit their Scratch projects to Raven. As detailed in Algorithm 1, Raven executes a Hybrid Logic-Video Evaluation pipeline. This process begins with static logic , Vol. 1, No. 1, Article . Publication date: April 2026.
10
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
You are a professional programming education assessment expert, specializing in logic-level evaluation of Scratch projects. You are able to precisely compare differences between a student’s implementation and the expected objectives, and provide fair and impartial scores and feedback. You must first obtain the necessary information through tools before answering. Do not guess or make assumptions. You are required to strictly follow the workflow below to generate a grading report: 1. Read the task description and logic grading rules below, and learn the task requirements and the logic grading criteria. <Task Description> + <Logic Grading Rules> 2. Logic evaluation: Strictly follow the grading criteria and perform the evaluation as follows: - Read and understand the Preprocessed Outputs below and the analyzed states of all sprites and backgrounds. <Preprocessed Outputs> - For each grading item, perform logic-level verification on the student Scratch project. - If the verification meets the expected requirement, the grading item receives full marks; otherwise, it receives 0 points. - List the score for each logic grading item in the following format: - Logic Grading Rule 1: [Grading Rule Description] + [Scoring Evidence (must describe in detail: which code blocks or states were verified and whether all conditions were met)] + Scoring Result (xx points) - Logic Grading Rule 2: [Grading Rule Description] + [Scoring Evidence] + Scoring Result (xx points) - Logic Grading Rule 3: [Grading Rule Description] + [Scoring Evidence] + Scoring Result (xx points) - ... - Total Score: xx points Carefully list the score for each grading item and compute the total score. Please note that your response must not include tool usage details. Only include the final analysis results. Your response must contain only numbers, punctuation marks and text. Do not include any other content, including emojis, images, tables, code or formatting.
Fig. 7. Prompt for the logic_checker
inspection to verify code structures against task requirements, followed by the video generation based on instructor-specified rules. These videos undergo robust visual analysis using a multi-run strategy to mitigate VLM stochasticity before being synthesized with logic checks to produce the final grading results. A Scratch project is typically provided in .sb3 format, which is unpacked to obtain the JSON-based source code and associated media assets. Preprocessing. Since large language models do not automatically extract salient semantic features from Scratch code, Raven first performs a preprocessing step. The system parses the project’s JSON source code to extract background and sprite descriptions, initial states, and structural information. These extracted features are packaged together with the original JSON code and serve as shared input to both logic-based and video-based evaluation modules. Logic Evaluation. For logic evaluation, Raven feeds the preprocessed outputs, logic grading rules, and task description into a logic_checker. The logic checker uses the state-of-the-art large , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
11
language model Qwen3-max [64] to compare the student implementation against the intended functionality and produces a logic score. Fig. 7 displays the prompt for the logic_checker. Video Evaluation. Video evaluation addresses the long-standing difficulty of automated testing in Scratch, where high randomness, complex event-triggered behaviors, and diverse implementation strategies make it challenging to define reliable assertions. Instead of reasoning over abstract system states, Raven evaluates correctness directly from execution videos.
You are a professional programming education assessment expert, specializing in event-based testing and grading of Scratch projects. You are able to precisely compare differences between a student’s implementation and the expected objectives, and provide fair and impartial scores and feedback. Video Frame List: <Video Frame List> Brief project analysis: <Preprocessed Outputs> You should only focus on the grading rules provided below: <Video Grading Rules> Do not guess or make assumptions. You are required to strictly follow the event testing criteria to generate the evaluation results: - For each grading rule, if the behaviors and events in the video fully comply with all conditions of the criteria, the item receives full marks; otherwise, it receives 0 points. - Summarize the evaluation results in the following format: - Grading Rule 1: [Grading Rule Description] + [Scoring Evidence (must describe in detail: which frames were checked, what specific phenomena were observed, and whether all conditions were met)] + Scoring Result (xx points / Full points) - Grading Rule 2: [Grading Rule Description] + [Scoring Evidence] + Scoring Result (xx points / Full points) - ... - Total Score: xx points / Full points Your response must contain only numbers, punctuation marks, and text. Do not include any other content, including emojis, images, tables, code, or special formatting markers.
Fig. 8. Prompt for the video_checker
The video generation module executes the student project according to the instructor-specified video generation rules and records one or more videos. These videos serve as the sole evidence for video-based grading. To handle interactive ask blocks, where traditional tools rely on random or heuristic responses, Raven incorporates a lightweight vision model Qwen-vl-plus [5] that interprets on-screen content and generates human-like responses, enabling effective human– computer interaction during automated execution. For each video test case, instructors only need to provide simple arrays specifying event sequences, such as keyboard and mouse actions, mouse movements, and responses to ask blocks. This design allows even instructors without technical testing expertise to configure executable test scenarios. The video_checker evaluates whether the recorded videos satisfy the video grading rules using the multimodal video analysis model Qwen3-vl-plus [4]. For each video, Raven performs multiple repeated evaluations to reduce the risk of hallucinations. A grading rule is awarded full credit only if it is satisfied across all repeated evaluations; otherwise, no credit is given. Fig. 8 displays the prompt for the video_checker. The video checker is provided with three inputs: (1) preprocessed outputs to support semantic understanding of sprites and interactions, (2) instructor-defined video grading rules to guide , Vol. 1, No. 1, Article . Publication date: April 2026.
12
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
Assignment
Difficulty
B
S
ICC
#Submissions
Core Blocks
Underwater World Colorful Fireworks Magic Show Draw Five-Pointed Star Ball Goes Up Big Fish Eats Small Fish Pea Shooter (Pests) Thinking Graphics Catch the Eggs Hit the Bats Mental Math Race Math Pea Shooter Rock–Paper–Scissors
Easy Easy Easy Easy Easy Medium Medium Medium Medium Hard Hard Hard Hard
21 40 24 22 41 34 197 43 65 139 85 50 82
3 1 3 1 2 2 11 4 2 5 5 3 7
7 6 11 4 3 12 44 10 18 36 30 13 21
10 10 11 10 12 18 12 6 11 9 13 11 13
looks, loops, motion special effects, sound, loops next costume, events, sound pen, loops, randomness looks, motion, events looks, events, motion, collisions looks, wait, randomness, collisions interaction, control, operators events, clone, collisions, timing events, randomness, timing interaction, timing, variables interaction, collisions, broadcast variables, randomness, broadcast
Table 2. Characteristics of the Scratch assignments used in evaluation. B = number of blocks; S = number of scripts; ICC = inter-procedural cyclomatic complexity. Core Blocks showcase the core components in each Scratch assignment.
attention, and (3) execution videos, which are sampled at 10 FPS to balance recognition accuracy and token efficiency. 4.3
Result Checking and Merging
Finally, Raven performs a consistency check between logic-based and video-based evaluation results using Qwen-plus [64], a large language model balancing capability and cost. If conflicts arise, video-based results take precedence, as they directly reflect observable runtime behavior. The system then merges the two evaluation results and computes the final grade for the student submission. 5
Evaluation
We empirically evaluate Raven with respect to the following research questions: • RQ1: Can Raven accurately assess the correctness of Scratch programs? • RQ2: Is Raven scalable and deployable in real classroom environments? 5.1
RQ1: Correctness of Automated Assessment
Dataset and Experimental Setup. We collect data from 13 Scratch assignments used in real-world classroom instruction at an after-school education center, covering multiple difficulty levels, grade ranges, and instructional topics. In total, the dataset contains 146 student submissions. Table 2 summarizes the characteristics of each assignment, including difficulty level, number of blocks (B), number of scripts (S), inter-procedural cyclomatic complexity (ICC), number of student submissions, and the core Scratch blocks involved. To evaluate the grading quality of Raven, we recruited five Scratch instructors, each with over two years of Scratch teaching experience, to independently grade the same set of student submissions using identical task descriptions and grading rules. The full score for each assignment is 100, with raw scores given based on grading rules described in section 4.1. , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
13
We define binary correctness in such a way that a submission is labeled as correct only if the majority of the five instructors assigns a full score of 100, and incorrect otherwise. This serves as the ground truth for evaluating the grading quality of Whisker and Raven. Next, we use the reference solution of each assignment as the seed program for Whisker and apply its default test generation mechanism to produce test cases. These test cases are then executed on the remaining student submissions. In parallel, Raven evaluates the same submissions using instructor-provided assignment descriptions, grading rules, and video generation rules. We report accuracy, precision, recall, and F1 score for both tools. Accuracy (𝐴𝑐𝑐) is defined as the fraction of submissions whose grades match the ground truth labels. Precision (𝑃𝑟𝑒𝑐) is defined as the fraction of submissions graded as incorrect by the tool that are truly incorrect. Recall (𝑅𝑒𝑐) is defined as the fraction of truly incorrect submissions that are correctly identified by the tool. The F1 score is the harmonic mean of precision and recall. Binary Correctness Grading Results: Raven vs. Whisker. Table 3 summarizes the results across all assignments from the binary correctness perspective. Across all the 13 assignments, Raven achieves near-perfect performance, significantly outperforming Whisker over all the metrics. This demonstrates Raven’s outstanding assessment capabilities handling complex real-world teaching projects that involve visual effects and human-computer interactions. The only imperfect performance on the assignment “Magic Show” is due to a false positive on a single submission where the student’s depicted hat sprite displays visual effects inconsistent with the physical world while the instructors think it’s only a minor detail for that assignment. We find that Whisker overall exhibits very low precision, which implies that it is too harsh when grading the submissions such that many correct submissions are misclassified as incorrect. In contrast, Raven significantly reduced false positives. This improvement is particularly important in educational settings, where false positives undermine grading fairness and instructional trust [24]. Additionally, Whisker exhibits very limited support for ask blocks questions that involve user interaction. Of the 30 submissions requiring user interaction, Whisker correctly triggers the intended conditional logic in only two cases. Raven
Assignment
Whisker
Acc(%)
Prec
Rec
F1
Acc(%)
Prec
Rec
F1
Underwater World Colorful Fireworks Magic Show Draw Five-Pointed Star Ball Goes Up Big Fish Eats Small Fish Pea Shooter (Pests) Thinking Graphics Catch the Eggs Hit the Bats Mental Math Race Math Pea Shooter Rock–Paper–Scissors
100 100 90.9 100 100 100 100 100 100 100 100 100 100
1.00 1.00 0.86 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
1.00 1.00 0.92 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00
60 60 63.6 50 25 16.7 16.7 100 100 44.4 76.9 81.8 84.6
0.00 0.56 0.60 0.67 0.18 0.12 0.09 1.00 1.00 0.38 0.75 0.86 0.83
0.00 1.00 1.00 0.57 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.86 1.00
0.00 0.71 0.75 0.62 0.31 0.21 0.17 1.00 1.00 0.55 0.86 0.86 0.91
Average
99.3
0.99
1.00
0.99
56.8
0.53
0.89
0.66
Table 3. Binary correctness comparison between Raven and Whisker on a curated dataset consisting of submissions collected from real Scratch classes.
, Vol. 1, No. 1, Article . Publication date: April 2026.
14
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
Agreement between Raven and Human Instructor Scores (146 submissions)
40 20 0
0
20
40
60
Raven Score
80
100
Instructor 2
100
Pearson r=0.974, p < 0.001 Kendall =0.936, p < 0.001
60 40 20 0
0
20
40
60
Raven Score
80
80
Instructor 3
100
Pearson r=0.952, p < 0.001 Kendall =0.906, p < 0.001
60 40 20
100
0
0
20
40
60
Raven Score
80
100
80
Instructor 4
100
Pearson r=0.933, p < 0.001 Kendall =0.893, p < 0.001
Instructor 5 Score
60
80
Instructor 4 Score
100
Pearson r=0.953, p < 0.001 Kendall =0.906, p < 0.001
Instructor 3 Score
80
Instructor 2 Score
Instructor 1 Score
100
Instructor 1
60 40 20 0
0
20
40
60
Raven Score
80
100
80
Instructor 5
Pearson r=0.946, p < 0.001 Kendall =0.895, p < 0.001
60 40 20 0
0
20
40
60
Raven Score
80
100
Fig. 9. Agreement between Raven and human instructor scores as shown in scatterplots with 146 submissions. A jittering effect with a magnitude of 1.12% of the data range is applied on overlapping data points for visualization. Each scatterplot (Raven against each instructor) exhibits strong correlations with fitted line close to y=x.
Raw Scores Grading Results: Raven vs. Human Instructors. Compared to Whisker, Raven additionally enables the provision of detailed scores based on custom grading rules. To assess the agreement in the raw scores provided by Raven and human instructors, we compute the Pearson correlation (linear relationship) and Kendall’s tau (non-parametric) between Raven and each one of the five human instructors over the same submission dataset, as shown in the scatterplots in Fig. 9. We see that the scores provided by Raven strongly positively correlate with the scores given by the instructors, and are well calibrated. While there can be minor grading inconsistencies between different instructors, these comparisons show that Raven overall aligns very well with the grading criteria of experienced Scratch instructors. Summary for RQ1. In an evaluation dataset constructed from real Scratch classes, Raven achieves very high grading accuracy, significantly outperforming existing the state-of-the-art Scratch assessment tool. The raw graded scores it provides closely align with experienced human instructors. Unlike prior tools that only support binary grading results, Raven provides detailed grading evidence and score breakdowns, addressing a key limitation of existing Scratch assessment tools. 5.2
RQ2: Classroom Deployability
Classroom Study Design. To assess Raven’s real-world deployability, we conducted a classroom study led by 10 experienced Scratch instructors at the same after-school education center. The study spanned five class sessions, with in total 30 student participants aged 8–12 with varying Scratch experience (0–3 years). The instructors selected and configured two assignments aligned with their teaching goals (“Ball Goes Up” and “Big Fish Eats Small Fish”). The instructors illustrate the end-to-end process, including assignment description, goal of the assignments, use of tool, grading criteria, as well as clarifications of questions on the survey. Students submitted their projects via USB drives, and all the grading was performed live during class. After receiving the grading results, students completed the survey. The instructors completed their survey after all the sessions were finished. When designing the questionnaires, we adopt the Technology Acceptance Model (TAM) [14] that evaluates perceived usefulness, perceived ease of use, and willingness of adoption towards Raven. TAM is a widely accepted framework for modeling user adoption of information technologies and has been widely applied in the evaluation of educational systems [1, 10, 17, 19, 51]. TAM characterizes tool acceptance through two major axes. Perceived usefulness (PU) measures the extent to which users believe that a system improves their task performance (e.g., “Raven helps me , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
Instructor Perspectives on Instructional Efficacy and Acceptance (N=10) 1.Intention to Use
2.Perceived Usefulness
10%
90%
4.90
10%
90%
4.90
80%
4.80
Q4: Grading time saved
10%
90%
4.90
Q5: Logic/Creativity insights
30% 10%
Q6: Classroom efficiency
60%
4.40
100% 5.00
Q8: Ease of configuration
10% 10%
Q9: Interface clarity
10%
Q3: Peer Recommendation 2.Perceived Usefulness
Q12: Scoring consistency
80%
20% 1
Strongly disagree
4.60 4.80
80%
2
3
4
Somewhat disagree
90%
4.83
Q6: Simple Submission
90%
4.83
13.3%
4.67
90.1%
4.77 4.83
96.7% 4.97
Q11: Perceived Fairness
96.7% 1
Somewhat agree
80.1%
86.7%
10%
Q10: Scoring Accuracy
5
Neutral
4.80
Q5: Criteria Clarity
Q7: Result Interpretation 3. Perceived Ease of Use Q8: Smooth Process
4. Output Quality
4.77
90.1%
96.7% 4.97
4.80
70%
4.30
83.3%
Q4: Diagnostic Feedback
4.00
70%
20% 20%
76.6%
10%
Q9: Quick Onboarding
Q10: Low learning curve Q11: Non-technical usability 10%
4. Output Quality
4.50
80%
20% 10%
Q1: Iterative Improvement 1.Intention Q2: Continued Usage to Use
4.50
80%
20%
Q7: Educational value
3. Perceived Ease of Use
Student Perspectives on Learning Experience and Satisfaction (N=30)
Q1: Future use intention Q2: Willingness to recommend Q3: Evaluation efficiency
20%
15
2
3
4
4.87 5
Strongly agree
Fig. 10. Survey results collected after a live classroom study conducted at an after-school education center. 𝑁 denotes the number of participants. The diverging stacked bar charts illustrate instructors’ (left) and students’ (right) feedback collected by questionnaires designed according to the Technology Acceptance Model (TAM). The questions cover Intention to Use, Perceived Usefulness, Perceived Ease of Use, and Output quality. Users respond on a Likert scale of 1-5, with the average score for each question shown on the right. The full list of questions and responses can be found in the dataset that we release.
evaluate student Scratch projects more efficiently,” instructor questionnaire Q3). Perceived ease of use (PEOU) captures the extent to which users perceive the system as requiring minimal effort (e.g., “The entire process of submission and viewing results is smooth, with no technical obstacles,” student questionnaire Q8). We also included survey questions measuring the Intention to Use and Output Quality, similar to the ones in [17]. Additionally, we added open feedback questions at the end of the questionnaires in case there are critical missing points. The full list of questions and responses from the surveys can be found in the dataset we release. Instructor Feedback. Raven receives high average survey scores across all the dimensions, as shown on the left of Fig. 10. Notably, they all agree that Raven brings educational value and are willing to use the tool in the future. All the instructors agree that the system improves grading efficiency, and five instructors highlight in open feedback questions that Raven is particularly helpful in large classrooms over 20 students, justifying the potential of scalability in real-world teaching environments. All the instructors also affirm the consistency and fairness of the grading results. In open feedback questions eight instructors express strong intent to adopt the system once available. Three instructors recommend further simplifying the task configuration and providing a web-based deployment, which we consider as an important direction for future work to improve user experience. Student Feedback. The student survey results after using Raven are shown on the right of Fig. 10, which again show overall high ratings across all the dimensions. Raven receives a high average score of 4.90/5 in the Perceived Usefulness dimension. Specifically, all the students agree that the grading results are reliable and that the detailed feedback help them understand the strengths and weaknesses of their submissions. It also receives a high average score of 4.78/5 in the Perceived Ease of Use dimension. 96.7% of the students suggest that Raven is a fairer system compared to manual grading, and 93.3% of them indicate willingness to recommend it to their peers. These responses suggest strong acceptance and practical viability in real classrooms. We analyze the few 1-point ratings to identify potential weaknesses. Based on the suggestions in the open feedback questions on the student survey, the dissatisfaction is primarily due to the fact that currently submissions , Vol. 1, No. 1, Article . Publication date: April 2026.
16
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
can only be graded by Raven on teacher’s computer which results in longer waiting time, and that the feedback generated is usually text-heavy. To address these user experience concerns from the product perspective, future versions will transition to a more intuitive, visual-centric reporting system and a web-based platform for convenient grading. Summary for RQ2. Classroom survey results collected from both instructors and students indicate that Raven overall receives strong positive feedback in live Scratch instruction environments, showing deployment potential in real-world Scratch programming education. 6 6.1
Discussion Sources of Discrepancy Between Raven and Human Graders
Although Raven achieves high correlation with instructor grading, a small gap remains relative to inter-instructor agreement. We attribute this gap to two main factors. LLM Variability. Both logic- and video-based evaluation in Raven rely on large language models and thus exhibit inherent variability. Logic-based evaluation may occasionally miss subtle errors due to overreliance on surface code patterns [27], while video-based evaluation can hallucinate visual events not present in execution [63]. Such effects lead to sporadic misjudgments. Grading Philosophy in Creative Tasks. Discrepancies also arise from differing grading criteria for creative assignments. For instance, in the “Magic Show” task, Raven enforces precise visual-state rules, whereas instructors adopt a more outcome-oriented view, prioritizing correct interaction behavior over exact visual fidelity. This philosophical difference naturally limits alignment with automated grading. 6.2
Comparison with ViScratch
ViScratch [55] is a feedback generation system that catches Scratch bugs by aligning runtime videos with logic bugs in code. Although both Raven and ViScratch analyze Scratch programs using video-based signals, they address different problems and serve distinct goals. Scope and Objectives. ViScratch assumes programs are buggy and focuses on diagnosing faults in incorrect submissions. Raven makes no such assumption: submissions may be correct or incorrect, and the system evaluates correctness for grading. As a result, ViScratch supports debugging scenarios, whereas Raven is designed for formative and summative assessment. Execution and Outputs. Raven actively executes programs using instructor-defined interaction sequences and evaluates runtime behavior through generated videos, enabling fine-grained, rubricbased scoring with partial credit. In contrast, ViScratch does not rely on systematic executiondriven video generation and primarily produces binary or categorical bug diagnoses. 6.3
Design Implications and Future Work
The core value of Raven lies in shifting the focus of automated assessment from how to design systemlevel assertions to how to evaluate observed behavioral outcomes. This shift enables automated grading to better serve educational objectives rather than rigid technical specifications. By simplifying and customizing the instructor-facing task configuration process through video generation rules, and by providing detailed grading evidence, Raven substantially lowers the technical barrier for educators. In addition, the lightweight video understanding module effectively supports human– computer interaction scenarios, addressing the long-standing challenge posed by ask blocks in Scratch programs. , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
17
Future work will focus on three directions. First, we plan to develop more efficient video generation and analysis techniques to further reduce assessment latency. Second, we aim to extend the system to better support the evaluation of creative and open-ended projects that go beyond fixed grading rubrics. Third, we plan to build a cloud-based assessment platform that provides a large collection of reusable tasks for instructors and practice opportunities for students, supporting online task customization and submission, and fostering a scalable educational assessment ecosystem. We believe that Raven not only provides a practical solution for Scratch education, but also opens new avenues for automated assessment in other graphical programming environments, ultimately contributing to the quality and accessibility of computational thinking education. 7
Threats to Validity
We discuss potential threats to the validity of our evaluation and how we mitigate them. Dataset Size and Task Complexity. A potential threat to external validity is whether our dataset adequately represents the diversity and complexity of Scratch programs encountered in real classrooms. To address this concern, our evaluation dataset consists of 13 assignments collected from real-world instructional settings, spanning multiple grade levels, difficulty levels, and programming constructs. The assignments provide comprehensive coverage of fundamental Scratch categories, including Motion, Looks (e.g., costumes and special effects), Sound, Events (e.g.,broadcasting), Control (e.g., loops and cloning), Sensing (e.g., timer, user interaction, and collisions), Operators (e.g., variables, randomness, and logic), and specialized extensions such as the Pen module. This diversity reduces the risk that our results are biased toward a narrow class of programs. Nevertheless, we acknowledge that the dataset size is limited, and future work should evaluate Raven on larger-scale datasets. Model Stability and Nondeterminism. Another threat to internal validity stems from the nondeterministic behavior of large language models used in both logic-based and video-based evaluation. We mitigate this threat in two ways. First, the video checker runs each evaluation three times and considers a grading rule as satisfied only if all the runs produce the same result, reducing the impact of sporadic hallucinations. Second, Raven employs a lightweight check, repair, and merge procedure to reconcile logic- and video-based results: when conflicts occur, video judgments take precedence, reflecting the primacy of observable runtime behavior in Scratch assessment and improving robustness to isolated model errors. Human Grading Variability. There are inevitably variabilities of grading outcomes by different human instructors even on the same student submission, as each grader can have their own interpretation over the same grading rules. Our evaluation assumes grading rules accurately reflect instructional intent; poorly specified rules could limit grading validity. In our studies, we recruited multiple experienced instructors to construct the evaluation dataset. The ground truth labels in the evaluation dataset are based on a majority vote of these instructors, which ensures the quality of reference. We showed in RQ1 that Raven provides grading results highly consistent with multiple experienced Scratch teachers. Also, the introduction of Raven may help reduce human grading variabilities in teaching, as illustrated by the high scores Raven received on scoring accuracy and perceived fairness in the student survey in RQ2. 8
Related Work
Automated Assessment for Scratch Programs. Automated assessment for Scratch programs has attracted increasing attention as block-based programming has become widely adopted in Computer Science education. The assessment of Scratch programs has evolved from static code analysis to assertion-based testing. Early approaches, such as Hairball [8] and Dr. Scratch [34], pioneered , Vol. 1, No. 1, Article . Publication date: April 2026.
18
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
static analysis by parsing the Abstract Syntax Tree (AST) to detect code smells, bad practices, or missing blocks [23, 26]. While effective for assessing coding style and computational thinking concepts [41, 49, 61], static methods fail to capture runtime logic and functional correctness, often leading to false positives in semantically complex projects [47]. To address this, assertion-based testing [40] techniques were introduced to evaluate functional behavior. For example, Whisker [58] models Scratch programs as event-driven systems and evaluates correctness by simulating user interactions (e.g., key presses and mouse events) and checking state assertions. Subsequent work has focused on improving test generation for Whisker, including search-based testing [16], model-based testing [22], and automated oracle generation [15]. While these techniques reduce the burden of manual test authoring, they fundamentally inherit the limitations of assertion-based evaluation, including sensitivity to execution order, reliance on precise state specifications, and difficulty handling visual outcomes and diverse implementations. Other program analysis tools target specific tasks such as bug detection, program repair [17, 52, 57, 59, 67, 70], or feedback generation [18, 21, 44, 48, 50, 68, 69], often assuming the presence of a reference implementation or known buggy submissions. These systems typically focus on debugging rather than grading and do not address the challenge of assessing open-ended, visually defined correctness in real classroom settings. Visual and Behavior-Based Program Assessment. Since Scratch projects are inherently visual and interactive, their assessment shares challenges with Graphical User Interface (GUI) testing and automated game validation. Traditional GUI testing tools, such as Sikuli [66] and GUITAR [38], utilize screenshot matching and widget-tree analysis to verify interface correctness [7, 12, 29]. In the domain of game development, automated agents have been employed to explore game states [3, 71] or detect glitches via visual artifacts [30]. However, most existing techniques still depend on predefined assertions or task-specific heuristics and are not designed for open-ended educational programs with high implementation diversity. Recent work has begun to explore video as an analysis artifact for understanding program execution [54], particularly in robotics and embodied agents [35, 60]. These efforts demonstrate that video captures aspects of system behavior that are difficult to express through program states alone. Nevertheless, existing visual analysis tools are primarily designed for crash detection or debugging [6, 37, 42, 55] rather than pedagogical grading. they do not address instructor-facing configuration or real-world educational deployment. Large Language Models in Educational Assessment. Large language models (LLMs) have recently been applied to educational tasks such as automated feedback generation [13, 43], shortanswer grading [11, 20], and programming assistance [25]. In the context of programming education, LLMs have been used to explain code [28, 36], identify bugs [31], or guide students through debugging steps [56]. These approaches primarily operate on source code or textual artifacts and assume well-defined specifications or reference solutions. More recently, multimodal LLMs like GPT-4V [39] and Qwen-VL [4] have shown promise in reasoning over visual inputs, including diagrams and videos [62, 65]. However, their application to automated assessment of graphical, event-driven programs remains largely unexplored. In particular, existing work does not leverage LLMs to reason over execution videos as a first-class grading artifact, nor does it integrate video-based reasoning with instructor-defined grading rules. 9
Conclusion
This work introduces Raven, a significant advance in automated assessment for Scratch programs. By integrating large language models with video-based analysis, Raven addresses key limitations of prior approaches, enabling evaluation of visual behaviors, accommodation of diverse student implementations, and support for human–computer interaction assessment. Our evaluation shows , Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
19
that Raven produces accurate, fine-grained assessments that align closely with instructor grading, and in-classroom deployments show that Raven is practical and usable in real-world settings. These results suggest that Raven provides a strong foundation for scalable and grounded assessment in block-based programming environments. References [1] Dimah Al-Fraihat, Mike Joy, Ra’ed Masa’deh, and Jane Sinclair. 2020. Evaluating E-learning systems success: An empirical study. Computers in Human Behavior 102 (2020), 67–86. doi:10.1016/j.chb.2019.08.004 [2] Kirsti M Ala-Mutka. 2005. A Survey of Automated Assessment Approaches for Programming Assignments. Computer Science Education 15, 2 (2005), 83–102. doi:10.1080/08993400500150747 [3] Sinan Ariyurek, Aysu Betin-Can, and Elif Surer. 2021. Automated Video Game Testing Using Synthetic and Humanlike Agents. IEEE Transactions on Games 13, 1 (2021), 50–67. doi:10.1109/TG.2019.2947597 [4] Shuai Bai, Yuxuan Cai, Ruizhe Chen, et al. 2025. Qwen3-VL Technical Report. arXiv:2511.21631 [cs.CV] https: //arxiv.org/abs/2511.21631 [5] Shuai Bai, Keqin Chen, Xuejing Liu, et al. 2025. Qwen2.5-VL Technical Report. arXiv:2502.13923 [cs.CV] https: //arxiv.org/abs/2502.13923 [6] Mohammad Bajammal, Andrea Stocco, Davood Mazinanian, and Ali Mesbah. 2022. A Survey on the Use of Computer Vision to Improve Software Engineering Tasks. IEEE Transactions on Software Engineering 48, 5 (2022), 1722–1742. doi:10.1109/TSE.2020.3032986 [7] Ishan Banerjee, Bao Nguyen, Vahid Garousi, and Atif Memon. 2013. Graphical user interface (GUI) testing: Systematic mapping and repository. Information and Software Technology 55, 10 (2013), 1679–1694. doi:10.1016/j.infsof .2013.03.004 [8] Bryce Boe, Charlotte Hill, Michelle Len, Greg Dreschler, Phillip Conrad, and Diana Franklin. 2013. Hairball: lintinspired static analysis of scratch projects. In Proceeding of the 44th ACM Technical Symposium on Computer Science Education (Denver, Colorado, USA) (SIGCSE ’13). Association for Computing Machinery, New York, NY, USA, 215–220. doi:10.1145/2445196.2445265 [9] Janet Carter, Kirsti Ala-Mutka, Ursula Fuller, Martin Dick, John English, William Fone, and Judy Sheard. 2003. How shall we assess this?. In Working Group Reports from ITiCSE on Innovation and Technology in Computer Science Education (Thessaloniki, Greece) (ITiCSE-WGR ’03). Association for Computing Machinery, New York, NY, USA, 107–123. doi:10.1145/960875.960539 [10] Cecilia Ka Yuk Chan and Wenjie Hu. 2023. Students’ voices on generative AI: perceptions, benefits, and challenges in higher education. International Journal of Educational Technology in Higher Education 20, 1 (July 2023), 43. doi:10.1186/ s41239-023-00411-8 [11] Li-Hsin Chang and Filip Ginter. 2024. Automatic Short Answer Grading for Finnish with ChatGPT. Proceedings of the AAAI Conference on Artificial Intelligence 38, 21 (Mar. 2024), 23173–23181. doi:10.1609/aaai.v38i21.30363 [12] Tsung-Hsiang Chang, Tom Yeh, and Robert C. Miller. 2010. GUI testing using computer vision. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Atlanta, Georgia, USA) (CHI ’10). Association for Computing Machinery, New York, NY, USA, 1535–1544. doi:10.1145/1753326.1753555 [13] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, and et al. 2021. Evaluating Large Language Models Trained on Code. arXiv:2107.03374. https://arxiv.org/abs/2107.03374 [14] Fred D. Davis. 1989. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Q. 13, 3 (Sept. 1989), 319–340. doi:10.2307/249008 [15] Adina Deiner, Patric Feldmeier, Gordon Fraser, Sebastian Schweikl, and Wengran Wang. 2023. Automated Test Generation for Scratch Programs. Empirical Software Engineering 28, 1 (2023), 79. doi:10.1007/s10664-022-10255-x [16] Adina Deiner, Christoph Frädrich, Gordon Fraser, Sophia Geserer, and Niklas Zantner. 2020. Search-based Testing for Scratch Programs. CoRR abs/2009.04115 (2020). arXiv:2009.04115 https://arxiv.org/abs/2009.04115 [17] Adina Deiner and Gordon Fraser. 2024. NuzzleBug: Debugging Block-Based Programs in Scratch. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE ’24). 1—2. doi:10.1145/3597503.3623331 [18] Benedikt Fein, Florian Obermüller, and Gordon Fraser. 2022. CATNIP: An Automated Hint Generation Tool for Scratch. In Proceedings of the 27th ACM Conference on on Innovation and Technology in Computer Science Education Vol. 1 (Dublin, Ireland) (ITiCSE ’22). Association for Computing Machinery, New York, NY, USA, 124–130. doi:10.1145/3502718.3524820 [19] Andrina Granić and Nikola Marangunić. 2019. Technology acceptance model in educational context: A systematic literature review. British Journal of Educational Technology 50, 5 (2019), 2572–2593. doi:10.1111/bjet.12864 [20] Christian Grévisse. 2024. LLM-based automatic short answer grading in undergraduate medical education. BMC Medical Education 24, 1 (Sept. 2024), 1060. doi:10.1186/s12909-024-06026-5
, Vol. 1, No. 1, Article . Publication date: April 2026.
20
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
[21] Jialiang Gu, Keren Zhou, Daming Li, Hanyuan Shi, and Jialu Zhang. 2026. Context-Aware Feedback Compression in Online Judge Programming with LLMs. In Proceedings of the 34th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE Companion ’26) (Montreal, QC, Canada, 5–9 July 2026) (FSE Companion ’26). Association for Computing Machinery, New York, NY, USA. doi:10.1145/3803437.3805565 Model-based Testing of Scratch Programs. [22] Katharina Götz, Patric Feldmeier, and Gordon Fraser. 2022. arXiv:2202.06271 [cs.SE] https://arxiv.org/abs/2202.06271 [23] Felienne Hermans and Efthimia Aivaloglou. 2016. Do code smells hamper novice programming? A controlled experiment on Scratch programs. In 2016 IEEE 24th International Conference on Program Comprehension (ICPC). 1–10. doi:10.1109/ICPC.2016.7503706 [24] Silas Hsu, Tiffany Wenting Li, Zhilin Zhang, Max Fowler, Craig Zilles, and Karrie Karahalios. 2021. Attitudes Surrounding an Imperfect AI Autograder (CHI ’21). Association for Computing Machinery, New York, NY, USA, Article 681, 15 pages. doi:10.1145/3411764.3445424 [25] Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, et al. 2024. CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 650, 20 pages. doi:10.1145/3613904.3642773 [26] Hieke Keuning, Bastiaan Heeren, and Johan Jeuring. 2017. Code Quality Issues in Student Programs. In Proceedings of the 2017 ACM Conference on Innovation and Technology in Computer Science Education (ITiCSE ’17). Association for Computing Machinery, New York, NY, USA, 110–115. doi:10.1145/3059009.3059061 [27] Hieke Keuning, Johan Jeuring, and Bastiaan Heeren. 2019. A Systematic Literature Review of Automated Feedback Generation for Programming Exercises. ACM Transactions on Computing Education 19, 1 (2019), 3:1–3:43. doi:10.1145/ 3231711 [28] Juho Leinonen, Paul Denny, Stephen MacNeil, et al. 2023. Comparing Code Explanations Created by Students and Large Language Models. In Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1 (ITiCSE 2023). Association for Computing Machinery, New York, NY, USA, 124–130. doi:10.1145/3587102.3588785 [29] Yuanchun Li, Ziyue Yang, Yao Guo, and Xiangqun Chen. 2017. DroidBot: a lightweight UI-Guided test input generator for android. In 2017 IEEE/ACM 39th International Conference on Software Engineering Companion (ICSE-C). 23–26. doi:10.1109/ICSE-C.2017.8 [30] Xiaoyun Liang, Jiayi Qi, Yongqiang Gao, Chao Peng, and Ping Yang. 2023. AG3: Automated Game GUI Text Glitch Detection Based on Computer Vision. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (San Francisco, CA, USA) (ESEC/FSE 2023). Association for Computing Machinery, New York, NY, USA, 1879–1890. doi:10.1145/3611643.3613867 [31] Qianou Ma, Hua Shen, Kenneth Koedinger, and Sherry Tongshuang Wu. 2024. How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging. Springer Nature Switzerland, 265–279. doi:10.1007/978-3-03164302-6_19 [32] John Maloney, Mitchel Resnick, Natalie Rusk, Brian Silverman, and Evelyn Eastmond. 2010. The Scratch Programming Language and Environment. 10, 4, Article 16 (Nov. 2010), 15 pages. doi:10.1145/1868358.1868363 [33] Marcus Messer, Neil C. C. Brown, Michael Kölling, and Miaojing Shi. 2024. Automated Grading and Feedback Tools for Programming Education: A Systematic Review. 24, 1, Article 10 (Feb. 2024), 43 pages. doi:10.1145/3636515 [34] Jesús Moreno-León and Gregorio Robles. 2015. Dr. Scratch: a Web Tool to Automatically Evaluate Scratch Projects. In Proceedings of the Workshop in Primary and Secondary Computing Education (London, United Kingdom) (WiPSCE ’15). Association for Computing Machinery, New York, NY, USA, 132–133. doi:10.1145/2818314.2818338 [35] Tushar Nagarajan and Kristen Grauman. 2021. Shaping embodied agent behavior with activity-context priors from egocentric video. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 29794–29805. https://proceedings.neurips.cc/ paper_files/paper/2021/file/f8b7aa3a0d349d9562b424160ad18612-Paper.pdf [36] Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024. Using an LLM to Help With Code Understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (Lisbon, Portugal) (ICSE ’24). Association for Computing Machinery, New York, NY, USA, Article 97, 13 pages. doi:10.1145/ 3597503.3639187 [37] Seong-Guk Nam and Yeong-Seok Seo. 2023. GUI Component Detection-Based Automated Software Crash Diagnosis. Electronics 12, 11 (2023). doi:10.3390/electronics12112382 [38] Bao N. Nguyen, Bryan Robbins, Ishan Banerjee, and Atif Memon. 2014. GUITAR: an innovative tool for automated testing of GUI-driven software. Automated Software Engineering 21, 1 (March 2014), 65–105. doi:10.1007/s10515-0130128-9 [39] OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, et al. 2024. GPT-4 Technical Report. arXiv:2303.08774 [cs.CL] https://arxiv.org/abs/2303.08774
, Vol. 1, No. 1, Article . Publication date: April 2026.
Raven: Rethinking Automated Assessment for Scratch Programs via Video-Grounded Evaluation
21
[40] Carlos Pacheco, Shuvendu K. Lahiri, Michael D. Ernst, and Thomas Ball. 2007. Feedback-Directed Random Test Generation. In 29th International Conference on Software Engineering (ICSE’07). 75–84. doi:10.1109/ICSE.2007.37 [41] José Carlos Paiva, José Paulo Leal, and Álvaro Figueira. 2022. Automated Assessment in Computer Science Education: A State-of-the-Art Review. ACM Trans. Comput. Educ. 22, 3, Article 34 (June 2022), 40 pages. doi:10.1145/3513140 [42] Raphael Pham, Helge Holzmann, Kurt Schneider, and Christian Brüggemann. 2014. Tailoring video recording to support efficient GUI testing and debugging. Software Quality Journal 22, 2 (June 2014), 273–292. doi:10.1007/s11219-013-9206-2 [43] Tung Phung, José Cambronero, Sumit Gulwani, Tobias Kohn, Rupak Majumdar, Adish Singla, and Gustavo Soares. 2023. Generating High-Precision Feedback for Programming Syntax Errors using Large Language Models. Proceedings of the 16th International Conference on Educational Data Mining (EDM 2023) (2023), 370–377. doi:10.5281/zenodo.8115653 [44] Thomas W. Price, Yihuan Dong, and Dragan Lipovac. 2017. iSnap: Towards Intelligent Tutoring in Novice Programming Environments. In SIGCSE. 483–488. doi:10.1145/3017680.3017762 [45] Mitchel Resnick. 2017. Lifelong Kindergarten: Cultivating Creativity through Projects, Passion, Peers, and Play. MIT Press, Cambridge, MA. https://mitpress.mit.edu/9780262037297/lifelong-kindergarten/ [46] Mitchel Resnick, John Maloney, Andrés Monroy-Hernández, Natalie Rusk, Evelyn Eastmond, Karen Brennan, Amon Millner, Eric Rosenbaum, Jay Silver, Brian Silverman, and Yasmin Kafai. 2009. Scratch: programming for all. Commun. ACM 52, 11 (Nov. 2009), 60–67. doi:10.1145/1592761.1592779 [47] Zachary P. Reynolds, Abhinandan B. Jayanth, Ugur Koc, et al. 2017. Identifying and Documenting False Positive Patterns Generated by Static Code Analysis Tools. In 2017 IEEE/ACM 4th International Workshop on Software Engineering Research and Industrial Practice (SER&IP). 55–61. doi:10.1109/SER-IP.2017..20 [48] Kelly Rivers and Kenneth R. Koedinger. 2017. Data-Driven Hint Generation in Vast Solution Spaces: A Self-Improving Python Programming Tutor. International Journal of Artificial Intelligence in Education 27, 1 (2017), 37–64. doi:10.1007/ s40593-015-0070-z [49] Marcos Román-González, Jesús Moreno-León, and Gregorio Robles. 2017. Complementary Tools for Computational Thinking Assessment. https://www.researchgate.net/publication/ 318469859_Complementary_Tools_for_Computational_Thinking_Assessment [50] Mark Santolucito, Jialu Zhang, Ennan Zhai, Jürgen Cito, and Ruzica Piskac. 2022. Learning CI Configuration Correctness for Early Build Feedback. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). 1006–1017. doi:10.1109/SANER53432.2022.00118 [51] Ronny Scherer, Fazilat Siddiq, and Jo Tondeur. 2019. The technology acceptance model (TAM): A meta-analytic structural equation modeling approach to explaining teachers’ adoption of digital technology in education. Computers & Education 128 (2019), 13–35. doi:10.1016/j.compedu.2018.09.009 [52] Sebastian Schweikl and Gordon Fraser. 2025. RePurr: Automated Repair of Block-Based Learners’ Programs. Proc. ACM Softw. Eng. 2, FSE, Article FSE067 (June 2025), 24 pages. doi:10.1145/3715786 [53] Scratch Foundation. 2026. Scratch Statistics - Scratch Imagine, Program, Share. https://scratch.mit.edu/statistics/. Accessed: 2026-01-29. [54] Yuan Si, Simeng Han, Daming Li, Hanyuan Shi, and Jialu Zhang. 2026. ScratchEval : A Multimodal Evaluation Framework for LLMs in Block-Based Programming. arXiv:2602.00757 [cs.SE] https://arxiv.org/abs/2602.00757 [55] Yuan Si, Daming Li, Hanyuan Shi, and Jialu Zhang. 2025. ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch. arXiv:2509.11065 [cs.SE] https://arxiv.org/abs/2509.11065 [56] Yuan Si, Kyle Qi, Daming Li, Hanyuan Shi, and Jialu Zhang. 2025. Stitch: Step-by-step LLM Guided Tutoring for Scratch. arXiv:2510.26634 [cs.SE] https://arxiv.org/abs/2510.26634 [57] Yuan Si, Ming Wang, Daming Li, Hanyuan Shi, and Jialu Zhang. 2026. EcoScratch: Cost-Effective Multimodal Repair for Scratch Using Execution Feedback. arXiv:2603.29624 [cs.SE] https://arxiv.org/abs/2603.29624 [58] Andreas Stahlbauer, Marvin Kreis, and Gordon Fraser. 2019. Testing scratch programs automatically. In Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (Tallinn, Estonia) (ESEC/FSE 2019). Association for Computing Machinery, New York, NY, USA, 165–175. doi:10.1145/3338906.3338910 [59] Niko Strijbol, Robbe De Proft, Klaas Goethals, Bart Mesuere, Peter Dawyndt, and Christophe Scholliers. 2024. Blink: An educational software debugger for Scratch. SoftwareX 25 (2024), 101617. doi:10.1016/j.softx.2023.101617 [60] Shao-Hua Sun, Hyeonwoo Noh, Sriram Somasundaram, and Joseph Lim. 2018. Neural Program Synthesis from Diverse Demonstration Videos. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80), Jennifer Dy and Andreas Krause (Eds.). PMLR, 4790–4799. https://proceedings.mlr.press/ v80/sun18a.html [61] Xiaodan Tang, Yue Yin, Qiao Lin, Roxana Hadad, and Xiaoming Zhai. 2020. Assessing computational thinking: A systematic review of empirical studies. Computers & Education 148 (2020), 103798. doi:10.1016/j.compedu.2019.103798 [62] Peng Wang, Shuai Bai, Sinan Tan, et al. 2024. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv:2409.12191 [cs.CV] https://arxiv.org/abs/2409.12191
, Vol. 1, No. 1, Article . Publication date: April 2026.
22
Donglin Li, Daming Li, Hanyuan Shi, and Jialu Zhang
[63] Yuxuan Wang, Yueqian Wang, Dongyan Zhao, Cihang Xie, and Zilong Zheng. 2024. VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models. arXiv:2406.16338 [cs.CV] https://arxiv.org/ abs/2406.16338 [64] An Yang, Anfeng Li, Baosong Yang, et al. 2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] https://arxiv.org/ abs/2505.09388 [65] Zhengyuan Yang, Linjie Li, Kevin Lin, et al. 2023. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision). arXiv:2309.17421 [cs.CV] https://arxiv.org/abs/2309.17421 [66] Tom Yeh, Tsung-Hsiang Chang, and Robert C. Miller. 2009. Sikuli: using GUI screenshots for search and automation. In Proceedings of the 22nd Annual ACM Symposium on User Interface Software and Technology (Victoria, BC, Canada) (UIST ’09). Association for Computing Machinery, New York, NY, USA, 183–192. doi:10.1145/1622176.1622213 [67] Jialu Zhang, José Pablo Cambronero, Sumit Gulwani, Vu Le, Ruzica Piskac, Gustavo Soares, and Gust Verbruggen. 2024. PyDex: Repairing Bugs in Introductory Python Assignments using LLMs. Proc. ACM Program. Lang. 8, OOPSLA1 (2024), 1100–1124. doi:10.1145/3649850 [68] Jialu Zhang, Jialiang Gu, Wangmeiyu Zhang, José Pablo Cambronero, John Kolesar, Ruzica Piskac, Daming Li, and Hanyuan Shi. 2025. A Systematic Study of Time Limit Exceeded Errors in Online Programming Assignments. arXiv:2510.14339 [cs.SE] https://arxiv.org/abs/2510.14339 [69] Jialu Zhang, De Li, John Charles Kolesar, Hanyuan Shi, and Ruzica Piskac. 2023. Automated Feedback Generation for Competition-Level Code. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (Rochester, MI, USA) (ASE ’22). Association for Computing Machinery, New York, NY, USA, Article 13, 13 pages. doi:10.1145/3551349.3560425 [70] Jialu Zhang, Todd Mytkowicz, Mike Kaufman, Ruzica Piskac, and Shuvendu K. Lahiri. 2022. Using pre-trained language models to resolve textual and semantic merge conflicts (experience paper). In ISSTA ’22: 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, South Korea, July 18 - 22, 2022, Sukyoung Ryu and Yannis Smaragdakis (Eds.). ACM, 77–88. doi:10.1145/3533767.3534396 [71] Yan Zheng, Xiaofei Xie, Ting Su, Lei Ma, et al. 2019. Wuji: Automatic Online Combat Game Testing Using Evolutionary Deep Reinforcement Learning. In 2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE). 772–784. doi:10.1109/ASE.2019.00077
, Vol. 1, No. 1, Article . Publication date: April 2026.