arXiv:2607.19209v1 [cs.CY] 21 Jul 2026
Assessment in Team Problem-Solving Exercises in Computing Education Valdemar Švábenský
Jan Vykopal
Sukrit Leelaluk
Faculty of Informatics Masaryk University Brno, Czech Republic [email protected] 0000-0001-8546-280X
Faculty of Informatics Masaryk University Brno, Czech Republic [email protected] 0000-0002-3425-0951
Center for ICT Infrastructure Yamaguchi University Yamaguchi, Japan [email protected] 0009-0003-3493-6878
Pavel Čeleda
Fumiya Okubo
Atsushi Shimada
Faculty of Informatics Masaryk University Brno, Czech Republic [email protected] 0000-0002-3338-2856
Graduate School of Information Science and Electrical Engineering, Kyushu University Fukuoka, Japan [email protected] 0000-0002-0077-9072
Graduate School of Information Science and Electrical Engineering, Kyushu University Fukuoka, Japan [email protected] 0000-0002-3635-9336
Abstract—This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable learner teams to prepare for workplace tasks and practice crisis responses, such as resolving cybersecurity incidents. While assessment is essential for determining how well teams achieve learning objectives, the complex, openended nature of TTXs often leads to delayed or incomplete feedback. TTX learning platforms can record teams’ actions and communication; yet, leveraging these data to assess performance is underexplored. To address this gap, we compared two post-TTX team assessment methods—clustering and large language models (LLMs)—using an original dataset from 81 participants across two countries. We evaluated these methods against instructorassigned scores based on standardized rubrics. Clustering grouped teams that approached TTX tasks similarly, enabling instructors to deliver faster, targeted feedback to teams within a cluster. This method was valid and reliable, with low computational requirements. LLMs used the standardized rubrics to assess teams’ communication. While GPT-4o frequently disagreed with instructor scores, GPT-5.2 demonstrated considerably lower error. The researched methods have been integrated into INJECT, an open-source TTX learning platform, to support scalability and teaching practice. To encourage community adoption, we publicly share all datasets, software tools, and a full-fledged TTX scenario. Index Terms—collaborative learning, cyber security, cyber exercise, TTX, incident response, INJECT, clustering, LLM
I. I NTRODUCTION Hands-on learning through tabletop exercises (TTXs) is gaining traction in educational applications [1], [2], [3], [4]. In TTXs, which are grounded in simulation-based learning [5], small teams of students collaboratively solve complex tasks This research was supported by the Open Calls for Security Research 2023– 2029 (OPSEC) program granted by the Ministry of the Interior of the Czech Republic under No. VK01030007 – Intelligent Tools for Planning, Conducting, and Evaluating Tabletop Exercises.
within a limited time frame. All teams are seated in one room at separate tables, and instructors present shared tasks that each team addresses independently of other teams. Solutions are presented orally or submitted digitally, using tools ranging from simple online forms to specialized learning platforms. TTXs are suitable for teaching cybersecurity: a computing discipline [6] that integrates technology, people, information, and processes to ensure operations in the face of adversaries [7]. TTXs simulate cyber incidents: crisis scenarios that disrupt IT operations in an organization, such as a data breach. Team members collaborate to resolve the incident, tackling technical and non-technical challenges. As a result, TTXs foster competencies in realistic, workplace-like settings, preparing learners for the multifaceted demands of this rapidly evolving field. Due to these benefits, as well as the alignment with computing curricula [6], [8], TTXs have been applied in various cybersecurity training scenarios [9], [10], [11], [12], [13], [14]. A. Background, Current Gaps, and Problem Statement In a TTX, each team’s objective is to address incidents by completing specific tasks in accordance with appropriate procedures. In the customer data breach example, this involves actions such as blocking remote access to the database and contacting the system administrator. While each team faces the same incident scenario, their approaches can differ substantially in the activities they perform, their sequence, and timing. Therefore, to increase the TTXs’ educational impact, it is essential to assess team performance by evaluating which activities were completed and how well. Such an assessment can serve various educational purposes, including: 1) grading teams in formal educational settings, 2) supporting an immediate post-TTX debriefing (so-called hot wash [15], [16]), and
©2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. Cite this article as follows: V. Švábenský, J. Vykopal, S. Leelaluk, P. Čeleda, F. Okubo, A. Shimada. Assessment in Team Problem-Solving Exercises in Computing Education. In Proceedings of the 56th IEEE Frontiers in Education Conference (FIE ’26). Paphos, Cyprus, 2026. DOI: TODO add after the proceedings publication.
3) providing feedback that highlights team strengths and 4) Post-Exercise, Not Intermediate: During a TTX, instrucareas for improvement, including how the teams deviated tors are busy facilitating and cannot divide their attention from the procedures expected by the instructor (which are between intermediate assessments. However, timely postnot necessarily the only possible correct solution). exercise feedback is crucial [20], [21]. Assessment is thus However, assessing teams in TTXs poses many challenges. conducted right after the TTX, delivering instant results within Teams may solve TTX tasks in various ways, complicating minutes (instead of hours or days). 5) Fully Automated, Not Manual: To overcome the limevaluation. Unlike established open-ended assessment tasks itations of manual assessment (see Section I-A), we utilize (e.g., essay grading [17], [18]), TTX assessment lacks standardautomated methods. Software tools process TTX data without ized frameworks. Lastly, team activity data vary in structure, human intervention, ensuring efficiency. Manual assessment is format, and length, since teams submit written responses and only used as a benchmark for evaluation in this study. supporting documents. In some TTXs, additional data include observer notes, audio/video recordings, or even IT system logs. C. Goals and Contributions of This Study Traditionally, teams have been assessed manually through Prior research on assessment in complex exercises (see questionnaires and observations – a practice still used toSection II) indicates that clustering and LLMs were effective day [19]. While manual data collection and analysis can in related educational contexts. Therefore, these methods may yield valuable insights, it suffers from two major drawbacks. be viable for addressing the challenges in TTX assessment. First, TTX participants have to wait several days or even However, the evaluation of the methods in the TTX context weeks to receive assessment-based feedback, which reduces remains unexplored. To address this gap, the research aspect its educational impact [20]. By the time feedback is delivered, of our study is framed by two research questions (RQ): participants may have forgotten details of their actions or lack RQ1 [Quantitative] How does team assessment using (a) the time to attend a separate feedback session. Furthermore, clustering and (b) LLMs align with instructor scores? the time constraints force instructors to strike a difficult balance between speed and depth of assessment. Providing RQ2 [Qualitative] Which insights for team assessment can be derived from TTX data using (a) clustering and (b) LLMs? faster feedback often leads to superficial, generic assessments that fail to address individual teams’ needs, while thorough Regarding the practice aspect of this research-to-practice paper, feedback requires a lot of the instructor’s time. This difficulty we bring various contributions and insights: of delivering assessments that are both timely and meaningful • Based on our literature review (Section II), this paper adhas also been observed in other cybersecurity exercises [21]. vances prior research and practice by evaluating automated We aim to improve the assessment by automatically anTTX assessment using learning analytics methods. alyzing TTX data using learning analytics methods. Using • We conducted two different cybersecurity TTXs in aueducational logs to provide insights into learning processes and thentic teaching environments (see Section III), involving outcomes has proven useful in various contexts [22], [23], [24]. 24 teams of 81 students in two countries (see Section IV). However, these methods have not yet been validated in the TTX These multi-national, multi-institutional study settings context. Our paper addresses this gap by empirically evaluating address many of the limitations of previous work [26]. automated assessment of teams and sharing practical experience • We demonstrate and evaluate the assessment methods from real-world TTXs. The proposed approach aims to speed by analyzing student datasets. We compare the methods’ up the assessment and deliver insights into team performance alignment with manual scoring, their advantages and that are difficult to obtain manually. limitations, and their educational use cases (see Section V). • All materials and tools are available, enabling researchers B. Definition of Scope of Student Assessment in This Study and educators to adopt or adapt them (see Section VI). We focus on assessment as characterized below, grounded II. R ELATED W ORK IN T EAM A SSESSMENT in assessment theory and computing education frameworks [8]. 1) Teams, Not Individuals: Our study examines overall team Effective team-based problem-solving is critical in a wide performance, since a team is the fundamental unit. A team range of domains [27], especially in high-stakes contexts succeeds only when its members collectively resolve the tasks. such as healthcare [28], military [29], and cybersecurity [13], Analyzing individual contributions within a team, as in other [30], [31], [32], [33], [34]. TTXs simulate crisis scenarios in non-TTX work [25], is outside the scope of the current paper. these contexts, making team assessment vital for evaluating 2) Summative, Not Formative: Team success is measured performance and learning outcomes. We examine prior research by completing tasks marked by predefined checkpoints called in this area, aligned with our RQs, to show how assessment milestones (e.g., “blocked the malicious website”). Milestones challenges have been addressed in non-TTX contexts. serve as indicators of the learning objectives the team has met, A. Clustering in Team Assessment and our assessment aims to summarize overall achievement. 3) Exercise-Specific, Not Longitudinal: The assessment Clustering is an unsupervised machine learning technique relies only on data collected during a TTX. We do not widely applied in education [23] to analyze team behavior, acincorporate historical participant data to ensure independence tivities, and performance [35], as well as to provide educational of external factors and support student data privacy. feedback [36].
For example, [37] leveraged logs from software development TABLE I S IMILARITIES AND DIFFERENCES BETWEEN OUR EXERCISES . group projects to enhance teamwork. They classified seven teams into three clusters based on activities such as commit TTX Cyber incident # Milestones # Tools (with representative frequency, ticket management, and wiki updates. The clusters’ topic (important + examples in parentheses) secondary) analysis uncovered diverse work strategies to inform teaching EXF Data exfiltration 17 (5 + 12) 11 (IP geolocation, IP blocking, interventions. In another programming context, [38] clustered by malware log search, . . . ) written discussion data from 34 teams to analyze collaborative PHI Phishing attacks 18 (14 + 4) 11 (DNS lookup, log search, behaviors. The analysis revealed three clusters reflecting campaign email sender blocking, . . . ) varying levels of collaboration, identifying teams with notably less interaction. Similarly, [39] examined team orientation in terms of goals, roles, and processes within a projectTTXs in our study simulate real-world cybersecurity inbased learning environment. Clustering 23 teams showed that cident handling, offering an immersive learning experience. teams with balanced orientations achieved better academic The exercises mirror the tasks performed by a Computer performance compared to those with unbalanced orientations. Security Incident Response Team (CSIRT) in medium to large These studies demonstrate the utility of clustering when clear organizations. Participants must prioritize and address multiple assessment objectives guide the evaluation. We advance prior issues within a limited time frame (90 minutes), reflecting the work by extending this method to the open-ended TTX context, pressures and dynamics of realistic cyber incident response. where such objectives are not explicitly defined. By clustering The TTXs in this research have the same learning objectives, TTX data, we assess similarities and differences among teams. which align with the skills defined for the Incident Response role in the well-established NICE Cybersecurity Workforce B. Large Language Models in Team Assessment Framework [52]. In addition, the TTXs aim to develop The recent surge in the use of LLMs in education has dispositions [6], [8] such as teamwork, decision-making, and introduced new benefits and challenges [40], [41]. Among their professional communication – both within the CSIRT and with many applications, the use of LLMs to analyze language data stakeholders involved in the simulated incidents. is relevant to our study. [42] used GPT-3.5 to give feedback B. Exercise Format and Content on student project reports. While the researchers cautioned Participants assume the role of CSIRT members tasked with about potentially unreliable outputs, model fine-tuning reduced handling reports of incidents affecting the IT infrastructure of hallucinations in the generated feedback. However, [43] evalua simulated organization. Table I details the incident scenarios. ated GPT-3.5 and GPT-4 by coding data from interviews, and The teams’ responsibilities include: (1) discussing the reports neither model achieved the expected reliability. to determine appropriate actions, (2) requesting information or [44] used LLMs to assess open-text responses from tutors in collaboration from internal or external actors, and (3) using educational scenarios. While this domain shares the ill-defined simulated tools for incident response. Although the TTX themes nature of TTX assessment, its objectives differ. Next, [45] differ, they both address the overarching learning objectives assessed student programming collaboration by summarizing defined in Section III-A. Therefore, we study both TTXs jointly. conversations of nine student pairs. [46] compared student The TTXs were conducted in-person, using the open-source perceptions of feedback generated by a teacher vs. an LLM. Lastly, [47] provided evidence that traditional machine learning INJECT Exercise Platform (IXP) [53]. Prior to the TTX, participants completed a brief tutorial on using the platform and models outperformed LLMs in short answer grading. were assigned to teams (see Section IV-A2). After a briefing, Compared to prior work, we integrate rubrics from estabparticipants could ask questions and designate team roles, such lished cybersecurity curricular guidelines into the LLM-based as appointing a team member to interact with the IXP interface. assessment. The rubric aims to mitigate prior challenges and Once the TTX began, teams interacted solely with their ensure that the TTX assessment aligns with learning objectives. teammates and the IXP, which delivered messages based on III. E DUCATIONAL C ONTEXT OF E XERCISES IN O UR S TUDY a predefined scenario. Instructors role-played as TTX actors who communicated with teams via email in the IXP. The This section outlines the properties of TTXs we used. TTX intentionally provided minimal guidance to simulate the uncertainty of a cyber incident. Participants received context A. Exercise Goals and Learning Objectives (e.g., “Your organization is experiencing phishing attacks.”) From an educational-theory perspective, TTXs connect to the and were thrown in the middle of the scenario. They had to broader landscape of simulation-based learning [5]. Through triage written queries from stakeholders (e.g., “I opened this simulations of authentic scenarios, students learn to address email attachment, and my computer stopped working. What practical problems in a safe learning environment where failure should I do?”), evaluate the situation, prioritize actions, devise has no negative consequences. Each team progresses with solutions, and write up relevant responses. minimal to no guidance from the instructor, employing selfEach TTX included activities aligned with the learning regulated learning (SRL) strategies. SRL is a well-studied objectives. These activities could sometimes be tackled in any framework [48], [49], [50], [51], but not in the TTX context. order, reflecting the non-linear nature of real-world incident
response. To further enhance authenticity, the IXP did not provide immediate feedback on teams’ progress; this was discussed later in the debriefing. After the TTX, teams received a scenario summary, and instructors supplemented it by explaining the suitable procedures. To provide timely, in-depth support to instructors and teams, our paper evaluates assessment methods that enhance the post-exercise phase.
participants in 23 teams in the dataset. Our group size aligns with or exceeds that of comparable state-of-the-art research; e.g., [25] analyzed data from 15 teams of four, [54] used data from 10 teams of three, and [32] had 4 teams of four to five. The TTXs were in English, which was a second language for all teams, though they were sufficiently proficient. Verbal communication within the team (out of scope of this study) was in the team’s respective native languages. IV. R ESEARCH AND A SSESSMENT M ETHODS 3) Data Content: The IXP logged each team’s activity, Figure 1 provides an overview of our study, illustrating the e.g., milestone achievements, tool usage, and email responses. important components and their interconnections. Communication between teams and instructors occurred only via the IXP. The data required filtering to use only teams’ = open tools activity, ensuring the analysis was relevant to team assessment. and materials We excluded data such as pre-scripted messages by the IXP, Analytical tools 13 teams × 2–3 people EXF exercise IXP logs (milestones, Czech Rep., university (practicing skills in activity, tool use, emails...) (Python scripts) which were uniform across teams and would have caused noise. All 13 teams consented incident handling) 4) Instructor-Assigned Scores: To establish a benchmark for Web interface automated assessment, two expert human instructors manually of the INJECT Exercise Platform compare assessed the teams. This manual scoring is used solely for (IXP) 11 teams × 3–5 people PHI exercise Estonia, workshop (practicing skills in Scores assigned manually Clusters and research evaluation in this paper to address RQ1. However, in by instructors (benchmark) LLM scores 10 / 11 teams consented incident handling) a practical classroom setting, this manual scoring would not Fig. 1. Conceptual illustration of the research methods. Two distinct groups of occur, which aligns with the objectives in Section I-A. teams independently completed the two TTXs in the same learning platform. The instructors used the IXP logs to assign numerical scores Based on data collected from this platform, an assessment was conducted both to each team. Two independent aspects of team performance manually and automatically, and the scores and outputs were compared. were assessed – milestone completion and communication. In both cases, the instructors followed pre-defined scoring rubrics A. Data Collection From Authentic Teaching Contexts (see below) to ensure consistency. Both rubrics assigned scores We focus on computing students (including, but not limited on a 3-step ordinal scale from 0 to 2 to ensure comparability. to, those in dedicated cybersecurity programs) in higher First, a simple script developed by the instructor scored each and formal education. To support assessment generalizability, team’s achievement of TTX milestones to obtain team scores. we gathered diverse data from various TTXs, students, and For each team and each milestone, we assigned a score of countries, as advised by computing education researchers [26]. 2 for completing an important milestone, 1 for completing a 1) Research Ethics: We received a waiver from the ethics secondary (optional) milestone, and 0 for missing a milestone. review board, i.e., the committee did not require explicit The final assessment for each team was an m-dimensional approval. All TTXs were conducted primarily for educational vector of these scores, where m is the number of milestones purposes. In addition, the instructor explained to the student in that TTX, as indicated in Table I (i.e., 17 or 18). teams that anonymized TTX activity data could be used for Second, both instructors assessed whether a team’s written educational research. Each team indicated their consent as communication aligns with standards for cybersecurity work a whole via the IXP; declining consent had no negative roles. We based the assessment criteria on well-established consequences. This gave students the most control: if at least cybersecurity curricular guidelines [7]. These refer to a U.S. one team member was uncomfortable, the data were not stored. government-endorsed competency framework [55] with a rubric 2) Participant Population: In the EXF exercise (see Table I), for assessing writing in professional contexts. For each of 36 cybersecurity students from a Czech university participated. the rubric’s three aspects (see Table II later), the instructors The TTX was integrated at the end of a semester-long course assigned a score of 2 if an aspect was satisfied, 1 if it on cyber incident handling, allowing students to apply the was partially satisfied, and 0 if it was not satisfied. So, the knowledge they had gained. The instructor divided the students assessment for each team was a 3-dimensional score vector. into 13 teams to ensure balance based on their skill levels. While the course had 39 enrolled students, three absences B. Automated Assessment of TTX Teams To address our RQs, we evaluated automated methods. One resulted in 10 teams of three and 3 teams of two. PHI exercise was part of an event in Estonia to foster author implemented them in Python (see code in Section VI-B), cybersecurity skills and collaboration between the two countries. another conducted a code review, and the next validated the Participants were 11 teams – 5 teams of university students with results. For each method, we explain the use cases, algorithms, a security background (three per team, self-selected), 5 teams and input data. Then, we explain the method evaluation. 1) Clustering: In TTXs, where performance is reflected of vocational school students (five seniors per team, selected in the collaborative problem-solving process, behavior-based by their teachers), and 1 team of five computing educators. In total, 81 participants in 24 teams completed the TTXs. methods, such as clustering, are particularly well-suited for However, one team declined consent for research, leaving 76 assessment. Our goal was to assess how similar or different the
teams’ approaches were in solving the TTX tasks. The resulting clusters help instructors identify teams that approached the tasks similarly, enabling them to provide assessment-based feedback to the whole cluster and to determine which approaches were more effective. Since setting the suitable number of clusters in advance is difficult, we used density-based clustering with the DBSCAN algorithm. To reduce the impact of randomness caused by different seeds, we applied Unsupervised Consensus Clustering [56], averaging results from 50 iterations. Features were based on IXP logs of which milestones teams reached (milestone IDs) and when (timestamps), including the sequence of actions and tool usage. To quantitatively evaluate the outputs (RQ1), we examined team similarity within each cluster using manually assigned score vectors. We used the score vectors based on the milestones (not communication). Since the vectors have numeric components (0, 1, 2), we chose the Root Mean Squared Error (RMSE) because it captures both (a) score pattern closeness, indicating similar strengths and weaknesses, and (b) score magnitude, i.e., the performance on the important milestones matters more. We computed a pairwise RMSE between each member of a cluster (indicating the distance between two teams), then averaged the results for the whole cluster. This metric determines the validity of the automated output. As a comparative benchmark, our baseline was the average RMSE across all possible combinations of cluster groupings. 2) LLM-based Approach: An LLM mimicked the human assessment of a team’s written communication. To guide the LLM-based assessment, we used the same rubric as in Section IV-A4, ensuring comparability with the human-based scores. This rubric was integrated into a contextualized prompt that has been improved through multiple revisions. Then, we selected GPT-4o as the latest model at the time of our pilot study and GPT-5.2 for our validation study and provided the prompt via an API, tasking the model to assess how well the teams’ communication adhered to the standards. The input data consisted of the teams’ incident response email messages. We tried several prompts, and the first few yielded results that differed from the expected format. Initially, the LLM produced a summary paragraph of high-level feedback for each team, rather than separately considering the three rubric criteria. After iteratively updating our strategy, the prompt that provided a desired output format is detailed below (shortened version to save space; see details in the code linked in Section VI-B): You are given email texts written by teams of learners during a tabletop exercise with the topic of cybersecurity incident handling. These teams practice their cybersecurity competencies in a professional context. For each team, evaluate their email communication with the stakeholders using the following three criteria, assessing whether the team: 1) Communicates information in a succinct and organized manner. 2) Produces written information appropriate for the intended audience.
3) Uses correct English grammar, punctuation, and spelling. For each of the three criteria, output your assessment on the scale: Yes, No, Partially. To quantitatively evaluate the LLM output (RQ1), we examined the error (i.e., the “disagreement”) between the instructor’s and the LLM’s ratings in the communication-based score vectors. We chose RMSE as the metric for the same reasons as for clustering, thereby also enabling comparison. V. R ESULTS AND T HEIR D ISCUSSION Figure 2 and Table II show example outputs of the automated assessment. Section V-A provides an overall statistical comparison of the results (RQ1). Then, Section V-B dives deeper into the specifics of each method (RQ2). Section V-C discusses limitations, and Section V-D focuses on educational implications. Results are discussed separately for each exercise, highlighting differences across participant groups.
Fig. 2. Clustering in EXF. All 13 teams are identified by 3-digit IDs and grouped by activity similarity. The axes of the figure represent abstract dimensions in the clustering space, so their meaning is not directly interpretable.
A. Comparison of Manual and Automated Assessment (RQ1) Table III shows the evaluation results. For clustering, we assessed the cohesion of cluster membership according to the instructor’s scores. For the LLM-based approach, we assessed its error compared to the instructor. 1) Clustering: For n teams, the number of possible nonempty cluster groupings is the n-th Bell number Bn . EXF had 13 teams, giving B13 = 27,644,437 options. PHI had 10 teams, giving B10 = 115,975 options. Based on Table III, the baseline RMSE (average error across all these options) is much larger than for our clustering results. Therefore, our assessment has better internal cohesion of the clusters with respect to the manually assigned scores. We also note that this result applies across both TTXs. While the scenarios differ (leading to different baseline RMSEs), the clustering method is effective within each TTX. At the same time, some variance in the RMSE is essential. If the RMSE was 0, then all teams would be the same. So, to answer our RQ1(a), clustering sufficiently aligns with instructor scores and outperforms the baseline.
TABLE II G RADING OF COMMUNICATION OF SELECTED TEAMS ( OUT OF ALL 23), DEMONSTRATING EXAMPLES AND DIFFERENCES BETWEEN THE FINAL SELECTED INSTRUCTORS ’ SCORE AND LLM’ S ASSESSMENT. G RADING SCALE : Y ES ( ○ = 2), PARTIALLY ( = 1), N O ( ○ = 0). T HE LAST COLUMN ALSO ILLUSTRATES THE EFFECTS ON THE RMSE METRIC BASED ON THE GRADE VALUES IN EACH ROW. T HIS ALSO PROVIDES CONTEXT TO TABLE III. Note: The quotes in this column do not represent the entire grading decision. The conversations are much longer, and a broader context needs to be taken into account. The examples serve only illustrative purposes. Team Example quote from the written communication 333 The suspicious traffic was generated from user with compromised account 328 We have serious reason to believe this is malware trying to spread. 329 Please check if this scr ip is not in your competancy. 319 !Check your email password and what you have sent!!!!!!!!!!!
TABLE III D IFFERENCES BETWEEN MANUAL AND AUTOMATED ASSESSMENT, EXPRESSED AS AVERAGE RMSE ACROSS TEAMS PER TTX AND PER METHOD . T HE BASIS FOR MEASURING THESE VALUES IS DEFINED IN S ECTION IV-B. S INCE THE GRADING SCALE RANGED FROM 0 TO 2, THE RMSE RANGE IS THE SAME ; A LOWER RMSE INDICATES LOWER ERROR .
TTX EXF PHI Avg
Clustering: Assessing milestone completion Baseline RMSE 0.53 (27%) 0.45 (22%) 1.12 (56%) 0.79 (40%) 0.83 (41%) 0.62 (31%)
GPT-4o: Assessing communication RMSE 1.03 (51%) 1.18 (59%) 1.10 (55%)
GPT-5.2: Assessing communication RMSE 0.73 (37%) 0.68 (34%) 0.71 (35%)
2) LLM-based Approach: We first compared the two instructors to establish the ground truth. In both exercises, the graders’ RMSE was 0.32 (16%), which represents differences in 10% of assigned scores. In all cases, the differences were at most ±1 grade level. After one round of discussion, a final human score was established, favoring either instructor. LLMs analyzed teams’ email texts in natural language. However, since the GPT-4o’s disagreement with the humanassigned grades was 55%, the output essentially matched chance. Even though the LLM was guided by a clear prompt based on the standardized rubric criteria [55], it often yielded surprising results (see, e.g., the last row in Table II). GPT5.2 performed better, representing a substantial upgrade to 35% error. Therefore, we answer RQ1(b): while a newer LLM performs better, the assessments are not sufficiently aligned with the instructor’s judgment in the cyber TTX context. B. Qualitative Insights From Automated Assessment (RQ2) Next, we evaluated the qualitative utility of both methods. We link the obtained insights to educational use cases and show their practical benefits and limitations. 1) Clustering: Clustering grouped similar teams by compressing their data into two dimensions while preserving the relationships among features. Given the relatively small number of teams (13 and 10 in our data), it is expected to obtain few clusters. We observed from three to four clusters, each highlighting different patterns of team behavior and activity. For example, the EXF data yielded four clusters based on teams’ milestone achievements, see Figure 2. Cluster A included the highest-scoring teams that took a balanced approach, thoroughly addressing both technical and non-technical
Information succinct and organized Human GPT-4o
○ ○
○ ○ ○ ○
Appropriate for the intended audience Human GPT-4o
○ ○ ○
○ ○ ○
Correct grammar, spelling, punctuation Human GPT-4o
○ ○ ○
○ ○ ○ ○
RMSE 0.0000 0.5774 0.8165 2.0000
aspects. Cluster B also included high-scoring teams that excelled at procedural tasks but overlooked smaller technical interventions and secondary stakeholders. Cluster C was formed by the lowest-scoring teams. Although their performance was decent, they sometimes missed important milestones, did not use the correct tools, and lacked communication with TTX actors. Cluster D contained average-scoring teams that used technical tools but forgot to contact some key stakeholders. This information can be used by instructors to enable clusterdifferentiated feedback, for example, by recommending that teams in Cluster D contact actors such as the data protection officer, who is an important entity in the exercise context. Another reason why clustering is valuable for instructors is that it provides a visual overview of teams that approached the tasks similarly. Moreover, outliers that do not belong to any cluster may indicate teams with unique approaches. For example, in PHI, only two teams were successful and formed a separate cluster. It was the only one to follow the entire incident response protocol, achieving even easy-to-miss milestones. Next, our analysis revealed that the utility of time-based features, such as the sequence of tool usage, depends on the TTX content. In TTXs centered around a single incident (like in our study), these features provided relevant information for clustering. However, they may be less informative if the task completion order is more arbitrary. Finally, we note a limitation of clustering. Features derived from email texts, such as common words and n-grams, were not useful for distinguishing between the teams. Many phrases – mainly those related to cybersecurity processes – were equivalent across most teams, resulting in a single cluster. 2) LLM-based Approach: Since the teams’ written communication is in natural language, we assumed that an LLM would be well-suited for uniform assessment of teams’ email content. However, GPT-4o’s insights were not useful due to significant disagreement with the instructor; GPT-5.2 performed better. We attempted multiple prompts (see Section IV-B), but the first few produced unusable results. Throughout iterative improvements, we arrived at a prompt that appears to provide comprehensive coverage, and although the output looks reasonable at first glance, it often differs from the human expert’s assessment. We also experimented with categorizing the emails’ emotional tone. However, the general LLM was biased to almost always interpret messages as exhibiting only anxiety-related
emotions. Most likely, this stemmed from the prevalence of incident words (e.g., “cyber attack” or “data loss”) in the dataset. These results show that to obtain valid results, it is vital to frame the assessment within the TTX context or to use a domain-specific model, such as CyLLM [57] for cybersecurity. Another limitation of the assessment using public (non-local) LLMs is that it depended on querying an external service – unlike clustering, which was executed locally. If the LLM became unavailable or its version changed, the assessment’s validity and reliability would change as well. C. Limitations of This Study Regarding the teaching context, our study focused on students in cybersecurity or computing education programs who had at least a basic knowledge of cybersecurity. So, when using the same TTXs, our research findings may only reliably generalize to the same population type. Nevertheless, the TTX format may be used in many other teaching contexts (see Section I), potentially enabling the transfer of our methods to other classrooms and to entirely different curricula. Regarding the research methods, one could argue that the sample size was relatively small; however, Section IV-A2 shows that our team size aligns with or exceeds those in comparable recent research in learning analytics. Nevertheless, a limitation in the methods is the restriction to specific LLMs. Since we do not have access to the LLM’s internal workings, it is difficult to determine the reason the assessment sometimes differed. In a broader sense, this is a limitation of all black box methods.
nuances in this domain-specific context, even when guided by a standardized rubric in the prompt. 4) Consider Ethical Implications of LLM-based Assessment: As with all automated decisions, the output must be reviewed by a human instructor to ensure accountability and fairness. Our clustering considers only teams’ activity logs, not any external factors, unlike a (potentially biased) black box like an LLM. Moreover, since clustering runs locally, learners’ data are not shared with an external service, preserving privacy. 5) Generalize the Principles to Other Team Learning Scenarios: Our assessment methods are applicable not only in TTXs but can also be adapted to related computing education contexts. These include Capture the Flag (CTF) [33], [58], Cyber Defense Exercise (CDX), and other complex collaborative problemsolving exercises in cybersecurity, e.g., in cyber ranges [31], [32], [34] and beyond [59]. Although neither of these contexts is entirely similar to TTX, they are related in that they provide limited information to the student and may involve open-ended solutions, complicating assessment. VI. C ONCLUSIONS AND F UTURE W ORK
This study evaluated the automated assessment of student teams who practiced cybersecurity incident resolution. While clustering enhanced assessment, LLMs were less effective in this task. Even a relatively recent model struggled to assess the teams’ texts using a structured prompt, despite our revisions and inclusion of task context. Moreover, the LLM’s opaque internal mechanisms limited our ability to diagnose the root cause. These findings highlight that, despite growing D. Implications for Educational Research and Practice enthusiasm around LLMs, their applicability may be limited in In TTXs, student cohort sizes are constrained (tens of certain domain-specific scenarios. Should LLMs be used for teams), but manual assessment remains slow and insufficient. assessment, their outputs need to be carefully reviewed. Furthermore, it is complicated by the fact that teams can vary Next, this study contributed the clustering framework to in size and come from diverse backgrounds, as evidenced by advance the state of the art of assessment in TTXs. We our dataset from real classrooms. Based on our study and evaluated our methods on a unique dataset comprising two teaching experience, we provide recommendations to inform exercises and teams from two countries, thereby supporting educators and help improve TTX assessment practices. the generalizability of the findings. Even though the scale of 1) Design TTXs With Assessment in Mind: For automated TTXs may seem relatively small (tens of teams), the instructor assessment to work, a thorough design of TTX learning workload using this teaching method is high, limiting manual objectives and activities is crucial. Otherwise, the log data assessment. Our study identified and demonstrated this gap in may be less accurate in capturing the learning process. TTX the TTX literature and provided classroom-based evidence of creators should consider assessment goals during the design clustering’s applicability in supporting team assessment. phase to enable higher-quality analytics in the post-TTX phase. 2) Clustering Makes Assessment Feasible: Section V-A A. Open Research Challenges showed that there are millions of options for team similarity Future research can investigate combining the two methods, in TTXs, making it impossible for the instructor to discover potentially including other assessment methods, into a single relevant groupings manually. Clustering provides a quick ensemble framework. Here, aggregation (e.g., by voting) indicator of team similarity based on activity logs. For each can enhance generalizability. Next, future work can focus TTX, all analytics were computed in under 2 minutes on a on collecting audio or video data from TTXs to support standard laptop, addressing the assessment challenges outlined assessment. Lastly, future work can develop cybersecurityin Section I-A. Subsequently, the instructor can provide tailored LLMs [57] to mitigate the limitations of general LLMs feedback to the whole cluster more quickly, saving time when and focus on custom TTX assessment. focusing on the personalized specifics of each team. B. Practical Tools and Resources for the Community 3) Be Cautious About LLM-based Assessment in TTXs: Although professional communication is an important skill in To support the adoption of TTXs, we release numerous TTXs, a general-purpose LLM sometimes failed to recognize materials under a free, public, open-source license:
• The PHI exercise definition and others [60], which can
be deployed in the free INJECT Exercise Platform [53]. • The research dataset from both TTXs [61], usable to research student behavior and task strategies. • Comprehensive results and visualizations (including those omitted from this paper due to space limitations) [62]. • The analytical toolset (Python code) to process the dataset and derive the results in this paper [62]. The IXP team has since incorporated these assessment methods into the latest version of the platform; see https://inject.muni.cz/. ACKNOWLEDGMENT This research was supported by the Open Calls for Security Research 2023–2029 (OPSEC) program granted by the Ministry of the Interior of the Czech Republic under No. VK01030007 – Intelligent Tools for Planning, Conducting, and Evaluating Tabletop Exercises. We also thank Tomáš Hájek for his help with establishing the ground truth for the LLM assessment. R EFERENCES [1] A. Frégeau, A. Cournoyer, M.-A. Maheu-Cadotte, M. Iseppon, N. Soucy, J. S.-C. Bourque, S. Cossette, V. Castonguay, and R. Fleet, “Use of tabletop exercises for healthcare education: a scoping review protocol,” BMJ open, vol. 10, no. 1, 2020. [Online]. Available: https://doi.org/10.1136/bmjopen-2019-032662 [2] A. M. Wendelboe, J. Amanda Miller, D. Drevets, L. Salinas, E. Miller, D. Jackson, A. Chou et al., “Tabletop exercise to prepare institutions of higher education for an outbreak of covid-19,” Journal of emergency management, vol. 18, no. 2, 2020. [Online]. Available: https://doi.org/10.5055/jem.2020.0463 [3] C. Husna, H. Kamil, M. Yahya, T. Tahlil, and D. Darmawati, “Does tabletop exercise enhance knowledge and attitude in preparing disaster drills?” Nurse Media Journal of Nursing, vol. 2, no. 10, pp. 182–190, 2020. [Online]. Available: https://doi.org/10.14710/nmjn.v10i2.29117 [4] B. E. Sandström, H. Eriksson, L. Norlander, M. Thorstensson, and G. Cassel, “Training of public health personnel in handling cbrn emergencies: A table-top exercise card concept,” Environment International, vol. 72, pp. 164–169, 2014. [Online]. Available: https://doi.org/10.1016/j.envint.2014.03.009 [5] O. Chernikova, N. Heitzmann, M. Stadler, D. Holzberger, T. Seidel, and F. Fischer, “Simulation-based learning in higher education: A metaanalysis,” Review of educational research, vol. 90, no. 4, pp. 499–541, 2020. [Online]. Available: https://doi.org/10.3102/0034654320933544 [6] The Joint Task Force on Computer Science Curricula, Computing Curricula 2023. New York, NY, USA: ACM, 2024. [Online]. Available: https://doi.org/10.1145/3664191 [7] Joint Task Force on Cybersecurity Education, “Cybersecurity curricular guideline,” 2017. [Online]. Available: http://cybered.acm.org [8] R. Raj, M. Sabin, J. Impagliazzo, D. Bowers, M. Daniels, F. Hermans, N. Kiesler, A. N. Kumar, B. MacKellar, R. McCauley, S. W. Nabi, and M. Oudshoorn, “Professional competencies in computing education: Pedagogies and assessment,” in Working Group Reports on Innovation and Technology in Computer Science Education. ACM, 2022, p. 133–161. [Online]. Available: https://doi.org/10.1145/3502870.3506570 [9] J. Vykopal, P. Čeleda, V. Švábenský, M. Hofbauer, and M. Horák, “Research and Practice of Delivering Tabletop Exercises,” in 29th Conference on Innovation and Technology in Computer Science Education, ser. ITiCSE ’24. New York, NY, USA: ACM, 2024, pp. 220–226. [Online]. Available: https://doi.org/10.1145/3649217.3653642 [10] G. Angafor, I. Yevseyeva, and L. Maglaras, “Malaware: A tabletop exercise for malware security awareness education and incident response training,” Internet of Things and Cyber-Physical Systems, vol. 4, 2024. [Online]. Available: https://doi.org/10.1016/j.iotcps.2024.02.003 [11] S. Hays and J. White, “Using LLMs for Tabletop Exercises within the Security Domain,” 2024. [Online]. Available: https: //arxiv.org/abs/2403.01626
[12] L. Müller, “Tabletop exercise for ransomware negotiations,” in Augmented Cognition, vol. 14695. Springer, 2024, pp. 166–184. [Online]. Available: https://doi.org/10.1007/978-3-031-61572-6_12 [13] K. Nakayama, I. Koshijima, and K. Watanabe, “Analyzing important factors in cybersecurity incidents using table-top exercise,” Human Factors in Cybersecurity, vol. 127, pp. 105–114, 2024. [Online]. Available: https://doi.org/10.54941/ahfe1004770 [14] J. Kävrestad, S. Johansson, and E. Bergström, “Using tabletop exercises to raise cybersecurity awareness of decision-makers,” in Critical Information Infrastructures Security. Springer Nature Switzerland, 2025, pp. 231– 248. [Online]. Available: https://doi.org/10.1007/978-3-031-84260-3_14 [15] M. Gafic, S. Tjoa, P. Kieseberg, O. Hellwig, and G. Quirchmayr, “Cyber exercises in computer science education,” in Proceedings of the 8th International Conference on Information Systems Security and Privacy, 2022. [Online]. Available: https://doi.org/10.5220/0010845800003120 [16] U.S. Environmental Protection Agency, “Cybersecurity,” 2022. [Online]. Available: https://ttx.epa.gov/CyberSecurity7.html [17] C. Xiao, W. Ma, Q. Song, S. X. Xu, K. Zhang, Y. Wang, and Q. Fu, “Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs,” in Proceedings of the 15th International Learning Analytics and Knowledge Conference, New York, NY, USA, 2025, p. 293–305. [Online]. Available: https://doi.org/10.1145/3706468.3706507 [18] K. Seßler, M. Fürstenberg, B. Bühler, and E. Kasneci, “Can ai grade your essays? a comparative analysis of large language models and teacher ratings in multidimensional essay scoring,” in Proceedings of the 15th International Learning Analytics and Knowledge Conference. New York, NY, USA: Association for Computing Machinery, 2025, p. 462–472. [Online]. Available: https://doi.org/10.1145/3706468.3706527 [19] R. Elvegård and N. Andreassen, “Exercise design for interagency collaboration training: The case of maritime nuclear emergency management tabletop exercises,” Journal of Contingencies and Crisis Management, vol. 32, no. 1, 2024. [Online]. Available: https://doi.org/10.1111/1468-5973.12517 [20] J. Jeuring, H. Keuning, S. Marwan, D. Bouvier, C. Izu, N. Kiesler, T. Lehtinen, D. Lohr, A. Peterson, and S. Sarsa, “Towards giving timely formative feedback and hints to novice programmers,” in Proceedings of the ITiCSE 2022 Working Group Reports. ACM, 2022. [Online]. Available: https://doi.org/10.1145/3571785.3574124 [21] J. Mirkovic, A. Aggarwal, D. Weinman, P. Lepe, J. Mache, and R. Weiss, “Using terminal histories to monitor student progress on hands-on exercises,” in ACM Technical Symposium on Computer Science Education, ser. SIGCSE. New York, NY, USA: ACM, 2020, p. 866–872. [Online]. Available: https://doi.org/10.1145/3328778.3366935 [22] C. Lang, G. Siemens, A. Wise, D. Gašević, and A. Merceron, The Handbook of Learning Analytics, 2nd ed. Canada: SoLAR, 2022. [Online]. Available: https://doi.org/10.18608/hla22 [23] C. Romero, S. Ventura, M. Pechenizkiy, and R. S. Baker, Handbook of educational data mining. USA: CRC Press, 2010. [Online]. Available: https://doi.org/10.1201/b10274 [24] J. C. Paiva, J. P. Leal, and A. Figueira, “Automated assessment in computer science education: A state-of-the-art review,” ACM Trans. Comput. Educ., vol. 22, no. 3, 2022. [Online]. Available: https://doi.org/10.1145/3513140 [25] S. Feng, L. Yan, L. Zhao, R. M. Maldonado, and D. Gašević, “Heterogenous network analytics of small group teamwork: Using multimodal data to uncover individual behavioral engagement strategies,” in Proceedings of the 14th Learning Analytics and Knowledge Conference, ser. LAK ’24. New York, NY, USA: ACM, 2024, p. 587–597. [Online]. Available: https://doi.org/10.1145/3636555.3636918 [26] M. Guzdial and B. du Boulay, “The history of computing education research,” in The Cambridge Handbook of Computing Education Research. Cambridge University Press, 2019, ch. 1, pp. 11–39. [Online]. Available: https://doi.org/10.1017/9781108654555 [27] M. Taylor, A. Barthakur, A. Azad, S. Joksimovic, X. Zhang, and G. Siemens, “Quantifying collaborative complex problem solving in classrooms using learning analytics,” in Learning Analytics and Knowledge. New York, NY, USA: ACM, 2024, p. 551–562. [Online]. Available: https://doi.org/10.1145/3636555.3636913 [28] L. Zhao, V. Echeverria, Z. Swiecki, L. Yan, R. Alfredo, X. Li, D. Gasevic, and R. Martinez-Maldonado, “Epistemic network analysis for end-users: Closing the loop in the context of multimodal analytics for collaborative team learning,” in 14th Learning Analytics and Knowledge Conference. ACM, 2024, p. 90–100. [Online]. Available: https://doi.org/10.1145/3636555.3636855
[29] J. Pande, W. Min, R. D. Spain, J. D. Saville, and J. Lester, “Robust team communication analytics with transformer-based dialogue modeling,” in Artificial Intelligence in Education, vol. 13916. Springer, 2023, pp. 639– 650. [Online]. Available: https://doi.org/10.1007/978-3-031-36272-9_52 [30] T. OConnor, A. Schmith, C. Stricklan, M. Carvalho, and S. Sudhakaran, “Pwn lessons made easy with docker: Toward an undergraduate vulnerability research cybersecurity class,” in Tech. Symposium on Comp. Sci. Educ., ser. SIGCSE. New York, NY, USA: ACM, 2024, p. 986–992. [Online]. Available: https://doi.org/10.1145/3626252.3630911 [31] S. Narain, P. Rayavaram, C. Morales-Gonzalez, M. Harper, M. Abbasalizadeh, K. Vellamchety, and X. Fu, “Practical cybersecurity education: A course model using experiential learning theory,” in 56th ACM Tech. Symposium on Comp. Sci. Educ. ACM, 2025, p. 819–825. [Online]. Available: https://doi.org/10.1145/3641554.3701922 [32] M. Won, L. R. Carrington, D. M. Espinoza, M. H. Ali, and D. Dasgupta, “A cybersecurity summer camp for high school students using autonomous r/c cars,” in Tech. Symposium on Comp. Sci. Educ., ser. SIGCSE. New York, NY, USA: ACM, 2024, p. 1435–1441. [Online]. Available: https://doi.org/10.1145/3626252.3630758 [33] G. Costa, S. De Francisci, M. Renieri, and S. Valiani, “Tackling the gender gap in cybersecurity education,” in Proceedings of the 56th ACM Technical Symposium on Computer Science Education, ser. SIGCSE. New York, NY, USA: ACM, 2025, p. 234–240. [Online]. Available: https://doi.org/10.1145/3641554.3701807 [34] C. Gough, C. Mann, C. Ficke, M. Namukasa, M. Carroll, and T. OConnor, “Remote controlled cyber: Toward engaging and educating a diverse cybersecurity workforce,” in Proceedings of the 55th ACM Technical Symposium on Computer Science Education. ACM, 2024, p. 394–400. [Online]. Available: https://doi.org/10.1145/3626252.3630917 [35] A. Dutt, M. A. Ismail, and T. Herawan, “A systematic review on educational data mining,” IEEE Access, vol. 5, pp. 15 991–16 005, 2017. [Online]. Available: https://doi.org/10.1109/access.2017.2654247 [36] V. Švábenský et al., “Student Assessment in Cybersecurity Training Automated by Pattern Mining and Clustering,” Educ. and Inf. Tech., 2022. [Online]. Available: https://doi.org/10.1007/s10639-022-10954-4 [37] D. Perera, J. Kay, I. Koprinska, K. Yacef, and O. R. Zaïane, “Clustering and sequential pattern mining of online collaborative learning data,” IEEE Transactions on knowledge and Data Engineering, vol. 21, no. 6, pp. 759– 772, 2008. [Online]. Available: https://doi.org/10.1109/TKDE.2008.138 [38] F. C. Serçe, K. Swigger, F. N. Alpaslan, R. Brazile, G. Dafoulas, and V. Lopez, “Online collaboration: Collaborative behavior patterns and factors affecting globally distributed team performance,” Computers in human behavior, vol. 27, no. 1, pp. 490–503, 2011. [Online]. Available: https://doi.org/10.1016/j.chb.2010.09.017 [39] A. Jaiswal, T. Karabiyik, P. Thomas, and A. J. Magana, “Characterizing team orientations and academic performance in cooperative project-based learning environments,” Education Sciences, vol. 11, no. 9, p. 520, 2021. [Online]. Available: https://doi.org/10.3390/educsci11090520 [40] E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser et al., “ChatGPT for good? On opportunities and challenges of large language models for education,” Learning and individual differences, vol. 103, 2023. [Online]. Available: https://doi.org/10.1016/j.lindif.2023.102274 [41] M. Pankiewicz and R. S. Baker, “Enhancing student focus and problem-solving with real-time llm feedback on compiler errors,” in Two Decades of TEL. From Lessons Learnt to Challenges Ahead. Cham: Springer Nature Switzerland, 2025, pp. 412–426. [Online]. Available: https://doi.org/10.1007/978-3-032-03870-8_28 [42] Q. Jia, J. Cui, R. Xi, C. Liu, P. Rashid, R. Li, and E. Gehringer, “On Assessing the Faithfulness of LLM-generated Feedback on Student Assignments,” in Proceedings of the 17th International Conference on Educational Data Mining. MA, USA: IEDMS, 2024, pp. 491–499. [Online]. Available: https://doi.org/10.5281/zenodo.12729868 [43] R. Garg, J. Han, Y. Cheng, Z. Fang, and Z. Swiecki, “Automated discourse analysis via generative artificial intelligence,” in Proceedings of the 14th Learning Analytics and Knowledge Conference. ACM, 2024, pp. 814–820. [Online]. Available: https://doi.org/10.1145/3636555.3636879 [44] S. Kakarla, C. Borchers, D. Thomas, S. Bhushan, and K. R. Koedinger, “Comparing few-shot prompting of gpt-4 llms with bert classifiers for open-response assessment in tutor equity training,” 2025. [Online]. Available: https://arxiv.org/abs/2501.06658 [45] C. Snyder, N. M. Hutchins, C. Cohn, J. H. Fonteles, and G. Biswas, “Analyzing Students Collaborative Problem-Solving Behaviors in Synergistic STEM+C Learning,” in 14th Learning Analytics and
Knowledge Conference. New York, NY, USA: ACM, 2024, p. 540–550. [Online]. Available: https://doi.org/10.1145/3636555.3636912 [46] S. Rüdian, J. Podelo, J. Kužílek, and N. Pinkwart, “Feedback on feedback: Student’s perceptions for feedback from teachers and few-shot llms,” in Proc. of the 15th Intl. Learning Analytics and Knowledge Conf. ACM, 2025. [Online]. Available: https://doi.org/10.1145/3706468.3706479 [47] R. Ferreira Mello, C. Pereira Junior, L. Rodrigues, F. D. Pereira, L. Cabral, N. Costa, G. Ramalho, and D. Gasevic, “Automatic Short Answer Grading in the LLM Era: Does GPT-4 with Prompt Engineering beat Traditional Models?” in Proceedings of the 15th International Learning Analytics and Knowledge Conference. ACM, 2025, p. 93–103. [Online]. Available: https://doi.org/10.1145/3706468.3706481 [48] T. Li, Y. Fan, N. Srivastava, Z. Zeng, X. Li, H. Khosravi, Y.-S. Tsai, Z. Swiecki, and D. Gašević, “Analytics of planning behaviours in selfregulated learning: Links with strategy use and prior knowledge,” in Proc. of the 14th Learning Analytics and Knowledge Conf. ACM, 2024, p. 438–449. [Online]. Available: https://doi.org/10.1145/3636555.3636900 [49] L. Zhao, M. Raković, E. B. Cloude, X. Li, D. Gašević, and L. Bardach, “The effect of sequential transition of self-regulated learning processes on performance: Insights from ordered network analysis,” in Proc. of the 15th International Learning Analytics and Knowledge Conf. ACM, 2025, p. 516–526. [Online]. Available: https://doi.org/10.1145/3706468.3706534 [50] Y. Cheng, R. Guan, T. Li, M. Raković, X. Li, Y. Fan, F. Jin, Y.-S. Tsai, D. Gašević, and Z. Swiecki, “Self-regulated learning processes in secondary education: A network analysis of trace-based measures,” in Proceedings of the 15th International Learning Analytics and Knowledge Conference. ACM, 2025, p. 260–271. [Online]. Available: https://doi.org/10.1145/3706468.3706502 [51] J. Saint, Y. Fan, and D. Gasevic, “Analytics of scaffold compliance for self-regulated learning,” in Proceedings of the 14th Learning Analytics and Knowledge Conference. New York, NY, USA: Association for Computing Machinery, 2024, p. 326–337. [Online]. Available: https://doi.org/10.1145/3636555.3636887 [52] National Initiative for Cybersecurity Careers and Studies (NICCS), “Incident response,” 2020. [Online]. Available: https://niccs.cisa.gov/ tools/nice-framework/work-role/incident-response [53] V. Švábenský, J. Vykopal, M. Horák, M. Hofbauer, and P. Čeleda, “From Paper to Platform: Evolution of a Novel Learning Environment for Tabletop Exercises,” in Innovation and Technology in Computer Science Education. New York, NY, USA: ACM, 2024, pp. 213–219. [Online]. Available: https://doi.org/10.1145/3649217.3653639 [54] M. Bradford, I. Khebour, N. Blanchard, and N. Krishnaswamy, “Automatic detection of collaborative states in small groups using multimodal features,” in Artificial Intelligence in Education. Springer, 2023, pp. 767–773. [Online]. Available: https://doi.org/10.1007/ 978-3-031-36272-9_69 [55] U.S. Office of Personnel Management, “Competency model,” 2011. [Online]. Available: https://www.opm.gov/chcoc/transmittals/ 2011/competency-model-cybersecurity_02-16-2011_508.pdf [56] N. Nguyen and R. Caruana, “Consensus clusterings,” in 7th IEEE International Conference on Data Mining. IEEE, 2007, pp. 607–612. [Online]. Available: https://doi.org/10.1109/ICDM.2007.73 [57] K. Mai, R. Beuran, and N. Inoue, “CyLLM-DAP: Cybersecurity Domain-Adaptive Pre-Training Framework of Large Language Models,” in 11th Int. Conf. on Inf. Systems Security and Privacy. SCITEPRESS, 2025, pp. 24–35. [Online]. Available: https://www.scitepress.org/Papers/ 2025/130948/130948.pdf [58] C. Nelson and Y. Shoshitaishvili, “Dojo: Applied cybersecurity education in the browser,” in Technical Symposium on Computer Science Education, ser. SIGCSE. New York, NY, USA: ACM, 2024, p. 930–936. [Online]. Available: https://doi.org/10.1145/3626252.3630836 [59] A. Zamecnik, V. Kovanovíc, S. Joksimovíc, G. Grossmann, D. Ladjal, R. Marshall, and A. Pardo, “Using online learner trace data to understand the cohesion of teams in higher education,” Journal of Computer Assisted Learning, vol. 39, no. 6, pp. 1733–1750, 2023. [Online]. Available: https://onlinelibrary.wiley.com/doi/abs/10.1111/jcal.12829 [60] INJECT Team, “Available Exercise Definitions,” 2026. [Online]. Available: https://docs.inject.muni.cz/INJECT_process/available-definitions [61] Vykopal, Jan and Švábenský, Valdemar and Čeleda, Pavel, “Dataset From Cybersecurity Tabletop Exercises in the INJECT Platform,” 2026, v. 1.1.0. [Online]. Available: https://doi.org/10.5281/zenodo.21396276 [62] Paper authors, “Results, visualizations, and analytical toolset,” 2026. [Online]. Available: https://gitlab.fi.muni.cz/inject/papers/2026-fie-assessment