RegKT: Interpretable and Robust Deep Knowledge Tracing With IRT-Regularizer Samuel Girard1,* , Juan D. Pinto2 , Jill-Jênn Vie1 and Amel Bouzeghoub3 1
Inria-Saclay University of Illinois Urbana-Champaign, Urbana, IL, USA 3 Telecom-Sud Paris 2
Abstract As deep learning models continue to advance, knowledge tracing models have achieved higher accuracy. However, these gains come at the cost of reduced interpretability, which is crucial for practitioners in educational settings to adopt new methodologies. Additionally, deep learning models are prone to overfitting, particularly when dealing with the small datasets that are common in educational applications. In this paper, we propose a novel regularization technique designed to enhance the robustness of deep-learning-based knowledge tracing models, while simultaneously improving their interpretability. Our method addresses both the interpretability and overfitting challenges, making it more feasible for real-world educational applications.
Keywords Knowledge tracing, Item Response Theory, Interpretability, Deep learning
1. Introduction
knowledge tracing models and regularization techniques for educational applications. We then describe our methods and present the results of our approach. We also touch on the implications that our work has on interpretability. Finally, we discuss future research directions.
Deep learning (DL) has significantly advanced the field of knowledge tracing, enabling more accurate predictions of student performance over time. Knowledge tracing models have evolved from early methods, such as Bayesian knowledge tracing (BKT), to more sophisticated approaches like deep knowledge tracing (DKT), which leverage the power of recurrent neural networks. These advancements have led to substantial improvements in predictive accuracy, making knowledge tracing models indispensable tools for adaptive learning platforms and personalized education. Yet, despite these improvements, DL-based knowledge tracing models suffer from a critical limitation: they often lack interpretability [1, 2, 3]. That is, while they can accurately predict students’ success on the next problem, they lack the interpretable parameters of BKT or item response theory (IRT). In educational settings, where teachers and instructional designers rely on clear, understandable results, the opaque nature of deep models limits their adoption [4, 5]. This also makes such models unusable for tasks such as open learner modeling [6]. Additionally, DL models are susceptible to overfitting, especially in contexts where educational datasets are small or sparse [7]. To address these issues, we introduce a novel hybrid model, RegKT, which integrates some of the interpretability of IRT with the temporal modeling capabilities of DKT [8, 9]. By incorporating a regularization term based on IRT into the DKT framework, we can control the trade-off between accuracy and interpretability using a hyper-parameter 𝜖. This hybrid approach allows our model to achieve high accuracy without sacrificing explainability, making it a better fit for real-world educational applications. The remainder of this paper is structured as follows: we first review related work on interpretability in DL-based
2. Related Work The task of estimating students’ latent knowledge states is one that has been studied at depth using a variety of approaches. In this section we review some past approaches that are relevant to our integrated RegKT model.
2.1. Item Response Theory IRT [10] is one approach based on logistic regression and used primarily in the field of assessment and psychometrics. IRT models the relationship between a student’s latent ability and their performance on test items, making it highly interpretable and less prone to overfitting compared to more complex models. The basic IRT model, known as the one-parameter logistic model or Rasch model, is given by: 𝑃 (𝑋𝑖𝑗 = 1|𝜃𝑖 , 𝑏𝑗 ) =
1 1 + exp −(𝜃𝑖 − 𝑏𝑗 )
Where 𝑃 (𝑋𝑖𝑗 = 1) represents the probability that student 𝑖 correctly answers item 𝑗, 𝜃𝑖 denotes the latent ability of student 𝑖, and 𝑏𝑗 is the difficulty parameter of item 𝑗. IRT produces simple, interpretable curves that illustrate the probability of a correct response based on student ability. This framework’s simplicity helps prevent overfitting, allowing it to generalize well, even with smaller datasets. However, its limitation is that it assumes a student’s ability is static, whereas knowledge evolves over time.
HEXED’25: 2nd Human-Centric eXplainable AI in Education Workshop, 20 July, 2025, Palermo, Italy Corresponding author. $ [email protected] (S. Girard); [email protected] (J. D. Pinto); [email protected] (J. Vie); [email protected] (A. Bouzeghoub) https://jdpinto.com (J. D. Pinto); https://jill-jenn.net (J. Vie); https://amel.wp.imtbs-tsp.eu (A. Bouzeghoub) 0009-0009-8205-7484 (S. Girard); 0000-0002-2972-485X (J. D. Pinto); 0000-0002-9304-2220 (J. Vie); 0000-0003-4890-9005 (A. Bouzeghoub)
2.2. Deep Knowledge Tracing
*
Knowledge tracing, as another family of methods, emphasizes knowledge growth over time—something IRT does not account for. The most well-known models for knowledge tracing include BKT [11], a hidden Markov model approach, and DKT [12], built on recurrent neural networks and their long short-term memory (LSTM) variants.
1
Samuel Girard et al. CEUR Workshop Proceedings
1–7
DKT and other DL-based knowledge tracing algorithms have become popular in the literature due to their improvements in accuracy and their lack of reliance on predefined assumptions about knowledge structure. In DKT, for each student interaction, the model updates its hidden state ℎ𝑡 , which encodes the student’s evolving knowledge. The model’s forward pass can be described as: ℎ𝑡 = tanh(𝑊ℎℎ𝑥 𝑥𝑡 + 𝑊ℎℎℎ ℎ𝑡−1 + 𝑏ℎ )
(1)
𝑦𝑡 = 𝜎(𝑊𝑦ℎ ℎ𝑡 + 𝑏𝑦 )
(2)
capabilities while incorporating an IRT-based regularization term to enhance interpretability.
3.1. Model Traditional recurrent neural networks are designed to map a sequence of input vectors 𝑥1 , . . . , 𝑥𝑇 to a corresponding sequence of output vectors 𝑦1 , . . . , 𝑦𝑇 . This is accomplished by maintaining a sequence of hidden states ℎ1 , . . . , ℎ𝑇 , where each hidden state ℎ𝑡 encodes relevant information from previous time steps t. Essentially, the hidden state ℎ𝑡 represents the student’s evolving knowledge over time as they interact with items. In RegKT, we do not only have one sequence of output vectors but also a sequence of scalar outputs 𝑣1 , . . . , 𝑣𝑇 that correspond to the prediction of the latent ability parameter of the given student. The variables are related using a simple network defined by the equations:
DKT’s flexibility comes from its ability to learn complex patterns directly from student interaction data. However, this flexibility comes at the cost of interpretability due to the interplay between flexibility, model complexity, and interpretability. Furthermore, DKT models are prone to overfitting, particularly when training on small datasets.
2.3. Integrating Approaches To our knowledge, [13] is the only prior work that attempts to integrate the advantages of IRT and DKT into a single model. Rather than the original DKT, they rely on the dynamic key-value memory network (DKVMN) variant [14], which creates a memory matrix that maps student knowledge states to implicitly derived knowledge components (though the accuracy of this implicit mapping and the interpretability of such an approach has been previously questioned—see [2]). They use this matrix’s vector representations to infer student ability and difficulty level, which they then plug into the IRT model as 𝜃 and 𝑏 to predict student success on the next item. While their model obtains promising results, their reliance on DKVMN increases the data requirements of such an approach. Our approach differs from this in that we aim to constrain DKT via an added penalty term to its loss function. A similar constraints-based approach was followed by [15], in which a convolutional neural network trained to predict gaming-the-system behavior was constrained in order for the network to learn binary rather than continuous weights in its convolutional layer. As with our approach, the researchers achieved this using an added regularization term to their model’s loss function. Their rationale for binarizing weights was that it would allow the model to better fit the nature of the features with the goal of improving interpretability. Similarly, [16] proposed introducing a penalty term to the loss function of a DKT model to ensure that the knowledge components learned would follow sensible learning curves. However, this was a theoretical proposal, so the details and results of the implementation are not available. While there are similarities between our approach and these prior works, our method is unique in that it seeks to integrate specific interpretable parameters from a different model with the temporal modeling capabilities of DKT. This allows us to maintain some of the interpretability of IRT while benefiting from the flexibility of DKT.
ℎ𝑡 = tanh(𝑊ℎℎ𝑥 𝑥𝑡 + 𝑊ℎℎℎ ℎ𝑡−1 + 𝑏ℎ )
(3)
𝑦𝑡 = 𝜎(𝑊𝑦ℎ ℎ𝑡 + 𝑏𝑦 )
(4)
𝑣𝑡 = 𝑊𝑣ℎ ℎ𝑡 + 𝑏𝑣
(5)
Equations (3) and (4) are exactly the same as in DKT and equation (5) is a simple perceptron linking the hidden state to the latent ability estimation. The architecture of the model can be visualized in figure 1. As with DKT, the input 𝑥𝑡 is the one-hot encoding of the student interaction tuple 𝑞𝑡 , 𝑎𝑡 that represents the combination of which exercise was answered and if the exercise was answered correctly, so 𝑥𝑡 ∈ {0, 1}2𝑀 with M the number of exercise in the corpus.
3.2. Training Objective The training is the most important part of our model as it is here that it differs from DKT. We first start by splitting the dataset into a train set 𝒟𝑡𝑟𝑎𝑖𝑛 and test set 𝒟𝑡𝑒𝑠𝑡 by students i.e no student can appear in both the train set and test set. From this we fit a simple IRT-1PL model with only 𝒟𝑡𝑟𝑎𝑖𝑛 to estimate the ability parameter 𝜃𝑖 of every student 𝑖 in the train set. The values 𝜃 will serve as a target for the predictions 𝑣1 , . . . , 𝑣𝑇 during the training procedure. The training objective is the negative log likelihood of the observed sequence of student responses under the model. Let 𝛿(𝑞𝑡+1 ) be the one-hot encoding of which exercise is answered at time 𝑡 + 1, and let 𝑙 be binary cross entropy. The loss for a given prediction is 𝑙(𝑦 𝑇 𝛿(𝑞𝑡+1 ), 𝑎𝑡+1 ), and the loss for a single student 𝑖 is: 𝐿𝑖 =
𝑇𝑖 ∑︁ 𝑡
⏟
𝑙(𝑦 𝑇 𝛿(𝑞𝑡+1 ), 𝑎𝑡+1 ) + 𝜖 · 𝐿irt (𝑓 ((𝑣𝑗 )𝑗≤𝑇𝑖 ) , 𝜃𝑖 ) ⏟ ⏞ IRT-based regularization term ⏞ DKT loss objective
Here 𝑓 (.) is a simple linear form that maps the vector (𝑣𝑗 )𝑗≤𝑇𝑖 into a scalar (we chose to use 𝑓 ((𝑣𝑗 )𝑗≤𝑇𝑖 ) = 𝑣𝑇𝑖 as our last prediction should have converged better to fixed parameter 𝜃𝑖 ) and 𝐿irt is a classical error signal such as the MSE or the L1 distance. The key innovation in RegKT is the introduction of a regularization term based on IRT. This term ensures that the model’s predictions align with the IRT-based estimates of student ability 𝜃𝑖 .
3. Methodology To address the limitations of both IRT and DKT, we propose a hybrid model, RegKT, which integrates the key properties of both approaches. RegKT uses DKT’s temporal modeling
2
Samuel Girard et al. CEUR Workshop Proceedings
1–7
Figure 1: The connection between variables in RegKT. The input (𝑥𝑡 ) is either a one-hot encoding or compressed representation of student action, and the prediction (𝑦𝑡 ) is a probability vector for answering each problem correctly, with (𝑣𝑡 ) as the student’s proficiency parameter estimate.
Dataset
We hypothesize that this regulator allows for more interpretable weights during gradient descent: the gradient loss from 𝐿irt is backpropagated into the weights of the RNN that are used in the calculation of the hidden states ℎ1 , . . . , ℎ𝑇 and the prediction sequences 𝑦1 , . . . , 𝑦𝑇 . This allows the hidden states to be more explainable and to carry less noise in the data—since they are inputs of a simple 1-layer perceptron they cannot vary too much as 𝑣𝑇𝑖 should converge toward 𝜃𝑖 . Additionally, the scaling weight 𝜖 allows for interpolation between weights in the RNN that are optimized to reduce 𝐿irt and weights that are optimized to capture more complex patterns in the data. This is useful in a setting where the data is scarce, as it implies that we should sometime put more trust in the loss signal of an underfitted model rather than a complex one that could capture unwanted noise.
Fraction
RoboMission
ASSISTments2009
Synthetic BKT
Synthetic M-IRT
Model DKT RegKT IRT DKT RegKT IRT DKT RegKT IRT DKT RegKT IRT DKT RegKT IRT
AUC 0.879 0.889 0.656 0.835 0.838 0.786 0.741 0.772 0.666 0.708 0.727 0.681 0.712 0.732 0.612
Accuracy 0.805 0.812 0.615 0.898 0.9015 0.890 0.696 0.723 0.634 0.731 0.743 0.719 0.664 0.677 0.584
Table 1 Results in AUC and accuracy (Acc.) for all datasets.
4. Results We have also looked at the last hidden state ℎ𝑇 ∈ R2 for every student, which is essential for practitioners to interpret, as it encodes the student’s knowledge derived from the entirety of their learning trace. Before running the experiment, we calculated the student’s proficiency parameter using a simple 1-PL IRT model with the full dataset. This allowed us to better understand the visualizations, as we noticed that the clusters of the hidden states aligns with the results from IRT (see figure 3). We notice that our method allows for much better clusters that makes more use of the 2-D plane (Figure 3), hence giving more room for interpretation between different types of knowledge states.
In this section, we present two sets of results that demonstrate the effectiveness of our proposed method. The first set of results centers on the model’s explainability. We provide visual interpretations of the learned weights. These visualizations can be valuable for practitioners, as they make the model’s decision-making process more transparent and accessible, potentially guiding instructional decisions. The second set focuses on the model’s ability to mitigate overfitting, particularly in small dataset scenarios. By comparing performance across different dataset sizes, we show that our approach can achieve better or similar accuracy than existing methods when data is limited, highlighting its robustness. To provide more technical details, the results in the table 1 are from a grid search to find the best hyperparameters.
4.2. RoboMission The RoboMission dataset consists of over 20,000 students, each answering questions in a corpus of 85 exercises. In this larger-scale setting, we maintained model interpretability by using classical DKT and RegKT models with a hidden state dimension of 2 and 5 layers. Similarly, we analyzed the final hidden state ℎ𝑇𝑖 ∈ R2 for student 𝑖. Figure 4 shows that our method leads to more distinct clusters in the 2-D projection of the hidden states, but there are no better separations where it comes to the visualizations of the weights linking to item correctness 𝑊𝑦ℎ . Concurrently we do not observe considerable improvement in accuracy.
4.1. Fractions Fractions is a very simple dataset consisting of 535 users answering 20 exercises on fractions. In our setup, we want a model with low complexity to allow visualization hence our choice to have,for both classical DKT and RegKT, the hidden state of dimension 2. We also chose the number of layers in our LSTM to be 5 (the number of LSTM units stacked together) as it gave better accuracy. In figure 2 we observe that our model allows for better clusters of items.
3
Samuel Girard et al. CEUR Workshop Proceedings
1–7
In classical DKT In classical DKT
In our model (RegKT) Figure 4: Visualization of the hidden-state clusters in classical DKT (top) vs. RegKT (bottom) in RoboMission. In our model (RegKT) Figure 2: Visualization of the weights linking the hidden state to item correctness, with each label associated with the ordering of difficulties.
Figure 5: Alignment of the hidden state with its proficiency parameter in ASSISTments: ℎ𝑇𝑖 associated to higher proficiency parameter 𝑣𝑇𝑖 are on the top right of the plot In classical DKT
4.4. Synthetic Datasets To better understand the benefits of RegKT, we tested our model on synthetic data where we could control the underlying dimensionality. In our synthetic M-IRT experiments, we generated learning traces for 200 students using a corpus containing a variable number 𝐾 of exercises and a model with dimensionality 𝐷. In this setting, each student is characterized by a 𝐷-dimensional ability vector, while each item is defined by a 𝐷-dimensional discrimination vector and an associated difficulty parameter. Each student responds to a random subset of available items. RegKT outperforms DKT when the data is generated from a simpler underlying model, characterized by lower values of 𝐾 and 𝐷. This suggests that high-capacity deep learning models, like DKT, tend to capture not only meaningful patterns but also more noise when no regularization is applied. In contrast, RegKT appears to mitigate this issue, leading to better generalization in lower-complexity settings, as illustrated in figure 6. Experiments have also been made with a synthetic dataset with learning traces generated from the BKT-student model.
In our model (RegKT) Figure 3: Visualization of the hidden state clusters in classical DKT (top) vs. RegKT (bottom) in fraction. We select subsets of top, middle, and low-performing students (based on their IRT ability)
4.3. ASSISTments We used the ASSISTments 2009-2010 dataset, where we have only looked at the 500 most played items with about 2675 distinct users. In this data set, we have had better accuracy with ℎ𝑡 ∈ R5 as it allows us to capture more complex patterns in the data. In order to allow visualizations, we decide to run a PCA with two components on the hidden states.
5. Enhancing interpretability The diagrams above demonstrate how our approach improves the clustering of hidden states, making it easier for
4
Samuel Girard et al. CEUR Workshop Proceedings
1–7
approaches that make indirect predictions of the model’s internal workings [19]. Following the XAI framework described by [20], an explanation is an inference made from interpreting specific evidence. In our case, we can identify 𝑣𝑡𝑖 as the evidence, with its alignment with ℎ𝑡𝑖 providing its interpretation as corresponding to the model’s conception of student proficiency. Depending on the specifics of the model’s architecture, its inputs and outputs, and the parameters used, this framework would thus allow us to calculate 𝑣𝑡𝑖 ’s explanatory potential. Furthermore, following the explanation evaluation framework described by [21], one of the desirable qualities of useful explanations that is enhanced when using the model’s internal parameters is faithfulness, referring to how well an explanation aligns with the model’s internal state. Another important quality is plausibility, which in our use case can be improved when students’ learning trajectories align with human intuition.
Figure 6: Illustration of the AUC curve, where students interact with exercises generated under the M-IRT model. Here 𝜖 = 0 implies the DKT model
6. Discussion
practitioners to understand how the model makes decisions. Clearer clusters help identify states that correspond to similar knowledge levels or learning patterns, allowing for a more interpretable analysis of a student’s evolving “knowledge state” over time. This provides insights into how the model perceives mastery progression after each question. However, our approach’s most important contribution to enhancing interpretability lies in its association of each hidden state ℎ𝑡𝑖 with a proficiency scalar parameter 𝑣𝑡𝑖 . This makes it possible to use students’ proficiency estimates in ways similar to more traditional KT algorithms. For example, we can calculate learning curves [17], which allow us to determine whether the model perceives the student as improving. Such curves can be used to test the whether the model’s understanding of a student is sound or to identify potential issues in the model. We can observe in figure 7 that both outputs of our model align properly. Notably, the IRT estimation in RegKT converges well toward the ground truth1 proficiency parameter (red dotted line), even for students in the test set, reinforcing the model’s reliability in assessing student learning.
The preliminary results of our approach, as presented here, demonstrate its usefulness for interpretable knowledge tracing, particularly when using smaller datasets on which other models may overfit. One promising area for future research is the exploration of alternative regularization techniques to further enhance both the performance and interpretability of the model. While our current approach uses IRT-based regularization, other models such as performance factors analysis (PFA) [22] or BKT [11] could also serve as effective regularizers. Additionally, investigating how to modify the architecture of more advanced knowledge tracing models that move away from recurrent neural networks and instead leverage mechanisms like neural attention [23] or temporal point processes [24] offers an exciting path forward. To broaden the applicability of the RegKT model, future work should focus on its scalability and computational efficiency when applied to larger and more complex datasets. While the current experiments were conducted on deliberately small datasets, real-world educational systems often involve millions of students and diverse learning tasks. A key question is whether RegKT can maintain its interpretability while achieving even greater accuracy in such large-scale settings. Another interesting direction is to analyze the impact of the regularizer on the model’s weight distribution in more detail, potentially uncovering deeper insights into how different regularization techniques influence model behavior. Lastly, incorporating more complex student behaviors, such as emotional engagement and collaboration, could provide a richer understanding of the learning process [25]. A potential approach is to model these behavioral data using a simple, interpretable framework and integrate it as a regularizer to capture more intricate patterns in the knowledge tracing task.
Figure 7: Evolution of the proficiency parameter for a student in the test set. The red dotted line represents a ground-truth proficiency parameter for the student
Students’ proficiency parameter can also be used to create dashboard visualizations for teachers and students, providing an explanation for their learning trajectory [18]. Deriving such explanations directly from RegKT’s parameters provides certain advantages over post-hoc explainability 1
7. Conclusion In this work, we introduced the RegKT model, which combines DKT and IRT to strike a balance between predictive accuracy and interpretability. While RegKT may not be
There is no actual parameter theta. This is simply an estimation calculated on the full dataset.
5
Samuel Girard et al. CEUR Workshop Proceedings
1–7
the most accurate model available, it bridges the gap between two of the most widely used approaches for latent knowledge estimation: DKT and IRT. By integrating IRT as a regularization mechanism, we mitigate the overfitting challenges typically encountered in deep learning models. At the same time, the inclusion of IRT encourages the model’s weights to align more closely with interpretable signals derived from simpler statistical models. More broadly, this approach demonstrates that using an underfitted model as a regularization tool can effectively combat overfitting, especially in data-scarce environments. This opens up interesting possibilities for improving model robustness in educational settings with limited data.
Modeling the acquisition of procedural knowledge, User Modelling and User-Adapted Interaction 4 (1995) 253–278. doi:10/b4wwjx. [12] C. Piech, J. Spencer, J. Huang, S. Ganguli, M. Sahami, L. Guibas, J. Sohl-Dickstein, Deep knowledge tracing, 2015. URL: https://arxiv.org/abs/1506.05908. arXiv:1506.05908. [13] C.-K. Yeung, Deep-IRT: Make deep learning based knowledge tracing explainable using item response theory, in: Proceedings of The 12th International Conference on Educational Data Mining (EDM 2019), 2019. [14] J. Zhang, X. Shi, I. King, D.-Y. Yeung, Dynamic keyvalue memory networks for knowledge tracing, in: Proceedings of the 26th International Conference on World Wide Web, WWW ’17, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 2017, pp. 765–774. doi:10.1145/3038912.3052580. [15] J. D. Pinto, L. Paquette, N. Bosch, Interpretable neural networks vs. expert-defined models for learner behavior detection, in: Companion Proceedings of the 13th International Conference on Learning Analytics & Knowledge Conference (LAK23), 2023, pp. 105–107. [16] Y. Shi, Interpretable code-informed learning analytics for CS education, in: Companion Proceedings 13th International Conference on Learning Analytics & Knowledge (LAK23), 2023, pp. 180–185. [17] B. van de Sande, Learning curves for problems with multiple knowledge components, in: Proceedings of the 9th International Conference on Educational Data Mining, 2016. [18] D. Hooshyar, M. Pedaste, K. Saks, Ä. Leijen, E. Bardone, M. Wang, Open learner models in supporting selfregulated learning in higher education: A systematic literature review, Computers & Education 154 (2020) 103878. doi:10.1016/j.compedu.2020.103878. [19] Q. Liu, J. D. Pinto, L. Paquette, Applications of explainable AI (XAI) in education, in: D. Kourkoulou, A.-O. Tzirides, B. Cope, M. Kalantzis (Eds.), Trust and Inclusion in AI-Mediated Education: Where Human Learning Meets Learning Machines, Springer Nature Switzerland, Cham, 2024, pp. 93–109. doi:10.1007/ 978-3-031-64487-05. [20] M. Rizzo, A. Veneri, A. Albarelli, C. Lucchese, M. Nobile, C. Conati, A theoretical framework for AI models explainability with application in biomedicine, in: 2023 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), IEEE, Eindhoven, Netherlands, 2023, pp. 1–9. doi:10.1109/CIBCB56990.2023.10264877. [21] J. D. Pinto, L. Paquette, Towards a unified framework for evaluating explanations, in: Joint Proceedings of the Human-Centric eXplainable AI in Education and the Leveraging Large Language Models for Next Generation Educational Technologies Workshops (HEXED-L3MNGET 2024), volume 3840, CEUR-WS, Atlanta, Georgia, USA, 2024. [22] P. I. Pavlik, H. Cen, K. R. Koedinger, Performance factors analysis –a new alternative to knowledge tracing, in: Proceedings of the 2009 Conference on Artificial Intelligence in Education: Building Learning Systems That Care: From Knowledge Representation to Affective Modelling, IOS Press, NLD, 2009, pp. 531–538. [23] A. Ghosh, N. Heffernan, A. S. Lan, Context-aware at-
References [1] X. Ding, E. C. Larson, Why deep knowledge tracing has less depth than anticipated, in: Proceedings of The 12th International Conference on Educational Data Mining (EDM 2019), 2019, pp. 282–287. [2] X. Ding, E. C. Larson, On the interpretability of deep learning based models for knowledge tracing, 2021. URL: https://arxiv.org/abs/2101.11335. arXiv:2101.11335. [3] C.-Q. Huang, Q.-H. Huang, X. Huang, H. Wang, M. Li, K.-J. Lin, Y. Chang, Xkt: Towards explainable knowledge tracing model with cognitive learning theories for questions of multiple knowledge concepts, IEEE Transactions on Knowledge and Data Engineering (2024) 1–18. doi:10.1109/TKDE.2024.3418098. [4] Y. Lu, D. Wang, Q. Meng, P. Chen, Towards interpretable deep learning models for knowledge tracing, 2020. URL: https://arxiv.org/abs/2005.06139. arXiv:2005.06139. [5] J. D. Pinto, L. Paquette, Deep learning for educational data science, 2024. arXiv:2404.19675. [6] J. Kay, K. Bartimote, K. Kitto, B. Kummerfeld, D. Liu, P. Reimann, Enhancing learning by Open Learner Model (OLM) driven data design, Computers and Education: Artificial Intelligence 3 (2022) 100069. doi:10.1016/j.caeai.2022.100069. [7] T. Gervet, K. Koedinger, J. Schneider, T. Mitchell, When is deep learning the best approach to knowledge tracing?, Journal of Educational Data Mining 12 (2020). URL: https://jedm.educationaldatamining. org/index.php/JEDM/article/view/451. doi:10.5281/ zenodo.4143614. [8] J. Sun, F. Yu, Q. Wan, Q. Li, S. Liu, X. Shen, Interpretable knowledge tracing with multiscale state representation, in: Proceedings of the ACM Web Conference 2024, WWW ’24, Association for Computing Machinery, New York, NY, USA, 2024, pp. 3265– 3276. URL: https://doi.org/10.1145/3589334.3645373. doi:10.1145/3589334.3645373. [9] H. Zhang, Z. Liu, C. Shang, D. Li, Y. Jiang, A questioncentric multi-experts contrastive learning framework for improving the accuracy and interpretability of deep sequential knowledge tracing models, ACM Trans. Knowl. Discov. Data 19 (2025). URL: https://doi.org/10. 1145/3674840. doi:10.1145/3674840. [10] G. Rasch, Probabilistic Models for Some Intelligence and Attainment Tests, MESA Press, Chicago, IL, 1993. [11] A. T. Corbett, J. R. Anderson, Knowledge tracing:
6
Samuel Girard et al. CEUR Workshop Proceedings
1–7
tentive knowledge tracing, in: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD ’20, Association for Computing Machinery, New York, NY, USA, 2020, pp. 2330–2339. URL: https://doi.org/10.1145/3394486. 3403282. doi:10.1145/3394486.3403282. [24] C. Wang, W. Ma, M. Zhang, C. Lv, F. Wan, H. Lin, T. Tang, Y. Liu, S. Ma, Temporal cross-effects in knowledge tracing, in: Proceedings of the 14th ACM International Conference on Web Search and Data Mining, WSDM ’21, Association for Computing Machinery, New York, NY, USA, 2021, pp. 517–525. URL: https://doi.org/10.1145/3437963.3441802. doi:10. 1145/3437963.3441802. [25] R. Baker, S. DMello, M. Rodrigo, A. Graesser, Better to be frustrated than bored: The incidence and persistence of affect during interactions with three different computer-based learning environments, International Journal of human-computer studies 68 (2010) 223–241.
7