RECALL: Recovery Experience Collection for Active Lifelong Learning in Vision-Language-Action Models
arXiv:2606.23617v1 [cs.RO] 22 Jun 2026
Ulas Berk Karli Department of Computer Science Yale University United States [email protected]
Tesca Fitzgerald Department of Computer Science Yale University United States [email protected]
Abstract: Vision-Language-Action (VLA) models are commonly fine-tuned through passive imitation learning, where additional demonstrations are collected for tasks where the policy performs poorly. This approach incurs several downsides: it requires the robot to fail before data collection is triggered, provides little guidance about which states require supervision, and wastes demonstrator effort on redundant parts of the task where the policy already performs well. In this paper, we propose an active, continual learning paradigm for VLAs. We demonstrate that active, uncertainty-guided data collection leads to more efficient fine-tuning than when using passively-collected demonstrations. However, we also find that fine-tuning only on actively-collected recovery data leads to catastrophic forgetting. We evaluate techniques for continual learning, including replay-based data mixing and elastic weight consolidation, and identify tradeoffs between plasticity to uncertainty-guided recovery data and retention of previously learned behaviors. Overall, our work contributes an empirical study of active continual learning for autoregressive VLAs, establishing that uncertainty-guided recovery demonstrations can improve adaptation efficiency while also revealing open challenges when targeted new data is incorporated into large robot policies. Keywords: VLA Models; Active Learning; Continual Learning
1
Introduction
Recent Vision-Language-Action (VLA) models such as RT-1 [1], RT-2 [2], Gemini Robotics [3], π0 -FAST [4], and π0.5 [5] have shown that large-scale robot learning can produce broadly capable manipulation policies. However, despite their generality, VLAs still require fine-tuning when deployed to new robots, environments, or task distributions. In practice, this adaptation is usually performed through passive imitation learning, where an expert (i) collects additional demonstrations starting from task initial states, (ii) fine-tunes the model, (iii) evaluates performance, and (iv) repeats if the policy remains unreliable. This process is inefficient for several reasons. First, it only reveals that the model needs more data after failures occurred, which can be time-consuming and unsafe. Second, it gives little guidance about how much additional data is needed, which can lead to underor over-collection of data. Third, passive recollection does not identify where supervision would be most useful. Demonstrators may therefore spend most of their efforts on states that are already well represented in the original training data; yet, failures may arise from a small set of difficult substeps or recovery from distribution-shifted intermediate states. We propose an active, continual learning paradigm for fine-tuning VLAs. Our central hypothesis is that collecting demonstrations from high-uncertainty states will lead to more informative (and thus, efficient) data collection compared to typical passive data collection. We use a state-of-the-art technique for uncertainty quantification [6] and utilize it to collect data that more directly addresses the failure modes limiting downstream task success. We then evaluate techniques for continual learning, with the aim of fine-tuning a VLA with uncertainty-guided data without catastrophic forgetting of previously-learned tasks. Our contributions are as follows:
1. An active data collection framework for autoregressive VLAs that uses INSIGHT [6] to select uncertain states to prompt new demonstrations and provide the way to best utilize that data. 2. An evaluation of active vs passive data collection, showing that demonstrations provided from high-uncertainty states lead to better model performance than unguided demonstrations. 3. A comparison of online- and offline-style data collection, showing that collecting data from the first uncertain state in a rollout is just as effective (and thus, more efficient) compared to collecting data from all uncertain states. 4. Application and analysis of continual-learning techniques (regularization- and replay-based) for VLA models, showing that replay-based data mixing is currently the most reliable strategy. 5. Empirical design lessons and open research questions for continual, active learning in VLAs.
2
Related Works
Active Learning for Robot Policies. Active learning involves identifying states, trajectories, tasks, or environments where additional demonstrations would be most informative. Interactive imitation learning methods such as DAgger [7] and HG-DAgger [8] reduce covariate shift by collecting supervision on states induced by the learned policy. Uncertainty-guided active learning further targets this process by querying the expert only when the policy is uncertain, for example using Monte Carlo dropout to select states for supervision [9]. However, most uncertainty-aware aggregation methods [10, 11, 12, 13] have been studied for smaller, task-specific policies, often in driving [8, 12] or classical imitation-learning settings [7], rather than for large pretrained VLAs. Uncertainty Quantification in VLAs. Prior work on uncertainty quantification (UQ) has explored ensemble methods [14], Bayesian approximations [15], conformal prediction [16], and learned intervention or failure predictors. For example, KnowNo [17] uses conformal prediction to decide when a VLM planner should ask for help, but operates over high-level plan rather than low-level VLA actions. INSIGHT [6] is an inference-time introspection method for UQ in autoregressive VLAs that uses token-level probability features to predict when the robot should ask for help. INSIGHT learns a model from token-level uncertainty features extracted during action decoding, including entropy, log-probability, and Dirichlet-based aleatoric and epistemic uncertainty estimates. Continual Learning. Continual learning studies how models acquire new knowledge without catastrophic forgetting, where training on new data degrades performance on earlier tasks or distributions [18]. This issue has been studied in LLMs [19], suggesting that large pretrained sequence models such as VLAs may be similarly vulnerable when fine-tuning on new data. Common solutions include replay-based methods, which mix new data with examples from prior datasets [20, 21], and regularization-based methods, which constrain changes to parameters important for prior behavior. Elastic weight consolidation (EWC) is a representative regularization method that uses a Fisher information matrix to penalize changes to parameters important for previous tasks [22].
3
Enabling Active, Continual Learning from Uncertainty-Guided Data
Active learning in VLAs remains unsolved and under-explored. In this work, we propose an active, continual learning paradigm for fine-tuning VLAs. We focus on autoregressive VLAs, particularly π0 -FAST [4], which tokenizes continuous robot actions and predicts them via a next-token objective. Toward enabling active learning in VLAs, we leverage INSIGHT for uncertainty quantification [6] and utilize it by interpreting its “help predictions” as a proxy for high-uncertainty states from which to collect new demonstrations. We study whether uncertainty-guided data collection enables effective active learning, leading to our first research question RQ-1: Compared to passive data collection, do demonstrations originating from high-uncertainty states result in better policy performance? We next consider the effect of online vs offline active learning in VLA performance. In online learning, the robot requests a demonstration in real-time, which originates at the first uncertain state 2
Figure 1: Overview of uncertainty-guided active continual learning. An initial π0 -FAST policy is rolled out on LIBERO-10 while INSIGHT identifies high-uncertainty states. We compare passive collection with online recovery from the first high-uncertainty state and offline recovery from all high-uncertainty states, then integrate the resulting data using replay, or EWC regularization.
within a rollout. In contrast, an offline learning setting involves the robot performing a full rollout, recording every uncertain state to prompt for demonstrations at a later. We pose RQ-2: Is onlinestyle data collection sufficient, or do we need dense offline collection from all uncertain states? Toward continual active learning for VLAs, we evaluate RQ-3: Can autoregressive VLAs be fine-tuned only on newly collected recovery data? This is in contrast to the dataset aggregation assumed for RQ-1&2, which combine the robot’s previous dataset and the new uncertainty-guided data. Finally, we apply and evaluate standard techniques for continual learning to address RQ-4: Do data replay and/or regularization methods mitigate catastrophic forgetting while still enabling adaptation to new training data? To evaluate these RQs, we propose a pipeline (Figure 1) consisting of an active learning component (to identify uncertain states to prompt new data collection) and a continual learning component (fine-tuning the model on newly-collected data without degrading prior capabilities). 3.1
Active Learning Pipeline
We begin with a π0 -FAST policy trained on LIBERO [23] via the π0 -FAST authors’ recipe for 30k steps, using the final checkpoint as the initial policy πθ0 . π0 -FAST tokenizes continuous action chunks and trains with a next-token prediction objective. Given language instruction ℓ, visual observation ot , and robot state st , the policy predicts a future action-token chunk. This policy serves both as the LIBERO-10 baseline and as the rollout policy used to induce states for active data collection. During each rollout τi , we apply INSIGHT [6] at each policy step to create a set of candidate recovery timesteps Ti = {∀t ∈ τi | H(ot , st , ℓ, πθ0 ) = 1} where Ht = 1 indicates that INSIGHT predicts the policy requires assistance in that state. We evaluate Strong INSIGHT, which is trained with step-wise labels indicating when the policy requires help during real robot executions, and Weak INSIGHT, which is trained from episode-level success/failure labels on the LIBERO dataset. These terms describe the model training strategy, not downstream task difficulty. While Weak INSIGHT is more aligned with the LIBERO tasks in our evaluation, we prioritize Strong INSIGHT throughout our experiments due to its generalization across real-world and simulated settings (as shown in [6]). We consider two active recovery datasets. The online dataset collects one recovery demonstration from the first high-uncertainty state in each rollout, tonline = min{t : ht = 1}, simulati ing deployment-time intervention. The offline dataset collects from every high-uncertainty state, Tioffline = Ti , simulating offline data collection from completed rollouts. We also collect passive start-state baselines that match the task distribution and number of demonstrations of the online and offline datasets, isolating the value of collecting from high-uncertainty states rather than merely adding more task-relevant data. To collect new demonstration data, we reset the simulator to each state within each dataset and record an expert recovery trajectory to task completion, forming 3
datasets Doffline , Donline , and Dpassive . After data collection, we fine-tune with the same autoregresPT sive cross-entropy objective used for the initial policy, LCE (θ) = − t=1 log pθ (xt | x<t ), where x1:T includes prompt, observation-conditioned context, and action tokens. 3.2
Continual Learning Strategies
Replay-based data mixtures. Let Dnew denote the newly collected recovery dataset (either Doffline , Donline , or Dpassive ). We compare new-only fine-tuning on Dnew , full replay on Dold ∪ Dnew , LIBERO-10 replay on DLIBERO10 ∪ Dnew , and targeted replay on Dcollected tasks ∪ Dnew . Colllected tasks refer to the tasks determined to have low accuracy and used in active learning based data collection. These mixtures test whether recovery data alone is sufficient, whether full prior-data replay is necessary, and whether replay restricted to targeted tasks preserves non-collected tasks. Elastic weight consolidation (EWC). We evaluate EWC as a regularization-based continuallearning method. EWC penalizes movement away from the initial parameters P θ0 using a diagonal Fisher estimate F computed on a reference dataset: LEWC (θ) = LCE (θ) + λ j Fj (θj − θ0,j )2 . The coefficient λ controls the stability-plasticity tradeoff: larger values better preserve parameters important for prior behavior, but may reduce adaptation to recovery demonstrations. Learning-rate ablations. Finally, we test whether smaller updates alone reduce forgetting. Standard fine-tuning uses the original learning-rate schedule from the baseline VLA training setup, consisting of a linear warmup to α = 2.5 × 10−5 followed by cosine decay to a high learning rate of α = 2.5 × 10−6 . We compare this against a low constant learning rate, α = 2.5 × 10−8 , separating the effect of reduced update magnitude from explicit continual-learning regularization.
4
Experiment Overview and General Setup
We perform 5 experiments, corresponding to our 4 research questions and 1 follow-up experiment. Each experimental involves configuring the following independent variables: • State Selection: passive start-state collection vs active learning (guided either by Strong INSIGHT or Weak INSIGHT, and producing either online or offline datasets). • Training Data Buffer: full replay vs new-only vs LIBERO-10 replay vs targeted replay • Training Loss Function: Standard cross-entropy (CE) vs EWC regularization • Learning Rate: Standard vs low, constant LR Each experiment involves evaluating a set of policies (corresponding to various experimental conditions) with 50 rollouts per task in LIBERO-10, for a total of 500 rollouts at each checkpoint. Unless otherwise stated, replay-based active-learning experiments use new normalization statistics (recomputed on the exact dataset being used to train the model). Appendix D provides best-checkpoint summaries, pairwise statistical tests, normalization-statistics ablations, and additional EWC sweeps. We report three metrics: overall success across all 10 tasks, collected-task success on the 5 tasks used for recovery-data collection, and retained-task success on the remaining 5 tasks. In all bar plots, statistical comparisons are computed with two-sided two-proportion z-tests over rollout success counts. Significance markers indicate p < 0.05 (*), < 0.01 (**), and < 0.001 (***). In all line graphs, we denote baseline performance (dashed line) and its 95% Wilson CI (shaded band).
5
Experiment 1: Active Collection vs. Passive Start-State Collection
We address RQ-1: compared to passive data collection, do demonstrations from high-uncertainty states lead to better policy performance? We compare performance after fine-tuning the baseline VLA on passive data vs active, online data collected via either Strong or Weak INSIGHT. We train the model on a full replay buffer using the standard LR schedule and CE loss. 4
Figure 2: Online recovery collection improves over passive start-state collection. We compare fine-tuned policy performance after adding matched passive or active data. Active collection is guided by Strong INSIGHT (left) or Weak INSIGHT (right). Results: Figure 2 shows that online recovery collection outperforms matched passive collection for both INSIGHT variants. In the Strong INSIGHT setting, online recovery reaches 72.4% overall success, compared to 60.2% for matched passive collection and 59.8% for the original baseline. This improvement over passive collection is statistically significant (p = 0.0001). The Weak INSIGHT comparison shows the same qualitative trend under the same new-normalization setting. Discussion: These results show that the benefit of active collection is not simply due to adding more demonstrations from the same tasks. When task distribution and data volume are controlled, demonstrations from high-uncertainty states provide more useful supervision than demonstrations from task initial states. Appendix B reports old-normalization active-versus-passive comparisons for both Strong and Weak INSIGHT. These runs show more muted improvement, consistent with old normalization statistics acting as a stabilizing constraint that can preserve prior behavior but limit adaptation to the recovery-data distribution. This motivates our use of old normalization statistics in the new-only experiments, where we test recovery-data-only adaptation under a conservative setting.
6
Experiment 2: Online vs. Offline Recovery Collection
We address RQ-2: is online-style collection from the first high-uncertainty state sufficient, or is offline collection from all high-uncertainty states necessary? We compare performance using online vs offline datasets. Both datasets are informed by Strong INSIGHT, and use a full replay buffer and the standard CE loss and LR schedule. Results: Figure 3 compares online and offline recovery collection. Online collection reaches 72.4% overall success at its best checkpoint, while offline collection reaches 68.4%. The two methods achieve nearly identical collected-task performance, with online collection reaching 49.2% and offline collection reaching 48.4%. The overall difference is not statistically significant (p = 0.1659). Discussion: We interpret this as a demonstration-efficiency result rather than strict superiority. Online collection reaches comparable performance while requiring fewer recovery demonstrations.
Figure 3: Online recovery collection is as effective as offline collection despite using fewer demonstrations. Left: training curves for online and offline Strong INSIGHT recovery datasets mixed with prior demonstrations. Right: best-checkpoint comparison across overall, collected-task, and retained-task success. 5
This suggests that the first reliable high-uncertainty state may mark the point where the rollout begins to diverge from successful behavior, so a corrective demonstration from that state can supervise both the uncertain state and the remainder of the task. Dense offline collection adds more data, but those additional states do not translate into proportionally better adaptation.
7
Experiment 3: New-Only Fine-Tuning
We address RQ-3: can autoregressive VLAs be fine-tuned only on newly collected recovery data? We compare performance after fine-tuning on full replay data vs new-only data. New data is collected using Strong INSIGHT online, Strong INSIGHT offline, and Weak INSIGHT online. We reuse old normalization statistics from the original LIBERO dataset for new-only training as a regularizer: old norms partially constrain distribution shift, so collapse under this setting provides stronger evidence that recovery data alone is insufficient.
Figure 4: New-only fine-tuning causes catastrophic forgetting. We compare training on newonly data (solid lines) against replay-based training that combines old and new data (dashed lines). Results: Figure 4 shows that new-only fine-tuning causes severe forgetting. Across Strong and Weak INSIGHT recovery datasets, retained-task performance collapses as training progresses and overall success falls far below the original baseline. In contrast, replay-based all-mix training preserves retained-task performance while improving collected tasks. For example, the best Strong INSIGHT online new-only checkpoint reaches only 28.4% overall success, with 0.4% collected-task success and 56.4% retained-task success. The corresponding replay-based online model reaches 72.4% overall success, with 49.2% collected-task success and 95.6% retained-task success. Discussion: Recovery demonstrations are informative precisely because they focus on difficult, policy-induced states, but this also makes the dataset distributionally narrow. New-only fine-tuning does not provide enough coverage to preserve behaviors that the original policy already performs well. Thus, active learning for VLAs cannot be reduced to selecting high-value states; the adaptation procedure must also preserve prior competence.
8
Experiment 4: Continual-Learning Regularization
We address RQ-4: do learning-rate reduction and regularization mitigate catastrophic forgetting while still enabling adaptation to recovery data? Using new-only data, we compare CE with standard LR vs CE with constant LR vs EWC with constant LR and multiple coefficient values. Full replay results are included as a reference. Results: Figure 5 compares representative regularization settings against low-learning-rate finetuning and replay. Low-learning-rate new-only training reaches 62.8% overall success and preserves retained-task performance at 91.2%, but collected-task success remains only 34.4%. EWC with λ = 1012 reaches 61.4% overall success, with 32.0% collected-task success and 90.8% retainedtask success. These methods preserve prior behavior better than standard new-only fine-tuning, but do not match full replay, which reaches 72.4% overall success and 49.2% collected-task success. The full EWC coefficient sweep and filtered-Fisher sweep are reported in Appendix C. They show the same pattern: stronger regularization can keep retained-task success near the original baseline, 6
Figure 5: Low learning rates (α) and EWC reduce forgetting but limit adaptation to new data. Compared to full replay, low learning-rate fine-tuning and EWC better preserve retained-task performance than standard new-only fine-tuning, but produce smaller gains on collected tasks.
but collected-task performance remains limited and overall success does not improve. Filtering the Fisher reference data changes the stability-plasticity behavior but does not remove the core tradeoff. Discussion: Low learning rates and EWC solve only part of the active continual-learning problem. Weak constraints allow more adaptation but risk forgetting, while strong constraints preserve retained-task behavior but suppress useful learning from recovery demonstrations. In our setting, replay provides a stronger stability-plasticity tradeoff than parameter regularization alone.
9
Experiment 5: Replay Scope
Motivated by the effectiveness of replay, we investigate how much prior data should be replayed. Using Strong INSIGHT online recovery data, we compare full replay vs LIBERO-10 replay vs targeted replay(containing the subset of LIBERO-10 replay data relevant to the collected tasks ). We also investigate a filtered EWC variant where the Fisher matrix excludes the collected tasks. Results: Figure 6 shows that full replay and LIBERO-10 replay both maintain strong retained-task performance while improving collected tasks. LIBERO-10 replay reaches 68.6% overall success, close to the 72.4% achieved by full replay; this difference is not significant (p = 0.1877). In contrast, targeted replay reaches 63.2% overall success and significantly reduces retained-task performance relative to LIBERO-10 replay, dropping from 90.0% to 83.6% retained-task success (p = 0.0345). Figure 7 further tests whether EWC can stabilize targeted replay. It does not consistently solve the retention problem: some variants briefly preserve performance early in training, but many later degrade substantially. Thus, when replay coverage is too narrow, parameter regularization cannot reliably compensate for missing retained-task data. Discussion: Replay does not necessarily need to include the full original dataset, but it must cover the behaviors that should be retained. Restricting replay to LIBERO-10 preserves most of the benefit of full replay, while targeted replay biases adaptation toward the collected tasks and allows non-collected task performance to degrade.
Figure 6: Replay coverage controls the tradeoff between adaptation and retention. Left: results for Strong INSIGHT online recovery data mixed with different replay subsets. Right: bestcheckpoint comparison. 7
Figure 7: Regularization does not consistently prevent degradation, indicating that replay coverage is critical. We evaluate EWC and filtered-EWC variants when using targeted replay data.
10
Summary and Conclusion
We studied active continual learning for autoregressive VLAs by using INSIGHT predictions to identify high-uncertainty, states and collect recovery demonstrations from those states. Across experiments, five findings emerge. First, high-uncertainty recovery demonstrations are more informative than matched passive start-state demonstrations in both the Strong INSIGHT and the Weak INSIGHT settings. Second, online recovery collection is more demonstration-efficient than offline collection: dense offline collection does not provide a statistically significant overall advantage despite requiring more demonstrations. Third, recovery data alone is insufficient for fine-tuning; newonly training causes catastrophic forgetting even under old normalization statistics, which partially constrain distribution shift. Fourth, low learning rates and EWC reduce forgetting but expose a stability-plasticity tradeoff, preserving old behavior at the cost of reduced adaptation. Fifth, replaybased mixing provides the strongest overall solution, but replay coverage matters: LIBERO-10 replay approaches full replay, while targeted replay fails to preserve retained-task performance. Together, these results support the central claim that active learning for VLAs must be formulated as active continual learning, where robots not only detect when they need help, but also use those moments to collect informative experience and integrate it without erasing prior skills. INSIGHT-style uncertainty predictors identify informative states for recovery collection, but the resulting data is narrow and distribution-shifted. Effective adaptation therefore requires both targeted data collection and mechanisms for preserving previously learned behaviors.
11
Limitations
Our study has several limitations. First, all experiments are in LIBERO-10 simulation. This enables controlled evaluation of active data collection and continual adaptation, but physical deployment introduces sensor noise, execution variability, safety constraints, imperfect resets, and higher demonstration costs. Our recovery-collection procedure also assumes access to simulator resets at high-uncertainty intermediate states; on real robots, analogous data would likely require runtime human intervention or demonstrations from naturally reached failure-adjacent states. Thus, the online setting is closer to deployment than offline collection from all high-uncertainty states, but both abstract away real-world collection costs. Second, our results depend on the quality of the help predictor. INSIGHT can produce false positives, which waste demonstrations, and false negatives, which miss useful recovery states; future work should study better calibrated and more risk-sensitive uncertainty estimates. Third, we evaluate one autoregressive VLA family on one benchmark. Other VLA architectures, especially diffusion-based or hybrid action heads, may require different uncertainty signals and recovery-state selection mechanisms. Fourth, our continual-learning methods are simple: replay, low-learning-rate fine-tuning, and EWC. More sophisticated methods such as LoRA, adapters, selective freezing, gradient projection, data reweighting, rehearsal buffers, or distillation may improve the stability-plasticity tradeoff. Finally, we focus primarily on task success; future evaluations should also measure intervention cost, demonstration count, recovery length, action smoothness, safety violations, uncertainty calibration, and repeated active-learning cycles. 8
Acknowledgments If a paper is accepted, the final camera-ready version will (and probably should) include acknowledgments. All acknowledgments go at the end of the paper, including thanks to reviewers who gave useful comments, to colleagues who contributed to the ideas, and to funding agencies and corporate sponsors that provided financial support.
References [1] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich. Rt-1: Robotics transformer for real-world control at scale, 2023. URL https://arxiv.org/abs/2212.06817. [2] A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, L. Lee, T.-W. E. Lee, S. Levine, Y. Lu, H. Michalewski, I. Mordatch, K. Pertsch, K. Rao, K. Reymann, M. Ryoo, G. Salazar, P. Sanketi, P. Sermanet, J. Singh, A. Singh, R. Soricut, H. Tran, V. Vanhoucke, Q. Vuong, A. Wahid, S. Welker, P. Wohlhart, J. Wu, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich. Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023. URL https://arxiv.org/abs/2307.15818. [3] G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, S. Bohez, K. Bousmalis, A. Brohan, T. Buschmann, A. Byravan, S. Cabi, K. Caluwaerts, F. Casarini, O. Chang, J. E. Chen, X. Chen, H.-T. L. Chiang, K. Choromanski, D. D’Ambrosio, S. Dasari, T. Davchev, C. Devin, N. D. Palo, T. Ding, A. Dostmohamed, D. Driess, Y. Du, D. Dwibedi, M. Elabd, C. Fantacci, C. Fong, E. Frey, C. Fu, M. Giustina, K. Gopalakrishnan, L. Graesser, L. Hasenclever, N. Heess, B. Hernaez, A. Herzog, R. A. Hofer, J. Humplik, A. Iscen, M. G. Jacob, D. Jain, R. Julian, D. Kalashnikov, M. E. Karagozler, S. Karp, C. Kew, J. Kirkland, S. Kirmani, Y. Kuang, T. Lampe, A. Laurens, I. Leal, A. X. Lee, T.-W. E. Lee, J. Liang, Y. Lin, S. Maddineni, A. Majumdar, A. H. Michaely, R. Moreno, M. Neunert, F. Nori, C. Parada, E. Parisotto, P. Pastor, A. Pooley, K. Rao, K. Reymann, D. Sadigh, S. Saliceti, P. Sanketi, P. Sermanet, D. Shah, M. Sharma, K. Shea, C. Shu, V. Sindhwani, S. Singh, R. Soricut, J. T. Springenberg, R. Sterneck, R. Surdulescu, J. Tan, J. Tompson, V. Vanhoucke, J. Varley, G. Vesom, G. Vezzani, O. Vinyals, A. Wahid, S. Welker, P. Wohlhart, F. Xia, T. Xiao, A. Xie, J. Xie, P. Xu, S. Xu, Y. Xu, Z. Xu, Y. Yang, R. Yao, S. Yaroshenko, W. Yu, W. Yuan, J. Zhang, T. Zhang, A. Zhou, and Y. Zhou. Gemini robotics: Bringing ai into the physical world, 2025. URL https://arxiv.org/abs/2503.20020. [4] K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv.org/abs/2501.09747. [5] P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky. π0.5 : a vision-language-action model with open-world generalization, 2025. URL https://arxiv.org/abs/2504.16054. 9
[6] U. B. Karli, Z. Shangguan, and T. FItzgerald. Insight: Inference-time sequence introspection for generating help triggers in vision-language-action models, 2025. URL https://arxiv. org/abs/2510.01389. [7] S. Ross and D. Bagnell. Efficient reductions for imitation learning. In Y. W. Teh and M. Titterington, editors, Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9 of Proceedings of Machine Learning Research, pages 661–668, Chia Laguna Resort, Sardinia, Italy, 13–15 May 2010. PMLR. URL https: //proceedings.mlr.press/v9/ross10a.html. [8] M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer. Hg-dagger: Interactive imitation learning with human experts. In 2019 International Conference on Robotics and Automation (ICRA), page 8077–8083. IEEE Press, 2019. doi:10.1109/ICRA.2019.8793698. URL https://doi.org/10.1109/ICRA.2019.8793698. [9] Y. Cui, D. Isele, S. Niekum, and K. Fujimura. Uncertainty-aware data aggregation for deep imitation learning, 2019. URL https://arxiv.org/abs/1905.02780. [10] K. Menda, K. Driggs-Campbell, and M. J. Kochenderfer. Ensembledagger: A bayesian approach to safe imitation learning, 2019. URL https://arxiv.org/abs/1807.08364. [11] M. Zhao, R. Simmons, H. Admoni, A. Ramdas, and A. Bajcsy. Conformalized interactive imitation learning: Handling expert shift and intermittent feedback, 2025. URL https:// arxiv.org/abs/2410.08852. [12] C. Wang and Y. Wang. Uncertainty-driven data aggregation for imitation learning in autonomous vehicles. Information, 15(6), 2024. ISSN 2078-2489. doi:10.3390/info15060336. URL https://www.mdpi.com/2078-2489/15/6/336. [13] S.-W. Lee, X. Kang, and Y.-L. Kuo. Diff-dagger: Uncertainty estimation with diffusion policy for robotic manipulation, 2025. URL https://arxiv.org/abs/2410.14868. [14] B. Lakshminarayanan, A. Pritzel, and C. Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles, 2017. URL https://arxiv.org/abs/1612.01474. [15] Y. Gal and Z. Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In M. F. Balcan and K. Q. Weinberger, editors, Proceedings of The 33rd International Conference on Machine Learning, volume 48 of Proceedings of Machine Learning Research, pages 1050–1059, New York, New York, USA, 20–22 Jun 2016. PMLR. URL https://proceedings.mlr.press/v48/gal16.html. [16] A. N. Angelopoulos and S. Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification, 2022. URL https://arxiv.org/abs/2107. 07511. [17] A. Z. Ren, A. Dixit, A. Bodrova, S. Singh, S. Tu, N. Brown, P. Xu, L. Takayama, F. Xia, J. Varley, Z. Xu, D. Sadigh, A. Zeng, and A. Majumdar. Robots that ask for help: Uncertainty alignment for large language model planners. In Proceedings of the Conference on Robot Learning (CoRL), 2023. [18] M. McCloskey and N. J. Cohen. Catastrophic interference in connectionist networks: The sequential learning problem. volume 24 of Psychology of Learning and Motivation, pages 109–165. Academic Press, 1989. doi:https://doi.org/10.1016/S0079-7421(08)60536-8. URL https://www.sciencedirect.com/science/article/pii/S0079742108605368. [19] Y. Luo, Z. Yang, F. Meng, Y. Li, J. Zhou, and Y. Zhang. An empirical study of catastrophic forgetting in large language models during continual fine-tuning, 2025. URL https: //arxiv.org/abs/2308.08747. 10
[20] S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert. icarl: Incremental classifier and representation learning, 2017. URL https://arxiv.org/abs/1611.07725. [21] D. Lopez-Paz and M. A. Ranzato. Gradient episodic memory for continual learning. In I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL https://proceedings.neurips.cc/paper_files/paper/ 2017/file/f87522788a2be2d171666752f97ddebb-Paper.pdf. [22] J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526, Mar. 2017. ISSN 1091-6490. doi: 10.1073/pnas.1611835114. URL http://dx.doi.org/10.1073/pnas.1611835114. [23] B. Liu, Y. Zhu, C. Gao, Y. Feng, Q. Liu, Y. Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning, 2023. URL https://arxiv.org/abs/2306. 03310.
11
A
Compute Infrastructure
All models were trained on an institutional high-performance computing cluster using NVIDIA H200 GPUs. Most evaluations were also run on the same cluster, with some additional local testing on an NVIDIA RTX 6000 Ada workstation. In total, the project used approximately 2,154 H200 GPU-hours of cluster compute.
B
Normalization-Statistics Ablations
The main active-versus-passive comparison uses new normalization statistics, which better match the state and action distribution of the aggregated replay-plus-recovery dataset. Here, we report additional active-versus-passive comparisons using old normalization statistics. These runs are useful because old normalization statistics partially constrain the fine-tuned policy toward the original data distribution. As a result, old normalization can act like an implicit regularizer: it can help preserve prior behavior, but it can also limit adaptation to the newly collected recovery demonstrations. Figure 8 shows the Strong INSIGHT active-versus-passive comparison with old normalization statistics. Compared to the main new-normalization result in Figure 5, the improvement from online recovery collection is more muted. This supports the interpretation that old normalization statistics stabilize the model but reduce the plasticity needed to exploit recovery data.
Figure 8: Strong INSIGHT active-versus-passive comparison with old normalization statistics. Old normalization statistics partially constrain the adapted policy toward the original data distribution. This can preserve retained-task behavior, but it also limits the improvement obtained from high-uncertainty recovery demonstrations. Figure 9 shows the same old-normalization comparison using Weak INSIGHT. Weak INSIGHT provides a complementary in-domain or LIBERO-aligned help signal, while Strong INSIGHT is the cleaner transfer setting used in the main paper. The old-normalization weak result follows the same qualitative pattern: recovery data can improve collected-task performance, but old normalization limits the magnitude of adaptation. These ablations also motivate our use of old normalization statistics in the new-only fine-tuning experiments. Since old normalization acts as a conservative stabilizing choice, catastrophic forgetting under old normalization provides stronger evidence that recovery data alone is insufficient for stable adaptation. 12
Figure 9: Weak INSIGHT active-versus-passive comparison with old normalization statistics. This comparison complements the Strong INSIGHT old-normalization ablation and shows that old normalization statistics similarly constrain adaptation under the Weak INSIGHT recovery dataset.
Figure 10: Full EWC coefficient sweep with low learning rate. Stronger regularization stabilizes retained-task performance, but does not provide enough plasticity to fully exploit uncertainty-guided recovery data.
C
Additional EWC Sweeps
Figure 10 reports the full EWC coefficient sweep using the low learning rate α = 2.5×10−8 . Across coefficients, stronger regularization generally improves retention, but collected-task performance remains limited and overall success does not match replay-based mixing. Figure 11 reports the corresponding sweep using filtered Fisher estimates. Filtering the Fisher reference data changes the detailed training dynamics, but does not remove the stability-plasticity tradeoff observed in the main paper.
D
Additional Experimental Results
This appendix reports the best-checkpoint summary, pairwise statistical tests, and raw percheckpoint task success counts. Best checkpoints are selected by highest overall success over evaluated checkpoints for each method. Full per-checkpoint results are included to make checkpoint selection transparent. 13
Figure 11: Filtered Fisher EWC sweep. Computing Fisher information on filtered prior data changes the stability-plasticity behavior, but regularization alone still does not match replay-based adaptation. Table 1: Best-checkpoint summary for key experimental conditions. We report overall success, collected-task success, and retained-task success. Each LIBERO-10 task is evaluated with 50 rollouts. Collected tasks are tasks 2, 3, 5, 8, and 9; retained tasks are tasks 0, 1, 4, 6, and 7. Experiment
Step
Overall
Overall %
Collected
Collected %
Retained
Retained %
Original baseline Strong First All-Mix New Norm Strong First Baseline All-Mix New Norm Strong All All-Mix New Norm Weak First All-Mix New Norm Weak First All-Mix Old Norm Weak First Baseline All-Mix Old Norm Strong First Only Old Norm Strong All Only Old Norm Weak First Only Old Norm Strong First Only Low LR Fisher 1e12 Low LR Strong First Libero 10 Mix Strong First Libero10 Collected Mix
0 9000 3000 5000 6000 500 1000 1000 1000 500 100 9999 7500 1000
299/500 362/500 301/500 342/500 354/500 329/500 311/500 142/500 81/500 195/500 314/500 307/500 343/500 316/500
59.8 72.4 60.2 68.4 70.8 65.8 62.2 28.4 16.2 39.0 62.8 61.4 68.6 63.2
80/250 123/250 83/250 121/250 134/250 98/250 86/250 1/250 0/250 10/250 86/250 80/250 118/250 107/250
32.0 49.2 33.2 48.4 53.6 39.2 34.4 0.4 0.0 4.0 34.4 32.0 47.2 42.8
219/250 239/250 218/250 221/250 220/250 231/250 225/250 141/250 81/250 185/250 228/250 227/250 225/250 209/250
87.6 95.6 87.2 88.4 88.0 92.4 90.0 56.4 32.4 74.0 91.2 90.8 90.0 83.6
D.1
Best Checkpoint Summary
D.2
Pairwise Statistical Tests
Table 2: Pairwise statistical comparisons for key claims. We use twoproportion z-tests over rollout success counts. Significance markers denote ∗ p < 0.05, ∗∗ p < 0.01, ∗∗∗ p < 0.001, and n.s. otherwise. Comparison
Metric
RQ1 strong new norm active vs passive overall Overall RQ1 strong new norm active vs passive collected Collected RQ1 strong new norm active vs passive retained Retained RQ1 strong new norm passive vs baseline overall Overall RQ1 strong new norm passive vs baseline collected Collected RQ1 strong new norm passive vs baseline retained Retained RQ1 strong new norm active vs baseline overall Overall RQ1 strong new norm active vs baseline collected Collected RQ1 strong new norm active vs baseline retained Retained Appendix RQ1 strong old norm active vs passive overall Overall Appendix RQ1 strong old norm active vs passive collected Collected Appendix RQ1 strong old norm active vs passive retained Retained Appendix RQ1 strong old norm passive vs baseline overall Overall Appendix RQ1 strong old norm passive vs baseline collected Collected Appendix RQ1 strong old norm passive vs baseline retained Retained
Method A
Method B
z
p
362/500 123/250 239/250 301/500 83/250 218/250 362/500 123/250 239/250 308/500 86/250 222/250 303/500 74/250 229/250
301/500 83/250 218/250 299/500 80/250 219/250 299/500 80/250 219/250 303/500 74/250 229/250 299/500 80/250 219/250
4.08 3.63 3.35 0.13 0.29 -0.13 4.21 3.92 3.22 0.32 1.15 -1.05 0.26 -0.58 1.47
0.0000∗∗∗ 0.0003∗∗∗ 0.0008∗∗∗ 0.8973 n.s. 0.7747 n.s. 0.8928 n.s. 0.0000∗∗∗ 0.0001∗∗∗ 0.0013∗∗ 0.7457 n.s. 0.2500 n.s. 0.2924 n.s. 0.7961 n.s. 0.5611 n.s. 0.1429 n.s.
Continued on next page
14
Table 2: Pairwise statistical comparisons for key claims continued. Comparison
Metric
Method A
Method B
z
p
Appendix RQ1 strong old norm active vs baseline overall Appendix RQ1 strong old norm active vs baseline collected Appendix RQ1 strong old norm active vs baseline retained Appendix RQ1 weak old norm active vs passive overall Appendix RQ1 weak old norm active vs passive collected Appendix RQ1 weak old norm active vs passive retained Appendix RQ1 weak old norm passive vs baseline overall Appendix RQ1 weak old norm passive vs baseline collected Appendix RQ1 weak old norm passive vs baseline retained Appendix RQ1 weak old norm active vs baseline overall Appendix RQ1 weak old norm active vs baseline collected Appendix RQ1 weak old norm active vs baseline retained RQ1 weak new norm active vs passive overall RQ1 weak new norm active vs passive collected RQ1 weak new norm active vs passive retained RQ1 weak new norm passive vs baseline overall RQ1 weak new norm passive vs baseline collected RQ1 weak new norm passive vs baseline retained RQ1 weak new norm active vs baseline overall RQ1 weak new norm active vs baseline collected RQ1 weak new norm active vs baseline retained RQ2 online vs offline overall, best RQ2 online vs offline collected, best RQ2 online vs offline retained, best RQ2 online vs baseline overall, best RQ2 online vs baseline collected, best RQ2 online vs baseline retained, best RQ2 offline vs baseline overall, best RQ2 offline vs baseline collected, best RQ2 offline vs baseline retained, best RQ5 full replay vs baseline overall, best RQ5 full replay vs baseline collected, best RQ5 full replay vs baseline retained, best RQ5 full replay vs LIBERO-10 replay overall, best RQ5 full replay vs LIBERO-10 replay collected, best RQ5 full replay vs LIBERO-10 replay retained, best RQ5 full replay vs Targated replay overall, best RQ5 full replay vs Targated replay collected, best RQ5 full replay vs Targated replay retained, best RQ5 LIBERO-10 replay vs Targated replay overall, best RQ5 LIBERO-10 replay vs Targated replay collected, best RQ5 LIBERO-10 replay vs Targated replay retained, best RQ5 LIBERO-10 replay vs baseline overall, best RQ5 LIBERO-10 replay vs baseline collected, best RQ5 LIBERO-10 replay vs baseline retained, best RQ5 Targated replay vs baseline overall, best RQ5 Targated replay vs baseline collected, best RQ5 Targated replay vs baseline retained, best
Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained Overall Collected Retained
308/500 86/250 222/250 329/500 98/250 231/250 311/500 86/250 225/250 329/500 98/250 231/250 354/500 134/250 220/250 308/500 79/250 229/250 354/500 134/250 220/250 362/500 123/250 239/250 362/500 123/250 239/250 342/500 121/250 221/250 362/500 123/250 239/250 362/500 123/250 239/250 362/500 123/250 239/250 343/500 118/250 225/250 343/500 118/250 225/250 316/500 107/250 209/250
299/500 80/250 219/250 311/500 86/250 225/250 299/500 80/250 219/250 299/500 80/250 219/250 308/500 79/250 229/250 299/500 80/250 219/250 299/500 80/250 219/250 342/500 121/250 221/250 299/500 80/250 219/250 299/500 80/250 219/250 299/500 80/250 219/250 343/500 118/250 225/250 316/500 107/250 209/250 316/500 107/250 209/250 299/500 80/250 219/250 299/500 80/250 219/250
0.58 0.57 0.42 1.19 1.11 0.95 0.78 0.57 0.85 1.96 1.68 1.79 3.08 4.97 -1.33 0.58 -0.10 1.47 3.65 4.88 0.14 1.39 0.18 2.97 4.21 3.92 3.22 2.83 3.74 0.28 4.21 3.92 3.22 1.32 0.45 2.42 3.11 1.44 4.40 1.80 0.99 2.11 2.90 3.47 0.85 1.10 2.50 -1.27
0.5601 n.s. 0.5688 n.s. 0.6775 n.s. 0.2357 n.s. 0.2658 n.s. 0.3436 n.s. 0.4366 n.s. 0.5688 n.s. 0.3949 n.s. 0.0497∗ 0.0927 n.s. 0.0736 n.s. 0.0021∗∗ 0.0000∗∗∗ 0.1836 n.s. 0.5601 n.s. 0.9235 n.s. 0.1429 n.s. 0.0003∗∗∗ 0.0000∗∗∗ 0.8913 n.s. 0.1659 n.s. 0.8580 n.s. 0.0030∗∗ 0.0000∗∗∗ 0.0001∗∗∗ 0.0013∗∗ 0.0046∗∗ 0.0002∗∗∗ 0.7831 n.s. 0.0000∗∗∗ 0.0001∗∗∗ 0.0013∗∗ 0.1877 n.s. 0.6545 n.s. 0.0154∗ 0.0019∗∗ 0.1511 n.s. 0.0000∗∗∗ 0.0717 n.s. 0.3227 n.s. 0.0345∗ 0.0037∗∗ 0.0005∗∗∗ 0.3949 n.s. 0.2692 n.s. 0.0126∗ 0.2027 n.s.
D.3
Raw Per-Checkpoint Results
Tables below report raw task success counts. T0–T9 denote LIBERO-10 tasks. Each task is evaluated with 50 rollouts.
15
D.3.1 Active Collection and Replay Baselines Table 3: Raw per-checkpoint results for Active Collection and Replay Baselines. Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
16
Strong All All-Mix New Norm 1000 Strong All All-Mix New Norm 3000 Strong All All-Mix New Norm 5000 Strong All All-Mix New Norm 7000 Strong All All-Mix Old Norm 1000 Strong All All-Mix Old Norm 3000 Strong All All-Mix Old Norm 5000 Strong First All-Mix New Norm 500 Strong First All-Mix New Norm 1000 Strong First All-Mix New Norm 1500 Strong First All-Mix New Norm 2000 Strong First All-Mix New Norm 2500 Strong First All-Mix New Norm 3000 Strong First All-Mix New Norm 3500 Strong First All-Mix New Norm 4000 Strong First All-Mix New Norm 4500 Strong First All-Mix New Norm 5000 Strong First All-Mix New Norm 5500 Strong First All-Mix New Norm 6000 Strong First All-Mix New Norm 6500 Strong First All-Mix New Norm 7000 Strong First All-Mix New Norm 7500 Strong First All-Mix New Norm 8000 Strong First All-Mix New Norm 8500 Strong First All-Mix New Norm 9000 Strong First All-Mix New Norm 9500 Strong First All-Mix New Norm 9999 Strong First All-Mix Old Norm 1000 Strong First All-Mix Old Norm 3000 Strong First All-Mix Old Norm 5000 Strong First Baseline All-Mix New Norm 1000
0 42 41 45 42 47 36 45 47 45 40 46 50 42 46 41 40 42 43 46 43 48 44 45 48 43 45 44 48 46 45
43 47 45 47 46 48 50 48 50 49 44 47 47 48 46 49 49 46 48 43 48 48 47 47 48 50 48 49 49 0 50
2 6 17 22 11 3 7 0 4 15 14 20 28 19 21 33 6 28 28 15 16 18 26 26 28 19 23 11 15 24 3
20 34 38 41 25 21 18 15 42 40 31 30 38 37 34 45 26 31 32 31 29 39 34 34 36 42 36 15 29 22 22
34 45 42 41 46 48 39 43 44 32 43 45 45 39 43 41 37 43 43 41 39 46 42 36 48 45 47 47 39 43 42
24 46 45 45 20 37 19 19 28 25 36 35 28 29 36 22 25 27 31 18 32 30 35 35 29 28 31 28 23 15 26
30 34 46 30 36 39 43 41 45 36 39 43 40 45 44 45 33 38 36 42 34 35 37 41 46 43 39 39 38 36 29
46 41 47 50 48 45 47 46 47 43 46 47 44 46 48 46 43 46 46 48 48 49 48 49 49 46 48 49 48 47 47
0 1 10 3 0 0 0 0 4 6 8 12 5 9 2 8 7 5 4 10 0 13 10 4 7 6 7 0 0 0 2
3 21 11 14 10 6 1 2 22 15 13 11 10 23 15 12 11 21 23 22 28 25 22 26 23 14 23 10 19 15 10
40.4 63.4 68.4 67.6 56.8 58.8 52.0 51.8 66.6 61.2 62.8 67.2 67.0 67.4 67.0 68.4 55.4 65.4 66.8 63.2 63.4 70.2 69.0 68.6 72.4 67.2 69.4 58.4 61.6 49.6 55.2
19.6 43.2 48.4 50.0 26.4 26.8 18.0 14.4 40.0 40.4 40.8 43.2 43.6 46.8 43.2 48.0 30.0 44.8 47.2 38.4 42.0 50.0 50.8 50.0 49.2 43.6 48.0 25.6 34.4 30.4 25.2
61.2 83.6 88.4 85.2 87.2 90.8 86.0 89.2 93.2 82.0 84.8 91.2 90.4 88.0 90.8 88.8 80.8 86.0 86.4 88.0 84.8 90.4 87.2 87.2 95.6 90.8 90.8 91.2 88.8 68.8 85.2
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
17
Strong First Baseline All-Mix New Norm 3000 Strong First Baseline All-Mix New Norm 5000 Strong First Baseline All-Mix New Norm 5500 Strong First Baseline All-Mix New Norm 6000 Strong First Baseline All-Mix New Norm 6500 Strong First Baseline All-Mix New Norm 9999 Strong First Baseline All-Mix Old Norm 1000 Strong First Baseline All-Mix Old Norm 3000 Strong First Baseline All-Mix Old Norm 5000 Weak First All-Mix New Norm 500 Weak First All-Mix New Norm 1000 Weak First All-Mix New Norm 1500 Weak First All-Mix New Norm 2000 Weak First All-Mix New Norm 2500 Weak First All-Mix New Norm 3000 Weak First All-Mix New Norm 3500 Weak First All-Mix New Norm 4000 Weak First All-Mix New Norm 4500 Weak First All-Mix New Norm 5000 Weak First All-Mix New Norm 5500 Weak First All-Mix New Norm 6000 Weak First All-Mix New Norm 6500 Weak First All-Mix New Norm 7000 Weak First All-Mix New Norm 7500 Weak First All-Mix New Norm 8000 Weak First All-Mix New Norm 8500 Weak First All-Mix New Norm 9000 Weak First All-Mix New Norm 9500 Weak First All-Mix New Norm 9999 Weak First All-Mix Old Norm 500 Weak First All-Mix Old Norm 1000 Weak First All-Mix Old Norm 1500 Weak First All-Mix Old Norm 2000 Weak First All-Mix Old Norm 2500 Weak First All-Mix Old Norm 3000
37 41 41 42 38 47 45 0 37 45 44 38 44 42 38 38 23 44 39 43 46 43 49 41 40 31 40 39 43 43 45 46 41 45 43
48 49 49 49 48 47 50 48 48 50 45 47 49 47 44 46 47 46 45 45 47 47 45 46 46 46 48 47 49 50 46 49 46 47 48
4 3 3 15 7 3 10 11 2 0 25 25 17 18 23 25 28 21 38 26 26 32 36 21 32 29 29 29 28 25 2 20 11 11 22
33 23 28 14 17 21 24 16 24 3 30 38 36 44 31 37 39 34 41 40 29 32 36 38 41 42 38 36 38 22 19 36 23 23 36
45 34 41 37 44 47 46 45 42 39 44 43 42 45 46 36 35 36 44 34 42 36 41 44 45 49 41 43 47 48 43 42 45 40 42
24 31 24 30 29 28 29 15 17 32 41 41 36 40 29 24 33 39 31 33 35 40 32 42 21 35 31 39 28 28 39 29 29 24 24
43 37 37 42 35 35 39 39 45 40 46 33 33 31 39 35 37 36 38 38 38 34 43 42 37 39 34 40 38 41 47 39 39 41 35
45 4 18 49 0 10 48 0 18 45 0 14 46 1 14 49 2 8 49 0 11 48 1 20 48 1 8 43 1 0 44 5 5 42 2 23 41 5 23 49 9 20 46 4 19 45 3 11 43 6 12 38 3 23 46 5 20 43 6 18 47 11 33 46 3 31 44 4 22 47 2 28 47 1 24 46 7 24 47 8 28 48 5 28 46 4 16 49 3 20 48 3 21 43 0 11 49 0 12 45 0 17 48 0 7
60.2 55.4 57.8 57.6 55.8 57.4 60.6 48.6 54.4 50.6 65.8 66.4 65.2 69.0 63.8 60.0 60.6 64.0 69.4 65.2 70.8 68.8 70.4 70.2 66.8 69.6 68.8 70.8 67.4 65.8 62.6 63.0 59.0 58.6 61.0
33.2 26.8 29.2 29.2 27.2 24.8 29.6 25.2 20.8 14.4 42.4 51.6 46.8 52.4 42.4 40.0 47.2 48.0 54.0 49.2 53.6 55.2 52.0 52.4 47.6 54.8 53.6 54.8 45.6 39.2 33.6 38.4 30.0 30.0 35.6
87.2 84.0 86.4 86.0 84.4 90.0 91.6 72.0 88.0 86.8 89.2 81.2 83.6 85.6 85.2 80.0 74.0 80.0 84.8 81.2 88.0 82.4 88.8 88.0 86.0 84.4 84.0 86.8 89.2 92.4 91.6 87.6 88.0 87.2 86.4
18
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First All-Mix Old Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix New Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm
3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9500 9999 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 7500 8000 8500 9000 9500 9999 500 1000 1500
46 41 43 44 44 37 44 44 48 42 45 47 43 36 46 45 42 46 46 45 42 46 47 41 41 45 47 42 47 45 42 48 46 45 45
49 50 50 48 48 46 48 49 47 49 47 49 48 47 49 50 50 48 50 47 47 47 49 48 48 46 49 50 48 48 49 49 49 48 49
24 18 17 16 10 13 12 22 24 16 19 17 20 26 0 0 1 1 4 6 1 3 15 2 5 12 8 4 5 10 9 9 8 17 5
13 16 25 16 33 15 31 21 20 23 22 25 27 34 15 20 13 21 12 30 29 14 16 22 15 28 19 20 33 15 10 15 24 25 26
37 43 46 41 43 38 44 45 43 44 44 40 39 47 34 47 40 48 47 39 47 46 42 43 44 41 46 47 47 45 44 42 45 46 46
27 29 38 24 28 23 18 28 32 35 29 20 22 23 19 21 13 27 25 26 17 13 17 32 18 28 21 24 28 23 30 31 26 27 17
34 37 38 40 45 38 37 35 34 42 39 41 38 37 36 41 28 37 32 39 36 43 38 41 36 43 33 42 40 43 41 39 41 41 39
48 49 46 47 46 49 46 48 43 45 43 48 47 46 41 43 47 41 45 41 46 44 47 46 43 44 47 47 47 47 49 47 48 45 46
2 0 0 0 1 0 2 2 2 1 1 1 0 1 0 2 1 0 2 0 1 0 1 0 1 2 3 0 3 0 0 1 0 1 1
15 11 8 11 16 15 15 25 18 22 9 19 23 14 5 10 9 13 6 13 8 8 14 1 10 16 15 14 10 18 6 5 14 16 12
59.0 58.8 62.2 57.4 62.8 54.8 59.4 63.8 62.2 63.8 59.6 61.4 61.4 62.2 49.0 55.8 48.8 56.4 53.8 57.2 54.8 52.8 57.2 55.2 52.2 61.0 57.6 58.0 61.6 58.8 56.0 57.2 60.2 62.2 57.2
32.4 29.6 35.2 26.8 35.2 26.4 31.2 39.2 38.4 38.8 32.0 32.8 36.8 39.2 15.6 21.2 14.8 24.8 19.6 30.0 22.4 15.2 25.2 22.8 19.6 34.4 26.4 24.8 31.6 26.4 22.0 24.4 28.8 34.4 24.4
85.6 88.0 89.2 88.0 90.4 83.2 87.6 88.4 86.0 88.8 87.2 90.0 86.0 85.2 82.4 90.4 82.8 88.0 88.0 84.4 87.2 90.4 89.2 87.6 84.8 87.6 88.8 91.2 91.6 91.2 90.0 90.0 91.6 90.0 90.0
19
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm Weak First Baseline All-Mix Old Norm
2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9500 9999
39 42 35 45 45 42 45 40 43 43 42 41 43 42 47 45 44
49 7 10 43 22 33 44 48 9 9 47 28 41 48 49 2 18 47 24 40 47 50 11 18 41 20 39 46 48 14 15 45 15 44 45 49 5 13 44 22 42 46 49 9 17 46 23 38 44 49 4 15 45 24 40 43 49 3 16 39 24 45 44 50 2 10 44 15 39 46 50 4 17 46 25 41 47 48 8 15 40 15 36 49 50 5 10 48 18 43 48 47 2 17 42 23 37 46 48 2 23 43 17 37 46 47 2 22 36 20 35 46 48 6 12 46 23 48 48
0 1 0 0 0 0 0 1 0 0 0 0 0 0 1 1 0
16 20 14 14 13 18 22 9 7 5 8 8 9 9 23 10 17
52.6 58.6 55.2 56.8 56.8 56.2 58.6 54.0 54.0 50.8 56.0 52.0 54.8 53.0 57.4 52.8 58.4
22.0 26.8 23.2 25.2 22.8 23.2 28.4 21.2 20.0 12.8 21.6 18.4 16.8 20.4 26.4 22.0 23.2
83.2 90.4 87.2 88.4 90.8 89.2 88.8 86.8 88.0 88.8 90.4 85.6 92.8 85.6 88.4 83.6 93.6
D.3.2 New-Only Fine-Tuning Table 4: Raw per-checkpoint results for New-Only Fine-Tuning. Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
20
Strong All Only New Norm 1000 Strong All Only New Norm 3000 Strong All Only New Norm 5000 Strong All Only Old Norm 1000 Strong All Only Old Norm 3000 Strong All Only Old Norm 5000 Strong First Only New Norm 1000 Strong First Only New Norm 3000 Strong First Only New Norm 5000 Strong First Only Old Norm 1000 Strong First Only Old Norm 3000 Strong First Only Old Norm 5000 Weak First Only Old Norm 500 Weak First Only Old Norm 1000 Weak First Only Old Norm 1500 Weak First Only Old Norm 2000 Weak First Only Old Norm 2500 Weak First Only Old Norm 3000 Weak First Only Old Norm 3500 Weak First Only Old Norm 4000 Weak First Only Old Norm 4500 Weak First Only Old Norm 5000 Weak First Only Old Norm 5500
0 0 0 15 1 0 0 0 0 34 25 2 46 47 44 42 19 20 8 4 4 0 0
0 0 0 7 0 2 0 0 0 32 13 6 45 40 32 19 9 21 15 16 10 6 4
0 0 0 0 0 0 0 0 0 0 0 0 5 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 1 0 5 5 4 16 5 3 11 8 12 12 9 5
0 0 0 11 0 0 0 0 0 16 0 0 30 18 11 7 1 2 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0 0 6 2 0 0 0 2 0 1 0
0 0 0 18 0 0 0 0 0 25 0 0 26 28 20 5 1 2 0 0 0 0 0
0 0 0 30 3 0 0 0 0 34 4 0 38 42 33 31 13 11 10 3 1 2 1
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
0.0 0.0 0.0 16.2 0.8 0.4 0.0 0.0 0.0 28.4 8.4 2.6 39.0 35.8 32.4 22.2 9.2 13.4 8.2 7.4 5.4 3.6 2.0
0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.4 0.0 2.0 4.0 1.6 8.8 2.8 1.2 4.4 3.2 5.6 4.8 4.0 2.0
0.0 0.0 0.0 32.4 1.6 0.8 0.0 0.0 0.0 56.4 16.8 3.2 74.0 70.0 56.0 41.6 17.2 22.4 13.2 9.2 6.0 3.2 2.0
D.3.3 Low-Learning-Rate and EWC Sweeps Table 5: Raw per-checkpoint results for Low-Learning-Rate and EWC Sweeps. Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
21
Filtered Fisher 1e13 Low LR 500 Filtered Fisher 1e13 Low LR 1000 Filtered Fisher 1e13 Low LR 1500 Filtered Fisher 1e13 Low LR 2000 Filtered Fisher 1e13 Low LR 2500 Filtered Fisher 1e13 Low LR 3000 Filtered Fisher 1e13 Low LR 3500 Filtered Fisher 1e13 Low LR 4000 Filtered Fisher 1e13 Low LR 4500 Filtered Fisher 1e13 Low LR 5000 Filtered Fisher 1e13 Low LR 5500 Filtered Fisher 1e13 Low LR 6000 Filtered Fisher 1e13 Low LR 6500 Filtered Fisher 1e13 Low LR 7000 Filtered Fisher 1e13 Low LR 7500 Filtered Fisher 1e13 Low LR 8000 Filtered Fisher 1e13 Low LR 8500 Filtered Fisher 1e13 Low LR 9000 Filtered Fisher 1e13 Low LR 9999 Filtered Fisher 1e14 Low LR 500 Filtered Fisher 1e14 Low LR 1000 Filtered Fisher 1e14 Low LR 1500 Filtered Fisher 1e14 Low LR 2000 Filtered Fisher 1e14 Low LR 2500 Filtered Fisher 1e14 Low LR 3000 Filtered Fisher 1e14 Low LR 3500 Filtered Fisher 1e14 Low LR 4000 Filtered Fisher 1e14 Low LR 4500 Filtered Fisher 1e14 Low LR 5000 Filtered Fisher 1e14 Low LR 5500 Filtered Fisher 1e14 Low LR 6000
44 44 42 43 41 40 42 45 45 43 41 45 45 43 44 46 46 42 41 43 41 42 40 44 44 44 44 46 42 43 43
49 48 49 49 48 48 49 49 48 49 49 50 48 49 49 49 49 49 49 49 48 49 48 49 48 49 48 49 48 48 49
19 14 14 17 13 18 19 16 11 11 15 17 14 16 16 19 21 14 16 14 16 8 17 16 18 11 13 14 15 16 6
21 21 16 18 17 18 14 16 23 21 13 16 16 17 16 18 17 19 19 21 22 17 23 20 18 16 19 18 13 15 15
41 47 46 44 45 46 43 46 46 44 46 42 48 45 44 47 46 47 46 46 44 44 44 42 45 46 48 45 45 45 46
25 29 24 28 22 26 22 29 30 23 27 20 26 28 30 26 25 21 24 26 28 27 27 24 24 26 22 30 29 22 26
38 43 37 40 41 36 41 37 43 39 39 44 39 41 39 43 41 38 39 40 42 37 36 38 41 40 42 40 38 37 43
46 47 45 45 44 47 46 44 46 48 48 47 47 43 46 48 48 49 47 47 46 44 48 47 45 46 48 45 45 43 46
0 0 0 0 0 2 0 0 1 0 0 1 1 0 0 0 0 0 0 1 0 0 0 0 0 1 0 0 1 0 0
14 13 9 13 17 21 12 18 8 17 16 23 16 15 15 12 13 15 15 15 13 7 11 15 12 11 9 12 18 10 17
59.4 61.2 56.4 59.4 57.6 60.4 57.6 60.0 60.2 59.0 58.8 61.0 60.0 59.4 59.8 61.6 61.2 58.8 59.2 60.4 60.0 55.0 58.8 59.0 59.0 58.0 58.6 59.8 58.8 55.8 58.2
31.6 30.8 25.2 30.4 27.6 34.0 26.8 31.6 29.2 28.8 28.4 30.8 29.2 30.4 30.8 30.0 30.4 27.6 29.6 30.8 31.6 23.6 31.2 30.0 28.8 26.0 25.2 29.6 30.4 25.2 25.6
87.2 91.6 87.6 88.4 87.6 86.8 88.4 88.4 91.2 89.2 89.2 91.2 90.8 88.4 88.8 93.2 92.0 90.0 88.8 90.0 88.4 86.4 86.4 88.0 89.2 90.0 92.0 90.0 87.2 86.4 90.8
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
22
Filtered Fisher 1e14 Low LR 6500 Filtered Fisher 1e14 Low LR 7000 Filtered Fisher 1e14 Low LR 7500 Filtered Fisher 1e15 Low LR 500 Filtered Fisher 1e15 Low LR 1000 Filtered Fisher 1e15 Low LR 1500 Filtered Fisher 1e15 Low LR 2000 Filtered Fisher 1e15 Low LR 2500 Filtered Fisher 1e15 Low LR 3000 Filtered Fisher 1e15 Low LR 3500 Filtered Fisher 1e15 Low LR 4000 Filtered Fisher 1e15 Low LR 4500 Filtered Fisher 1e15 Low LR 5000 Filtered Fisher 1e15 Low LR 5500 Filtered Fisher 1e15 Low LR 6000 Filtered Fisher 1e15 Low LR 6500 Filtered Fisher 1e15 Low LR 7000 Filtered Fisher 1e15 Low LR 7500 Filtered Fisher 1e15 Low LR 8000 Filtered Fisher 1e15 Low LR 8500 Filtered Fisher 1e15 Low LR 9000 Filtered Fisher 1e15 Low LR 9999 Filtered Fisher 1e8 Low LR 500 Filtered Fisher 1e8 Low LR 1000 Filtered Fisher 1e8 Low LR 1500 Filtered Fisher 1e8 Low LR 2000 Filtered Fisher 1e8 Low LR 2500 Filtered Fisher 1e8 Low LR 3000 Filtered Fisher 1e8 Low LR 3500 Filtered Fisher 1e8 Low LR 4000 Filtered Fisher 1e8 Low LR 4500 Filtered Fisher 1e8 Low LR 5000 Filtered Fisher 1e8 Low LR 5500 Filtered Fisher 1e8 Low LR 6000 Filtered Fisher 1e8 Low LR 6500
42 45 46 40 41 45 44 43 43 46 41 42 42 46 48 45 44 45 42 40 43 42 39 45 41 42 40 42 40 41 43 41 44 41 42
49 49 49 48 50 49 49 49 49 50 49 48 48 49 49 49 49 48 49 49 49 48 49 48 49 49 49 50 49 50 49 49 48 49 50
19 15 13 19 10 12 17 15 15 18 12 10 15 17 14 19 8 12 21 11 10 13 13 9 5 2 2 2 4 2 2 8 4 3 8
25 20 18 22 15 18 19 16 15 16 17 17 20 21 13 23 20 18 21 16 20 19 18 26 22 18 18 14 14 13 16 10 12 9 8
46 47 43 43 44 45 44 45 46 45 40 46 45 45 49 47 44 46 46 45 47 45 48 42 44 46 46 45 47 47 47 46 48 46 42
24 26 30 22 28 25 25 22 27 22 28 28 23 28 22 24 25 28 20 27 26 27 32 30 25 28 23 29 24 28 25 27 23 28 23
45 38 38 39 44 44 39 41 36 41 38 38 44 40 36 39 41 43 43 39 44 38 37 37 43 38 37 40 40 37 41 38 43 41 39
46 45 44 45 47 43 46 48 42 45 46 48 47 49 47 46 45 48 46 45 45 46 48 46 46 48 48 48 50 46 47 50 48 48 47
0 2 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 1 0 0 2 0 1 0 1 0 1 1 3 2 3 2 2 0 3
10 14 11 10 12 15 11 10 11 11 11 13 15 14 11 14 10 10 13 15 16 15 15 11 6 4 3 2 0 2 0 1 0 0 0
61.2 60.2 58.4 57.6 58.2 59.4 58.8 57.8 56.8 58.8 56.4 58.0 59.8 61.8 57.8 61.2 57.2 59.8 60.2 57.4 60.4 58.6 60.0 58.8 56.4 55.0 53.4 54.6 54.2 53.6 54.6 54.4 54.4 53.0 52.4
31.2 30.8 28.8 29.2 26.0 28.4 28.8 25.2 27.2 26.8 27.2 27.2 29.2 32.0 24.0 32.0 25.2 27.6 30.0 27.6 29.6 29.6 31.6 30.4 23.6 20.8 18.8 19.2 18.0 18.8 18.4 19.2 16.4 16.0 16.8
91.2 89.6 88.0 86.0 90.4 90.4 88.8 90.4 86.4 90.8 85.6 88.8 90.4 91.6 91.6 90.4 89.2 92.0 90.4 87.2 91.2 87.6 88.4 87.2 89.2 89.2 88.0 90.0 90.4 88.4 90.8 89.6 92.4 90.0 88.0
23
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Filtered Fisher 1e8 Low LR Filtered Fisher 1e8 Low LR Filtered Fisher 1e8 Low LR Filtered Fisher 1e8 Low LR Filtered Fisher 1e8 Low LR Filtered Fisher 1e8 Low LR Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e10 Low LR Fisher 1e12
7000 7500 8000 8500 9000 9999 500 1000 1500 2000 2500 3000 3500 4000 4500 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9999 500
46 43 46 43 44 44 42 44 44 41 0 0 0 0 0 43 40 46 37 46 45 42 44 48 43 45 44 45 45 42 38 42 44 44 42
50 49 50 47 50 49 49 50 50 49 0 0 0 0 0 48 49 49 49 49 48 48 49 49 49 50 49 49 49 49 49 48 49 48 49
1 4 5 1 6 6 8 7 8 6 0 0 0 0 0 12 11 10 12 5 5 9 13 6 4 6 4 1 6 8 3 6 5 4 19
10 7 8 6 8 8 19 18 14 17 0 0 0 0 0 17 24 16 18 17 18 23 16 23 18 19 29 14 18 18 21 16 15 19 17
42 45 39 42 43 44 46 45 47 43 0 0 0 0 0 46 45 45 48 47 45 46 45 46 44 45 45 47 44 43 45 47 46 47 45
26 25 20 19 21 21 17 24 29 33 0 0 0 0 0 23 24 29 29 26 32 27 32 23 22 27 28 28 26 24 31 28 29 26 28
43 38 45 42 40 37 36 39 40 46 0 0 0 0 0 42 39 42 42 42 38 43 41 41 37 41 41 39 41 40 41 39 44 39 35
50 46 42 47 46 44 47 42 46 45 0 0 0 0 0 46 46 46 47 48 46 47 48 47 47 46 46 46 46 47 48 47 46 47 43
3 1 0 0 1 1 1 2 0 0 0 0 0 0 0 1 0 3 1 1 0 1 2 1 0 0 0 1 1 2 0 3 0 1 0
1 0 0 0 0 0 14 16 10 3 0 0 0 0 0 17 11 14 9 9 14 17 11 11 10 9 6 14 13 13 18 8 17 9 12
54.4 51.6 51.0 49.4 51.8 50.8 55.8 57.4 57.6 56.6 0.0 0.0 0.0 0.0 0.0 59.0 57.8 60.0 58.4 58.0 58.2 60.6 60.2 59.0 54.8 57.6 58.4 56.8 57.8 57.2 58.8 56.8 59.0 56.8 58.0
16.4 14.8 13.2 10.4 14.4 14.4 23.6 26.8 24.4 23.6 0.0 0.0 0.0 0.0 0.0 28.0 28.0 28.8 27.6 23.2 27.6 30.8 29.6 25.6 21.6 24.4 26.8 23.2 25.6 26.0 29.2 24.4 26.4 23.6 30.4
92.4 88.4 88.8 88.4 89.2 87.2 88.0 88.0 90.8 89.6 0.0 0.0 0.0 0.0 0.0 90.0 87.6 91.2 89.2 92.8 88.8 90.4 90.8 92.4 88.0 90.8 90.0 90.4 90.0 88.4 88.4 89.2 91.6 90.0 85.6
24
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR
1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9500 9999 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000
41 42 47 42 45 43 43 44 46 44 43 43 41 39 45 41 45 42 45 48 43 45 42 46 42 43 43 43 43 44 44 43 43 44 42
48 49 49 49 50 49 47 48 49 49 49 49 49 50 48 50 49 49 48 48 49 49 48 48 48 49 49 49 49 48 50 49 49 49 49
11 19 15 15 19 17 17 17 13 14 17 16 19 15 17 13 18 16 18 13 17 15 16 21 13 11 15 19 13 14 12 16 13 17 15
20 18 17 20 12 17 21 17 14 23 19 20 17 24 23 22 17 19 19 15 21 18 17 18 19 16 19 16 17 20 16 15 16 16 17
46 46 47 45 47 45 45 47 46 46 43 46 43 44 45 46 47 43 46 45 47 45 45 45 47 44 46 46 43 44 45 45 45 44 43
26 21 20 22 25 22 24 31 23 21 26 26 28 23 29 28 27 25 27 26 25 27 23 25 24 26 32 26 30 27 27 29 27 26 20
43 38 39 39 42 36 36 42 40 42 38 40 41 39 38 38 39 37 41 40 39 42 35 39 42 42 41 41 38 43 43 44 41 41 39
47 47 47 44 45 47 45 45 47 50 45 46 45 48 48 47 45 45 47 43 45 46 46 44 45 47 46 46 46 45 46 45 45 47 49
0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 1 0 0 2 0 0 0
10 19 14 13 12 12 16 14 17 9 12 17 12 13 13 15 9 9 13 14 15 16 10 10 9 16 13 11 18 12 14 11 9 12 11
58.4 59.8 59.0 58.0 59.4 57.6 58.8 61.0 59.0 59.6 58.4 60.6 59.0 59.0 61.2 60.0 59.2 57.0 60.8 58.4 60.2 60.6 56.6 59.2 57.8 58.8 60.8 59.4 59.6 59.4 59.4 59.8 57.6 59.2 57.0
26.8 30.8 26.4 28.4 27.2 27.2 31.2 31.6 26.8 26.8 29.6 31.6 30.4 30.0 32.8 31.2 28.4 27.6 30.8 27.2 31.2 30.4 26.8 29.6 26.0 27.6 31.6 28.8 31.6 29.2 27.6 29.2 26.0 28.4 25.2
90.0 88.8 91.6 87.6 91.6 88.0 86.4 90.4 91.2 92.4 87.2 89.6 87.6 88.0 89.6 88.8 90.0 86.4 90.8 89.6 89.2 90.8 86.4 88.8 89.6 90.0 90.0 90.0 87.6 89.6 91.2 90.4 89.2 90.0 88.8
25
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e12 Low LR Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR
8500 9000 9999 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9999 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500
40 43 43 43 44 45 44 38 43 43 45 42 47 42 45 43 42 45 45 45 42 42 42 43 42 45 45 43 44 45 44 41 42 46 43
49 48 48 49 49 48 49 49 49 49 49 49 49 49 49 49 49 47 49 48 49 49 48 49 49 49 48 49 47 50 49 49 48 49 49
16 8 15 13 18 21 16 15 18 14 17 13 13 14 16 12 16 12 16 17 15 16 16 9 20 23 12 18 11 14 21 16 17 13 18
19 22 22 22 17 25 19 20 21 21 18 14 15 18 13 21 19 22 20 20 22 22 26 15 16 17 22 16 22 15 18 22 14 17 19
46 46 49 47 45 44 46 43 48 45 47 45 48 45 46 44 47 46 47 42 44 46 43 42 46 46 47 45 44 45 46 44 45 46 46
24 21 27 31 28 30 21 28 24 24 25 24 25 31 22 26 29 24 26 25 26 25 27 25 28 26 25 19 31 27 27 26 29 25 25
44 36 43 38 46 41 37 43 46 44 38 41 44 40 36 38 42 42 39 38 39 43 45 41 42 39 37 43 37 40 35 40 39 42 43
46 44 44 47 43 46 46 45 46 44 46 46 46 47 43 47 46 47 46 46 46 49 45 46 44 45 43 45 48 47 45 44 47 46 41
0 1 0 1 0 0 1 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 1
12 14 16 9 11 8 13 15 13 16 12 12 10 12 13 10 12 17 12 12 14 14 11 14 10 8 8 13 13 12 13 10 9 13 9
59.2 56.6 61.4 60.0 60.2 61.6 58.4 59.2 61.6 60.0 59.4 57.2 59.4 59.6 56.8 58.0 60.4 60.4 60.0 58.6 59.4 61.2 60.6 56.8 59.4 59.6 57.4 58.4 59.4 59.0 59.6 58.4 58.0 59.4 58.8
28.4 26.4 32.0 30.4 29.6 33.6 28.0 31.2 30.4 30.0 28.8 25.2 25.2 30.0 26.0 27.6 30.4 30.0 29.6 29.6 30.8 30.8 32.0 25.2 29.6 29.6 26.8 26.8 30.8 27.2 31.6 29.6 27.6 27.2 28.8
90.0 86.8 90.8 89.6 90.8 89.6 88.8 87.2 92.8 90.0 90.0 89.2 93.6 89.2 87.6 88.4 90.4 90.8 90.4 87.6 88.0 91.6 89.2 88.4 89.2 89.6 88.0 90.0 88.0 90.8 87.6 87.2 88.4 91.6 88.8
26
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e15 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e6 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e8 Low LR Fisher 1e9 Low LR
7000 7500 8000 8500 9000 9999 1000 1500 2000 2500 3000 3500 4000 4500 4999 500 1000 1500 2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000 8500 9000 9999 500
45 44 45 44 43 46 43 44 45 39 44 42 45 47 41 43 37 45 46 44 44 42 42 40 44 43 41 44 42 47 44 47 45 40 41
48 49 49 49 49 48 48 50 49 50 50 48 50 48 48 49 49 50 49 49 48 50 49 50 50 49 50 49 50 50 49 49 48 48 49
18 12 9 15 15 13 9 8 1 1 3 4 0 2 3 14 7 4 2 0 4 2 4 3 5 3 6 1 6 4 3 11 7 3 20
17 20 15 20 11 15 22 19 20 18 19 16 18 11 7 22 20 18 21 21 20 13 16 13 12 10 12 8 8 5 10 9 10 8 16
45 45 46 45 44 43 44 45 46 46 45 45 48 45 45 44 46 43 46 47 47 46 47 45 45 47 46 45 46 43 39 41 37 45 46
28 19 25 28 29 30 27 29 21 21 27 25 22 25 25 26 27 25 29 29 25 28 23 26 23 27 27 27 24 24 31 24 17 21 19
39 39 41 40 35 40 41 34 41 37 36 37 36 42 41 41 41 37 35 41 37 44 40 40 35 39 38 43 36 40 41 38 37 38 39
44 46 47 46 47 45 46 49 48 48 47 48 49 45 45 45 47 46 48 49 49 48 47 49 47 47 49 46 48 45 49 49 47 45 45
1 0 0 0 0 0 0 1 1 5 2 2 0 1 3 1 0 1 0 2 3 4 1 0 0 3 2 1 4 4 2 1 1 1 1
12 10 9 13 17 12 11 8 4 3 0 1 3 0 1 11 9 7 5 2 0 1 0 1 0 0 0 0 0 0 0 1 0 0 14
59.4 56.8 57.2 60.0 58.0 58.4 58.2 57.4 55.2 53.6 54.6 53.6 54.2 53.2 51.8 59.2 56.6 55.2 56.2 56.8 55.4 55.6 53.8 53.4 52.2 53.6 54.2 52.8 52.8 52.4 53.6 54.0 49.8 49.8 58.0
30.4 24.4 23.2 30.4 28.8 28.0 27.6 26.0 18.8 19.2 20.4 19.2 17.2 15.6 15.6 29.6 25.2 22.0 22.8 21.6 20.8 19.2 17.6 17.2 16.0 17.2 18.8 14.8 16.8 14.8 18.4 18.4 14.0 13.2 28.0
88.4 89.2 91.2 89.6 87.2 88.8 88.8 88.8 91.6 88.0 88.8 88.0 91.2 90.8 88.0 88.8 88.0 88.4 89.6 92.0 90.0 92.0 90.0 89.6 88.4 90.0 89.6 90.8 88.8 90.0 88.8 89.6 85.6 86.4 88.0
27
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Fisher 1e9 Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR Strong First Only Low LR
1000 1500 2000 2500 3000 3500 4500 4999 5500 6000 6500 7000 7500 8000 8500 9000 9999 100 200 300 400 500 600 700 800 900 1000 1500 2000 2500 3000 3500 4000 4500
43 46 43 41 45 45 41 44 43 45 46 46 42 42 44 44 42 44 42 43 42 42 45 45 40 44 43 44 43 43 43 42 46 43
49 49 49 49 49 50 50 48 49 50 50 49 49 49 49 48 50 50 50 49 49 48 50 50 49 49 50 49 49 49 49 50 48 48
13 6 6 7 4 5 3 1 2 4 6 5 10 6 3 6 5 20 16 17 17 16 12 16 11 10 14 5 4 1 0 2 3 2
19 21 19 20 20 17 12 27 15 16 17 19 17 15 13 13 14 27 16 16 27 31 21 20 20 19 19 23 25 25 24 20 11 12
48 47 43 49 45 44 47 47 44 42 47 45 43 46 45 45 43 47 46 40 46 44 43 47 45 47 46 45 45 46 46 46 46 43
25 30 25 25 25 23 24 17 28 27 22 22 29 25 27 23 25 27 21 23 28 28 25 31 29 33 30 25 26 29 28 25 24 24
34 39 42 39 38 39 36 38 40 37 39 36 41 39 42 43 37 41 38 43 44 43 37 41 41 44 41 41 37 32 38 40 39 38
48 48 47 48 48 48 48 47 48 47 48 49 48 47 49 49 47 46 47 49 43 42 44 50 48 46 47 46 49 47 47 47 50 47
1 1 0 1 1 3 3 2 1 2 3 2 1 2 3 1 2 1 0 0 0 0 0 2 0 2 3 0 2 1 2 3 0 3
7 5 5 1 3 4 2 5 3 7 1 4 5 3 2 1 2 11 12 17 7 15 13 8 14 13 7 8 5 3 2 2 1 0
57.4 58.4 55.8 56.0 55.6 55.6 53.2 55.2 54.6 55.4 55.8 55.4 57.0 54.8 55.4 54.6 53.4 62.8 57.6 59.4 60.6 61.8 58.0 62.0 59.4 61.4 60.0 57.2 57.0 55.2 55.8 55.4 53.6 52.0
26.0 25.2 22.0 21.6 21.2 20.8 17.6 20.8 19.6 22.4 19.6 20.8 24.8 20.4 19.2 17.6 19.2 34.4 26.0 29.2 31.6 36.0 28.4 30.8 29.6 30.8 29.2 24.4 24.8 23.6 22.4 20.8 15.6 16.4
88.8 91.6 89.6 90.4 90.0 90.4 88.8 89.6 89.6 88.4 92.0 90.0 89.2 89.2 91.6 91.6 87.6 91.2 89.2 89.6 89.6 87.6 87.6 93.2 89.2 92.0 90.8 90.0 89.2 86.8 89.2 90.0 91.6 87.6
D.3.4 LIBERO-10 Replay Scope Table 6: Raw per-checkpoint results for LIBERO-10 Replay Scope. Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
28
Strong First Libero 10 Mix 500 Strong First Libero 10 Mix 1000 Strong First Libero 10 Mix 1500 Strong First Libero 10 Mix 2000 Strong First Libero 10 Mix 2500 Strong First Libero 10 Mix 3000 Strong First Libero 10 Mix 3500 Strong First Libero 10 Mix 4000 Strong First Libero 10 Mix 4500 Strong First Libero 10 Mix 5000 Strong First Libero 10 Mix 5500 Strong First Libero 10 Mix 6000 Strong First Libero 10 Mix 6500 Strong First Libero 10 Mix 7000 Strong First Libero 10 Mix 7500 Strong First Libero 10 Mix 8000 Strong First Libero 10 Mix 8500 Strong First Libero 10 Mix 9000 Strong First Libero 10 Mix 9500 Strong First Libero 10 Mix 9999
45 43 41 45 44 45 45 42 47 42 42 43 46 44 47 48 46 44 46 45
49 47 48 49 47 49 46 47 47 46 46 49 48 50 47 50 46 48 48 47
21 22 20 16 16 25 21 24 30 22 26 18 22 19 29 20 17 22 18 17
14 12 21 20 27 29 19 21 22 28 17 27 25 21 29 31 14 25 20 19
46 42 44 44 44 46 43 45 42 39 40 44 44 40 41 45 44 50 46 45
31 29 28 29 31 34 25 24 22 33 22 31 27 21 34 27 24 34 27 23
35 40 44 43 41 42 34 34 32 36 39 38 38 40 43 43 42 38 41 30
48 48 46 49 43 48 45 46 46 47 48 43 47 49 47 48 46 46 45 45
1 0 0 0 2 1 0 0 1 0 0 0 1 2 0 3 2 1 1 0
24 23 15 19 19 17 26 13 23 19 21 18 18 17 26 25 21 21 10 22
62.8 61.2 61.4 62.8 62.8 67.2 60.8 59.2 62.4 62.4 60.2 62.2 63.2 60.6 68.6 68.0 60.4 65.8 60.4 58.6
36.4 34.4 33.6 33.6 38.0 42.4 36.4 32.8 39.2 40.8 34.4 37.6 37.2 32.0 47.2 42.4 31.2 41.2 30.4 32.4
89.2 88.0 89.2 92.0 87.6 92.0 85.2 85.6 85.6 84.0 86.0 86.8 89.2 89.2 90.0 93.6 89.6 90.4 90.4 84.8
D.3.5 Collected-Task Replay and EWC Table 7: Raw per-checkpoint results for Collected-Task Replay and EWC. Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
29
Strong First Libero10 Collected Mix 500 Strong First Libero10 Collected Mix 1000 Strong First Libero10 Collected Mix 1500 Strong First Libero10 Collected Mix 2000 Strong First Libero10 Collected Mix 2500 Strong First Libero10 Collected Mix 3000 Strong First Libero10 Collected Mix 3500 Strong First Libero10 Collected Mix 4000 Strong First Libero10 Collected Mix 4500 Strong First Libero10 Collected Mix 5000 Strong First Libero10 Collected Mix 5500 Strong First Libero10 Collected Mix 6000 Strong First Libero10 Collected Mix 6500 Strong First Libero10 Collected Mix 7000 Strong First Libero10 Collected Mix 7500 Strong First Libero10 Collected Mix 8000 Strong First Libero10 Collected Mix 8500 Strong First Libero10 Collected Mix 9000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 1000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 1500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 2000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 2500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 3000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 3500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 4000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 4500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 5000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 5500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 6000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 6500
44 44 47 47 46 42 48 35 42 40 37 24 22 16 16 14 16 14 44 44 45 42 41 40 15 14 31 34 37 42 36
46 39 32 24 20 29 28 26 22 21 18 22 18 15 13 20 13 16 49 48 49 50 50 49 2 20 34 32 26 22 23
18 24 15 19 20 17 14 16 20 14 34 31 32 24 22 32 29 24 4 6 4 5 17 13 0 0 0 0 0 3 4
18 31 18 22 30 26 28 24 30 18 23 24 23 21 22 19 28 25 17 22 17 19 20 10 0 0 1 5 8 7 4
41 43 41 31 10 10 8 3 3 1 2 0 0 0 0 0 0 0 47 45 46 44 46 42 14 13 26 28 35 31 32
23 33 24 32 34 18 27 36 22 27 26 24 28 23 29 27 28 26 30 25 26 26 30 25 0 2 1 2 5 5 11
37 35 42 38 35 28 30 19 27 19 15 18 9 6 7 4 3 6 36 42 40 41 34 40 18 8 14 15 18 16 21
47 48 43 43 37 39 31 37 39 23 33 22 24 14 14 18 13 14 48 45 45 47 46 49 12 7 17 18 22 16 15
0 0 1 1 2 0 2 3 0 4 1 0 1 1 1 1 0 0 0 2 0 2 0 1 0 0 0 0 0 1 0
16 19 13 17 16 15 16 18 24 14 17 24 21 22 18 19 20 20 10 13 10 18 7 1 0 0 0 0 0 0 0
58.0 63.2 55.2 54.8 50.0 44.8 46.4 43.4 45.8 36.2 41.2 37.8 35.6 28.4 28.4 30.8 30.0 29.0 57.0 58.4 56.4 58.8 58.2 54.0 12.2 12.8 24.8 26.8 30.2 28.6 29.2
30.0 42.8 28.4 36.4 40.8 30.4 34.8 38.8 38.4 30.8 40.4 41.2 42.0 36.4 36.8 39.2 42.0 38.0 24.4 27.2 22.8 28.0 29.6 20.0 0.0 0.8 0.8 2.8 5.2 6.4 7.6
86.0 83.6 82.0 73.2 59.2 59.2 58.0 48.0 53.2 41.6 42.0 34.4 29.2 20.4 20.0 22.4 18.0 20.0 89.6 89.6 90.0 89.6 86.8 88.0 24.4 24.8 48.8 50.8 55.2 50.8 50.8
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
30
Strong First Libero10 Collected Mix Filtered Fisher 1e10 7000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 7500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 8000 Strong First Libero10 Collected Mix Filtered Fisher 1e10 8500 Strong First Libero10 Collected Mix Filtered Fisher 1e10 9000 Strong First Libero10 Collected Mix Filtered Fisher 1e9 500 Strong First Libero10 Collected Mix Filtered Fisher 1e9 1000 Strong First Libero10 Collected Mix Filtered Fisher 1e9 1500 Strong First Libero10 Collected Mix Filtered Fisher 1e9 2000 Strong First Libero10 Collected Mix Filtered Fisher 1e9 2500 Strong First Libero10 Collected Mix Filtered Fisher 1e9 3000 Strong First Libero10 Collected Mix Filtered Fisher 1e9 3500 Strong First Libero10 Collected Mix Filtered Fisher 1e9 4000 Strong First Libero10 Collected Mix Filtered Fisher 1e9 4500 Strong First Libero10 Collected Mix Fisher 1e10 500 Strong First Libero10 Collected Mix Fisher 1e10 1000 Strong First Libero10 Collected Mix Fisher 1e10 1500 Strong First Libero10 Collected Mix Fisher 1e10 2000 Strong First Libero10 Collected Mix Fisher 1e10 2500 Strong First Libero10 Collected Mix Fisher 1e10 3000 Strong First Libero10 Collected Mix Fisher 1e10 3500 Strong First Libero10 Collected Mix Fisher 1e10 4000 Strong First Libero10 Collected Mix Fisher 1e10 4500 Strong First Libero10 Collected Mix Fisher 1e10 5000 Strong First Libero10 Collected Mix Fisher 1e10 5500 Strong First Libero10 Collected Mix Fisher 1e10 6000 Strong First Libero10 Collected Mix Fisher 1e10 6500 Strong First Libero10 Collected Mix Fisher 1e10 7000 Strong First Libero10 Collected Mix Fisher 1e10 7500 Strong First Libero10 Collected Mix Fisher 1e10 8000 Strong First Libero10 Collected Mix Fisher 1e10 8500 Strong First Libero10 Collected Mix Fisher 1e10 9000 Strong First Libero10 Collected Mix Fisher 1e8 500 Strong First Libero10 Collected Mix Fisher 1e8 1000 Strong First Libero10 Collected Mix Fisher 1e8 1500
43 39 33 40 40 45 42 42 44 44 45 47 40 42 46 40 45 47 46 42 39 19 40 42 40 41 40 42 42 40 41 41 48 47 17
30 29 28 32 28 47 49 49 46 42 40 39 33 32 49 49 49 49 49 48 37 13 24 35 33 31 22 20 25 28 23 15 49 49 15
3 5 3 5 1 11 7 14 13 9 4 11 10 10 9 15 12 4 5 5 2 0 0 4 2 5 4 4 5 8 14 8 18 1 0
13 4 6 9 6 21 22 23 21 27 17 12 8 14 19 15 16 16 22 21 5 0 1 5 6 8 5 7 12 14 12 11 18 12 0
31 28 28 27 30 43 43 43 42 38 43 42 37 37 44 46 47 48 46 44 27 7 28 35 37 32 32 32 30 28 28 32 47 46 6
12 15 8 13 13 27 23 25 25 30 25 20 22 20 26 24 22 27 33 30 1 2 2 2 8 1 14 9 8 11 13 12 27 26 1
22 20 20 18 15 40 47 43 43 37 36 28 17 16 40 39 38 44 48 43 38 5 20 22 25 20 23 31 24 23 23 23 42 38 1
18 15 9 17 16 48 48 46 44 48 45 42 43 43 45 41 37 42 49 45 36 3 12 22 27 34 27 28 28 19 21 10 46 43 3
0 0 1 0 0 2 0 1 6 0 3 3 2 3 0 1 0 1 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0
2 1 0 3 1 10 8 6 5 5 6 11 9 11 12 12 14 9 6 2 0 0 0 0 0 0 0 0 1 1 0 1 17 2 0
34.8 31.2 27.2 32.8 30.0 58.8 57.8 58.4 57.8 56.0 52.8 51.0 44.2 45.6 58.0 56.4 56.0 57.4 60.8 56.2 37.0 9.8 25.4 33.4 35.6 34.4 33.4 34.6 35.0 34.4 35.0 30.6 62.4 52.8 8.6
12.0 10.0 7.2 12.0 8.4 28.4 24.0 27.6 28.0 28.4 22.0 22.8 20.4 23.2 26.4 26.8 25.6 22.8 26.4 23.6 3.2 0.8 1.2 4.4 6.4 5.6 9.2 8.0 10.4 13.6 15.6 12.8 32.0 16.4 0.4
57.6 52.4 47.2 53.6 51.6 89.2 91.6 89.2 87.6 83.6 83.6 79.2 68.0 68.0 89.6 86.0 86.4 92.0 95.2 88.8 70.8 18.8 49.6 62.4 64.8 63.2 57.6 61.2 59.6 55.2 54.4 48.4 92.8 89.2 16.8
Experiment
Step T0 T1 T2 T3 T4 T5 T6 T7 T8 T9 Overall % Collected % Retained %
Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8 Strong First Libero10 Collected Mix Fisher 1e8
2000 2500 3000 3500 4000 4500 5000 5500 6000 6500 7000 7500 8000
9 5 8 2 5 1 12 2 20 5 23 9 8 1 14 0 17 3 5 1 11 0 30 14 15 3
0 0 0 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 1 0 1 0 1 1
4 0 3 0 3 0 1 0 1 5 6 11 1 1 0 2 1 5 0 2 0 5 0 9 0 12
0 1 1 0 3 1 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0
0 0 0 0 0 0 0 0 0 0 0 0 0
3.6 2.8 2.0 3.0 6.8 10.0 2.2 3.4 5.2 1.8 3.2 10.8 6.2
0.0 0.0 0.0 0.0 2.0 4.4 0.4 1.2 2.0 1.2 2.0 4.0 5.2
7.2 5.6 4.0 6.0 11.6 15.6 4.0 5.6 8.4 2.4 4.4 17.6 7.2
31