Boosting Brain-to-Image Decoding with TRIBE v2 Data Augmentation Yohann Benchetrit , Marlène Careil , Simon Dahan , Hubert Banville , Stéphane d’Ascoli , Jean-Rémi King
arXiv:2606.06345v1 [cs.AI] 4 Jun 2026
Meta AI Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes. To address this issue, we investigate whether and when brain decoding can be boosted by augmenting small fMRI datasets with synthetic data generated by a pretrained model of fMRI responses to stimuli. We use TRIBE v2, a large encoding model pretrained on more than 1000 hours of fMRI responses to video, audio and language. For each dataset, we evaluate systematic grids that show how the performance of image decoders varies with the amount of synthetic data used for training. Our results, based on two datasets (the 7T fMRI Natural Scenes Dataset and 3T fMRI BOLD5000), show up to 68% improvement in Top-10 image-retrieval accuracy compared to decoders trained only on real data. Importantly, the proportion of augmented data required to reach a given image decoding performance needs to be adjusted depending on the data source. Surprisingly, image decoders trained exclusively on synthetic fMRI can perform above chance in some settings, suggesting that TRIBE v2 can support zero-shot brain-to-image decoding. Together, these results show how large-scale models of the fMRI responses to sight, sound and language may provide a foundation to improve the data efficiency for image decoding. Date: June 5, 2026 Correspondence: [email protected]
1
Introduction
Brain-to-image decoding has progressed rapidly thanks to the pairing of fMRI with powerful pretrained visual representations and generative models (Scotti et al., 2024; Ozcelik and VanRullen, 2023; Chen et al., 2023; Benchetrit et al., 2024; Shen et al., 2019). Yet its practical reach remains restricted by the amount (and quality) of brain imaging data: High-performing decoders typically require thousands of stimulus-response pairs from the same individual, collected over many scanning sessions (Banville et al., 2025). For example, one of the most popular datasets for fMRI-to-image reconstruction is the Natural Scenes Dataset (NSD), which contains 30-40 hours of 7T fMRI per subject (Allen et al., 2022). In comparison, most neuroscience studies are far smaller, noisier, and heterogeneous e.g. (Chang et al., 2019; Shen et al., 2019; Hebart et al., 2023). This discrepancy creates a main bottleneck: modern decoding methods are strongest precisely in the data regime that few laboratories can afford to collect. Data augmentation is the standard response to data scarcity in machine learning, but fMRI presents an unusual challenge. Classical augmentations such as cropping, flipping, or color jitter do not apply trivially to neural recordings, because neural data is inherently tied to the fixed biological architecture of the subject rather than the fluid, translation-invariant structure of a 2D image grid. In addition, ‘model-free’ alternatives like noise injection or trial-mixing strategies like MixCo (Scotti et al., 2023, 2024) can regularize a decoder and improve performance, but they operate on neural responses that have already been measured and therefore do not provide responses for new external stimuli. Here, we investigate an alternative – model-based and stimulus-conditioned – data augmentation: we leverage a pretrained encoding model to predict how the brain would respond to new stimuli, then use those predicted responses as additional training data for the inverse problem i.e., decoding. We benchmark this approach on fMRI-to-image decoding using TRIBE v2 (d’Ascoli et al., 2026), a stateof-the-art foundation model trained to predict fMRI responses to video, audio, and language. Although
1
TRIBE v2 was not trained specifically on static image stimuli, its visual pathway can be applied to images by treating them as short static videos. We use TRIBE v2 to synthesize fMRI responses for novel images, train fMRI-to-image decoders on controlled mixtures of real (i.e., measured) and synthetic fMRI, and evaluate retrieval performance on a held-out set of real fMRI data. Beyond investigating whether synthetic fMRI helps image decoding, we here aim to resolve a more practical question: under what conditions and data regimes can an fMRI-to-image decoder leverage synthetic data? By varying two quantities – the fraction of real fMRI and the amount of synthetic, TRIBE-generated fMRI added relative to that real subset – we show that using synthetic fMRI data can substantially improve image decoding in low-to-medium data regimes, but excessive synthetic data can saturate or even hurt performance. We find that the best real-to-synthetic ratio differs across datasets and decoders. This leads to three main contributions: 1. We introduce a model-based, image-conditioned data augmentation pipeline for fMRI-to-image decoding, using TRIBE v2’s visual pathway to synthesize fMRI responses to novel images. 2. We apply this protocol to the Natural Scenes Dataset and BOLD5000, and grid how linear and deep decoders respond to different real-to-synthetic training ratios. In low-data regimes, adding TRIBEsynthetic responses improves Top-10 image-retrieval accuracy by up to 68%. We further show that the same augmentation strategy can improve an fMRI-to-image reconstruction model (Careil et al., 2025). 3. We show that TRIBE augmentation is not plug-and-play: its benefits depend on the dataset and decoder capacity, and can saturate, making careful calibration essential for using synthetic fMRI effectively.
2
Related Work
Brain decoding. Brain decoding aims to infer perceived or imagined content from neural activity. Early fMRI work decoded simple visual features and natural image identity with linear and Bayesian models (Kamitani and Tong, 2005; Kay et al., 2008; Naselaris et al., 2009). Recent systems use pretrained visual representations and diffusion models to retrieve or reconstruct images from brain activity, including MindEye (Scotti et al., 2023), MindEye2 (Scotti et al., 2024), Brain-Diffuser (Ozcelik and VanRullen, 2023), MindVis (Chen et al., 2023), and DynaDiff (Careil et al., 2025). As shown in this prior work, these advances make the data bottleneck more visible: strong decoders usually require many subject-specific samples. Where MindEye2 addresses this challenge through multi-subject pretraining, we explore a complementary approach: using image-conditioned fMRI synthesis to expand the decoder’s training set without acquiring additional real fMRI. fMRI encoding models. Encoding models predict neural responses from stimuli and have long served as a bridge between computational neuroscience and machine learning (Kay et al., 2008; Naselaris et al., 2011; St-Yves et al., 2023). Modern encoding models use deep visual features and can predict responses across the visual cortex (St-Yves et al., 2023; Conwell et al., 2023). TRIBE v2 extends this line of research to a large-scale, tri-modal model trained on over 1,000 hours of fMRI from 720 subjects (d’Ascoli et al., 2026). Importantly, we use TRIBE v2 not as a forward model, but as a generator of synthetic fMRI responses for the inverse task, i.e., decoding. Synthetic data for neuroimaging and machine learning. Synthetic data augmentation improves data efficiency in image classification, medical imaging, and language modeling (Azizi et al., 2023; He et al., 2023; Fernandez et al., 2022; Scotti et al., 2023, 2024). In fMRI, augmentation typically relies on noise injection, trial manipulation, or generative models trained on the target data (Nguyen et al., 2023; Wang et al., 2023). Our setting is different: the synthetic samples are conditioned on external stimuli and generated by a pretrained brain model without fitting a new fMRI generator on the target decoding dataset.
2
Unseen COCO images
TRIBE v2 image → fMRI
Synthetic training pairs a × pN
Train decoder fMRI → DINOv2-small (1 + a)pN training pairs
Real training pairs pN
Held-out real fMRI Top-10 image retrieval
(a) TRIBE v2 image-conditioned fMRI augmentation protocol.
(b) DINOv2-small image embeddings predict cortical fMRI responses. Figure 1 TRIBE v2 image-conditioned fMRI augmentation and target-embedding neural signal. (a) For each subject and
dataset, we retain pN real image-fMRI training pairs and sample additional COCO images not seen by that subject. TRIBE v2 predicts synthetic fMRI responses to these images, yielding apN synthetic training pairs, where a is the augmentation factor. We train an fMRI-to-DINOv2-small decoder on the resulting mixture of real and synthetic pairs and evaluate it only on held-out real fMRI using Top-10 image retrieval. (b) Brain-encoding grids show Pearson correlation between DINOv2-small-predicted and real fMRI cortical activations on fsaverage5, averaged across subjects within each dataset, supporting the use of DINOv2-small as the decoded image representation.
3
Methods
3.1
TRIBE v2 as a Synthetic fMRI Generator
TRIBE v2 (d’Ascoli et al., 2026) predicts BOLD fMRI responses on the cortical surface (fsaverage5 (Fischl et al., 1999)) from video, audio, and language inputs. It uses frozen modality-specific backbones, including V-JEPA 2 (Assran et al., 2025) for visual input (videos), and integrates their features with a transformer encoder to predict responses of T = 100 s at a frequency of 1 Hz (i.e., TR=1 s), starting 5 s after stimulus onset (to account for hemodynamic delay). We use TRIBE v2 in its default mode, which predicts population-level responses without adapting to a target subject. For brevity, we refer to this model as “TRIBE” in the remainder of the paper. Synthesizing fMRI for images. TRIBE was trained on naturalistic video, audio, and text, and not on isolated still images. To apply it to static image datasets, we therefore convert each image into a short 3 s still video, where every frame shows the same image. Of note, this use of TRIBE is significantly out-of-distribution, as we feed it video that has no natural motion; a property absent from TRIBE’s training set. The resulting synthetic fMRI time-series of T seconds is resampled to match the target dataset acquisition frequency and truncated to a single timestep (i.e., 1 TR).
3.2
Data-Augmentation Operating Grids for Image Decoding
Image Decoding. Let D = Dtrain ∪ Dtest be a dataset of real fMRI responses to images for a given subject, V with disjoint train and test splits. That is, D = {(si , yi )}N i=1 , where si is an image stimulus and yi ∈ R is the corresponding fMRI brain response from the subject. The task of image decoding is to train a model to predict a representation (or reconstruction) of s from y.
3
Data Augmentation. Let fθ (s) denote the TRIBE-synthesized fMRI response to stimulus s. For a percentage of retained real-data p and TRIBE augmentation factor a ∈ R≥0 , we train a model M on Dp,a = {(si , yi ) : i ∈ Ip } ∪ {(s̃j , fθ (s̃j )) : j = 1, . . . , ⌊a|Ip |⌋},
(1)
where Ip is a random subset of Dtrain of size ⌊pN ⌋, and the s̃j are sampled from an image augmentation pool A disjoint from D. We evaluate M on Dtest , which is kept fixed and never augmented. The condition a = 0 means subsampling a subset of Dtrain of p% of its total size, and represents a setting where the availability of real data is limited, with no access to synthetic fMRI augmentation. We refer to this as the matched real-only setting (at p%). Synthetic-only Image Decoding. On the other hand, we train ’real-data zero-shot’, synthetic-only image decoders: the model is trained only on TRIBE-synthesized responses and evaluated on the real fMRI in Dtest . Operating Grids. By training decoders over representative values of p and a and evaluating them on the same held-out dataset Dtest of real data, we obtain an operating grid of how the decoding performance γ(p, a) of models varies (on Dtest ) with the amount of real data available (p) and the amount of synthetic TRIBE data added (a). Encoder-decoder Leakage. We use the output of the last layer of DINOv2-small (Oquab et al., 2023) as the decoding target. While it does not entirely rule out all forms of representational overlap with TRIBE’s visual encoder (V-JEPA 2), it intentionally reduces a direct leakage concern: that augmentation gains could arise simply from using the same visual features to synthesize fMRI and to score the decoder.
3.3
Reconstructing Images from BOLD fMRI
Beyond the task of decoding an image embedding, we use DynaDiff (Careil et al., 2025) to test whether synthesizing fMRI responses to images with TRIBE also helps reconstruct images from fMRI responses. We train DynaDiff on Dp,a (Section 3.2) for p = 100% and a > 0. For each data modality (real or synthetic), we add a new modality-specific linear layer at the brain module’s input. This informs the model of which kind of data it is processing and facilitates training convergence. All remaining layers in the brain module and diffusion model remain shared across modalities, so synthetic trials still update every shared parameter. Training details essentially follow the Dynadiff paper (Careil et al., 2025), additional details can be found in Appendix C.
3.4
Models and Evaluation
Linear decoders. Our primary decoder is a ridge regression model that maps fMRI inputs (i.e., cortical brain activations) to image embeddings. This pipeline normalizes fMRI features into standard units (across training samples and voxels) and fits a ridge regression with per-target regularization and Leave-One-Out Cross-Validation. Deep decoders. We additionally evaluate a MindEye-style residual MLP decoder (Scotti et al., 2023). This model maps fMRI cortical responses to image embeddings through a residual MLP; architecture and optimization details are provided in Appendix A. We train it with a CLIP contrastive retrieval loss as in (Benchetrit et al., 2024; Banville et al., 2025): for each batch, the predicted image embedding is matched to its paired target against the other batch targets using cross-entropy over dot-product similarities. Target embeddings are normalized in the loss, with fixed temperature τ = 1. The CLIP loss is one-way i.e., it is not symmetrized across the target-to-prediction direction. Both decoders are trained separately for each target subject. Decoding Metrics. The primary metric is Top-K image retrieval accuracy on the entire test set (made only of real data). Given a predicted image embedding, we rank all candidate test images by cosine similarity to their ground-truth embeddings and score whether the correct image appears among the top K. We also
4
report median retrieval rank in the Appendix, defined as the median of rank positions of the correct image in the same similarity-sorted retrieval list; lower ranks indicate better performance. Reconstruction Metrics. Following standard practice in fMRI-to-image reconstruction, we report PixCorr (pixel-wise correlation) to assess low-level image similarity as well as EfficientNet and SwAV to measure high-level semantic similarity.
4
Experiments and Results
4.1
Datasets
Hemodynamic Delay. We evaluate our approach on two image-fMRI datasets that are disjoint from TRIBE’s training data (d’Ascoli et al., 2026). For both datasets, we decode a one-TR response window at the expected BOLD peak, starting 5 s after stimulus onset (consistent with TRIBE predictions), and train separate decoders for each subject. The synthetic augmentation pool for a given subject is sampled from the COCO images dataset after excluding all stimuli seen by this subject. Natural Scenes Dataset. NSD (Allen et al., 2022) is a 7T dataset acquired with whole-brain gradient-echo EPI at 1.8 mm isotropic resolution and TR=1.6 s. It contains eight participants, each scanned for 30–40 one-hour sessions while viewing natural images from MS-COCO. Following common practice for image-decoding work, we use the four subjects who completed the full 40-session protocol and have fsaverage-space responses in our pipeline (subjects 1, 2, 5, and 7). Each of these subjects viewed 10,000 unique images, repeated three times; the standard split uses 9,000 subject-specific training images and a shared 1,000-image test set. Images were presented for 3 s with a 1-s blank interval. BOLD5000 BOLD5000 (Chang et al., 2019) is a smaller 3T image-fMRI dataset acquired with multiband T2*-weighted EPI at 2 mm in-plane resolution and TR=2 s. It includes four participants, each scanned across repeated 1.5-hour sessions, with one participant completing fewer sessions than the others. Stimuli were drawn from MS-COCO, ImageNet, and SUN scene images, giving the full-session subjects 4,916 unique images. A small repeated subset of 112 images forms the held-out test set, and each stimulus was shown for 1 s followed by a 9-s fixation interval. Single-trial responses. Some stimuli are presented multiple times to subjects in the training set, most notably in NSD where each image is shown three times. Before constructing the operating-grid conditions, we keep one randomly selected fMRI response per subject and image and discard all other repetitions, and apply this deduplication uniformly to all real-only, TRIBE, and control scenarios in both datasets. This avoids an imbalance between real data, which can contain multiple noisy repetitions of the same stimulus, and TRIBE predictions, which are deterministic and therefore provide only one synthetic response per image. fMRI preprocessing. Anatomical and functional MRI data are preprocessed with fMRIPrep (Esteban et al., 2019) using its default workflow, requesting cortical-surface outputs in fsaverage space. We retain the fsaverage5 time series from each run so that real fMRI and TRIBE predictions live on the same 20,484-vertex cortical mesh.
4.2
Experimental Setup
Operating Grids. We compute operating grids γ(p, a) for each dataset and each decoder family, using real-data percentages p ∈ {10, 30, 50, 70, 90, 100}% and augmentation factors a ∈ {0, 0.25, 0.5, 0.75, 1, 2, 4, 8, 16}. For each dataset D, a decoder is trained separately for each subject on Dp,a and is evaluated on the fixed held-out test set Dtest this subject (i.e., our decoders are single-subject). We pick Dtest as defined by the dataset itself, following previous Image Decoding studies (see Section A). The resulting grids summarize performance both as a percentage of the 100% real-data reference (Panel A in Figures 2 and 3) and as the relative gain over the matched real-only baseline (a = 0) at the same retained-real percentage p (Panel B in Figures 2 and
5
34%
39%
44%
43%
37%
30%
57%
57%
59%
61%
62%
66%
71%
72%
59%
10%
0%
+2%
+5%
+8% +18% +35% +53% +53% +30%
30%
0%
+1%
+4%
+7%
30
80 50%
76%
75%
77%
79%
80%
70%
90%
90%
91%
93%
94% 101% 103% 102% 83%
89%
88%
72%
90% 100%
96%
96%
40
99% 100% 103% 109% 111% 112% 95%
20 0
100% 100% 103% 104% 108% 114% 116% 115% 102%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
%
60
10 50%
0%
-0%
+2%
+4%
+6% +14% +18% +17%
0
-5%
−10 70%
0%
90%
0%
16x
-0%
+1%
+1%
+3%
+3%
+4%
+5% +12% +15% +13%
-7%
+7% +14% +16% +17%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
8x
−20 −30
-0%
55%
55%
58%
60%
68%
30%
65%
67%
72%
78%
80%
91% 104% 102% 93%
50%
79%
82%
85%
89%
70%
85%
89%
93%
97% 104% 118% 119% 121% 112%
90%
16x
60 40
10 8 6 4 2
0
0 0x
0.25x 0.5x 0.75x
1x
2x
4x
8x
30% real 50% real 70% real
)
al Re
90% real 100% real 100% real, unaugmented
94%
80
99% 106% 111% 121% 129% 120% 111%
4x
8x
+7%
0%
+3% +11% +20% +23% +40% +60% +57% +43%
50%
0%
+4%
+7% +12% +16% +36% +40% +35% +26%
70%
0%
+5%
+9% +14% +22% +39% +40% +42% +31%
90%
0%
-2%
+3% +10% +16% +26% +34% +25% +15%
+5% +11% +15% +32% +55% +68% +57%
20
0
0 −20
20
100% 93% 101% 112% 113% 122% 132% 126% 112%
2x
0%
30%
40
100
−40
16x
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
16x
TRIBE augmentation factor
D
160
20
140
18
120 100 80 60
15 12 10 8 5 2
40 0
16x
TRIBE augmentation factor 100% synthetic Chance 10% real
96%
10%
C
12
20
120
TRIBE augmentation factor
Normalized performance (%)
Top-10 Accuracy (%)
80
82%
92% 107% 111% 107% 100%
0x 0.25x 0.5x 0.75x 1x
D
100
87%
40
100%
14
120
81%
60
TRIBE augmentation factor
C Normalized performance (%)
20
+9% +16% +25% +28% +4%
B
52%
Top-10 Accuracy (%)
86%
Real data retained
Real data retained
100
A 10%
Relative Improvement over real-only (%)
31%
Real data retained
30%
%
29%
Real data retained
B
28%
Relative Improvement over real-only (%)
A 10%
ly
on
(1
0%
+
8x
(1 IBE
TR
100% real
0x
)
)
0%
0%
+
8x
0.25x 0.5x 0.75x
1x
2x
4x
8x
16x
TRIBE augmentation factor
(1 ise
100% synthetic Chance 10% real
Chance
(a) NSD
30% real 50% real 70% real
)
nly
lo
a Re
No
90% real 100% real 100% real, unaugmented
)
)
0%
(1
IBE
0% (1
+
8x
TR
100% real
ise
0%
+
8x
(1
No
Chance
(b) BOLD5000
Figure 2 Operating grids for TRIBE augmentation using Ridge decoders. A. Top-10 retrieval accuracy normalized to the
performance of a 100% real-data-only baseline (i.e., a ridge model trained on the full, unaugmented training set); values above 100% indicate that TRIBE augmentation surpasses this 100% real-data-only baseline. B. Relative performance improvement over the matched real-only condition (a = 0) at the same retained-real percentage p%. C. Augmentation scaling curves for each p% of real-data retained, including chance and synthetic-only baselines. Shading shows SEM across all subjects (four subjects for each dataset). D. Raw Top-10 retrieval at the selected real-data fraction p% against real-only (p = 100%) and noise-augmentation controls. Ridge models provide the cleanest operating-grid evidence: TRIBE augmentation improves low- and medium-data regimes. The best factor depends on dataset and the fraction of real data.
3). Scores are computed after averaging 5 seeds within subject, controlling dataset subsampling and train / validation splits. Models. We train our decoders to predict a DINOv2-small embedding of a stimulus image from an input of one TR of cortical fsaverage5 activations at stimulus onset 5 s (totalizing 20,484-vertex values across both hemispheres). Ridge regressions are tuned over 20 alphas evenly log-spaced from 1 to 108 . The deep models use a residual architecture as introduced by Scotti et al. (2023) (see Appendix A for details) and are trained using AdamW with learning rate 5 × 10−4 and weight decay 0.01. Optimization uses a step-wise OneCycleLR schedule with maximum learning rate 10−3 and warm-up over the first 10% of training steps. Models are trained for up to 40 epochs with early stopping on validation loss (20% validation split) and patience 10. Recording times. The percentages can also be read as approximate scan-time budgets (Appendix D). For NSD, the 100% real reference corresponds to 9,000 training images, or roughly 10 hours of image-task acquisition per subject (3 s presentation + 1 s blank interval). Thus p = {10, 30, 50, 70, 90, 100}% corresponds to about {1, 3, 5, 7, 9, 10} real hours. For BOLD5000, the approximately 4,804 non-test images are presented for 1 s followed by 9 s fixation, so the full real training set corresponds to roughly 13.3 hours per subject; the same p values correspond to about {1.3, 4.0, 6.7, 9.3, 12.0, 13.3} hours.
4.3
Operating Grids for Ridge decoders
Figure 2 shows operating grids, augmentation scaling curves and controls averaged across all subjects of each dataset (shading in Panel C and error bars in Panel D indicate SEM across subjects). On the Natural Scenes Dataset (NSD), augmenting data with TRIBE produces significant gains in low-data regimes. For example: around 90% of the full real-only dataset performance (i.e., p = 100%, a = 0) can be obtained with only 50% of retained real data (see p = 50%, a = 4). In scan-time terms, because the deduplicated NSD training set corresponds to about 10 hours of real fMRI per subject, this means reaching roughly 90% of full-data performance with about 5 real hours instead of 10. The full-real-only baseline’s performance can 6
26%
30%
35%
31%
27%
30%
44%
43%
45%
50%
52%
57%
57%
48%
40%
10%
0%
-3%
-8%
+11% +28% +13%
-2%
30%
0%
-2%
+2% +14% +19% +31% +32% +9%
-8%
-3%
-6%
60%
70%
81%
90%
82%
67%
85%
71%
87%
70%
89%
76%
91%
71%
82%
58%
68%
51%
60
60%
40
99% 103% 102% 104% 102% 94%
78%
71%
20
100% 108% 106% 107% 108% 108% 97%
80%
76%
0
8x
16x
96%
100%
62%
0x 0.25x 0.5x 0.75x 1x
2x
4x
10 50%
0%
+3% +12% +17% +16% +26% +18%
-16%
0 −10
70%
0%
+1%
+6%
+8% +10% +12% +2%
-15% -26%
−20 90%
0%
+4%
+8%
+7%
-2%
-19% -26%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
TRIBE augmentation factor
+7%
+6%
40%
47%
44%
46%
48%
49%
50%
47%
30%
58%
65%
63%
63%
65%
69%
74%
70%
66%
50% 70% 90%
16x
74%
40 20 0
74%
80%
73%
83%
84%
83%
78%
79%
60 80%
86%
81%
92% 100% 99%
93%
96%
86%
92%
95%
82%
86%
40
97% 107% 103% 91%
87%
20
100% 101% 105% 108% 110% 111% 111% 94%
91%
0
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
0.25x 0.5x 0.75x
1x
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
20
15
10
5
0%
-0%
+8%
-1%
+12% +14% +12% +6%
+6%
70%
0%
+8%
+1% +15% +7% +15% +18% +2%
+6%
90%
0%
+9%
+8%
-1%
-6%
8x
16x
-1%
-5%
-2%
+3%
+4%
+7%
+1%
10 5 0 −5
+4%
+6% +16% +12%
−10
0x 0.25x 0.5x 0.75x 1x
2x
4x
TRIBE augmentation factor
D 30
120 100 80 60
25 20 15 10 5
40
0
al Re
90% real 100% real 100% real, unaugmented
50%
C
16x
TRIBE augmentation factor
+7% +11% +18% +27% +20% +14%
16x
0 0x
+12% +8%
TRIBE augmentation factor
Normalized performance (%)
Top-10 Accuracy (%)
60
-15%
0%
−15
D
80
0%
30%
15
80
100%
120 100
10% 100
TRIBE augmentation factor
C Normalized performance (%)
-5%
B
47%
Top-10 Accuracy (%)
50%
Real data retained
20 80
%
Real data retained
100
A 10%
Relative Improvement over real-only (%)
26%
Real data retained
25%
%
26%
Real data retained
B
27%
Relative Improvement over real-only (%)
A 10%
ly
on
(3
%)
0
0%
+
4
x)
(3 IBE
TR
100% real
0%
+
0x
x)
4
0.25x 0.5x 0.75x
1x
2x
4x
8x
16x
TRIBE augmentation factor
(3 ise
100% synthetic Chance 10% real
Chance
(a) NSD
30% real 50% real 70% real
)
nly
lo
a Re
No
90% real 100% real 100% real, unaugmented
)
)
0%
(3
IBE
0% (3
+
4x
TR
100% real
ise
0%
+
4x
(3
No
Chance
(b) BOLD5000
Figure 3 Operating grids for TRIBE augmentation using Deep decoders. The operating grids as in Figure 2 are computed for the Deep Residual decoders. A. Top-10 retrieval accuracy normalized to the performance of a 100% real-data-only
baseline (i.e., a deep model trained on the full, unaugmented training set); values above 100% indicate that TRIBE augmentation surpasses this 100% real-data-only baseline. B. Relative performance improvement over the matched real-only condition (a = 0) at the same retained-real percentage p%. C. Augmentation scaling curves for each p% of real-data retained, including chance and synthetic-only baselines. Shading shows SEM across all subjects (four subjects for each dataset). D. Raw Top-10 retrieval at the selected real-data fraction p% against real-only (p = 100%) and noise-augmentation controls. The qualitative pattern observed for Ridge decoders remains: TRIBE helps in selected low- and medium-data regimes, but gains are less uniform than for Ridge and depend on both the dataset and the augmentation factor. This supports the claim that TRIBE contains useful decoding signal while emphasizing that synthetic fMRI must be calibrated to the decoder.
be reached with 70% of real data, or about 7 real hours, saving roughly 3 hours per subject. Furthermore, TRIBE provides relative improvements up to 53% even in the lowest regime p = 10%. At larger retained-real fractions, the relative improvements become smaller and can saturate, indicating that TRIBE augmentation is most useful when real data is scarce. Similarly, adding too much new data eventually saturates or hurts performance (topping around a = 8). Of note, the 100% synthetic-only data reference is at chance level and shows that real data remains necessary to decode NSD. The BOLD5000 dataset shows even stronger performance gains. The Ridge decoders benefit substantially from TRIBE at p = 10 . . . 50% of retained real data, reaching already approximately 90% of the full real-only baseline performance for p = 10% and 100% with only p = 30% of real data retained. In scan-time terms, this corresponds to about 4.0 hours of real BOLD5000 training data instead of 13.3 hours, a saving of roughly 9 hours per subject. The gain brought by adding TRIBE synthetic data remains consistent for all presented data regimes and saturates between augmentation factors a = 4 and a = 8, at a slightly smaller augmentation factor than NSD. Also, in contrast to NSD, where the 100%-synthetic-only baseline is roughly at chance, the corresponding BOLD5000 reference is slightly above, suggesting that TRIBE’s subject-agnostic synthetic responses carry some transferable visual signal in this specific dataset. Remarkably, despite the different acquisition protocols and absolute performance levels, the strongest relativeimprovement cell for Ridge appears at the same operating point in both datasets: p = 10% retained real data and augmentation factor a = 8. See Appendix B for per-subject operating grids for Ridge decoders for Top-10 accuracy and median rank.
4.4
Operating grids for Deep Decoders
Figure 3 extends the operating grids for Ridge models to deep decoders. On both datasets, the deeper architecture can benefit from TRIBE augmentation at selected operating points, and BOLD5000 again shows 7
stronger high-retention gains than NSD. The effects are more modest and noisier than for Ridge decoders: performance saturates, and can start decreasing, at smaller augmentation factors, typically around a = 4. This may be explained by the fact that deep decoders are already strong baselines on matched real-only conditions, which makes it more difficult to push performance further with TRIBE augmentation. Despite this noisier behavior, the same practical regimes appear in both datasets. Around 90% of the full real-only baseline performance can be reached with only p = 70% retained real data, using a = 2 for NSD and a = 0.75 for BOLD5000. Recovering the full real-only baseline requires more real data: both datasets reach it at approximately p = 90% with a small amount of TRIBE augmentation (a = 0.25). Here too, the best relative-improvement score aligns across datasets, but at a different regime than Ridge: for Deep decoders, the largest relative gain is observed at p = 30% and a = 4 in both NSD and BOLD5000. Finally, unlike the Ridge case in NSD, the full-synthetic reference remains above chance for Deep decoders in both datasets, suggesting that the Deep decoder can extract some usable visual signal from TRIBE predictions even without real target-dataset fMRI. See Appendix B for per-subject operating grids for Deep decoders for Top-10 accuracy and median rank.
4.5
Synthetic-Only Training and Controls
The operating grids also include two controls. First, the synthetic-only baseline remains above chance for Ridge models for BOLD5000 and for Deep Models for both datasets , (see Panel C in Figure 2 and 3). This is notable because TRIBE’s visual pathway was trained on natural videos rather than static-image fMRI responses, yet its predictions still carry visual information aligned with the DINOv2-small image embedding. Second, Panel D compares a selected TRIBE operating point (p, a) to a noise control, obtained by training the decoder exactly as for (p, a) but replacing the TRIBE-synthesized fMRI with random Gaussian noise. The TRIBE condition is generally much stronger than the noise control, showing that the gain is not explained by simply increasing the number of training samples or by adding high-dimensional variability to the fMRI input space; the synthetic responses must preserve stimulus-dependent structure that is useful for decoding.
4.6
Augmenting Image Reconstruction
Beyond visual embedding retrieval, DynaDiff (Careil et al., 2025) provides a natural test of whether TRIBE augmentation transfers to generative image reconstruction. We train DynaDiff on the NSD dataset for subject 5 (the best-performing subject in our retrieval experiments), retaining all available real fMRI (p = 100%) and varying the TRIBE augmentation factor a ∈ {0, 0.5, 1, 2}. The held-out test set is identical to the retrieval experiments: the original NSD test split, which contains no synthetic data. Figure 4 reports image reconstruction metrics (PixCorr, SwAV, and EfficientNet; see Section 3.4) as a function of a, with the no-augmentation baseline (a = 0) shown as a dashed line. TRIBE augmentation improves reconstruction across all three metrics and at every augmentation factor a > 0, with a consistent optimum at a = 1: PixCorr rises from 0.06 to 0.16 (∼2.5× on low-level pixel similarity), while the SwAV and EfficientNet distances drop from 0.39 to 0.35 and from 0.72 to 0.66, respectively. At a = 2, all three metrics regress toward the baseline but remain better than the unaugmented reference.
Figure 4 Image reconstruction metrics on NSD with Dynadiff. Similar to image decoding, TRIBE boosts results up to a
certain augmentation factor
8
5
Discussion
Synthetic fMRI augmentation only helps in specific data-regimes. Our results show that TRIBE v2(d’Ascoli et al., 2026) augmentation can improve brain-to-image decoding by up to 68% gain when fMRI data is scarce. As expected, such gains depend both on the amount of available data and on the proportion of synthetic data used; beyond a certain point, decoding performance saturates and can even deteriorate as more synthetic data are added. The practical implication is direct: with a limited scan-time budget, the synthetic-to-real data ratio should be carefully tuned. Out-of-distribution pretraining. It is quite remarkable that this approach works at all: TRIBE v2(d’Ascoli et al., 2026) is neither trained on brain responses to static images (or static image embedding) nor is it trained for decoding. Its visual pathway is trained on brain responses to long movies, jointly with audio and language. Here, we simply converted the static images into short movie clips. Consequently, the present results suggest that TRIBE captures patterns of visual responses that transfer from naturalistic multimodal stimulation to a more controlled rapid-serial-visual-presentation of images. At the same time, this mismatch may also explain why augmentation benefits saturate and why the optimal synthetic-to-real ratio must be calibrated. Why can augmentation exceed the full-real reference? Some configurations exceed the 100% real reference. This should not be read as synthetic fMRI containing more subject-specific information than real fMRI. A more plausible explanation is that TRIBE v2 adds stimulus diversity and population-level visual response structure that regularize the fMRI-to-embedding mapping, especially when the real training set is small or idiosyncratic. Because the augmentation pool excludes images from the target dataset and evaluation is always on real fMRI, these gains are not caused by overlap with evaluation stimuli or by testing on synthetic data. Subject-agnostic augmentation has limits. Here, TRIBE v2 predicts an average-subject brain response (i.e. there is no target-subject adaptation). This makes the augmentation broadly deployable, but it also means synthetic responses lack individual-specific anatomy and response idiosyncrasies. The synthetic-only reference captures this limitation: above-chance decoding is possible, but the strongest results still require real fMRI from the target subjects. Future work could combine the operating grid protocol with subject-specific TRIBE adaptation, learned synthetic-data weighting, or multi-subject decoder pretraining (Scotti et al., 2024). Limitations. First, although we replicate our findings on two datasets, the latter have a small number of subjects. Extension to additional datasets, including other modalities thus remains to be further evaluated. Given TRIBE v2 is multimodal, our approach should not require major changes. Second, the primary task focuses on image retrieval from a DINOv2-small embedding (Oquab et al., 2023); image reconstruction is evaluated only preliminarily through the DynaDiff experiments in Section 4.6. This choice was constrained by the fact that most available fMRI-to-image models (e.g. Scotti et al. (2023, 2024)) are based on timeindependent fMRI beta maps, while TRIBE v2 generates dynamic fMRI signals. It will be important to extend the present tests to other fMRI-to-image models, e.g., by deriving beta maps from TRIBE v2 with preprocessing matched to NSD (Allen et al., 2022) and BOLD5000 (Chang et al., 2019). Broader impact. If reliable, model-based fMRI augmentation could make image-decoding research less dependent on very large scan-time budgets and therefore easier to study across laboratories and populations. At the same time, the results should not be interpreted as enabling decoding without any real data: the best performance still relies on subject-specific real fMRI responses.
6
Conclusion
Synthetic fMRI augmentation could mark a shift toward more accessible neuroimaging by reducing the reliance on massive scan-time budgets. While TRIBE v2 demonstrates that naturalistic, multimodal pretraining can effectively regularize decoding in data-scarce regimes, the future of the field lies in balancing these general
9
population-level priors with subject-specific nuances. Ultimately, calibrating the synergy between synthetic diversity and real neural signals will be essential for building efficient decoders of brain activity.
References E. J. Allen, G. St-Yves, Y. Wu, J. L. Breedlove, J. S. Prince, L. T. Dowdle, M. Nau, B. Caron, F. Pestilli, I. Charest, et al. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence. Nature Neuroscience, 25(1):116–126, 2022. Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-JEPA 2: Self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985, 2025. S. Azizi, S. Kornblith, C. Saharia, M. Norouzi, and D. Fleet. Synthetic data from diffusion models improves ImageNet classification. Transactions on Machine Learning Research, 2023. Hubert Banville, Yohann Benchetrit, Stéphane d’Ascoli, Jérémy Rapin, and Jean-Rémi King. Scaling laws for decoding images from brain activity. arXiv preprint arXiv:2501.15322, 2025. Yohann Benchetrit, Hubert Banville, and Jean-Remi King. Brain decoding: toward real-time reconstruction of visual perception. In The Twelfth International Conference on Learning Representations, 2024. M. Careil, Y. Benchetrit, and J.-R. King. Single-stage decoding of images from continuously evolving fMRI. arXiv preprint arXiv:2505.14556, 2025. N. Chang, J. A. Pyles, A. Marcus, A. Gupta, M. J. Tarr, and E. M. Aminoff. BOLD5000, a public fMRI dataset while viewing 5000 visual images. Scientific Data, 6(1):49, 2019. doi: 10.1038/s41597-019-0052-3. Z. Chen, J. Qing, T. Xiang, W. L. Yue, and J. H. Zhou. Seeing beyond the brain: Conditional diffusion model with sparse masked modeling for vision decoding. In CVPR, 2023. C. Conwell, J. S. Prince, K. N. Kay, G. A. Alvarez, and T. Konkle. What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? bioRxiv, 2023. S. d’Ascoli, J. Rapin, Y. Benchetrit, T. Brookes, K. Begany, J. Raugel, H. Banville, and J.-R. King. A foundation model of vision, audition, and language for in-silico neuroscience. arXiv preprint, 2026. Oscar Esteban, Christopher J Markiewicz, Ross W Blair, Craig A Moodie, A Ilkay Isik, Asier Erramuzpe, James D Kent, Mathias Goncalves, Elizabeth DuPre, Madeleine Snyder, et al. fMRIPrep: a robust preprocessing pipeline for functional MRI. Nature methods, 16(1):111–116, 2019. V. Fernandez, W. H. Pinaya, P. Borges, P.-D. Tudosiu, M. S. Graham, T. Vercauteren, and M. J. Cardoso. Can segmentation models be trained with fully synthetically generated data? In MICCAI Workshop on Simulation and Synthesis in Medical Imaging, 2022. Bruce Fischl, Martin I Sereno, Roger BH Tootell, and Anders M Dale. High-resolution intersubject averaging and a coordinate system for the cortical surface. Human brain mapping, 8(4):272–284, 1999. R. He, S. Sun, X. Yu, C. Xue, W. Zhang, P. Torr, S. Bai, and X. Qi. Is synthetic data from generative models ready for image recognition? In ICLR, 2023. Martin N Hebart, Oliver Contier, Lina Teichmann, Adam H Rockter, Charles Y Zheng, Alexis Kidder, Anna Corriveau, Maryam Vaziri-Pashkam, and Chris I Baker. THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior. eLife, 12:e82580, feb 2023. ISSN 2050-084X. doi: 10.7554/eLife.82580. https://doi.org/10.7554/eLife.82580. Y. Kamitani and F. Tong. Decoding the visual and subjective contents of the human brain. Nature Neuroscience, 8(5): 679–685, 2005. K. N. Kay, T. Naselaris, R. J. Prenger, and J. L. Gallant. Identifying natural images from human brain activity. Nature, 452(7185):352–355, 2008. T. Naselaris, R. J. Prenger, K. N. Kay, M. Oliver, and J. L. Gallant. Bayesian reconstruction of natural images from human brain activity. Neuron, 63(6):902–915, 2009. T. Naselaris, K. N. Kay, S. Nishimoto, and J. L. Gallant. Encoding and decoding in fMRI. NeuroImage, 56(2):400–410, 2011.
10
Kevin P Nguyen, Vyom Raval, Abu Minhajuddin, Thomas Carmody, Madhukar H Trivedi, Richard B Dewey Jr, and Albert A Montillo. BLENDS: augmentation of functional magnetic resonance images for machine learning using anatomically constrained warping. Brain connectivity, 13(2):80–88, 2023. M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2023. F. Ozcelik and R. VanRullen. Natural scene reconstruction from fMRI signals using generative latent diffusion. Scientific Reports, 13(1):15666, 2023. Paul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin, Alex Nguyen, Ethan Cohen, Aidan J. Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth A. Norman, and Tanishq Mathew Abraham. Reconstructing the mind’s eye: fmri-to-image with contrastive learning and diffusion priors, 2023. https://arxiv.org/abs/2305.18274. Paul S. Scotti, Mihir Tripathy, Cesar Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A. Norman, and Tanishq Mathew Abraham. Mindeye2: Shared-subject models enable fmri-to-image with 1 hour of data, 2024. https://arxiv.org/abs/2403.11207. G. Shen, K. Dwivedi, K. Majima, T. Horikawa, and Y. Kamitani. End-to-end deep image reconstruction from human brain activity. Frontiers in Computational Neuroscience, 13:21, 2019. G. St-Yves, E. J. Allen, Y. Wu, K. Kay, and T. Naselaris. Brain-optimized deep neural network models of human visual areas learn non-hierarchical representations. Nature Communications, 14, 2023. doi: 10.1038/s41467-023-38674-4. Jiyao Wang, Nicha C Dvornek, Lawrence H Staib, and James S Duncan. Learning sequential information in task-based fMRI for synthetic data augmentation. In International Workshop on Machine Learning in Clinical Neuroimaging, pages 79–88. Springer, 2023.
11
Appendix A
MindEye-Style MLP Decoder Architecture
The deep decoder used in the operating grids follows the residual MLP architecture introduced by MindEye (Scotti et al., 2023), with the parameterization described below. The input is a one-TR fsaverage5 response with 20,484 vertices, and the output is a 384-dimensional DINOv2-small image embedding. Table 1 reports the architecture used in our experiments: hidden width 553, two residual blocks, an input projection from fsaverage5 vertices to the hidden width, and a 384-dimensional output head. We optimize this model with AdamW (learning rate 5 × 10−4 , weight decay 0.01) and a OneCycle learning-rate schedule with maximum learning rate 10−3 . Training runs for up to 40 epochs with early stopping on validation loss (patience 10) and batch size 128. Training and evaluation were run on a single NVIDIA V100 GPU with 16 GB of memory; the longest training runs took approximately 2 hours. Table 1 Residual MLP decoder architecture used for deep-decoder robustness experiments. Parameter counts are for one
per-subject decoder. All parameters are trainable. Component
Operation
Parameters
Vertex-to-hidden map Hidden mixing layer Post-TR block Residual MLP blocks Temporal aggregation Embedding readout Output head
Linear map, 20,484 → 553, no bias Linear map within the hidden space, 553 → 553 LayerNorm(553), GELU, dropout p = 0.5 2× [Linear 553 → 553, LayerNorm, GELU, dropout p = 0.15] Linear aggregation over the single retained TR Linear 553 → 384 Linear projection 384 → 384, dropout p = 0
11,327,652 306,362 1,106 614,936 2 212,736 147,840
Total
B
12,610,634
Additional Operating Grids
Figures 5–8 show the per-subject Top-10 operating grids underlying the subject-averaged results in the main text. These grids use the same normalization, retained-real percentages, augmentation factors, and panel layout as Figures 2 and 3, but avoid averaging across subjects. They make visible the heterogeneity of the augmentation effect: some subjects benefit from TRIBE across a broad range of low-data regimes, whereas others show narrower optima or saturation at smaller augmentation factors. This subject-level variability is expected due to the subject-agnostic nature of TRIBE v2 synthetic responses and motivates reporting the operating regime rather than a single global augmentation factor.
12
72%
35%
78%
80%
83%
84%
91%
94%
87%
70%
60 70% 90% 100%
95%
98%
94%
98%
94%
96%
99% 106% 108% 101% 79%
40
99% 100% 105% 115% 116% 111% 88%
20 0
100% 102% 106% 108% 110% 119% 120% 114% 96%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
%
79%
0%
30%
0%
+1%
+9% +20% +37% +55% +41% +16%
30 +0%
+5%
+7%
20
+8% +13% +20% +19% -13%
10 50%
0%
-1%
+1%
+4%
+6% +15% +18% +9%
0
-11%
−10 70%
0%
90%
0%
16x
-1%
-1%
-1%
+1%
+1%
+2%
+4% +11% +14% +6%
-17%
−20 −30
+7% +17% +18% +13% -10%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
8x
A
C
28%
29%
30%
31%
35%
39%
44%
48%
39%
30%
58%
59%
61%
63%
63%
70%
74%
79%
60%
50% 70% 90%
16x
60 40 20
12 10 8 6 4
2x
4x
8x
16x
)
TRIBE augmentation factor 100% synthetic Chance 10% real
30% real 50% real 70% real
93%
94%
88%
91%
91%
93%
96%
ly
on
al Re
(1
20
100% 105% 108% 109% 115% 119% 118% 121% 107%
0
8x
0%
+
8x
+
31%
37%
41%
43%
0%
56%
58%
60%
61%
65%
70%
72%
63%
50%
74%
74%
76%
77%
78%
82%
85%
94%
74%
70% 90% 100%
88%
93%
89%
96%
88%
99%
92%
91%
97% 101% 99%
84%
40
99% 100% 106% 110% 110% 98%
20 0
100% 99% 102% 101% 105% 114% 117% 112% 107%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
%
60
0%
30%
60 40
+3%
+6%
+9% +17% +18% +17%
90%
0%
+3%
+5% +10% +12% +17% +18% +17% +4%
0 −10
20
-9%
−20
2x
4x
8x
16x
12 10 8 6 4 2 0
0.25x 0.5x 0.75x
100% synthetic Chance 10% real
+4%
+9% +10% +16% +38% +57% +62% +40%
0%
+1%
+5%
+8% +10% +17% +26% +30% +14%
20 10
50%
0%
-0%
+4%
+4%
+6% +12% +15% +28% +1%
70%
0%
+1%
+0%
+4%
+3% +10% +15% +12%
0 −10
-5%
−20 −30
90%
0%
16x
+3%
+6%
+6%
+7% +14% +18% +18% +5%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
1x
2x
4x
8x
16x
)
30% real 50% real 70% real
nly
)
)
0%
(1
lo
a Re
IBE
0% (1
+
8x
ise
+
0%
8x
(1
No
TR
90% real 100% real 100% real, unaugmented
100% real
2x
4x
8x
A 10%
Chance
B
29%
29%
29%
30%
33%
37%
41%
40%
36%
100 30%
53%
51%
51%
54%
56%
59%
65%
66%
60%
50%
70%
70%
71%
72%
72%
78%
83%
83%
74%
70% 90% 100%
16x
85%
93%
85%
87%
86%
88%
94%
60
96% 102% 88%
91%
93%
92%
95%
100% 95%
95%
97%
99% 104% 109% 112% 98%
0x 0.25x 0.5x 0.75x 1x
40 20
99% 103% 110% 93%
2x
4x
8x
10%
0%
+1%
+1%
+4% +11% +27% +40% +37% +24%
30%
0%
-3%
-3%
+2%
+6% +12% +24% +25% +14%
50%
0%
-0%
+1%
+3%
+3% +11% +19% +19% +6%
70%
0%
-1%
+1%
+1%
+3% +10% +12% +19% +3%
90%
0%
-2%
-0%
-1%
+2%
20 80
TRIBE augmentation factor
C
+1%
TRIBE augmentation factor
30
80
0%
10
(b) Subject 2.
10%
Real data retained
Real data retained
55%
70%
20
TRIBE augmentation factor
80
Chance
100 30%
+8% +17% +19% +11% -13%
No
TR
100% real
37%
+4%
0x 0.25x 0.5x 0.75x 1x
%
29%
+2%
D
(1 ise
(1 IBE
Real data retained
29%
-0%
14
0x
Relative Improvement over real-only (%)
28%
0%
−30
100
8x
B
26%
50%
16x
(a) Subject 1. 10%
+9% +10% +22% +29% +37% +3%
120
)
)
0%
90% real 100% real 100% real, unaugmented
A
40
98% 101% 103% 108% 110% 115% 116% 115% 102%
4x
+7%
30
69%
99% 106% 107% 106% 83%
2x
+8% +10% +26% +39% +58% +70% +39%
+2%
0
0 1x
85%
+3%
0%
C
2
0
83%
0%
30%
TRIBE augmentation factor
Normalized performance (%)
Top-10 Accuracy (%)
80
81%
0x 0.25x 0.5x 0.75x 1x
D
100
79%
60
100%
14
0.25x 0.5x 0.75x
79%
10% 100 80
16
0x
B
10%
TRIBE augmentation factor
120
Normalized performance (%)
+1%
100
53%
80 50%
10%
Relative Improvement over real-only (%)
73%
43%
0
10
−10 −20
16x
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
D
0
+7% +11% +18%
-0%
2x
16x
4x
8x
Relative Improvement over real-only (%)
68%
47%
Real data retained
65%
42%
Top-10 Accuracy (%)
65%
37%
Real data retained
64%
33%
%
61%
31%
Real data retained
60%
31%
Real data retained
Real data retained
30%
B
31%
Relative Improvement over real-only (%)
A 10%
TRIBE augmentation factor
C
D
60 40 20
14 12 10 8 6 4 2
12
100
Top-10 Accuracy (%)
80
Normalized performance (%)
16 100
Top-10 Accuracy (%)
Normalized performance (%)
120
80 60 40 20
0.25x 0.5x 0.75x
1x
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
6 4
0
16x
TRIBE augmentation factor
8
2
0 0x
10
nly
lo
a Re
90% real 100% real 100% real, unaugmented
(1
%)
0
0%
+
) 8x
(1 IBE
TR
100% real
e ois
0%
+
0x
) 8x
0.25x 0.5x 0.75x
1x
2x
4x
8x
16x
TRIBE augmentation factor
(1
100% synthetic Chance 10% real
Chance
(c) Subject 5.
30% real 50% real 70% real
)
nly
lo
a Re
N
90% real 100% real 100% real, unaugmented
)
)
0%
(1
IBE
0% (1
+
8x
TR
100% real
ise
0%
+
8x
(1
No
Chance
(d) Subject 7.
Figure 5 Per-subject Top-10 operating grids for Ridge decoders on NSD. Each panel follows the same format as Figure 2,
but reports one subject before averaging across subjects.
13
51%
56%
75%
69%
85%
10%
0%
30%
0%
+3%
+3%
+3% +10% +22% +63% +50% +83%
50%
81%
80%
79%
82%
85%
70%
88%
90%
92%
95%
98% 108% 115% 126% 120%
62%
62%
69%
75%
85% 100% 112% 89%
100 80
98% 105% 100% 102%
60 40 90% 100%
99%
95%
100% 88%
93%
98% 102% 113% 128% 129% 116%
2x
4x
8x
40 0%
+1% +13% +22% +39% +63% +81% +45%
20 50%
0%
-1%
-2%
+2%
+5% +22% +30% +24% +26%
0
0 −20
70%
0%
+2%
+3%
+7% +10% +23% +30% +43% +36%
−40
20
92% 104% 104% 119% 135% 138% 119%
0x 0.25x 0.5x 0.75x 1x
Real data retained
62%
%
Real data retained
120 30%
60
90%
0%
16x
-5%
-6%
-2%
+2% +14% +29% +30% +17%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
8x
−60
A 10%
B
50%
56%
58%
63%
72%
83%
95%
82%
120 30%
62%
50%
77%
84%
90%
91%
70%
80%
87%
92%
96% 100% 115% 115% 128% 113%
65%
74%
83%
83%
88% 101% 91%
93%
100 80
95% 108% 110% 109% 99%
60 40 90% 100%
16x
88%
92%
95% 101% 105% 120% 127% 123% 117%
100% 86%
95% 109% 109% 118% 136% 127% 123%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
C
55%
2x
4x
8x
10%
0%
+13% +11% +16% +25% +44% +65% +91% +64%
30%
0%
+6% +19% +34% +34% +43% +63% +47% +50%
50%
0%
+8% +16% +18% +22% +40% +42% +41% +28%
70%
0%
+9% +15% +20% +25% +43% +44% +60% +41%
90%
0%
+4%
40 20
−40
16x
+7% +14% +19% +36% +44% +39% +33%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
D
0 −20
20 0
60
2x
4x
8x
−60
Relative Improvement over real-only (%)
48%
Real data retained
48%
%
48%
Real data retained
B
46%
Relative Improvement over real-only (%)
A 10%
16x
TRIBE augmentation factor
C
D
10
5
40 1x
2x
4x
8x
16x
)
TRIBE augmentation factor 100% synthetic Chance 10% real
30% real 50% real 70% real
100 80 60
ly
on
al Re
(1
0%
+
8x
0%
8x
1x
2x
4x
100% synthetic Chance 10% real
Chance
65%
30%
68%
76%
81%
88%
86% 105% 113% 105% 98%
50%
87%
88%
91% 102% 105% 124% 129% 121% 102%
79%
89%
63%
10%
89%
80 60
94% 104% 109% 116% 140% 137% 129% 118%
40 90%
100% 103% 118% 124% 133% 136% 135% 126% 118%
100%
100% 116% 121% 132% 132% 142% 142% 132% 126%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
+15% +4% +21% +19% +35% +44% +63% +15%
40
% 70%
0%
120 100
30%
0%
+11% +18% +29% +26% +54% +65% +54% +43%
20 50%
0%
+1%
70%
0%
+5% +16% +22% +29% +56% +53% +45% +32%
0 −20 −40
90%
0%
16x
+3% +18% +24% +33% +36% +35% +26% +18%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
30% real 50% real 70% real
2x
4x
8x
57%
61%
62%
63%
76%
30%
71%
68%
75%
75%
77%
88% 105% 100% 95%
50%
72%
77%
80%
82%
86% 102% 101% 100% 97%
80 60
0.25x 0.5x 0.75x
1x
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
+
0%
8x
(1
Chance
70% 90%
83%
97%
86%
88%
86%
10%
0%
-4%
+4%
+5%
+7% +29% +48% +70% +61%
30%
0%
-3%
+6%
+6%
+9% +25% +49% +42% +34%
50%
0%
+7% +12% +15% +21% +43% +41% +40% +35%
88% 104% 115% 112% 100% 96%
40
70%
0%
+4%
+4%
+6% +25% +38% +34% +20% +15%
94% 104% 111% 118% 124% 100% 91%
20
16x
0
100% 84% 100% 105% 111% 111% 116% 105% 79%
90%
0%
-9%
-3%
+8% +14% +22% +28% +3%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
16x
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
20 0 −20
14 12 10 8 6 4
2x
4x
8x
-7%
16x
TRIBE augmentation factor
C
D 18
120 100 80 60 40
15 12 10 8 5 2 0
16x
TRIBE augmentation factor
ise
No
40
0 0x
8x
−40
2
40
+
100
60
Normalized performance (%)
Top-10 Accuracy (%)
100
0% (1
100% real
87% 100% 95%
80
D
120
IBE
TR
B
59%
100%
16
140
)
)
0%
(1
90% real 100% real 100% real, unaugmented
10%
TRIBE augmentation factor
C Normalized performance (%)
+4% +17% +20% +42% +48% +39% +17%
20 0
)
nly
lo
a Re
Top-10 Accuracy (%)
74%
16x
A
Real data retained
66%
Relative Improvement over real-only (%)
57%
Real data retained
Real data retained
63%
8x
(b) Subject 2.
B
55%
5
TRIBE augmentation factor
(a) Subject 1. 10%
8
No
TR
100% real
A
10
0 0.25x 0.5x 0.75x
(1 ise
(1 IBE
90% real 100% real 100% real, unaugmented
+
12
2
0x
)
)
0%
15
Relative Improvement over real-only (%)
0.25x 0.5x 0.75x
120
40
0 0x
Top-10 Accuracy (%)
60
15
18
Real data retained
80
20
140
%
100
Normalized performance (%)
120
Top-10 Accuracy (%)
Normalized performance (%)
140
nly
lo
a Re
90% real 100% real 100% real, unaugmented
(1
%)
0
0%
+
) 8x
(1 IBE
TR
100% real
e ois
0%
+
0x
) 8x
0.25x 0.5x 0.75x
1x
2x
4x
8x
16x
TRIBE augmentation factor
(1
100% synthetic Chance 10% real
Chance
(c) Subject 3.
30% real 50% real 70% real
)
nly
lo
a Re
N
90% real 100% real 100% real, unaugmented
)
)
0%
(1
IBE
0% (1
+
8x
TR
100% real
ise
0%
+
8x
(1
No
Chance
(d) Subject 4.
Figure 6 Per-subject Top-10 operating grids for Ridge decoders on BOLD5000. Each panel follows the same format as
Figure 2, but reports one subject before averaging across subjects.
14
34%
36%
31%
30%
46%
42%
46%
54%
55%
64%
64%
51%
44%
10%
0%
-10%
-7%
-3%
+9% +16%
-0%
30%
0%
-10%
-1%
+16% +18% +37% +39% +9%
-6%
-18%
+3%
20
70% 90% 100%
67%
86%
67%
82%
70%
86%
73%
89%
74%
88%
81%
92%
73%
87%
98% 102% 107% 97% 108% 107% 96%
100% 115% 113% 109% 109% 105% 99%
0x 0.25x 0.5x 0.75x 1x
2x
4x
63%
74%
80%
52%
60
61%
40
70%
20
81%
76%
0
8x
16x
10 50%
0%
+1%
-6%
0
-22%
−10 70%
0%
-5%
-1%
+3%
+2%
+7%
+0%
-14% -29%
−20 90%
0%
+4%
+9%
-1%
+10% +9%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
-2%
-18% -28%
4x
8x
31%
27%
27%
28%
31%
39%
32%
27%
30%
46%
48%
49%
40%
51%
61%
64%
51%
44%
50% 70% 90%
16x
40 20 0
15
10
5
1x
2x
4x
8x
16x
)
TRIBE augmentation factor 100% synthetic Chance 10% real
30% real 50% real 70% real
ly
on
al Re
(3
60%
71%
73%
74%
82%
76%
59%
55%
60 67%
88%
90%
94%
96% 102% 90%
72%
99% 101% 104% 109% 100% 107% 103% 87%
100% 105% 109% 112% 109% 120% 107% 88%
2x
4x
8x
64%
0%
+
4x
+
0%
20
83%
0
24%
28%
33%
28%
50%
62%
65%
67%
70%
68%
70%
70%
54%
48%
70%
88%
82%
85%
85%
88%
83%
76%
65%
56%
43%
41%
51%
52%
52%
53%
46%
39%
80 60
100%
99% 100% 102% 100% 88%
70%
67%
20
100% 107% 105% 105% 105% 104% 92%
74%
72%
0
8x
16x
100% 98%
0x 0.25x 0.5x 0.75x 1x
2x
4x
0%
-10% -19%
30%
0%
+4%
+30% +33% +39% +42% +52% +33% +7%
-5%
90%
0%
+3%
50%
0%
60 40
−10
20
-8%
-3%
+12% +31% +15% +7%
30
-0%
+24% +25% +27% +28% +10%
-6%
+6%
0
+9% +14% +10% +14% +13% -13% -23%
0%
-7%
-3%
-3%
-0%
-5%
-14% -26% -36%
−20 90%
0%
-1%
-1%
+0%
+2%
+0%
-12% -30% -33%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
−30
10%
1x
2x
4x
8x
30% real 50% real 70% real
26%
25%
26%
25%
29%
34%
25%
53%
53%
61%
65%
65%
71%
65%
55%
48%
70%
80%
76%
81%
80%
82%
85%
78%
62%
57%
38%
43%
53%
50%
90%
52%
49%
43%
34%
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
16x
10
5
nly
)
)
0%
(3
IBE
0% (3
+
4x
ise
+
0%
4x
(3
No
TR
100% real
Chance
60 40
74%
67%
20
100% 103% 97% 102% 109% 105% 92%
80%
72%
0
8x
16x
96% 101% 104% 106% 96%
2x
4x
10%
0%
+10% +6% +10% +7% +26% +46% +9%
-9%
30%
0%
-5%
+7% +31% +22% +29% +22% +6%
-15%
50%
0%
-0%
+15% +21% +23% +33% +22% +2%
-11%
70%
0%
-6%
+0%
-2%
-23% -29%
90%
0%
+12% +19% +22% +25% +13% +7%
-13% -21%
15 10
10
)
lo
nly
a Re
(3
0%
+
4x
(3 IBE
TR
100% real
e ois
0%
+
4x
−20
80 60 40 20
4x
8x
−30
16x
18 15 12 10 8 5 2 0
0.25x 0.5x 0.75x
100% synthetic Chance 10% real
(c) Subject 5.
2x
D
1x
2x
4x
8x
16x
30% real 50% real 70% real
)
nly
lo
a Re
N
Chance
+6%
20
TRIBE augmentation factor
(3
+2%
TRIBE augmentation factor
100
0x
)
)
0%
0 −10
+0%
0x 0.25x 0.5x 0.75x 1x
0
16x
30 20
C
20
90% real 100% real 100% real, unaugmented
8x
TRIBE augmentation factor
Normalized performance (%)
Top-10 Accuracy (%) 1x
TRIBE augmentation factor
4x
)
80
91%
85%
0x 0.25x 0.5x 0.75x 1x
0 0.25x 0.5x 0.75x
2x
lo
21%
50%
16x
5
0x
0x 0.25x 0.5x 0.75x 1x
a Re
120
0
−20
B
23%
40%
D
20
-12% -20%
90% real 100% real 100% real, unaugmented
30%
100%
25
40
+4%
16x
100
120
60
+8%
15
A
TRIBE augmentation factor
80
+5% +10% +2%
0 0.25x 0.5x 0.75x
100% synthetic Chance 10% real
−10 70%
100
0
TRIBE augmentation factor
80
Chance
10
TRIBE augmentation factor
Normalized performance (%)
0%
TRIBE augmentation factor
20
C
70%
20 10
Top-10 Accuracy (%)
90%
40
Real data retained
41%
%
Real data retained
30%
-5%
(b) Subject 2.
10%
100
+4% +21% +26% +26% +40% +31% +1%
No
TR
100% real
26%
0%
16x
%
23%
50%
D
0x
Real data retained
20%
-6%
−30
100
4x
Relative Improvement over real-only (%)
22%
-7%
-5%
20
(3 ise
(3 IBE
B
25%
40
79%
(a) Subject 1. 10%
+6% +32% +10%
-13% +10% +32% +37% +9%
120
)
)
0%
90% real 100% real 100% real, unaugmented
A
-9%
+5%
0
0 0.25x 0.5x 0.75x
-7%
+3%
C Normalized performance (%)
Top-10 Accuracy (%)
60
+4%
0%
TRIBE augmentation factor
D
0x
58%
0x 0.25x 0.5x 0.75x 1x
20
80
0%
30%
30
80
100%
120 100
10% 100
TRIBE augmentation factor
C Normalized performance (%)
+4% +10% +10% +22% +9%
B
30%
Top-10 Accuracy (%)
50%
%
80
Real data retained
Real data retained
100
A 10%
Relative Improvement over real-only (%)
32%
Relative Improvement over real-only (%)
26%
Real data retained
30%
Real data retained
29%
%
28%
Real data retained
B
31%
Relative Improvement over real-only (%)
A 10%
90% real 100% real 100% real, unaugmented
)
)
0%
(3
IBE
0% (3
+
4x
TR
100% real
ise
0%
+
4x
(3
No
Chance
(d) Subject 7.
Figure 7 Per-subject Top-10 operating grids for Deep decoders on NSD. Each panel follows the same format as Figure 3,
but reports one subject before averaging across subjects.
15
78%
53%
81%
56%
79%
57%
80%
80%
70%
85%
93%
96% 101% 81%
72%
89%
86%
90%
80%
80%
98% 107% 84%
87%
60
100%
103% 108% 106% 110% 102% 116% 115% 92%
40
92%
20 0
100% 106% 102% 114% 123% 117% 117% 101% 93%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
30%
0%
-22% -23% -22% -18% -17%
-2%
-0%
20
50%
0%
70%
0%
90%
0%
16x
+11% +5%
-6%
-6%
+4%
-16%
-7%
+5%
+10% +14% +19%
+5%
+3%
+7%
-4%
-1%
+15% +19% +16% +5%
+1%
+6%
-6%
+16% +27%
-1%
-6%
+2%
+13% +12% -11% -11%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
8x
10
0
−10
−20
A 10%
B
41%
38%
C
47%
46%
46%
43%
44%
48%
39%
100 30%
49%
50%
72%
71%
79%
79%
76%
83%
81%
71%
74%
70%
81%
85%
71%
87%
92%
91%
94%
85%
72%
54%
90%
57%
59%
65%
62%
69%
72%
59%
80 60
100%
16x
40
89%
77%
20
100% 97% 103% 103% 111% 117% 103% 89%
89%
0
92%
91%
96%
88%
98% 104% 99%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
8x
10%
0%
-8%
30%
0%
+11% +16% +22% +32% +26% +41% +46% +22%
50%
0%
-1%
+10% +10% +6% +15% +12%
70%
0%
+6%
90%
0%
-2%
80 60 40
25 20 15 10 5
16x
1x
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
ly
on
al Re
(3
%)
0
0%
+
4
x)
+
0%
80 60 40
0.25x 0.5x 0.75x
1x
2x
4x
44%
50%
45%
48%
100% synthetic Chance 10% real
Chance
10%
63%
67%
66%
72%
71%
70%
64%
0%
-17% +13% +1%
+7% +22% +9% +17% +23%
20
68%
80 70%
70%
85%
86%
74%
89%
88%
75%
76%
82%
60 96%
76%
96%
93%
88%
87%
76%
92% 107% 100% 102% 95% 104% 107% 88%
93%
40
91%
20
30%
0%
50%
0%
70%
0%
+9% +15% +14% +25% +24% +22% +10% +19%
-2%
+23% +5% +26% +25% +7%
+13% -11% +13% +11% +4%
+4%
+8% +16%
-11% +11%
10 0 −10 −20
100%
100% 99% 113% 111% 105% 108% 120% 93%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
0
97%
90%
0%
16x
+15% +8% +10% +3% +12% +15%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
C
2x
4x
-5%
-1%
8x
16x
10%
30% real 50% real 70% real
80 60 40
44%
51%
47%
47%
56%
55%
49%
50%
66%
75%
73%
66%
75%
78%
85%
88%
79%
70%
70%
68%
80%
86%
74%
93%
89%
81%
93%
68%
56%
52%
59%
62%
76%
63%
67%
80 60
77%
96%
93%
80%
94%
88%
20
99% 102% 91%
84%
0
2x
16x
92% 102% 86%
100% 101% 103% 104% 97%
0x 0.25x 0.5x 0.75x 1x
40
4x
8x
1x
2x
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
4x
ise
+
0%
4x
(3
No
Chance
10%
0%
-11%
+5%
-5%
-5%
+14% +12%
30%
0%
+16%
-4%
-11%
+1%
+6% +29% +8% +14%
50%
0%
+13% +11%
-0%
+13% +18% +29% +34% +20%
70%
0%
-2%
90%
0%
+25% +20% +4% +19% +33% +12% +22% +14%
20 15 10
-14%
30 20
0 −10
+15% +23% +6% +33% +28% +16% +33%
−20
TRIBE augmentation factor
25
-0%
10
2x
4x
8x
−30
16x
TRIBE augmentation factor
C
D 25
120 100 80 60
20
15
10
5
40 0
16x
TRIBE augmentation factor
+
0x 0.25x 0.5x 0.75x 1x
0 0.25x 0.5x 0.75x
0% (3
100% real
42%
59%
100%
5
0x
IBE
TR
B
49%
30%
90%
Normalized performance (%)
100
)
)
0%
(3
90% real 100% real 100% real, unaugmented
100
D
Top-10 Accuracy (%)
Normalized performance (%)
)
nly
lo
a Re
30
120
−30
10
A
TRIBE augmentation factor
140
16x
4x
15
16x
Top-10 Accuracy (%)
90%
69%
%
50%
Real data retained
Real data retained
58%
8x
2x
(b) Subject 2.
100 30%
-17%
20
No
TR
100% real
51%
8x
%
42%
-3%
+6% +13% +8%
25
TRIBE augmentation factor
Real data retained
46%
-5%
5
0x
4
Relative Improvement over real-only (%)
34%
+4%
0
D
100
x)
B
41%
-10%
30
(a) Subject 1. 10%
+8% +14% +12% +17% +6%
TRIBE augmentation factor
120
(3 ise
(3 IBE
90% real 100% real 100% real, unaugmented
A
-12%
0
16x
TRIBE augmentation factor
+2%
0x 0.25x 0.5x 0.75x 1x
Real data retained
0.25x 0.5x 0.75x
-1%
−20
0 0x
30
35
Normalized performance (%)
Top-10 Accuracy (%)
100
-4%
−10
C
30
120
+6% +17%
10
TRIBE augmentation factor
D
+14% +12% +13% +4%
20
35
140
Normalized performance (%)
-7%
Top-10 Accuracy (%)
90%
%
85%
0%
100
71%
80 50%
10%
Relative Improvement over real-only (%)
63%
47%
Relative Improvement over real-only (%)
71%
47%
Real data retained
72%
44%
%
76%
43%
Real data retained
68%
44%
Real data retained
Real data retained
30%
B
57%
Relative Improvement over real-only (%)
A 10%
nly
lo
a Re
90% real 100% real 100% real, unaugmented
(3
%)
0
0%
+
) 4x
(3 IBE
TR
100% real
e ois
0%
+
0x
) 4x
0.25x 0.5x 0.75x
1x
2x
4x
8x
16x
TRIBE augmentation factor
(3
100% synthetic Chance 10% real
Chance
(c) Subject 3.
30% real 50% real 70% real
)
nly
lo
a Re
N
90% real 100% real 100% real, unaugmented
)
)
0%
(3
IBE
0% (3
+
4x
TR
100% real
ise
0%
+
4x
(3
No
Chance
(d) Subject 4.
Figure 8 Per-subject Top-10 operating grids for Deep decoders on BOLD5000. Each panel follows the same format as
Figure 3, but reports one subject before averaging across subjects.
16
Table 2 NSD scan-time interpretation of the operating grid. Each cell reports real scan hours + synthetic recording-
p
a=0
a = 0.25
a = 0.5
a = 0.75
a=1
a=2
a=4
a=8
a = 16
10% 30% 50% 70% 90% 100%
1+0 3+0 5+0 7+0 9+0 10+0
1+0.2 3+0.8 5+1.2 7+1.8 9+2.2 10+2.5
1+0.5 3+1.5 5+2.5 7+3.5 9+4.5 10+5
1+0.8 3+2.2 5+3.8 7+5.2 9+6.8 10+7.5
1+1 3+3 5+5 7+7 9+9 10+10
1+2 3+6 5+10 7+14 9+18 10+20
1+4 3+12 5+20 7+28 9+36 10+40
1+8 3+24 5+40 7+56 9+72 10+80
1+16 3+48 5+80 7+112 9+144 10+160
35%
39%
44%
45%
38%
10%
0%
+2%
+4%
+7% +10% +21% +29% +31% +18%
20
55%
50%
73%
55%
57%
59%
60%
66%
70%
69%
54%
80 74%
76%
78%
85%
88%
83%
63%
60 70% 90% 100%
86%
96%
85%
95%
87%
90%
92%
98% 100% 96%
74%
40
98% 101% 105% 112% 113% 106% 86%
20 0
100% 99% 102% 108% 112% 118% 120% 111% 95%
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
%
73%
Real data retained
Real data retained
100 30%
30%
0%
+0%
+7%
+9% +17% +22% +21%
-1%
10
50%
0%
+0%
+2%
+4%
+7% +14% +18% +13% -15%
0
70%
0%
-0%
+1%
+5%
+7% +13% +15% +11% -16%
−10
90%
0%
16x
-1%
+2%
+5%
+8% +14% +15% +9%
0x 0.25x 0.5x 0.75x 1x
TRIBE augmentation factor
2x
4x
−20
-11%
8x
30%
80%
50%
86%
70%
60 40
4x
8x
100% synthetic Chance 10% real
30% real 50% real 70% real
200
95%
88%
95%
84%
90%
86%
95%
91% 101% 108% 99% 100%
10%
0%
-1%
+0%
+2%
30%
0%
+1%
+5%
+7% +12% +21% +26% +19% +20%
50%
0%
+2%
+4%
+9% +11% +21% +26% +24% +22%
70%
0%
+1%
+4%
+8% +13% +19% +26% +24% +19%
90%
0%
+4%
+7% +10% +14% +22% +26% +27% +21%
+3% +13% +20% +15% +21%
20
100 80
97% 110% 117% 114% 111%
60
98% 103% 108% 116% 127% 124% 116%
100%
100% 102% 105% 112% 117% 129% 132% 141% 124%
2x
4x
8x
10
20 0
0 −10 −20
16x
0x 0.25x 0.5x 0.75x 1x
2x
4x
8x
16x
TRIBE augmentation factor
D 50
140
120
100
80
nly
(1
) 0%
lo
a Re
E RIB
0%
+
(1
T
100% real
e ois
0%
+
8x
0.25x 0.5x 0.75x
1x
2x
4x
8x
100% synthetic Chance 10% real
(a) NSD.
30% real 50% real 70% real
)
nly
lo
a Re
N
Chance
20
16x
TRIBE augmentation factor
(1
30
0 0x
)
)
8x
40
10
60
16x
90% real 100% real 100% real, unaugmented
88%
C
300
TRIBE augmentation factor
94%
160
0 2x
86%
TRIBE augmentation factor
20 1x
94%
81%
0x 0.25x 0.5x 0.75x 1x
100
0.25x 0.5x 0.75x
77%
96% 100% 103% 106% 112% 124% 130% 132% 122%
Normalized performance (%)
Median Rank
80
76%
90%
16x
400
100
75%
40
D
120
74%
120
500
140
0x
B
75%
TRIBE augmentation factor
C Normalized performance (%)
+4%
A 10%
Real data retained
34%
Median Rank
32%
%
32%
Real data retained
B
31%
Relative Improvement over real-only (%)
A 10%
Relative Improvement over real-only (%)
equivalent hours for the retained-real percentage p and augmentation factor a. Synthetic hours are not acquired; they express how much additional stimulus-conditioned fMRI the decoder sees if one TRIBE prediction is treated as one recording-equivalent image response.
90% real 100% real 100% real, unaugmented
)
)
0%
(1
IBE
0% (1
+
8x
TR
100% real
ise
0%
+
8x
(1
No
Chance
(b) BOLD5000.
Figure 9 Median-rank operating grids for Ridge decoders. Median rank is lower better in raw units; the normalized
operating grids invert this direction so that larger values indicate better performance, matching the Top-10 figures in the main text.
C
Image reconstruction with DynaDiff
To adapt DynaDiff to be trained on both synthetic and real fMRI data, we slightly modify the brain module architecture in the following way: i) We add an initial linear layer to process independently real or synthetic data (Modality-specific initial layer), and ii) As synthetic fMRI data contains only one TR, we only use the last temporal aggregation layer for real data and keep only the first TR layer for synthetic fMRI data (Modality-specific temporal aggregation). All other architectural choices and training hyperparameters follow the original DynaDiff recipe (Careil et al., 2025). Each run was trained on 8 NVIDIA H100 GPUs for approximately 2 days.
D
Scan-Time Interpretation of Operating Grids
Tables 2 and 3 translate the abstract operating grid coordinates into approximate scan-time budgets. For a dataset with real-fMRI training budget H, retaining p% real data corresponds to (p/100)H real scan hours, while augmentation factor a adds a(p/100)H synthetic recording-equivalent hours. Only the first quantity is acquired in the scanner; the second is generated by TRIBE. This framing separates acquisition cost from effective training-set scale.
17
Table 3 BOLD5000 scan-time interpretation of the operating grid. Each cell reports real scan hours + synthetic recording-
equivalent hours for the retained-real percentage p and augmentation factor a. Synthetic hours are not acquired; they express how much additional stimulus-conditioned fMRI the decoder sees if one TRIBE prediction is treated as one recording-equivalent image response. p
a=0
a = 0.25
a = 0.5
a = 0.75
a=1
a=2
a=4
a=8
a = 16
10% 30% 50% 70% 90% 100%
1.3+0 4.0+0 6.7+0 9.3+0 12.0+0 13.3+0
1.3+0.3 4.0+1.0 6.7+1.7 9.3+2.3 12.0+3.0 13.3+3.3
1.3+0.7 4.0+2.0 6.7+3.3 9.3+4.7 12.0+6.0 13.3+6.7
1.3+1.0 4.0+3.0 6.7+5.0 9.3+7.0 12.0+9.0 13.3+10.0
1.3+1.3 4.0+4.0 6.7+6.7 9.3+9.3 12.0+12.0 13.3+13.3
1.3+2.7 4.0+8.0 6.7+13.3 9.3+18.7 12.0+24.0 13.3+26.7
1.3+5.3 4.0+16.0 6.7+26.7 9.3+37.4 12.0+48.0 13.3+53.4
1.3+10.7 4.0+32.0 6.7+53.4 9.3+74.7 12.0+96.0 13.3+106.7
1.3+21.3 4.0+64.0 6.7+106.7 9.3+149.4 12.0+192.1 13.3+213.4
E
Dataset Licenses
We used the Natural Scenes Dataset (NSD) (Allen et al., 2022) and BOLD5000 (Chang et al., 2019) in accordance with their respective license terms and conditions. For both datasets, we accessed the data through the official distribution channels, agreed to the applicable terms of use, and used the data only for research purposes consistent with those terms.
18