SAMPLE-CONDITIONED REPRESENTATION SELECTION FOR AUDIO FEW-SHOT LEARNING Fengrui Liu1,2,∗,‡,§ , Ningxin Shen3,∗ , Yi Li1 , Yiwei Fu2 , Feng Liu4 , Senior Member, IEEE, Jiangmeng Li1 , †
arXiv:2609.17076v1 [cs.AI] 15 Sep 2026
1
National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences 2 School of Computer Science and Technology, East China Normal University 3 School of Computer Science, Nanjing University 4 School of Psychology, Shanghai Jiao Tong University ∗ Equal contribution. ‡ Project lead. † Corresponding author. § Work done while the first author was an intern at the Institute of Software, Chinese Academy of Sciences. ABSTRACT Few-shot audio classifiers may rely on foreground–background cooccurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10% of channels explain 82.80% of the nullcorrected shift contribution. We propose S AMPLE S ELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source classifier frozen. Training uses differentiable Gumbel Top-k selection with foreground classification and cross-background contrastive losses; inference uses deterministic Top-k masks and support-only linear adaptation. Across ResNet12 and Conv64 in 5-way 1-shot and 5-shot evaluation, S AM PLE S ELECT gives the best OOD accuracy among the compared methods and improves the matched full-representation control by 4.90–8.38 percentage points. Ablations and representation analyses further support the learned selection mechanism.Codes available at https://github.com/Cross-Innovation-Lab/SAMPLESELECT/ Index Terms— Few-shot audio classification, background shift, representation selection, contrastive learning 1. INTRODUCTION Few-shot audio classification recognizes novel sound classes from limited labeled examples, commonly through episodic adaptation or source-trained representations [1–5]. Real recordings also contain recurring foreground–background co-occurrences that models can exploit as shortcuts [6]. When those associations change, previously useful features can become unreliable, while the small support set offers little evidence for identifying which coordinates remain trustworthy. MetaAudio studies transfer across acoustic domains [7], whereas SpurAudio changes foreground–background co-occurrences between IID and OOD episodes [8]. Existing representation adaptation and feature reweighting methods improve few-shot transfer [9–13], but do not directly target input-varying background sensitivity. This motivates selecting features for each example rather than always exposing the full representation or one global subset. A frozen ResNet12 reveals that the shift is highly concentrated and class dependent: the top 10% of its 640 channels account for 82.80% of the corrected shift contribution, and the affected channels vary substantially across foreground classes. We therefore propose
Baseline:
Full Sample Representations SUPPORT �
Episodic Classification Space Potential OOD Misclassification
D channels
Blender
● D channels
Pig·IID Pig·OOD
★ IID query
◆ Pig class ●
OOD query → other class
Ours:
Selected Sample Representations SUPPORT �
●
◆ Other class
★ OOD query ×
Pig
QUERY �
●
●
Crow
Episodic Classification Space More Robust OOD Classification
D channels
Blender
●
Crow
●
◆ Other class
Pig
QUERY � Pig·IID Pig·OOD
● D channels
●
★ IID query
◆ Pig class ●
★ OOD query ✓
OOD query → Pig class
Fig. 1. Overview of few-shot classification under background shift. F ULL R EP passes all representation coordinates to a fresh episodic linear head; S AMPLE S ELECT independently masks each support and query example before fitting and prediction. The class-space plots are schematic, not measured embeddings.
S AMPLE S ELECT, a fixed-budget, sample-conditioned selector. A lightweight scorer predicts feature importance from each input while the encoder and source classifier remain frozen. Training uses a differentiable Gumbel relaxation with foreground classification and grouped cross-background contrastive supervision [14–18]; inference uses deterministic Top-k masking and a support-only linear head. On SpurAudio, across ResNet12 and Conv64 with 5-way 1-shot and 5-shot episodes, S AMPLE S ELECT attains the highest OOD accuracy among the compared methods in all four settings and improves the matched F ULL R EP control by 4.90–8.38 percentage points. Ablations and further analyses support learned selection, cross-background supervision, and the functional relevance of the selected channels. Our contributions are: (1) evidence that background co-occurrence shift is concentrated and class dependent in frozen audio representations; (2) an inductive sample-conditioned fixed-budget selector trained without changing the source representation; and (3) consistent OOD gains
Stage 1 Representation
Stage 2 Selector Training
Learn source-domain representations
Sample-Conditioned Channel Selector Gφ
Learn sample-conditioned channel selection Channel scores sᵢ GAP C×H×W→C
Illustrative grouped Log-Mel samples
fθ
Feature map Fi
Backbone
Log-Mel spectrogram 1×128×157
25-way logits
LCE
Cross-entropy loss
(frozen)
Backbone
Feature map Fi
...
Update φ only
Lcon
Lsel=LCE+λLcon
Grouped contrastive loss
F̃ ᵢ=Fᵢ⊙mᵢ Selected
feature map F̃ᵢ
airplane rain ←positive pair: same class, different backgrounds→ 2 of 4 foreground classes shown
25-way source classifier
Soft Gumbel Top-k
...
C
GAP C×H×W→C
Channel mask mᵢ mᵢ ∈ [0,1]C
Source Linear Head
Source label y
Scorer
C
sᵢ ∈ ℝ
fθ
church bells engine idling ←positive pair: same class, different backgrounds→
Class b:
HSrc
MLP
... zᵢ ∈ ℝ
Class a:
Flatten: C×H×W→CHW
Backpropagation updates θ and Hsrc
Channel descriptor zi
H
Src Flatten C×H×W→CHW (frozen)
LCE
25-way logits
Cross-entropy loss
Source Linear Head
25-way source classifier
Source label y
Stage 3 Few-shot evaluation Selected representations
Frozen fθ and Gφ · fresh episodic head H𝓔
Support features R̃ S
Illustrative target episode E = (S, Q)
...
Support set S 5-way K-shot, 2 classes illustrated
...
fθ
Class a
(frozen)
Class b
Backbone
Query set Q
Feature map Fi
GAP C×H×W→C
Gφ Selector (frozen)
Channel mask mᵢ ... mᵢ ∈ [0,1]C
F̃ ᵢ=Fᵢ⊙mᵢ
Selected feature map F̃ᵢ
Fit on S with yS (CE)
...
Flatten C×H×W→CHW
...
(fresh) ...
Query labels hidden during prediction
5-way logits
HE
Query features R̃ Q ...
Query prediction ŷQ
Predict on Q
Episodic Linear Head 5-way episodic classifier
... .. . Query 1
Query 2
Fig. 2. Three-stage S AMPLE S ELECT pipeline. The source encoder/head are trained then frozen; the selector learns sample-wise masks with classification and cross-background contrastive losses; inference uses deterministic Top-k masks and a fresh support-only episodic head.
across two backbones and two shot settings, supported by ablation and representation analyses. 2. METHODOLOGY Figure 2 summarizes S AMPLE S ELECT. We first learn a source representation, then freeze it and train only a sample-conditioned selector; few-shot evaluation uses deterministic masks and a fresh support-only head. 2.1. Source Representation Learning For input xi with foreground label yi , the selectable descriptor zi and classifier representation hi are ResNet12: Fi = fθ (xi ) ∈ R640×4×5 , zi = GAP(Fi ), hi = vec(Fi ) ∈ R12800 , Conv64: zi = hi = fθ (xi ) ∈ R1600 .
(1)
We train the encoder and source classifier Hψ on source foreground classes, N
1 X CE(Hψ (hi ), yi ) , θ,ψ N i=1
(θ⋆ , ψ ⋆ ) = arg min
(2)
The selected classifier representation is ( vec(mi ⊙c Fi ) , ResNet12, e hi = mi ⊙ hi , Conv64,
(4)
where ⊙c broadcasts a channel mask over spatial locations. Zero masking preserves coordinate alignment across examples. The frozen source head preserves foreground information through B 1 X CE Hψ⋆ (e hi ), yi . Lcls = (5) B i=1 For background group bi , define P (i) = {p ̸= i : yp = yi , bp ̸= bi }. Let qi = GAP(mi ⊙c Fi ) for ResNet12 and qi = e hi for Conv64, with normalized similarity uij = (qi /∥qi ∥2 )⊤ (qj /∥qj ∥2 ). For a valid anchor, X exp(uip /T ) 1 L(i) . (6) log P con = − |P (i)| a∈G(i)\{i} exp(uia /T ) p∈P (i)
where G(i) is its comparison group. We average over anchors with valid positives and minimize Lsel = Lcls + λcon Lcon , updating only ϕ. Thus classification preserves category information, while contrastive supervision favors same-class consistency across observed backgrounds. Background labels are used only in selector training.
and freeze both thereafter.
2.3. Inductive Few-Shot Evaluation
2.2. Sample-Conditioned Selector Learning
At inference, each example receives an independent deterministic mask
A scorer predicts one score per selectable coordinate, with retention ratio r: si = Gϕ (stopgrad(zi )) ∈ RD , k = ⌊rD⌋, (3) D mtr = RelaxedTopK (s , k) ∈ [0, 1] , i i τ where D = 640 for ResNet12 and 1600 for Conv64. Gϕ acts independently on each input. During training, sequential GumbelSoftmax draws without replacement provide a differentiable Top-k relaxation [14, 15]; exact cardinality is enforced at inference.
meval i,c = 1[c ∈ TopK(si , k)] ,
∥meval ∥0 = k. i
(7)
For a 5-way K-shot episode E = (S, Q), a fresh linear head is fitted only on masked support representations and then applied to each masked query: X 1 WE⋆ = arg min CE W e h(xj ), yj , W |S| (xj ,yj )∈S (8) h i ybq = arg max WE⋆ e h(xq ) . c
c
Table 1. Results on SpurAudio with ResNet12 and Conv64. Accuracy values are percentages. ∆ = IID − OOD. F ULL R EP denotes the matched no-selection control. Bold indicates the best IID or OOD accuracy among the compared methods. ResNet12 1-shot Method
IID
OOD
Conv64 5-shot
∆
IID
OOD
1-shot ∆
IID
OOD
5-shot ∆
IID
OOD
∆
Baseline++ (2019) 57.694 53.078 4.616 75.032 65.490 9.542 50.705 46.667 4.038 65.286 56.934 8.352 R2D2 (2019) 57.995 54.228 3.767 74.760 66.149 8.611 41.093 39.420 1.673 68.695 61.873 6.822 ANIL (2020) 54.129 50.893 3.236 66.015 58.352 7.663 49.522 46.499 3.023 64.763 56.009 8.754 BDCSN (2022) 61.241 57.961 3.280 73.564 66.691 6.873 44.992 42.374 2.618 58.951 55.031 3.920 PADDLE (2022) 52.196 46.842 5.354 70.380 59.299 11.081 50.016 44.267 5.749 66.767 55.912 10.855 Proto-LP (2023) 59.852 56.036 3.816 74.762 67.454 7.308 57.054 53.446 3.608 69.916 61.328 8.588 BPA (2024) 60.325 55.828 4.497 75.964 63.516 12.448 49.158 45.369 3.789 66.173 50.924 15.249 ECPE (2026) 61.397 56.391 5.006 74.890 65.642 9.248 55.836 51.974 3.862 69.038 59.536 9.502 F ULL R EP S AMPLE S ELECT
57.169 51.997 5.172 74.369 63.259 11.110 51.470 47.032 4.438 67.641 56.813 10.828 62.307 57.979 4.327 77.523 68.161 9.362 57.357 53.635 3.721 72.993 65.188 7.805
No query label, query-set statistic, or test-time background label is used for masking or adaptation. The matched F ULL R EP control uses the same support-only procedure with mi = 1. 3. EXPERIMENTS 3.1. Experimental Setup Datasets. We evaluate on SpurAudio [8], which contains 25 training, 5 validation, and 8 test foreground classes and controls foreground– background co-occurrence to construct IID and OOD few-shot episodes. We follow the standard 5-way 1-shot and 5-shot protocols with 10 query examples per class and report the mean over 1,000 episodes. We report IID accuracy, OOD accuracy, and the shift gap ∆shift = AccIID − AccOOD . OOD accuracy is the primary metric. Backbones and Baselines. We evaluate two representation architectures. ResNet12 performs channel-level selection over 640 channels, while Conv64 performs coordinate-level selection over a 1,600dimensional projected representation. We compare with representative methods reported under the SpurAudio protocol and include F ULL R EP as a matched control using the same frozen backbone and episodic classifier without representation selection, thereby isolating the effect of selection. Implementation Details. Inputs are standardized 1 × 128 × 157 log-Mel spectrograms. The source encoder is trained for 30 epochs and then frozen. ResNet12 uses retention ratio r = 0.7 with k = 448, while Conv64 uses r = 0.8 with k = 1280. Unless otherwise stated, selector training uses Gumbel temperature τ = 0.3, contrastive temperature T = 0.07, and λcon = 0.02. During few-shot evaluation, the encoder and selector remain frozen and only a support-only linear classifier is optimized for each episode. 3.2. Main Results Table 1 shows that S AMPLE S ELECT achieves the highest OOD accuracy among the compared methods in all four backbone and shot settings. It reaches 57.979% and 68.161% with ResNet12, and 53.635% and 65.188% with Conv64 for 1-shot and 5-shot evaluation, respectively. IID accuracy is also highest in each setting. Compared with the matched F ULL R EP control, S AMPLE S ELECT improves OOD accuracy by 5.982 and 4.902 percentage points for ResNet12, and by 6.603 and 8.375 points for Conv64. The gains are consistent
across both representation architectures and shot settings, showing that sample-conditioned selection improves the use of frozen representations under class–background co-occurrence shift. 3.3. Ablation and Sensitivity Analysis Figure 3(a) evaluates the main components of S AMPLE S ELECT. Replacing learned selection with Random Soft decreases OOD accuracy by 2.882 and 1.254 points in the 1-shot and 5-shot settings. Removing the Gumbel relaxation reduces accuracy by 2.458 and 2.322 points, while removing the cross-background contrastive objective gives drops of 1.477 and 0.968 points. A learnable global selector achieves 55.799±1.372 and 66.869±1.004 OOD accuracy in the 1shot and 5-shot settings, respectively, trailing S AMPLE S ELECT by 2.180 and 1.292 points. These results support learned, differentiable, sample-conditioned selection and cross-background supervision. Figures 3(b,c) study the two main hyperparameters. Performance changes only modestly for r ∈ [0.4, 0.9], and values near the default r = 0.7 give similar OOD gains. For λcon , every tested positive value improves the OOD mean over λcon = 0 in both shot settings. We use λcon = 0.02 as a shared setting rather than tuning it separately for each evaluation condition. 3.4. Further Analysis The previous experiments establish the performance benefit of S AM PLE S ELECT. We next examine the representation behavior that motivates sample-conditioned selection. 3.4.1. Concentrated and Class-Dependent Shift For foreground class y and ResNet12 channel c, we measure the IID-to-OOD distribution change using the 1-Wasserstein distance and correct it with a within-class permutation null: dy,c = W1 PbIID (zc | y), PbOOD (zc | y) , (9) [gy,c ]2+ gy,c = dy,c − Eπ [dπy,c ], ωy,c = P . 2 ′ c′ [gy,c ]+ + ϵ As shown in Fig. 4(a), the top 10%, 20%, and 30% of channels account for 82.80%, 92.85%, and 96.76% of the positive corrected shift mass, with a mean Gini coefficient of 0.887. The number of
Fig. 3. Ablation and sensitivity analysis on ResNet12. (a) OOD accuracy drop after removing individual components of S AMPLE S ELECT. (b) OOD gain over F ULL R EP for different retention ratios r. (c) OOD gain over F ULL R EP for different contrastive weights λcon . Dashed lines indicate the default settings used in the main experiments.
(a) Shift concentration
(b) Shift and OOD degradation
(c) Functional blocking
Fig. 4. Further analysis on ResNet12. (a) Cumulative contribution of channels with the largest null-corrected representation shifts. (b) Mean Spearman correlation across three seeds between episode-level shift mass and IID-to-OOD degradation for accuracy, confidence, and classification margin. (c) Accuracy drop after blocking the highest-scored, random, or lowest-scored 30% of channels while keeping the episodic classifier fixed.
significantly shifted channels also varies strongly across foreground classes. For example, sneezing and pig show broader affected subsets than blender and crackling fire. This class dependence supports input-dependent rather than globally fixed selection. This analysis is diagnostic rather than a causal identification result: IID and OOD routes contain different foreground recordings, and concentration refers specifically to the null-corrected statistic above. Neither testset channel rankings nor this statistic supervise the selector; training uses source-class labels and observed background variation only. 3.4.2. Shift and OOD Degradation We correlate episode-level representation shift with IID-to-OOD predictive degradation over 1,000 episodes. Figure 4(b) summarizes the mean correlations across three random seeds for accuracy, predicted probability, and classification margin. Across the three seeds, two shot settings, and three measures, all 18 Spearman correlations are positive, with ρ between 0.268 and 0.455. Episodes with larger representation shift therefore tend to exhibit larger predictive degradation, establishing an association between the representation diagnosis and OOD performance. 3.4.3. Selector Scores Reflect Functional Importance We rank the 640 ResNet12 channels by their selector scores and block the highest-scored, random, or lowest-scored 30% while keeping the episodic classifier fixed. Figure 4(c) shows that blocking the highestscored channels decreases accuracy by 3.033 percentage points on
average, compared with 1.064 points for random blocking, while removing the lowest-scored channels produces almost no change. The learned scores therefore reflect clear differences in the functional contribution of representation channels. Together, these analyses connect the empirical motivation in Sec. 1 with the behavior of the learned selector: background-induced changes are concentrated and class-dependent, larger representation shifts are associated with larger OOD degradation, and the selector assigns higher scores to channels with greater functional importance. 4. CONCLUSION We presented S AMPLE S ELECT, a sample-conditioned fixed-budget representation selector for few-shot audio classification under class– background co-occurrence shift. The method keeps the source encoder frozen, learns per-example feature scores using source-class discrimination and cross-background contrastive supervision, and performs deterministic Top-k selection at inference. Masks are perexample and support-only; no query label/statistic or test-time background label is used. On SpurAudio, S AMPLE S ELECT achieves the best OOD accuracy in all four backbone/shot settings and improves the matched F ULL R EP control by 5.982, 4.902, 6.603, and 8.375 percentage points. Ablation and further analysis show that the gains depend on learned selection and cross-background supervision and are consistent with a representation structure in which backgroundinduced changes are concentrated, class-dependent, and associated with predictive degradation.
5. REFERENCES [1] O. Vinyals, C. Blundell, T. Lillicrap, K. Kavukcuoglu, and D. Wierstra, “Matching networks for one shot learning,” in Advances in Neural Information Processing Systems, vol. 29, 2016. [2] J. Snell, K. Swersky, and R. S. Zemel, “Prototypical networks for few-shot learning,” in Advances in Neural Information Processing Systems, vol. 30, 2017. [3] W.-Y. Chen, Y.-C. Liu, Z. Kira, Y.-C. F. Wang, and J.-B. Huang, “A closer look at few-shot classification,” in International Conference on Learning Representations, 2019. [4] F. Liu, R. Huang, Q. Zheng, Y. Wang, and F. Liu, “From physics to representation: Audio learning with synthetic pre-training via procedural generation,” in Proceedings of the 2026 International Conference on Multimedia Retrieval, 2026, pp. 595–604. [5] F. Liu, R. Huang, Q. Zheng, Y. Wang, and F. Liu, “PRP: Procedural-to-real masked pre-training for transferable and interpretable audio representations,” IEEE Transactions on Audio, Speech and Language Processing, 2026. [6] R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann, “Shortcut learning in deep neural networks,” Nature Machine Intelligence, vol. 2, pp. 665– 673, 2020. [7] C. Heggan, S. Budgett, T. Hospedales, and M. Yaghoobi, “MetaAudio: A few-shot audio classification benchmark,” arXiv:2204.02121, 2022. [8] G. Abu Ayoub, M. Tukan, and L. Mualem, “SpurAudio: A benchmark for studying shortcut learning in few-shot audio classification,” arXiv:2605.13672, 2026. [9] H. Li, D. Eigen, S. Dodge, M. Zeiler, and X. Wang, “Finding task-relevant features for few-shot learning by category traversal,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 1–10. [10] N. Dvornik, C. Schmid, and J. Mairal, “Selecting relevant features from a multi-domain representation for few-shot classification,” in Computer Vision – ECCV 2020, 2020, pp. 769–786.
[11] W.-H. Li, X. Liu, and H. Bilen, “Cross-domain few-shot learning with task-specific adapters,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 7161–7170. [12] S. Lee, W. Moon, and J.-P. Heo, “Task discrepancy maximization for fine-grained few-shot classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 5331–5340. [13] J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7132–7141. [14] E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with Gumbel-Softmax,” in International Conference on Learning Representations, 2017. [15] W. Kool, H. Van Hoof, and M. Welling, “Stochastic beams and where to find them: The Gumbel-Top-k trick for sampling sequences without replacement,” in Proceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 3499–3508. [16] M. F. Balın, A. Abid, and J. Zou, “Concrete autoencoders: Differentiable feature selection and reconstruction,” in Proceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 444–453. [17] P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” in Advances in Neural Information Processing Systems, vol. 33, 2020. [18] C. Sgouropoulos, C. Nikou, S. Vlachos, V. Theiou, C. Foukanelis, and T. Giannakopoulos, “Prototypical contrastive learning for improved few-shot audio classification,” in 2025 IEEE 35th International Workshop on Machine Learning for Signal Processing, 2025.