Improving Sparse Autoencoder with Dynamic Attention
arXiv:2604.14925v1 [cs.LG] 16 Apr 2026
Dongsheng Wang, Jinsen Zhang, Dawei Su, Hui Huang* College of Computer Science and Software Engineering, Shenzhen University, China {dongshengwang,2400101100,2023110003}@szu.edu.cn, [email protected]
Abstract Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts. However, identifying the optimal level of sparsity for each neuron remains challenging in practice: excessive sparsity might lead to poor reconstruction, whereas insufficient sparsity harms interpretability. While existing activation functions such as ReLU and TopK provide certain sparsity guarantees, they typically require additional sparsity regularization or cherry-picked hyperparameters. We show in this paper that adaptive sparse attention mechanisms using sparsemax can bridge this trade-off, due to their ability to determine the number of concepts in a datadependent manner. Specifically, we first explore a new class of SAEs based on the cross-attention architecture with the latent features as queries and the learnable dictionary as the key and value matrices. To encourage sparse pattern learning, we employ a sparsemax-based attention strategy that automatically infers a sparse set of concepts according to the complexity of each neuron, resulting in a more flexible and efficient activation function. Through comprehensive evaluation and visualization, we show that our approach successfully achieves lower reconstruction loss while producing high-quality concepts. Moreover, the sparsity level automatically determined by our approach can serve as tuning guidance to improve existing SAEs. The code is available https://github.com/qyj-bkjx/Sparsemax-SAE.
1. Introduction The impressive gains in reasoning and accuracy of recent large-scale machine learning models like CLIP [6, 55] and GPT [1, 54] have generally come at the cost of a loss of transparency into their functioning. Typically, neurons in these models are polysemantic, and they respond to seemingly unrelated inputs simultaneously. This can be explained as superposition [17], where models learn more independent features than they have neurons by viewing each feature as a * Corresponding author.
/
RuLU
/
/ IMG 1
TopK IMG 2
/ Sparsemax
Figure 1. Comparisons of our proposed Sparsemax SAE with previous SAEs (left) and histogram of activation frequency of the learned concepts on the ImageNet dataset (right). ReLU-based SAEs often suffer from the feature shrinkage issue, while TopK-based models tend to produce dead concepts as they only keep K largest concepts. In contrast, our Sparsemax-based SAE dynamically select the number of concepts based on feature complexity, thereby discovering more concepts.
linear combination of neurons. Fortunately, sparse autoencoders (SAEs) [19, 27, 28, 32, 45] have emerged as a promising technique for addressing this fundamental challenge by learning an overcomplete yet sparse representation of neural activations, effectively disentangling these superimposed features into more interpretable concepts [20, 60, 63, 66]. Despite its successes in reversing the effects of superposition, determining the optimal sparse level for the latent features remains an open problem. For example, assigning too many concepts to each feature may compromise interpretability, while insufficient concepts can degrade reconstruction—both scenarios lead to suboptimal concept learning. Early SAEs [28, 34] adopt the ReLU as the activation function and combine it with L1 regularization to balance sparsity and reconstruction. However, the L1 penalty often leads to feature shrinkage, where all activations tend toward zero (Seen as Fig. 1). GatedReLU [57] suggests decomposing ReLU into direction selection and magnitude estimation functions via the gate mechanism. JumpReLU [58] finds that zeroing out activations below a positive threshold is a better option. However, these models require additional regularizations to prompt the sparsity, and the balancing coefficient needs to be carefully selected to achieve satisfactory
performance [27]. On the other hand, recent attempts propose to limit the number of concepts explicitly. For example, TopK SAEs [20] employ K-sparse autoencoder [37] to directly choose the K largest concepts and zero the rest. This approach eliminates the need for an explicit sparsity penalty but imposes a rigid constraint on the number of active concepts per sample. BatchTopK SAEs [5] relax the top-K constraint to a batch-level constraint, enabling the SAEs to represent each sample within a batch with a variable number of concepts, resulting in more flexible and efficient utilization of the concept dictionary. Unfortunately, both TopK and BatchTopK view K as a hyperparameter, and how to set K properly remains an unsolved problem. In this paper, we aim to improve SAEs with adaptive sparse attention mechanisms under the cross-attention framework. Moving beyond the traditional single-layer MLP-style encode-decode structure, we here explore a new class of SAEs based on the transformer architecture due to its successes in various tasks [49, 53, 65, 68]. Specifically, we first view the to-be-learned dictionary as a set of concept vectors, which will be used as the key and value matrices via the corresponding projections. For each latent feature in the neural networks, we view it as the query and apply the cross-attention operation to obtain the reconstructed feature. Intuitively, the calculation of attention weights can be viewed as the encoding stage of SAEs, whose output measures the relevance score between the query and concepts. Notably, this transformer-based SAE connects the encoding and decoding stages by sharing the same concept vectors, rather than viewing them as two independent MLPs, achieving coherent, high-quality concept learning. With the designed transformer-based architecture, it is natural to replace the softmax with any well-studied sparse attention operations [8, 14, 24, 50]. Here, we adopt sparsemax [39] as our solution. On one hand, sparsemax is differentiable everywhere, and can be easily applied with gradientbased optimization. On the other hand, sparsemax has the ability to assign exactly zero probability to some of its outputs, sharing the same motivation as SAEs. Most importantly, unlike previous works that require hyperparameter K to constrain the number of activations, sparsemax dynamically estimates a threshold function for each sample according to its complexity. This enables the model to output the most relevant concepts while truncating others to zero. As shown in Fig 1, TopK-based SAEs sometimes fail to correctly assign the number of concepts due to the incorrect K. In contrast, our sparsemax successfully assigns six concepts to complex images and two concepts to simple images. Our sparsemax can be viewed as a more precise version of BatchTopK, where we set K at the sample level rather than the batch level, thereby achieving greater flexibility and accuracy in sparse estimation.
To sum up, our contributions are as follows: • We formalize a novel transformer-based SAEs under the cross-attention framework, which bridges the gap between the encoder and decoder of SAEs by sharing the same concept vectors, resulting in more coherent concept learning. • A novel sparsemax function is developed to replace the original softmax function in the cross-attention operation. Sparsemax can determine the number of activations dynamically for each sample, without any regularization or hard TopK truncation. • We provide extensive validation across image and text tasks, demonstrating that our approach not only captures coherent concepts but also achieves superior reconstruction results.
2. Related Work Mechanistic interpretability with SAEs. Mechanistic interpretability aims to uncover and explain the black box characteristics, enabling models to understand input data and generate reasonable responses [56]. Recently, Sparse Autoencoder (SAEs) have been applied to language models due to their inherent ability to generate interpretable latent concepts [28, 59]. Building upon the standard architecture with ReLU activation [4], a series of studies have developed numerous improvements to the original design. For example, Gated SAE [57], Switch SAE [43] aim to design complex encoders to ensure sparse outputs; Focusing on the feature shrinkage issue [4], JumpReLU [58] introduces threshold parameters for each concept to truncate the concepts with small scores. Topk-based SAEs [5, 20] directly keep K concepts with large scores and zero out the others; In terms of the sparsity regularization, P-annealing [30] and mutual feature regularization (MFR) [38] are designed as alternatives to L0 and L1 loss. Inspired by the successful application of SAE in LLMs, PatchSAE and its variants explore to train SAEs on top of the CLIP and DiNOv2 [34, 46, 52, 62], showing great potential in interpreting visual concepts. Additionally, there has also been interest in steering the generation process of diffusion models via SAEs [11, 20]. More recently, several studies have increasingly applied SAEs to multimodal LLMs [47, 71], and show that SAE can learn shared concepts across the vision and text modalities. However these model primarily extend existing SAEs (such as ReLU and TopK) to the vision domain, without designing new SAE architectures. Sparse Attention Mechanisms. Traditional attention mechanisms commonly employ the softmax transformation to convert scores into probability distributions [65]. However, softmax produces dense distributions, assigning nonzero attention weights to all elements, which limits interpretability and efficiency [33]. SlidingWindow is a commonly used strategy that allows the query to compute at-
tention only within a fixed windlow [2, 7, 69]. However, these models rely on pre-defined sparse patterns, limiting their potential for application in SAEs. Top-K attention [25] shares similar ideas with TopK SAEs and selects a set of keys with the K largest scores. SeerAttention [21] separates queries and keys into spatial blocks and performs blockwise selection for sparsity. α-entmax attention [51] provides natural, input-dependent sparsity patterns with an exact and differentiable transformation, attracting increasing attention in recent research. For example, Adaptively sparse transformers [10] employ α-entmax attention where attention heads can learn α dynamically. SparseFinder [64] aims to address the efficiency issues of α-entmax by predicting a prior. Sparse Flash Attention [12, 23] combines the efficiency of GPU-optimized algorithms with the sparsity benefits of α-entmax, showing great runtime and memory efficiency. Sparsemax [40] can be viewed as a special case of α-entmax with α = 2. It maps inputs onto the probability simplex while allowing exact zeros in the output. Moreover, sparsemax yields piecewise-linear activations and a welldefined Jacobian, enabling efficient gradient computation. It has been successfully applied in multi-label classification and attention-based networks, producing more selective and interpretable attention maps without sacrificing differentiability [29, 40, 42]. In this paper, we aim to introduce a new transformerbased SAEs based on the sparsemax attention. This approach enables SAE to be trained solely with the reconstruction loss, eliminating the need for additional penalty regularization or hyperparameter tuning. The proposed SAEs can also be easily applied to both visual and textual domains.
ReLU-based σ. Early SAEs often employ the ReLU activation function to generate sparse weights due to its simplicity of implementation. To address the feature shrinkage issue in ReLU, where activations in z tend toward zero, GatedReLU and JumpRelu are proposed. Although effective in learning sparse z, these models require additional sparse regularizations, for example: L = ||x − x̂||22 + λS(z),
(2)
where S denotes the function penalizing non-sparse decompositions, e.g., L1 in ReLU and GatedReLU and L0 in JumpReLU. λ sets the trade-off between sparsity and reconstruction, requiring careful tuning to achieve a balance between the two. TopK-based σ. Gao et al. [20] suggests that the TopK is another option to learn sparse z, where only the K largest concepts are kept for each sample, with all others set to zero. Due to its explicit sparsity selection, TopK SAEs can be trained via the reconstruction loss. Built upon TopK, BatchTopK is further developed to replace the sample-level TopK operation with a batch-level BatchTopK function, where the top n × K activations across the entire batch of n samples are selected, while all others are set to zero. Compared to ReLU-based SAEs, TopK-based SAEs empirically show a better balance between the sparsity and reconstruction. However, how to choose the optimal K in these models remains an open problem. In this paper, we aim to further relax the constraints of BatchTopK based on the sparsemax function, which dynamically determines the number of activations according to the feature’s complexity, rather than setting a fixed K.
3. Methodology
3.2. Sparsemax SAE
In this section, we first briefly review SAEs, then introduce our proposed model in detail.
As illustrated in Fig. 2, we improve SAEs in two aspects: 1) a transformer-based architecture is designed to connect the encoder and decoder by sharing the same concept vectors; and 2) within the transformer-based framework, each sample has the ability to determine its sparse level dynamically via sparsemax attention. Below, let us introduce each of them in detail.
3.1. SAE and its variants To disentangle the polysemantic activations (or features) x ∈ Rd into a set of monosemantic and interpretable concepts C = {c1 , c2 , ..., cM } ∈ Rd×M , where M ≫ d represents the concept space dimension, SAEs typically represent x as a sparse linear combination of these concepts under the encoder-decoder framework: z = σ(Wenc (x − benc )), x̂ = Wdec z + bdec ,
(1)
where Wenc ∈ RM ×d and Wdec ∈ Rd×M denote the weight matrices of the single-layer encoder and decoder, respectively. benc , bdec ∈ Rd are two bias terms. The columns of Wdec are the to-be-learned concepts C, and the reconstruction x̂ is obtained by weighting these concepts via z. To prompt the sparse combination, various structures of σ have been developed in previous studies:
Transformer-based SAE. Inspired by the great successes of transformer-based structures in various fields [49, 53, 65, 67], we aim to explore a novel SAE with the cross-attention mechanism. Mathematically , we rewrite Eq. 1 as: Q = x T WQ ,
K = C T WK ,
QKT x̂ = σ( √ )V, d
V = C T WV , (3)
where WQ , WK , WV ∈ Rd×d denotes the query, key, and value projections, respectively. σ is the sparsemax function, which will be introduced later. Intuitively, Eq. 3 views the
Sparsemax Attention K
Step 1: Calculate Score
Step 2: Rank
τ
K
... Step 3: Select & Re-norm
K
Q
Patch Embedding Input Image
...
Concept Embeddings
Step 4: Attention
^
||x x ||22
V
...
V V
Figure 2. Framework of our Saprsemax SAE, which reconstructs the input feature under the transformer architectures, and the sparsemax attention is employed to dynamically determine the number of concepts by estimating the threshold τ .
input feature as a query and reconstructs it using a set of concepts via the cross-attention framework. Compared to MLP-based SAEs in Eq. 1 that directly output z via Wenc , our approach models the activation weights as the similarity score of the query and concepts explicitly, thereby enabling more precise weight estimation. For example, a higher z denotes a closer distance between the query and concepts in the embedding space. More importantly, unlike Eq. 1 views Wenc and Wdec as two independent learnable projections, both the key and value in Eq. 3 originate from the same concept C. This reinforces the synergy between the concept vector V and its weights during the weighting (decoding) stage, showing stronger reconstruction capabilities. Sparsemax Attention. One of the core ideas of SAEs is to assign a sparse set of concepts for each latent feature, while the widely-used softmax function in transformers often outputs dense activations [65]. To this end, we introduce the sparsemax attention into our transformer-based SAE, as it is capable of producing exactly zero value to low-scoring concepts. Let z = QKT ∈ RM denote the similarity score between the query and M concepts, sparsemax aims to project z onto the probability simplex with Euclidean distance: sparsemax(z) = arg min ∥p − z∥2 , p∈∆M −1
where ∆
M −1
:=
n
p∈R
M
pi ≥ 0 ,
Proposition 1. The closed-form solution of Eq. 4 is: sparsemax(z)m = max(z m − τ, 0),
(5)
where τ is a threshold calculated so the result sums to 1: P (z j − τ ) = 1 for every selected z j , e.g., S = {j : j∈S z j > τ }. Furthermore, the support set S (and hence τ ) can be computed efficiently via sorting: if we sort z in descending order as z(1) ≥ · · · ≥ z(M ) , define Pr 1 − i=1 z(i) >0 , k = max r ∈ {1, . . . , M } z(r) + r (6) then, Pk z(i) − 1 τ = i=1 . (7) k Proof. We provide a detailed derivation in the appendix. Unlike TopK-based algorithms that truncate z by setting a hard threshold, Prop. 1 suggests a dynamical τ by measuring the content complexity of the input feature. For example, if the query feature consists of multiple concepts, the z tends to contain many comparable values, resulting in a large set S. Conversely, if the query feature represents pure concepts, the resulting S becomes very small. We summarize the Sparsemax algorithm in Alg. 1.
(4)
4. Experiments o p = 1 is i i=1
PM
the (M − 1)-dimensional simplex. Sparsemax aims to find the point inside the simplex that is nearest to z. Therefore, for small coordinates of z, the closest point in the simplex will force them to zero, in which case sparsemax(z) becomes sparse. Fortunately, Eq. 4 can be solved with lineartime algorithms [16, 41].
In this section, we first outline the experimental setup, followed by the evaluation of Sparsemax SAEs. We conduct experiments in both visual and textual domains and compare our approach with recent advances in terms of image classification and reconstruction tasks. Finally, we visualize the learned concepts at both the image and patch levels, revealing clear and interpretable visual patterns behind these concepts.
Figure 3. Comparisions of zero-shot image classfication using top-n concepts on 11 datasets. All results are calculated as the mean value of three runs with different random seeds. CLIP denotes the results using the original image features in CLIP, and on 49152 denotes all concepts are used. Table 1. NMSE scores with various dictionary sizes M on the OpenWeb and WikiText-103 test datasets. OpenWeb
Method RELU JumpRELU Gated TopK BatchTopK Saprsemax SAE (Ours)
WikiText-103
M =3072
M =6144
M =12288
M =24576
M =3072
M =6144
M =12288
M =24576
0.064 0.051 0.078 0.014 0.014 0.005
0.064 0.050 0.092 0.059 0.061 0.038
0.064 0.050 0.129 0.010 0.060 0.004
0.059 0.051 0.489 0.055 0.060 0.039
0.064 0.058 0.088 0.024 0.024 0.008
0.064 0.056 0.106 0.063 0.064 0.046
0.064 0.055 0.196 0.018 0.018 0.007
0.064 0.567 0.527 0.061 0.062 0.045
4.1. Experimental Setup Datasets. For the visual modality, we use the CLIP model with an image encoder of ViT-B/16. This results in a CLS
and 14 × 14 image patch tokens as inputs. Following PatchSAE [34], we extract the ViT output from the residual stream of the second-last attention layer. Therefore, the training
Table 2. CE degradation with various dictionary sizes M on the OpenWeb and WikiText-103 test datasets. OpenWeb
Method
WikiText-103
M =3072
M =6144
M =12288
M =24576
M =3072
M =6144
M =12288
M =24576
Relu JumpRELU Gated TopK BatchTopK
-4.709 -0.586 -1.738 0.209 0.196
-4.656 -1.300 -1.422 -1.778 -1.742
-4.614 -0.928 -2.679 0.306 0.302
-1.562 -1.151 -1.309 -1.261 -1.740
-4.709 -3.272 -6.637 -0.898 -0.867
-4.656 -3.060 -5.791 -4.587 -4.659
-4.614 -2.980 -5.453 -0.569 -0.566
-4.615 -3.260 -5.557 -4.251 -4.356
Saprsemax SAE (Ours)
0.031
-1.516
0.012
-0.395
-0.106
-2.234
-0.113
-2.079
Figure 4. Visualization of top three concepts given the reference image. For each concept, we provide its masking map within the input image and top five reference image from the ImageNet dataset. Compared to BatchSAE, our Sparsemax SAE learns clearer and more interpretable concepts.
Algorithm 1 Sparsemax Attention 1: Input: z 2: Sort z as z(1) ≥ · · · ≥ z(M ) 3: Find k(z) := max
n
r ∈ [M ] z(r) +
o P 1− ri=1 z(i) >0 , r
where [M ] := {1, . . . , M} P k
z(i) −1
i=1 4: Define τ (z) = . kz 5: Output: p such that pi = max{0, zi − τ (z)}
visual data has a size of Nimg × Npatch × d, with Nimg and Npatch denoting the number of images and patches per image, respectively. We train all SAEs on the ImageNet dataset [13], and evaluate the performance on 11 classifica-
tion datasets: ImageNet [13] and Caltech101 [18] for generic object classification, OxfordPets [48], StanfordCars [31], Flowers102 [44], Food101 [3] and FGVCAircraft [36] for fine-grained image recognition, EuroSAT [26] for satellite image classification, UCF101 [61] for action classification, DTD [9] for texture classification, and SUN397 [70] for scene recognition. For the textual modality, we choose GPT2 Small [54] as our pre-trained model, and extract the hidden embeddings from the residual stream of the 8-th transformer layer. We train all SAEs on the training set of the OpenWebText dataset [22], which was processed into sequences of a maximum of 128 tokens for input into the language models. We report the reconstruction results on the test sets of the OpenWebText and WikiText-103 datasets.
Baselines We compare our Sparsemax SAE against a range of state-of-the art baselines, including 1) ReLU-based SAEs: ReLUSAE [28], a pioneering method that explains LLMs using SAE; PatchSAE [34] views the image patches as tokens and extracts interpretable concepts at both the image and patch levels; GateSAE [57] designs two encoders to model the position and coefficients of the concepts simultaneously; JumpReLU [58] introduces a learnable threshold into ReLU to alleviate the feature shrinkage issue. And 2) TopK-based SAEs: TopKSAE [63] directly keeps the K largest concepts and zeroes out the others; BatchTopK [5] relaxes TopKSAE by introducing the batch-level operation, where samples within a batch size share K × Nbatch activations. Unlike above models that require additional regularization loss or hyperparameter tuning, Our sparsemax SAE dynamically determines the optimal sparsity level based on feature complexity, demonstrating greater flexibility and interpretability. Implementation Details. Following prior research [34], for the CLIP model, we set the number of concepts to M = 49, 152, as a value 64 times the latent dimension of the ViT model. The batch size is 32, and the training continued until a total of 2, 621, 440 patches were feed. For the GPT-2 Small model, we conduct experiments with M = 3072, 6144, 12288 and 24576 to test the performance with different dictionary sizes. The batch size is 128 and training continued until a total of 1 × 109 tokens were processed. All models were trained using the Adam optimizer with a learning rate of 3 × 10−4 , β1 = 0.9 and β2 = 0.99. For all baselines, we load the suggested hyperparameters according to their papers (K = 32 for TopK-based SAEs and the sparsity weighting set to 1e−3 for ReLU-based SAEs).
4.2. Results Analysis 4.2.1. Zero-Shot Image Classification To measure the quality of the learned concepts of SAEs and explore whether the activated concepts per class capture the core class-level concepts, we conduct zero-shot image classification tasks by replacing the intermediate embeddings of ViT in CLIP with the SAE reconstruction embeddings using only the top n = 1, 5, 10, 50 concepts. The final prediction is calculated by the cosine similarity between textual prototypes and steered image features. To specify the subset concepts per class, we follow PatchSAE [34] and first collect SAE latent activations from the training set of each dataset, and then select the largest n concepts according to their activation frequency. These concepts are then act as masks to control the ViT’s reconstruction. Intuitively, the selected top-n concepts capture key information of that class, and higher classification performance denotes a higher quality of the learned concepts. Fig. 3 reports the classification results of our Sparsemax
SAE and baselines on 11 datasets. We also report the results with original ViT embeddings (CLIP) and results with all concepts (on 49152). From the results, we have the following interesting finds: 1) Overall, our Sparsemax SAE achieves the best average performance across 11 datasets on all top-n settings (top-left subfigure). Particularly at extremely small n values (i.e., n = 1, 5, 10), our method significantly outperformed the second-best model. these results demonstrate that our sparse attention-based SAE is much more effective than ReLU and TopK-based SAEs. Our approach successfully assigns the most relevant concepts to the latent features, improving the representation learning of the dictionary. 2) SAE-based models outperform CLIP on the EuroSAT and DTD datasets n = 10, 50 settings, whose images differ significantly from the pre-trained natural images. This demonstrates that the learned concepts are also generalizable and can be useful in zero-shot scenarios due to their monotonicity. 3) Intuitively, more concepts should yield better reconstruction quality and thus higher prediction scores. However,we observe performance declining from on 50 to on 49152 across multiple datasets. we attribute this to the denoising capability of SAEs, where the top-n concepts capture the clear visual patterns that aid in identifying object labels. 4.2.2. Reconstruction Results on Text Following previous works [5], we evaluate the reconstruction performance of our approach in terms of normalized mean squared error (NMSE) and cross-entropy (CE) degradation on the OpenWeb and WikiText-103 corpora, and report the results in Table. 1 and 2, respectively. Similarly, we replace the intermediate features in GPT-2 Small with the reconstruction output of SAEs and measure the difference with the original outputs. Experiments demonstrate that across all datasets and dictionary sizes, our proposed Sparsemax SAE not only achieves significantly lower NMSE scores than other methods but also exhibits reduced CE degradation. This proves that Sparsemax SAE not only decouples polysemantic features into interpretable concepts but also reconstructs input data with lower information loss. The dynamic sparse attention mechanism enables the model to leverage more conceptual representations for complex features, fully demonstrating its effectiveness. 4.2.3. Ablation Study In this section, we aim to ablate the impacts of our proposed modules: transformer-based architecture and sparsemax attention. Specifically, we mainly consider two variants: MLP-based SAEs with the sparsemax function and transformer-based SAEs with the ReLU function. For the latter, we employ the L1 regularization to promote sparsity. Table. 12 reports the image classification results on the ImageNet dataset. From the ablations, we find that both the introduced modules improve the performance of our base-
Figure 5. Visualizations of the top one concept. For the query image, we interprete its most relevant concept from the image-level (top row) and patch-level (bottom row), respectively. Table 3. Ablation results on the ImageNet dataset. on 1 ReLU SAE Transformer + ReLU MLP + Sparsemax Sparsemax SAE (Ours)
on 5
on 10
on 50
on 49152
3.12 15.83 3.86 16.85 7.91 29.87 10.93 33.47
22.17 24.08 39.73 42.13
34.87 36.33 55.32 59.95
63.67 63.94 64.74 65.09
Table 4. Zero-shot classification results of different K on the Food101 dataset. TopK (K = 32) BatchTopK (K = 32) TopK (K = 24) BatchTopK (K = 24) Sparsemax SAE (Ours)
on 1
on 5
on 10
on 50
on 49152
8.64 6.46 0.99 7.36 26.11
21.74 37.86 24.42 44.53 51.71
30.21 48.99 42.88 49.52 59.23
57.89 68.37 66.80 69.40 69.49
54.96 75.44 47.67 75.60 79.95
line. The transformer framework connects the encoder and decoder by sharing the same concept vectors, resulting in more coherent concepts. The sparsemax function enables the model to dynamically determine the number of activations according to the feature’s complexity, improving the trade-off between the sparsity and reconstruction. 4.2.4. Further Analysis Our sparsemax SAE is trained solely with reconstruction loss, with sparsity determined by the dataset. This enables optimal K selection for TopK-based SAE.Specifically, we first collect all activations of our sparsemax SAE on the ImageNet dataset, and calculate the average number of activated concepts per sample K ∗ = 24. We then re-train TopK-based SAE with the new K ∗ . Table. 4 reports the results of K = 24, 32. We find that the calculated K using our approach is a better choice for both TopK and BatchTopK SAEs in most cases. This improvement suggests that our sparsemax attention is able to estimate the level of sparsity. 4.2.5. Visualization Results In addition to the above quantitative analysis, we also provide visualizations of the learned concepts. First, we aim to evaluate the top three concepts activated for each given image. Specifically, we use the CLS token as the global representation of the images, and obtain the top three concepts with the largest activation scores. Fig. 4 shows the comparison results of BatchTopK and our Sparsemax SAE.
For each model, we also visualize the activation scores of all concepts. For each selected concept, we visualize its corresponding patches of the input image and the top five reference images from the ImageNet dataset, which provide an easy tool to understand the meaning of the concepts. From the results, we find that our approach successfully disentangles the core concepts from the input image[35],and the top three concepts show clear and specific visual patterns. For example, Sparsemax SAE successfully extracts concepts of the swimming pool, building, and tree from the input image. Compared to BatchTopK, we find that both models are able to extract the relevant concepts. However, BatchTopK sometimes produces redundant or unclear patterns. For example, the second and third concepts of the first image learned from BatchTopK share similar reference images, and the third concept of the second image contains both the wheat and bread patterns. We also provide the visualization of the top one concept at the image level (top row of Fig. 5) and patch level (bottom row of Fig. 5). We find that the top one concept successfully extracts the key information of the given image. For the patch-level visualization, our approach accurately localizes the dog’s nose and cat’s eyes in the reference images, which demonstrates the high-quality of the learned concepts. This suggests that the learned concepts exhibit internal semantic structure that could be further modeled with graph-based discovery methods [15].
5. Conclusion In this paper, we present a novel transformer-based SAE with sparsemax attention mechanism, where the input feature acts as the query and the concepts are modeled as the key and value matrices. The sparsemax function is then employed to steer the sparsity between the query and key attention. The transformer structure connects the encoder and decoder of SAEs by sharing the same concept vectors, while the sparsemax function enables the model to dynamically determine the sparsity level based on the feature’s complexity. This synergy enhances the concept learning and reconstruction performance of traditional SAEs. Extensive experiments across image classification, text reconstruction, and visualization validate the effectiveness of our model. We hope our sparsemax SAE will offer novel insights for secure artificial intelligence and interpretability research.
Acknowledgments This work was supported in part by National Key R&D Program of China (2024YFB3908500, 2024YFB3908502), NSFC (62506237, 62576215), ICFCRT (W2441020), Guangdong Basic and Applied Basic Research Foundation (2023B1515120026), Shenzhen Science and Technology Program (KQTD20210811090044003, KJZD2024 0903100022028, RCJC20200714114435012), and Scientific Development Funds from Shenzhen University.
References [1] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 1 [2] Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020. 3 [3] Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101–mining discriminative components with random forests. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VI 13, pages 446–461. Springer, 2014. 6 [4] Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. 2 [5] Bart Bussmann, Patrick Leask, and Neel Nanda. Batchtopk sparse autoencoders. In NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning. 2, 7 [6] Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2818–2829, 2023. 1 [7] Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019. 3 [8] Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J. Colwell, and Adrian Weller. Rethinking attention with performers. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. 2 [9] Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3606–3613, 2014. 6
[10] Gonçalo M Correia, Vlad Niculae, and André FT Martins. Adaptively sparse transformers. In 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, pages 2174–2184. Association for Computational Linguistics, 2019. 3 [11] Bartosz Cywiński and Kamil Deja. Saeuron: Interpretable concept unlearning in diffusion models with sparse autoencoders. arXiv preprint arXiv:2501.18052, 2025. 2 [12] Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. Flashattention: Fast and memory-efficient exact attention with io-awareness. Advances in neural information processing systems, 35:16344–16359, 2022. 3 [13] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 6 [14] Zhibin Duan, Dongsheng Wang, Bo Chen, Chaojie Wang, Wenchao Chen, Yewen Li, Jie Ren, and Mingyuan Zhou. Sawtooth factorial topic embeddings guided gamma belief network. In International Conference on Machine Learning, pages 2903–2913. PMLR, 2021. 2 [15] Zhibin Duan, Yishi Xu, Bo Chen, Dongsheng Wang, Chaojie Wang, and Mingyuan Zhou. Topicnet: Semantic graph-guided topic discovery. In Advances in Neural Information Processing Systems, 2021. 8 [16] John Duchi, Shai Shalev-Shwartz, Yoram Singer, and Tushar Chandra. Efficient projections onto the l 1-ball for learning in high dimensions. In Proceedings of the 25th international conference on Machine learning, pages 272–279, 2008. 4 [17] Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy models of superposition. Transformer Circuits Thread, 2022. 1 [18] Li Fei-Fei, Rob Fergus, and Pietro Perona. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories. In 2004 conference on computer vision and pattern recognition workshop, pages 178–178. IEEE, 2004. 6 [19] Thomas Fel, Ekdeep Singh Lubana, Jacob S Prince, Matthew Kowal, Victor Boutin, Isabel Papadimitriou, Binxu Wang, Martin Wattenberg, Demba E Ba, and Talia Konkle. Archetypal sae: Adaptive and stable dictionary learning for concept extraction in large vision models. In International Conference on Machine Learning, pages 16543–16572. PMLR, 2025. 1 [20] Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. 1, 2, 3 [21] Yizhao Gao, Zhichen Zeng, Dayou Du, Shijie Cao, Peiyuan Zhou, Jiaxing Qi, Junjie Lai, Hayden Kwok-Hay So, Ting Cao, Fan Yang, et al. Seerattention: Learning intrinsic sparse attention in your llms. arXiv preprint arXiv:2410.13276, 2024. 3
[22] Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. Openwebtext corpus, 2019. 6 [23] Nuno Gonçalves, Marcos V Treviso, and Andre Martins. Adasplash: Adaptive sparse flash attention. In Forty-second International Conference on Machine Learning. 3 [24] Nuno Gonçalves, Marcos V Treviso, and Andre Martins. Adasplash: Adaptive sparse flash attention. In Forty-second International Conference on Machine Learning, 2025. 2 [25] Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant. Memory-efficient transformers via top-k attention. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing, pages 39–52, 2021. 3 [26] Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217–2226, 2019. 6 [27] Sai Sumedh R. Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba E. Ba. Projecting assumptions: The duality between sparse autoencoders and concept geometry. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. 1, 2 [28] Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. 1, 2, 7 [29] Tao Jin and Jiaming Liu. A text classification method by integrating mobile inverted residual bottleneck convolution networks and capsule networks with adaptive feature channels. Scientific Reports, 15(1):855, 2025. 3 [30] Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Smith, Claudio Mayrink Verdun, David Bau, and Samuel Marks. Measuring progress in dictionary learning for language model interpretability with board game models, 2024. page 16. 2 [31] Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013. 6 [32] Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng. Efficient sparse coding algorithms. Advances in neural information processing systems, 19, 2006. 1 [33] Miaoge Li, Dongsheng Wang, Xinyang Liu, Zequn Zeng, Ruiying Lu, Bo Chen, and Mingyuan Zhou. Patchct: Aligning patch set and label set with conditional transport for multilabel image classification. In Proceedings of the IEEE/CVF international conference on computer vision, pages 15348– 15358, 2023. 2 [34] Hyesu Lim, Jinho Choi, Jaegul Choo, and Steffen Schneider. Sparse autoencoders reveal selective remapping of visual concepts during adaptation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. 1, 2, 5, 7
[35] Xinyang Liu, Dongsheng Wang, Miaoge Li, Bowei Fang, Yishi Xu, Zhibin Duan, Bo Chen, and Mingyuan Zhou. Patchprompt aligned bayesian prompt tuning for vision-language models. In Uncertainty in Artificial Intelligence, 2024. 8 [36] Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151, 2013. 6 [37] Alireza Makhzani and Brendan Frey. K-sparse autoencoders. arXiv preprint arXiv:1312.5663, 2013. 2 [38] Luke Marks, Alasdair Paren, David Krueger, and Fazl Barez. Enhancing neural network interpretability with feature-aligned sparse autoencoders. arXiv preprint arXiv:2411.01220, 2024. 2 [39] Andre Martins and Ramon Astudillo. From softmax to sparsemax: A sparse model of attention and multi-label classification. In International conference on machine learning, pages 1614–1623. PMLR, 2016. 2 [40] André F. T. Martins and Ramón Astudillo. From softmax to sparsemax: A sparse model of attention and multi-label classification. In Proceedings of the 33rd International Conference on Machine Learning (ICML), pages 1614–1623. PMLR, 2016. 3 [41] Christian Michelot. A finite algorithm for finding the projection of a point onto the canonical simplex of ∝ n. Journal of Optimization Theory and Applications, 50(1):195–200, 1986. 4 [42] Timur Mudarisov, Mikhail Burtsev, Tatiana Petrova, et al. Limitations of normalization in attention. In The Thirtyninth Annual Conference on Neural Information Processing Systems. 3 [43] Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt. Efficient dictionary learning with switch sparse autoencoders, 2024. 273233368. 2 [44] Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729. IEEE, 2008. 6 [45] Bruno A Olshausen and David J Field. Sparse coding with an overcomplete basis set: A strategy employed by v1? Vision research, 37(23):3311–3325, 1997. 1 [46] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jégou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. Dinov2: Learning robust visual features without supervision. Trans. Mach. Learn. Res., 2024, 2024. 2 [47] Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie, and Zeynep Akata. Sparse autoencoders learn monosemantic features in vision-language models. arXiv preprint arXiv:2504.02821, 2025. 2 [48] Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and CV Jawahar. Cats and dogs. In 2012 IEEE conference on computer vision and pattern recognition, pages 3498–3505. IEEE, 2012. 6
[49] William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023. 2, 3 [50] Ben Peters, Vlad Niculae, and André F. T. Martins. Sparse sequence-to-sequence models. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 1504–1519. Association for Computational Linguistics, 2019. 2 [51] Ben Peters, Vlad Niculae, and André F. T. Martins. Sparse sequence-to-sequence models. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1504–1519. ACL, 2019. 3 [52] Yao Qiang, Chengyin Li, Prashant Khanduri, and Dongxiao Zhu. Interpretability-aware vision transformer. arXiv preprint arXiv:2309.08035, 2023. 2 [53] Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018. 2, 3 [54] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 1, 6 [55] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021. 1 [56] Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, and Ziyu Yao. A practical review of mechanistic interpretability for transformer-based language models. arXiv preprint arXiv:2407.02646, 2024. 2 [57] Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda. Improving dictionary learning with gated sparse autoencoders. arXiv preprint arXiv:2404.16014, 2024. 1, 2, 7 [58] Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, János Kramár, and Neel Nanda. Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders. arXiv preprint arXiv:2407.14435, 2024. 1, 2, 7 [59] Lee Sharkey, Dan Braun, and Beren Millidge. Taking features out of superposition with sparse autoencoders. 2022. 2023. 2 [60] Wei Shi, Sihang Li, Tao Liang, Mingyang Wan, Guojun Ma, Xiang Wang, and Xiangnan He. Route sparse autoencoder to interpret large language models. arXiv preprint arXiv:2503.08200, 2025. 1 [61] Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012. 6 [62] Samuel Stevens, Wei-Lun Chao, Tanya Berger-Wolf, and Yu Su. Sparse autoencoders for scientifically rigorous interpretation of vision models. arXiv preprint arXiv:2502.06755, 2025. 2
[63] Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Transformer Circuits Thread, 2024. 1, 7 [64] Marcos Treviso, António Góis, Patrick Fernandes, Erick Fonseca, and André FT Martins. Predicting attention sparsity in transformers. In Proceedings of the Sixth Workshop on Structured Prediction for NLP, pages 67–81, 2022. 3 [65] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 2, 3, 4 [66] Dongsheng Wang, Dandan Guo, He Zhao, Huangjie Zheng, Korawat Tanwisuth, Bo Chen, and Mingyuan Zhou. Representing mixtures of word embeddings with mixtures of topic embeddings. In International Conference on Learning Representations, 2022. 1 [67] Dongsheng Wang, Miaoge Li, Xinyang Liu, MingSheng Xu, Bo Chen, and Hanwang Zhang. Tuning multi-mode tokenlevel prompt alignment across modalities. Advances in Neural Information Processing Systems, 36:52792–52810, 2023. 3 [68] Dongsheng Wang, Jiequan Cui, Miaoge Li, Wang Lin, Bo Chen, and Hanwang Zhang. Instruction tuning-free visual token complement for multimodal llms. In European Conference on Computer Vision, pages 446–462. Springer, 2024. 2 [69] Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023. 3 [70] Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. Sun database: Large-scale scene recognition from abbey to zoo. In 2010 IEEE computer society conference on computer vision and pattern recognition, pages 3485–3492. IEEE, 2010. 6 [71] Kaichen Zhang, Yifei Shen, Bo Li, and Ziwei Liu. Large multi-modal models can interpret features in large multimodal models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3650–3661, 2025. 2
Improving Sparse Autoencoder with Dynamic Attention Supplementary Material 6. Proof of Prop. 1 1. Problem formulation. Recall that sparsemax is the Euclidean projection of z ∈ RM onto the probability simplex ∆
M −1
= {p ∈ R
M
⊤
| 1 p = 1, p ≥ 0},
(8)
2 1 2 ∥p − z∥2 .
(9)
i.e. sparsemax(z) = arg min
p∈∆M −1
3. Determining τ and the active set. (support) be S = {i | pi > 0} = {i | zi > τ },
M
1X (pi −zi )2 M 2 i=1 p∈R
M X
subject to
pi ≥ 0 ∀i.
2. Lagrangian and KKT conditions. IntroducePthe Lagrange multiplier τ ∈ R for the equality constraint i pi = 1 and multipliers µi ≥ 0 for the inequality constraints pi ≥ 0. The (augmented) Lagrangian is M M M X X 1X (pi − zi )2 + τ pi − 1 − µi pi . 2 i=1 i=1 i=1 (11) Karush–Kuhn–TuckerP (KKT) conditions for optimality are: M (i) Primal feasibility: i=1 pi = 1, pi ≥ 0 for all i. (ii) Dual feasibility: µi ≥ 0 for all i. (iii) Stationarity:
L(p, τ, µ) =
∂L = pi −zi +τ −µi = 0 ∂pi
=⇒
pi = zi −τ +µi ,
∀i. (12)
(iv) Complementary slackness: ∀i.
µi pi = 0,
(13)
Given Eq. 12 and Eq. 13, we have: • If pi > 0, complementary slackness forces µi = 0. Hence and thus
=⇒
µi = τ − z i .
Since µi ≥ 0, we conclude τ − zi ≥ 0, i.e. zi ≤ τ . Combining both cases yields the compact expression (Eq. 4 in the main paper) pi = max(zi − τ, 0),
∀i,
P
i∈S zi − 1
k
.
(17)
To find S efficiently, sort the coordinates in descending order: z(1) ≥ z(2) ≥ · · · ≥ z(M ) . (18) If the optimal support corresponds to the top k indices (this is always the case: if some j ∈ S were not among the top k, there would be an index in the top k not in S with a larger z, contradicting zj > τ ), then Pk τk =
i=1 z(i) − 1
k
.
(19)
The correct k is the largest integer for which the k-th sorted element exceeds this threshold: Pk 1 − i=1 z(i) > 0. (20) z(k) > τk ⇐⇒ z(k) + k Therefore set Pr n o 1 − i=1 z(i) k = max r ∈ {1, . . . , M } z(r) + >0 , r (21) and then take τ = τk , which completes the proof.
7. More Visualization Results
zi − τ > 0.
• If pi = 0, complementary slackness allows µi ≥ 0 and stationarity gives 0 = zi − τ + µi
(15)
i∈S
τ=
i=1
(10)
pi = zi − τ
i∈S
Thus pi = 1,
k := |S|.
PM Summing pi = zi − τ over i ∈ S and using i=1 pi = 1 gives X X (zi − τ ) = 1 =⇒ zi − kτ = 1. (16)
Equivalently, we solve the constrained quadratic program min
Let the active set
(14)
Analysis of Top 3 Concepts. Fig. 6- 8 show the visualizations of the top three concepts of the test image. Overall, we find that our Sparsemax SAE successfully captures the key visual patterns. For example, the mixture of fruit, wooden background, and apples in Fig. 6. Moreover, our approach is able to understand visual features from different perspectives. For example, the first concept in Fig. 7 corresponds to a pig (the main object of the input image), the second concept is related to cartoon characters, and the third concept is about the dressing.
Table 5. Sparsity metric values.
Model
L0 ↑
FVU↓
CS↑
CKNNA↑
DO↓
ReLU TopK BatchTopK Ours
0.928 0.966 0.814 0.979
0.098 0.169 0.278 0.129
0.953 0.925 0.904 0.934
0.812 0.701 0.750 0.796
0.003 0.003 0.002 0.001
Table 6. Interpretability metrics.
Model ReLU TopK BatchTopK Ours
MEAN-MS
MAX-MS
0.1627 0.0548 0.1243 0.3484
0.9172 0.8751 0.9031 0.9575
Failure Cases. We also provide several failure cases of our Sparsemax SAE in Fig. 9- 11. On one hand, we find that the learned concepts sometimes contain unclear visual patterns (The third concept in Fig. 9 and the first concept in Fig. 11). This may stem from feature absorption issues, where the learned concepts fail to decompose into into their subconcepts. On the other hand, the learned concepts occasionally share the similar masking concent of the input image. We attribute this to the fine-grained feature of the learned concepts, where concepts capture the similar visual patterns while focusing on distinct dimensions. For example, the concepts of baby, cute girl, and playing girl in Fig. 10. Sparsity and Interpretability Analysis To further evaluate the quality of the learned concepts, we report sparsity metric values in Table. 5 and interpretability metrics in Table. 6. Following prior work [30, 38], we compute metrics including L0 (higher is better), FVU (fraction of variance unexplained, lower is better), CS (cosine similarity, higher is better), CKNNA (higher is better), and DO (dead concepts, lower is better). For interpretability, we follow the evaluation protocol in [47] and report both the mean and maximum monosemantic scores of the learned concepts. Our Sparsemax SAE achieves the best performance under most metrics, demonstrating that the dynamic attention mechanism not only produces sparse representations but also yields more interpretable and semantically coherent concepts.
8. Results of Zero-shot Image Classification We report the detailed zero-shot image classification results in Table. 7- 17.
Table 7. Zero-shot classification results on the Caltech101 dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TOPK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
32.403 32.403 32.403 32.403 32.403 32.403
7.816 3.578 5.894 5.94 11.567 20.121
17.591 12.840 14.054 16.632 23.144 26.195
21.033 18.594 22.847 23.096 27.314 29.322
28.671 27.845 28.847 30.654 30.582 31.325
31.672 28.974 30.158 29.645 31.284 31.875
Table 8. Zero-shot classification results on the DTD dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
44.840 44.840 44.840 44.840 44.840 44.840
13.457 4.825 5.989 6.117 17.730 34.804
27.074 22.884 15.048 15.372 30.372 41.791
34.326 30.520 21.495 22.624 36.383 46.986
44.468 44.495 47.187 46.365 49.007 50.16
41.596 40.367 37.365 35.851 40.957 44.025
Table 9. Zero-shot classification results on the EuroSAT dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
38.605 38.605 38.605 38.605 38.605 38.605
10.200 11.680 15.815 16.46 10.0 19.301
24.826 20.475 32.487 11.468 34.990 36.463
32.939 33.648 39.846 25.344 34.957 34.056
39.832 40.987 44.682 47.041 45.317 47.839
33.315 32.158 34.577 32.741 31.338 32.421
Table 10. Zero-shot classification results on the FGVC dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
23.137 23.137 23.137 23.137 23.136 23.137
2.371 2.647 1.574 1.02 0.990 4.8346
2.617 2.689 4.855 1.683 5.191 7.069
3.314 4.047 5.128 3.156 7.390 8.14871
7.431 7.846 7.748 9.18 10.805 10.894
19.803 17.358 11.547 7.59 14.362 19.5187
Table 11. Zero-shot classification results on the Food101 dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
83.195 83.195 83.195 83.195 83.194 83.195
17.991 19.475 7.458 8.644 6.459 26.1134
34.420 35.187 25.486 21.735 37.864 51.705
39.286 39.954 36.487 30.207 48.996 59.23
56.2719 67.257 60.875 57.888 68.365 69.49
78.285 75.481 77.584 54.956 75.440 79.954
Table 12. Zero-shot classification results on the ImageNet dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
67.294 67.294 67.294 67.294 67.294 67.294
3.116 3.548 5.568 3.421 1.58 10.929
15.829 17.558 12.875 8.2501 15.562 33.469
22.173 34.975 35.495 32.646 27.782 42.127
34.872 40.876 50.159 56.135 58.472 59.947
63.670 62.957 62.547 47.956 61.072 65.087
Table 13. Zero-shot classification results on the ImageNet-Sketch dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
47.851 47.851 47.851 47.851 47.851 47.851
2.798 3.847 1.782 0.957 0.3509 12.460
10.283 12.657 9.257 8.1269 7.349 25.321
13.451 15.984 13.576 14.712 16.103 29.678
23.479 26.581 38.712 36.942 39.128 39.328
31.009 30.257 32.157 30.991 34.178 42.314
Table 14. Zero-shot classification results on the OxfordPets dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
83.477 83.477 83.477 83.477 83.477 83.477
26.964 26.981 18.547 6.973 17.081 25.186
49.611 50.258 53.492 44.886 56.985 58.135
61.350 63.204 64.581 52.492 67.100 67.925
71.127 73.980 74.012 79.621 76.918 77.351
79.061 78.041 76.251 65.274 77.828 80.986
Table 15. Zero-shot classification results on the StandfordCars dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
31.976 31.976 31.976 31.976 31.976 31.976
0.828 1.568 0.540 0.610 1.248 3.421
3.307 3.964 3.541 2.735 3.811 8.25
5.383 5.421 6.034 3.263 6.786 9.337
9.778 9.157 10.652 6.897 10.1140 13.155
24.602 20.367 19.068 8.7586 18.987 27.245
Table 16. Zero-shot classification results on the SUN397 dataset. K = 32 in TopK and BatchTopK.
ReLU JumpReLU Gated TopK BatchTopK Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
66.257 66.257 66.257 66.257 66.257 66.257
4.543 5.962 6.058 5.2819 5.117 15.946
19.045 22.084 27.643 20.468 29.824 39.904
25.902 25.035 31.247 29.9 40.5 48.4794
44.859 48.258 52.796 54.277 57.203 56.973
60.704 58.367 58.947 45.027 60.431 63.788
Table 17. Zero-shot classification results on the UCF101 dataset. K = 32 in TopK and BatchTopK.
ReLU TopK BatchTopK JumpReLU Gated Saprsemax SAE (Ours)
no sae
on 1
on 5
on 10
on 50
on 49152
61.322 61.322 61.322 61.322 61.322 61.322
7.607 3.039 3.591 6.257 4.068 21.712
19.600 18.003 26.474 20.587 24.947 33.575
26.056 26.157 33.346 26.971 30.189 37.779
42.945 41.396 45.194 43.579 43.840 47.214
58.626 34.647 56.986 51.278 52.497 59.064
Figure 6. This picture successfully activates three different concepts: fruit, background and apple.
Figure 7. This picture is from the cartoon task image in the animation. It can be seen that it has the characteristics of pigs. Therefore, the features of pig are activated through sae, and other relevant features are activated according to its shape.
Figure 8. This picture is from the monster in the film and television image. Its wolf’s head features activate the wolf’s features through sae. In addition, his armor and weapons also activate the corresponding features.
Figure 9. Messi’s picture in this example activated the football. In addition, in this experiment, sae activated the ”crowd” feature through the background of the picture,but unfortunately there is a phenomenon of feature absorption in the third feature, where images of short sleeves and similar color structures are considered to be the same concept
Figure 10. In this sample, pictures activate children and children’s eating characteristics through sae.However,the concept of Apple was not recognized and the concept on the right is too similar.
Figure 11. This picture is from a football match. The picture activates the football feature, field , and the crowd .But in the first concept, there were pictures unrelated to football