Conceptio › Archive › arXiv CS
arXiv CSopen access

Samsone: A Family of Open Small Audio Language Models for On-Device Inference

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Samsone: A Family of Open Small Audio Language Models for On-Device Inference Piotr Masztalski1,2,∗ , Michal K. Grzeszczyk1,∗ , Olaf Sikorski1 1

2

Samsung R&D Institute Poland AGH University of Kraków, Poland

{p.masztalski, m.grzeszczyk, o.sikorski}@samsung.com

90

The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacypreserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone. Index Terms: small audio language models, on-device inference, audio question answering

1. Introduction Traditionally, audio tasks such as sound classification [1], audio captioning [2, 3], and speaker identification [4] were addressed by task-specific architectures [5]. Recently, the rapid advancement of Large Language Models (LLMs) has catalyzed the development of unified Large Audio Language Models (LALMs) [6, 7, 8, 9, 10, 11, 12, 13]. Using audio-text pairs to train in the Audio Question Answering (AQA) setting, these models achieve generalized performance across diverse tasks. While the audio domain initially faced data scarcity, the use of multimodal LMs to generate synthetic, high-reasoning datasets, such as OpenAQA [13], ReasonAQA [14], and AudioSkillsXL [8], has significantly boosted LALMs capabilities. While the performance gains of scaling LLMs are undeniable, they come with significant environmental and computational costs [15]. Furthermore, strict privacy requirements in domains such as healthcare [16], the need for offline processing, and the rise of agentic frameworks for specialized, repetitive tasks [17] have created a demand for small-scale models. Defining the boundary between small and large models remains ambiguous. In the text domain, where parameter numbers frequently exceed hundreds of billions (e.g., Qwen3-235B [18]), even multi-billion parameter architectures like Llama38B [19] are categorized as small. However, the audio domain operates on a different scale: state-of-the-art (SOTA) LALMs rarely exceed 10 billion parameters, with leading models such * These authors contributed equally.

SALMs

80 70 MMAU-Test Score

arXiv:2609.21666v1 [eess.AS] 18 Sep 2026

Abstract

LALMs Audio-Thinker Audio Flamingo 3

60

Samsone-134M Samsone-356M Audio Flamingo 2 Samsone-99M

50

Mellow

Previous SALM SOTA

Qwen2-Audio-Instruct

40 30 20 Audio Flamingo

10 0

GAMA LTU

Pengi

0.1

0.5 1 3 Model Size [Billions of Parameters]

8

Figure 1: Samsone establishes the new state of the art on the MMAU benchmark among SALMs. Both Samsone-99M and Samsone-134M beat the previous SOTA (Mellow) while containing less parameters. All of the Samsone variants are also competitive with much larger, multi-billion parameter LALMs.

as Audio Flamingo 3 (8.4B) [8], typically ranging between 3B and 9B. Consequently, in this paper, we define Small Audio Language Models (SALMs) as those with fewer than 1 billion parameters. Existing SALMs remain underrepresented. Pengi (323M) [20] pioneered treating audio tasks as text generation but struggled with complex reasoning due to limited training text diversity. This was partially addressed by Mellow (167M) [14], which utilized the ReasonAQA dataset. Despite their potential for on-device execution, actual mobile deployment and benchmarking of such models remain largely unexplored. In this paper, we introduce Samsone, a family of SALMs that achieve SOTA performance across various audio tasks (Fig. 1). Samsone adopts a standard ALM architecture, comprising an Audio Encoder (AE), a modality projector mapping audio embeddings to the text embedding space, and an LM text decoder. Through rigorous dataset curation and token pruning, we developed Samsone-134M a 134-million-parameter model that outperforms existing SALMs such as Pengi and Mellow on diverse audio benchmarks. We further explore the efficiency frontier by depth-pruning the LM to create Samsone-99M, a sub-100M parameter model that maintains competitive performance despite its compact size. Conversely, we scale up the architecture to create Samsone-356M, which leverages a larger LM backbone to deliver enhanced accuracy. Finally, we val-

Audio Signals SEP

caption the audio file...

LM Tokenizer

User prompt

Language Model (LM)

LM Output Layer

Projection

Audio Encoder

...

Model response the audio contains the sound of water...

Model export

LM Embedding Layer On-device execution

Figure 2: Samsone family of SALMs consisting of Audio Encoder, projector and Language Model with on-device execution example.

idate the practical utility of the Samsone family by exporting model weights and demonstrating real-time inference on a commodity smartphone. The main contributions of this work are: 1. We introduce Samsone-134M, the best performing ALM in its class. It achieves SOTA results within its size range and remains competitive with models orders of magnitude larger. 2. We present Samsone-99M and Samsone-356M to explore the scalability of SALMs across varying on-device computational constraints, providing insights into the trade-offs between model capacity and efficiency. 3. To promote reproducibility and accelerate research in the field, we open-source the Samsone Family. This includes the training code, server-side and on-device weights, as well as an Android application for inference on smartphones* .

2. Proposed Method In this section we introduce Samsone - its architecture, multiple size variants and design choices for efficient operation. 2.1. Model architecture Samsone is a multimodal language model (LM) that accepts audio and text inputs to produce text output (Fig. 2). It adopts a standard ALM architecture with an audio encoder, a modality projector that maps audio embeddings into the text embedding space, and a text backbone that processes the multimodal input. As our audio backbone, we extract the encoder part of Whisper [21], a transformer-based model originally designed for Automatic Speech Recognition. The audio input gets processed through the Whisper encoder and is subsequently passed to the non-linear modality projector. Its role is to align the feature dimension of raw AE outputs with the text embedding dimension expected by the LM. The sequence length of the projection module output is equal to the sequence length of its input, the audio embeddings. For the language modeling task, we choose the SmolLM2 [22] models in 135M and 360M sizes, depending on the Samsone variant. SmolLM2 is a decoder-only LM based on the LLaMA2 [23] architecture trained on a highly curated multi trillion token text dataset. The text inputs are initially passed through the SmolLM tokenizer, resulting in a sequence of tokens that is mapped into a corresponding sequence of text embeddings by the SmolLM2 embedding layer. The audio-text input to the SmolLM2 transformer blocks is constructed, by concatenating the projected audio embeddings * https://github.com/SamsungLabs/samsone

with the text embeddings. As Samsone is designed to handle an arbitrary number of audio signals, we additionally introduce a trainable SEP token embedding that separates the input audio embeddings from each other, and from the surrounding text embeddings. This enables Samsone to handle multiple variablelength audio inputs while preserving the information about the number of audio signals in each input sequence and to position the audio embedding sequences in any part of the audio-text input. The resulting multimodal input is then processed by the remaining part of the SmolLM2 model to produce a text output. 2.2. Size optimizations We reduce the model size by targeting two components: (i) the embedding matrix, and (ii) the number of Transformer layers. Vocabulary reduction (VR): In LMs, the input text is tokenized and each token is mapped to a continuous embedding vector. When the embedding dimension is relatively large (e.g., 576 in SmolLM2-135M) and the vocabulary is extensive (49152 tokens), the embedding matrix constitutes a significant fraction of the total parameters. In SmolLM2-135M, the embedding layer alone contains more than 28M parameters (approximately 21% of the total model size). To reduce this overhead, we constrain the training text to lowercase ASCII characters (as in [14]) and filter out rarely useful tokens. Specifically, we remove tokens with more than four whitespace characters or more than three special characters. This step enables us to eliminate 15042 tokens from the vocabulary, reducing the embedding matrix by 8.7M parameters while preserving coverage of the target domain. Depth pruning (DP): Our second optimization strategy focuses on reducing model depth [24, 25]. Modern LMs are composed of multiple stacked Transformer blocks, each contributing substantially to the total parameter count. In our setup, all Transformer blocks are trainable, which allows us to apply straightforward depth pruning. Removing a single Transformer block in SmolLM2-135M reduces the model size by approximately 3.5M parameters. This simple yet effective strategy provides a controllable trade-off between model capacity and footprint, enabling systematic scaling. 2.3. Model variants Leveraging the described size optimizations, we introduce the Samsone family, consisting of three distinct variants. All models share a common AE based on the Whisper-Tiny architecture [21]. The core of the family is Samsone-134M, which utilizes the SmolLM2-135M backbone [22]. By applying our VR strat-

Table 1: Parameter count comparison of Samsone size variants with vocabulary reduction (VR) and depth pruning (DP). Variant

VR

DP

Enc.

Proj.

LM

Total size

Samsone-99M Samsone-134M Samsone-356M

✓ ✓ ✓

✓ × ×

9M 9M 9M

0.5M 0.5M 1M

90M 125M 347M

99M 134M 356M

egy, we decrease the total parameter count to approximately 134 million. To explore the lower bounds of on-device efficiency, we introduce Samsone-99M. This variant combines VR with DP, truncating the LM from 30 to 20 Transformer blocks by removing the last 10 layers to achieve a sub-100M parameter footprint. Finally, we scale the architecture upward with Samsone-356M, which employs the SmolLM2-360M backbone to provide higher reasoning capacity for more complex audio queries. The detailed parameter distribution for each variant is summarized in Table 1.

3. Implementation details In this section, we describe the architecture of the Samsone family models, initialization strategy, datasets, training configuration, and the export procedure for on-device deployment. 3.1. Models We initialize the AE from openai/whisper-tiny, and the LM from HuggingFaceTB/SmolLM2-135M (and SmolLM2-360M for Samsone-356M), using the Transformers library [26]. To bridge the audio and language representations, we adopt the nonlinear projector architecture introduced in Mellow [14]. The projector consists of two linear layers with a GeLU activation in between, followed by a residual connection and a final layer normalization step. The AE produces frame-level audio embeddings, which we temporally average-pool to obtain a fixedlength representation of 50 audio tokens per sample. To explicitly separate audio and text modalities, we insert a trainable SEP token before, after, and between audio tokens. This design helps the LM distinguish modalities and improves their alignment. 3.2. Data We train all Samsone variants on two datasets: ReasonAQA [14] and AudioSkillsXL [8]. AudioSkillsXL is a large-scale dataset comprising 8 million question–answer pairs across sound, music, and speech. ReasonAQA is designed specifically for reasoning and contains 1 million QA pairs, many of which require reasoning over two audio inputs. We identified a strong class imbalance in the multiple-choice subset of ReasonAQA: option (b) is the correct answer in the majority of examples, exceeding the combined frequency of all other answers. This imbalance can bias the model toward favoring the second option. To mitigate this issue, we randomly permute answer choices during training to enforce a uniform distribution of correct answers. As SmolLM base models do not utilize prompt templates, we append a postfix string ” answer: ” to the prompt (before the answer) to explicitly delimit the model response. 3.3. Training setup We implement all models in PyTorch and use pytorch-lightning [27] to manage the training pipeline. Given the relatively small parameter count of Samsone models, all components except

the LM embedding layer are trainable. We experimented with multi-stage training strategies commonly used in LALMs [7, 8], but observed no performance gains. Therefore, we adopt a single-stage training procedure. Each model is trained for 100 epochs on a single NVIDIA RTX PRO 6000 Blackwell 96GB GPU, with each epoch consisting of 200,000 training examples. We use the AdamW optimizer [28] with a cosine annealing learning rate scheduler. The schedule includes 10 linear warmup epochs and a minimum learning rate of 1e-7. The learning rate is set to 3e-4 for Samsone-99M and Samsone-134M, and 1e-4 for Samsone-356M. Models are trained using standard token-level cross-entropy loss. After training, we export checkpoints with XNNPACK using ExecuTorch for on-device inference.

4. Experiments and results In this section we describe our experiments on various audio tasks including audio understanding, captioning and AQA. We present the ablation study of key components of Samsone-134M and benchmark our models on a Samsung Galaxy S25 Ultra. To ensure reproducible results, in all of the following experiments, we opt for the greedy text decoding strategy during inference. 4.1. Massive Multitask Audio Understanding (MMAU) We assess Samsone’s audio understanding abilities using the MMAU benchmark, which includes 10,000 human-annotated AQA pairs across three domains: speech, sound, and music [29]. MMAU is designed to require expert-level knowledge and complex reasoning. Evaluation results presented in Table 2 show a strong improvement of all Samsone variants over the previous SALM SOTA, Mellow, across all domains. Even our smallest model, Samsone-99M, consistently outperforms Mellow, while containing over 40% less parameters. It is worth mentioning, that we compare our solution to an improved version of Mellow (v0 s)† , that was made available after the original paper publication. Additionally, Samsone-134M remains competitive with much larger Audio Flamingo 2, beating the likes of 60x larger Qwen2-Audio or 100x larger SALMONN. To evaluate our models in the most challenging environment, we utilize MMAU-Pro [33], a benchmark with 5,305 AQA instances testing long-form audio comprehension, spatial audio reasoning and multi-audio understanding. We observe an even bigger improvement of Samsone-99M and Samsone134M models (34% - 36%) over Mellow (Table 2). We attribute this performance gain to our more diverse and extensive training data mix. Comparing Samsone to competing LALMs, we find that larger model size provides a slight advantage on this benchmark, which we argue, stems from greater inbuilt knowledge of LLMs, on which these LALMs are based on. Finally, there is a correlation of Samsone’s performance gain with its size increase, which is consistent with scaling laws [34]. 4.2. Question Answering, Captioning and Reasoning To obtain a full picture of Samsone’s performance, we test our models on audio captioning, simple AQA (yes/no and one word answers), and audio entailment tasks using Clotho and AudioCaps based datasets. In Table 3, we observe, that when it comes to simple AQA, all Samsone models outperform the competing SALMs. Similar holds true for the audio entailment tasks with a slight performance downgrade for Samsone-99M on Audio† https://github.com/soham97/mellow

Table 2: Comparison of Audio Language Models performance on MMAU and MMAU-Pro datasets. [29]. Sound

Music

Speech

Name

Size

Test-mini

LTU [13] GAMA [30] SALMONN [31] Qwen2-Audio-Instruct [32] GPT-4o Audio [10] Audio Flamingo 2 [7] Audio Flamingo 3 [8]

7B 7.4B 13B 8.4B – 3B 8.4B

20.42 31.83 41.14 67.27 64.56 71.47 79.58

Pengi [20] Mellow [14] Samsone-99M (ours) Samsone-134M (ours)

323M 167M 99M 134M

3.00 69.37 72.97 76.28

3.87 66.17 71.13 73.23

0.29 56.89 60.78 66.47

2.27 58.13 61.17 62.87

Samsone-356M (ours)

356M

75.98

74.27

70.34

65.83

Test

Test-mini

Test

Test-mini

Avg.

MMAU-Pro

Test

Test-mini

Test

15.33 16.97 28.77 55.37 69.33 44.87 66.97

17.44 20.82 34.90 59.60 62.50 62.40 73.30

17.23 21.68 36.23 57.40 60.82 61.06 72.42

33.46 33.20 39.60 45.41 52.50 42.60 51.70

3.30 30.63 37.84 46.25

4.77 35.73 42.10 47.90

2.20 52.30 57.20 63.00

3.63 53.34 58.13 61.33

28.61 27.50 36.83 37.57

44.74

45.90

63.70

62.00

40.67

Large Audio Language Models 20.67 30.73 42.10 61.17 63.20 68.13 75.83

15.97 17.71 37.13 56.29 56.29 70.96 73.95

15.68 17.33 37.83 55.67 49.93 70.20 74.47

15.92 12.91 26.43 55.26 66.67 44.74 66.37

Small Audio Language Models

Table 3: Comparison of SALMs on audio captioning task on AudioCaps [2] (AC) and Clotho [3] (CL) with SPICE [35] metric (sp.), on AQA task (ClothoAQA dataset [36] - (AQA)), and audio entailment task (Clotho CLE, AudioCaps ACE [37]). Model

AC sp.

CL AQA CLE ACE sp. acc. acc. acc.

Pengi (323M) [20] 12.7 7.0 Mellow (167M) [14] 17.8 9.4 Samsone-99M (ours) 14.4 11.1 Samsone-134M (ours) 14.4 11.6

63.6 71.4 71.8 73.8

37.3 91.2 92.4 93.4

38.7 89.7 89.1 93.7

Samsone-356M (ours)

74.5

93.7

93.5

14.9

11.7

Caps. These results prove that Samsone has a particularly good grounding in audio while using very few parameters. Samsone models are also competitive on audio captioning tasks, beating Mellow on the Clotho dataset. We attribute the inferior SPICE score on AudioCaps to the lower share of the AudioCaps dataset in the full training dataset, than originally used for Mellow. 4.3. Ablation Study We evaluate the key components of Samsone-134M architecture through an ablation study on MMAU benchmark. To ensure a controlled comparison, we utilize the LM without VP as the baseline, hence the 143M size (Samsone-134M-NP). We test the impact of replacing chosen modules with common alternatives: the AE with AST (Audio Spectrogram Transformer) [38], the projector with a linear layer, and the LM with GPT-2 [39]. Results in Table 4 show that all alternative configurations yield inferior performance. These findings validate our architectural choices for balancing size and audio understanding capability. Table 4: Ablation study of Samsone-134M. Model/change

Size

MMAUmini

MMAU

Samsone-134M-NP

143M

61.80

59.92

w/ AST AE w/ GPT-2 LM w/ Linear projector

221M 133M 142M

59.70 58.80 61.20

57.47 57.96 58.82

Table 5: On-device inference of the Samsone family models on Samsung Galaxy S25 Ultra. Variant

Audio [ms]

Query [ms]

Generation [tok/s]

Samsone-99M Samsone-134M Samsone-356M

667 757 1116

15 23 47

125 87 39

4.4. On-Device Inference We develop an Android application to benchmark Samsone’s inference latency on a Samsung Galaxy S25 Ultra CPU using 15 audio-query pairs. As shown in Table 5, the generation process involves prefilling the LM cache with audio and query embeddings followed by autoregressive token generation. The Samsone family achieves generation speeds between 39 and 125 tokens per second depending on model size, confirming their suitability for real-time edge applications. Notably, these results were achieved without hardware-level optimizations, suggesting significant potential for further speed improvements.

5. Conclusion In this paper, we introduced Samsone, a family of SALMs for on-device inference. Our core model, Samsone-134M, establishes a new SOTA for its size class, outperforming existing SALMs on multiple benchmarks, including MMAU (+15%) and MMAU-Pro (+36%), despite a 20% reduction in parameter count. Samsone-134M also exceeds the performance of much larger models, such as GAMA and LTU, on these tasks. To facilitate further research we open-source all project artifacts. These include the training code, checkpoints, mobile-optimized weights, and an Android application for real-time inference. Our work has limitations. First, extensive AQA finetuning causes the LM to lose general-purpose linguistic capabilities. Second, we focused on parameter count as a proxy for efficiency, however, recent Precision-Aware Scaling Laws [40] suggest that larger, quantized models may offer superior performance-to-memory trade-offs. Finally, we have not yet implemented hardware-specific optimizations for mobile GPUs or NPUs. This will be explored in future research.

6. Generative AI Use Disclosure The content of this paper was conceived, researched, and written entirely by the authors. Generative AI tools were utilized solely for editorial and grammatical refinement purposes, such as improving clarity, readability, and linguistic accuracy. These tools did not contribute to the generation of ideas, data analysis, interpretation of results, or creation of original content. The final manuscript reflects the full responsibility and intellectual effort of the authors.

7. References [1] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, pp. 776–780. [2] C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 119–132. [3] K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 736–740. [4] A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language, vol. 60, p. 101027, 2020. [5] B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5. [6] Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro, “Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,” in International Conference on Machine Learning. PMLR, 2024, pp. 25 125–25 148. [7] S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro, “Audio flamingo 2: An audiolanguage model with long-audio understanding and expert reasoning abilities,” in Forty-second International Conference on Machine Learning, 2025. [8] S. Ghosh, A. Goel, J. Kim, S. Kumar, Z. Kong, S. gil Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, and B. Catanzaro, “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [9] G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen et al., “Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,” arXiv preprint arXiv:2507.06261, 2025. [10] A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford et al., “Gpt-4o system card,” arXiv preprint arXiv:2410.21276, 2024. [11] G. Li, J. Liu, H. Dinkel, Y. Niu, J. Zhang, and J. Luan, “Reinforcement learning outperforms supervised fine-tuning: A case study on audio question answering,” arXiv preprint arXiv:2503.11197, 2025. [12] Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint audio and speech understanding,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2023. [13] Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=nBZBPXdJlC

[14] S. Deshmukh, S. Dixit, R. Singh, and B. Raj, “Mellow: a small audio language model for reasoning,” in The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [15] A. Singh, N. P. Patel, A. Ehtesham, S. Kumar, and T. T. Khoei, “A survey of sustainability in large language models: Applications, economics, and challenges,” in 2025 IEEE 15th Annual Computing and Communication Workshop and Conference (CCWC). IEEE, 2025, pp. 00 008–00 014. [16] J. C. L. Ong, S. Y.-H. Chang, W. William, A. J. Butte, N. H. Shah, L. S. T. Chew, N. Liu, F. Doshi-Velez, W. Lu, J. Savulescu et al., “Ethical and regulatory challenges of large language models in medicine,” The Lancet Digital Health, vol. 6, no. 6, pp. e428– e432, 2024. [17] P. Belcak, G. Heinrich, S. Diao, Y. Fu, X. Dong, S. Muralidharan, Y. C. Lin, and P. Molchanov, “Small language models are the future of agentic ai,” arXiv preprint arXiv:2506.02153, 2025. [18] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv et al., “Qwen3 technical report,” arXiv preprint arXiv:2505.09388, 2025. [19] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. AlDahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan et al., “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783, 2024. [20] S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” Advances in Neural Information Processing Systems, vol. 36, pp. 18 090–18 108, 2023. [21] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International conference on machine learning. PMLR, 2023, pp. 28 492–28 518. [22] L. B. allal, A. Lozhkov, E. Bakouch, G. M. Blazquez, G. Penedo, L. Tunstall, A. Marafioti, A. P. Lajarı́n, H. Kydlı́ček, V. Srivastav, J. Lochner, C. Fahlgren, X. S. NGUYEN, B. Burtenshaw, C. Fourrier, H. Zhao, H. Larcher, M. Morlon, C. Zakka, C. Raffel, L. V. Werra, and T. Wolf, “SmolLM2: When smol goes big — data-centric training of a fully open small language model,” in Second Conference on Language Modeling, 2025. [Online]. Available: https://openreview.net/forum?id=3JiCl2A14H [23] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “Llama: Open and efficient foundation language models,” 2023. [Online]. Available: https://arxiv.org/abs/2302.13971 [24] B.-K. Kim, G. Kim, T.-H. Kim, T. Castells, S. Choi, J. Shin, and H.-K. Song, “Shortened LLaMA: A simple depth pruning for large language models,” in ICLR 2024 Workshop on Mathematical and Empirical Understanding of Foundation Models, 2024. [Online]. Available: https: //openreview.net/forum?id=18VGxuOdpu [25] F. Sandri, E. Cunegatti, and G. Iacca, “2SSP: A two-stage framework for structured pruning of LLMs,” Transactions on Machine Learning Research, 2025. [Online]. Available: https://openreview.net/forum?id=Qd7LzJBg21 [26] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Online: Association for Computational Linguistics, Oct. 2020, pp. 38–45. [Online]. Available: https: //www.aclweb.org/anthology/2020.emnlp-demos.6 [27] W. Falcon and The PyTorch Lightning team, “PyTorch Lightning,” Mar. 2019. [Online]. Available: https://github.com/ Lightning-AI/lightning [28] I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations, 2019. [Online]. Available: https: //openreview.net/forum?id=Bkg6RiCqY7

[29] S. Sakshi, U. Tyagi, S. Kumar, A. Seth, R. Selvakumar, O. Nieto, R. Duraiswami, S. Ghosh, and D. Manocha, “MMAU: A massive multi-task audio understanding and reasoning benchmark,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https: //openreview.net/forum?id=TeVAZXr3yv [30] S. Ghosh, S. Kumar, A. Seth, C. K. R. Evuru, U. Tyagi, S. Sakshi, O. Nieto, R. Duraiswami, and D. Manocha, “Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 6288–6313. [31] C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=14rn7HpKVk [32] Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin et al., “Qwen2-audio technical report,” arXiv preprint arXiv:2407.10759, 2024. [33] S. Kumar, Š. Sedláček, V. Lokegaonkar, F. López, W. Yu, N. Anand, H. Ryu, L. Chen, M. Plička, M. Hlaváček et al., “Mmau-pro: A challenging and comprehensive benchmark for holistic evaluation of audio general intelligence,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 27, 2026, pp. 22 688–22 697. [34] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” arXiv preprint arXiv:2001.08361, 2020. [35] P. Anderson, B. Fernando, M. Johnson, and S. Gould, “Spice: Semantic propositional image caption evaluation,” in European conference on computer vision. Springer, 2016, pp. 382–398. [36] S. Lipping, P. Sudarsanam, K. Drossos, and T. Virtanen, “Clothoaqa: A crowdsourced dataset for audio question answering,” in 2022 30th European Signal Processing Conference (EUSIPCO). IEEE, 2022, pp. 1140–1144. [37] S. Deshmukh, S. Han, H. Bukhari, B. Elizalde, H. Gamper, R. Singh, and B. Raj, “Audio entailment: Assessing deductive reasoning for audio understanding,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 22, 2025, pp. 23 769–23 777. [38] Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” in Proc. Interspeech 2021, 2021, pp. 571–575. [39] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019. [40] T. Kumar, Z. Ankner, B. F. Spector, B. Bordelon, N. Muennighoff, M. Paul, C. Pehlevan, C. Re, and A. Raghunathan, “Scaling laws for precision,” in The Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https: //openreview.net/forum?id=wg1PCg3CUP

Record · ID 1006907 · SHA-256 ba4f8033fe2f8788
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.