Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt Yanfeng Shi1 , Pengfei Cai1 , Jun Liu1 , Qing Gu1 , Nan Jiang1 , Lirong Dai1 , Ian McLoughlin2 , Yan Song1,∗∗ 1
National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China, Hefei, China 2 ICT Cluster, Singapore Institute of Technology, Singapore
arXiv:2604.13715v1 [cs.SD] 15 Apr 2026
yanf [email protected], [email protected], [email protected], [email protected]
Abstract Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, these models still face challenges in temporal perception (e.g., inferring event onset and offset), leading to limited utility in fine-grained scenarios. To address this issue, we propose Audio-Side Time Prompt and leverage Reinforcement Learning (RL) to develop the TimeProRL framework for fine-grained temporal perception. Specifically, we encode timestamps as embeddings and interleave them within the audio feature sequence as temporal coordinates to prompt the model. Furthermore, we introduce RL following Supervised Fine-Tuning (SFT) to directly optimize temporal alignment performance. Experiments demonstrate that TimePro-RL achieves significant performance gains across a range of audio temporal tasks, such as audio grounding, sound event detection, and dense audio captioning, validating its robust effectiveness. Index Terms: large audio-language model, fine-grained temporal perception, reinforcement learning
1. Introduction Audio conveys a wealth of information, ranging from human speech to environmental events, and serves as a fundamental modality for perceiving the world[1, 2]. Large Audio-Language Models (LALMs) have significantly advanced general audio understanding by integrating the linguistic reasoning of Large Language Models (LLMs) with audio encoders [3, 4]. These models have demonstrated strong versatility across a wide range of applications, such as acoustic scene classification, audio captioning and audio question answering [5, 6, 7]. Based on largescale cross-modal pre-training, LALMs exhibit a remarkable ability to comprehend diverse acoustic content and interact with users through flexible natural language, providing a unified interface for audio-related tasks. Despite these advancements, research indicates that current LALMs still exhibit shortcomings in fine-grained temporal understanding [8]. In many practical applications, models are required not only to interpret acoustic content but also to capture the temporal structure or event boundaries within it[9, 10]. Although LALMs excel at semantic recognition, they often struggle to precisely infer the onset and offset timestamps of specific sound events. This gap exposes a key limitation in finegrained temporal perception, which is essential for temporally grounded tasks such as audio grounding[11] and sound event detection[12]. ** indicates the corresponding author.
Prior work has strengthened temporal perception through the construction of time-annotated datasets [13], yielding significant performance improvements. Other approaches introduce specialized time tokens [14] to represent the concept of time, enhancing temporal alignment during generative inference. While such advancements are promising, two aspects merit further attention: 1) The audio input in LALMs lacks the explicit modeling of physical temporal cues, limiting the precise alignment between semantic content and its actual temporal coordinates. 2) Supervised Fine-Tuning (SFT) primarily focuses on semantic correctness, lacking optimization signals to address time-boundary prediction deviations. Notably, Reinforcement Learning (RL) has shown potential in Video Temporal Grounding (VTG) [15, 16], where reward signals can be directly designed around temporal alignment metrics, offering valuable insights for analogous tasks in the audio domain. Motivated by these observations, we propose the TimeProRL framework to enhance the fine-grained temporal perception of LALMs via two primary modules: temporal information integration and training objective design. Specifically, we introduce Audio-Side Time Prompt to inject temporal coordinates into audio features, then employ SFT to instruct the model in utilizing these coordinates. Subsequently, we further develop RL post-training with an advantage-driven adaptive temporal reward to incorporate temporal alignment quality into the optimization objective. Through the synergy of these strategies, TimePro-RL achieves significant performance gains across a range of audio temporal tasks.
2. Method In this section, we elaborate on the TimePro-RL framework, the overall schematic of which is illustrated in Figure 1. We first present the infusion of temporal cues into LALM’s audio input (Section 2.1), and then describe the RL post-training paradigm tailored for audio temporal tasks, focusing on reward design (Section 2.2). 2.1. Audio-Side Time Prompt Most LALMs predominantly rely on position embeddings, such as RoPE [17], to capture sequential structure. While these models can be trained to extract temporal information from sequence inputs [18], directly inferring absolute timestamps remains challenging. For VTG task, a previous study has overlaid frame indices onto video frames to provide perceptible time references, thereby enhancing temporal localization performance [19]. The core of this approach lies in providing “time
(a) Audio-Side Time Prompt Audio Feature
A
(b) SFT & RL Post-Training {“a railroad crossing rings”: [[0.1, 4.4], [7.3, 10.0]]}
Audio Token
Timestamp Embedding
T
Timestamp Token
Text Embedding
T
Text Token
SFT Stage
Full Data LALM
T
T
···
T
Predict Text Tokens
Tokenize
T
T
···
T
Label Text Tokens
Label T
···
T
T
Large Language Model
RL Stage
Subset Data
···
A
0.08s
A
A
0.16s
T
A
0.24s
T
···
T
T
T
···
T
···
T
T
t
T
Group Predictions
When does “ a railroad crossing rings” happen?
Group Relative Policy Optimization
···
A
T
T
···
A
T
Group Size
Audio Encoder
Label
Text Embedding Layer LALM
Timestamp Embedding Layer
Ground-Truth Time Segments
Adaptive Temporal Reward
T
Extract Time Segments
T
Predicted Time Segments
R1 R2 ··· RG
Figure 1: Overview of our TimePro-RL framework. Timestamp Embeddings are interleaved within audio features as time prompt, followed by SFT and RL post-training.
prompt” at the input level, effectively reducing reasoning difficulty and mitigating hallucinations. Inspired by this, we propose Audio-Side Time Prompt (ASTP), in which timestamps are encoded into embedding vectors and interleaved within the audio feature sequence to constitute LALM’s audio input. Specifically, we first extend the tokenizer with a set of Timestamp Tokens (e.g., <0.04>), each corresponding to a specific time point in seconds. During preprocessing, the audio input is partitioned into an audio token sequence based on its duration and the output frame rate of the audio encoder. Since each position in this sequence maintains a fixed mapping to the timeline, Timestamp Tokens can be inserted into the sequence according to their corresponding time points, serving as explicit temporal coordinates to prompt the model. The following shows an example of the processed input at the maximum time resolution: <s><audio><AUDIO><0.04><AUDIO><0.08><AUDIO> <0.12><AUDIO><0.16>· · · </audio>When does "a railroad crossing rings" happen?</s>
where <s> and </s> indicate the start and end of the sequence, while <audio> and </audio> demarcate the audio segment. <AUDIO> serves as a placeholder token, which is then replaced by the corresponding audio frame feature. In this example, the audio encoder’s output frame rate is 25 Hz, and tokens such as <0.04> represent the inserted Timestamp Tokens. Following this, Timestamp Tokens in the sequence are mapped to vector representations via the Timestamp Embedding Layer. To ensure the stability of the embedding space, we adopt an initialization strategy based on semantic priors: each Timestamp Embedding is initialized as the mean of the subword embeddings derived from tokenizing its corresponding numerical string. Taking <0.04> as an example: E⟨0.04⟩ =
1 |T (0.04)|
X
eu
(1)
u∈T (0.04)
where E⟨0.04⟩ indicates the embedding vector of Timestamp
Token <0.04>, eu represents the pre-trained embedding of the token with ID u in the vocabulary, and T (0.04) is the set of token IDs obtained by tokenizing the numerical string “0.04”. Adding time prompt to LALM’s input aligns well with its autoregressive architecture. This enables the model to retrieve temporal information from neighboring Timestamp Embeddings when attending to event cues within audio frames. Since the input structure is altered, SFT is required to guide the model in correctly understanding and utilizing this prompt. The semantic initialization strategy facilitates this by transferring pre-trained knowledge from the original language model. Moreover, all E⟨t⟩ parameters are frozen during training to prevent semantic drift. 2.2. RL for audio temporal tasks While LALM’s temporal perception can be manifested in its ability to localize sound events, the standard SFT objective remains misaligned. Consider a ground-truth segment [5.0 s, 6.0 s], a reasonable prediction of [4.9 s, 5.9 s] would be heavily penalized by token-level cross-entropy loss, which risks overfitting and impairs generalization capability [20]. RL offers a viable path to address this problem by optimizing for evaluation metrics, and has been successfully applied in generation tasks like summarization by directly utilizing ROUGE or CIDEr as rewards [21, 22]. However, the potential of RL for audio temporal tasks has yet to be extensively explored. For further enhancing fine-grained temporal perception of LALMs, we extend the training process following SFT by adopting Group Relative Policy Optimization (GRPO) [23], which computes relative advantages through group sampling. To align the training objective with temporal alignment performance, we utilize the Event-based F1 score (Eb-F1) [24], an established metric in sound event detection, as the main reward (rmain ). However, we observe in our experiments that since EbF1 is a threshold-based discrete metric, the limited group size in GRPO may result in identical rewards among sampled predictions, which leads to advantage degeneration and diminishes data efficiency. To address this, we incorporate a continuous auxiliary re-
Table 1: Performance comparison of TimePro-RL and zero-shot / finetuned baseline models on audio temporal tasks. Bold indicates the best results in each category. Model
Audio Grounding
Scale [email protected]
Sound Event Detection mIoU
Zero-shot 11.9 27.7 Finetuned 19.0 43.3 36.5 57.8 34.6 69.6 34.1 69.9 34.5 70.6 3.3 10.6
Dense Audio Captioning
Eb-F1
METEOR
Eb-F1
3.4 13.7
11.2 10.5
3.0 10.4
Qwen2-Audio Qwen2.5-Omni
7B 7B
9.2 25.4
5.1 17.4
Audio-Flaming2 TimeAudio Qwen2-Audio Qwen2.5-Omni Kimi-Audio
3B 7B 7B 7B 7B
37.0 75.7 74.8 74.0 76.1
27.6 61.2 57.9 59.8 60.0
8.9 − 49.8 48.9 50.9
25.7 20.4 32.2 31.3 31.2
12.7 37.4 35.0 35.2 32.7
Qwen2-Audio Qwen2.5-Omni
7B 7B
78.8 80.1
Post-Trained with TimePro-RL (Ours) 64.0 38.1 72.9 58.4 66.3 39.8 74.4 57.6
35.3 33.9
39.8 40.7
ward (raux ), such as the mean Intersection over Union (mIoU), to provide smoother optimization signals. We develop an advantage-driven adaptive temporal reward mechanism: ( R=
rmain ⊙ raux , rmain ,
if Var(rmain ) < ϵ otherwise
(2)
where R indicates the group reward vector for advantage computation, Var(·) represents the variance operator, and ϵ is the variance threshold. If rmain lacks discriminability among predictions in a group, the algorithm adopts the element-wise product of rmain and raux as the fused reward. This strategy leverages the smoothness of raux to recover the advantage signal while utilizing rmain as a weight to regulate optimization intensity. By dynamically adjusting the reward calculation based on real-time data conditions during training, this mechanism improves data efficiency while maintaining high temporal alignment quality.
This task is constructed using the FTAR dataset [8], and outputs follow the format: onset-offset, description. To measure performance, we employ METEOR [26] to assess the linguistic quality of the generated captions and Eb-F1 to evaluate the accuracy of temporal localization. The sample sizes for the training and test sets of each task are detailed in Table 2. Table 2: Summary of data statistics across the three tasks. Task Audio Grounding Sound Event Detection Dense Audio Captioning
Training Size
Test Size
61,862 15,041 92,443
483 1,153 741
3.2. Implementation details
3. Experimental setup 3.1. Tasks and datasets We conduct experiments across three representative audio temporal tasks: Audio Grounding (AG) [11]. The audio grounding task is defined as localizing a specific sound event within an audio clip based on a descriptive natural language query. For this task, we utilize the FTAR dataset [8] and require the model to output the corresponding timestamps in the {"query": [onset, offset]} format. Performance is measured via Intersection over Union (IoU) metrics, where we report recall values at various thresholds including [email protected], [email protected], and [email protected] alongside the mean IoU (mIoU). Sound Event Detection (SED) [12]. Sound event detection involves identifying event categories from a predefined set along with their respective occurrence periods. We conduct this task on the DESED dataset [25], with model outputs formatted as {"event": [onset, offset]}. The evaluation metric of this task leverages the Eb-F1 score [24], and we set the boundary tolerance to 0.2 s. Dense Audio Captioning (DAC) [8]. For dense audio captioning, the model should generate descriptions for the sound events in the audio clip paired with their associated timestamps.
To evaluate the effectiveness of the TimePro-RL framework, we utilize Qwen2-Audio [27] and Qwen2.5-Omni [28] as base models. Both of them integrate Whisper [29] as the audio encoder with an output frame rate of 25 Hz. To support AudioSide Time Prompt at the maximum time resolution, we expand the tokenizer with 750 Timestamp Tokens, covering from 0 s to 30 s with a stride of 0.04 s. We conduct parameter-efficient fine-tuning using LoRA [30] with r = 8 and α = 32. The overall post-training pipeline consists of two stages: 1) SFT: The models are trained on the full dataset for 3 epochs with a learning rate of 1 × 10−5 . 2) RL: Following SFT, we take GRPO for only a single epoch using a subset of 10,200 samples. The group size is set to 4, and the learning rate is 1 × 10−6 . In the adaptive temporal reward mechanism, we employ Eb-F1 as rmain across all three tasks. For raux , we utilize mIoU for AG and SED, while METEOR for DAC. The variance threshold ϵ is specified as 1 × 10−6 .
4. Results In this section, we first evaluate TimePro-RL against baselines under zero-shot and finetuned settings. Then, we conduct ablation studies to analyze the contributions of components in TimePro-RL.
Table 3: Ablation study of different components. ASTP denotes Audio-Side Time Prompt; “random init” refers to random initialization of Timestamp Embeddings; RL(Eb-F1) indicates using only Eb-F1 as the reward. Method
Audio Grounding
Sound Event Detection
Dense Audio Captioning
(Qwen2.5-Omni)
mIoU
Eb-F1
METEOR
Eb-F1
SFT Baseline
74.0
59.8
34.1
69.9
48.9
31.3
35.2
w/ ASTP (random init) w/ ASTP w/ ASTP + RL (Eb-F1) w/ ASTP + RL
73.2 77.6 77.8 80.1
57.2 61.7 63.1 66.3
32.8 35.8 38.9 39.8
68.8 71.7 72.7 74.4
46.0 50.1 56.9 57.6
31.4 32.6 31.6 33.9
33.3 37.0 38.1 40.7
4.1. Main results As shown in Table 1, we first evaluate Qwen2-Audio and Qwen2.5-Omni under zero-shot conditions, where their performance on high-precision metrics (e.g., [email protected] and Eb-F1) is notably constrained, revealing the limitations of existing generalpurpose LALMs in fine-grained temporal perception. Subsequently, we compare models post-trained via TimePro-RL framework with several prominent LALMs adapted by SFT on the same dataset, including Qwen2-Audio, Qwen2.5-Omni, Audio-Flamingo2 [31] and Kimi-Audio [32]. For TimeAudio, which is only trained on the FTAR dataset, we adopt the results reported in its original publication [8]. The results show that TimePro-RL consistently outperforms these baselines across multiple evaluation metrics, and crucially, it demonstrates a strong competitive edge in high-precision localization. For instance, in AG task, Qwen2.5-Omni improves from 34.1 during the SFT stage to 39.8 on [email protected]. Similarly, in DAC task, its Eb-F1 score rises from 35.2 to 40.7, while Qwen2-Audio also achieves a 4.8-point improvement. These gains validate the effectiveness of TimePro-RL in enhancing the fine-grained temporal perception of LALMs. 4.2. Ablation study We conduct a series of ablation experiments based on Qwen2.5Omni to systematically investigate the contributions of each component in TimePro-RL framework, with the results summarized in Table 3. Compared to the SFT baseline, using random initialization in Audio-Side Time Prompt leads to performance regressions across most metrics, such as a 2.9-point drop in SED Eb-F1, because randomly initialized Timestamp Embeddings introduce extraneous noise into the audio feature sequence. In contrast, the semantic initialization strategy proposed in Section 2.1 allows the model to correctly interpret and leverage Timestamp Embeddings as temporal coordinates, offering performance gains including a 1.7-point increase in AG [email protected]. This underscores the critical importance of semantic initialization for Audio-Side Time Prompt. Furthermore, our results demonstrate that even a modest volume of RL training significantly boosts model performance, with AG [email protected] increasing from 35.8 to 39.8 compared to the SFT (w/ ASTP) stage. In terms of the reward design, the adaptive temporal reward mechanism in Section 2.2 provides more balanced overall gains compared to the Eb-F1 only configuration. While the latter causes an optimization imbalance where the METEOR score for DAC task declines from the 32.6 achieved in the SFT (w/ ASTP) stage to 31.6, our adaptive approach recovers this score to 33.9 and pushes the final Eb-F1 to 40.7, facilitating more comprehensive data utilization.
Figure 2: Visualization of attention weights on Timestamp Embeddings. For the audio grounding query “a train horn honking”, the mel-spectrogram (top) is aligned with the chronologically arranged attention weights assigned to each Timestamp Embedding (bottom).
Figure 2 provides a visual analysis of the attention weights assigned to Timestamp Embeddings to further elucidate the internal mechanism of our framework. we extract the attention weights of the generate tokens toward each Timestamp Embedding from the model’s final layer and plot them along the timeline. The attention map reveals that the model’s focus exhibits high-intensity activations that precisely align with the onset and offset boundaries of the sound events marked in the melspectrogram. This sharp concentration of attention on the start and end coordinates suggests that the model effectively leverages Timestamp Embeddings to capture precise temporal cues for sound events. Such clear alignment provides intuitive evidence for the interpretability of Audio-Side Time Prompt, confirming that it successfully guides the model in perceiving and utilizing fine-grained temporal information integrated in the audio feature sequence.
5. Conclusion In this paper, we propose TimePro-RL, a framework to enhance the fine-grained temporal perception of LALMs. TimePro-RL interleaves Timestamp Embeddings into audio feature sequence to provide temporal cues, and introduces RL post-training with an adaptive temporal reward designed for temporal alignment to further strengthen temporal capabilities. Experiments show that TimePro-RL achieves significant performance gains across a range of audio temporal tasks, validating the effectiveness of the synergy between Audio-Side Time Prompt and RL posttraining. Future research will explore the application of our method in complex reasoning scenarios, such as Chain-ofThought (CoT), where fine-grained temporal cues serve as critical intermediate evidence.
6. References [1] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017, pp. 776–780. [2] Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880–2894, 2020. [3] S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 36. Curran Associates, Inc., 2023, pp. 18 090–18 108. [4] C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: towards generic hearing abilities for large language models,” in The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, 2024. [5] Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in The Twelfth International Conference on Learning Representations (ICLR). OpenReview.net, 2024. [6] Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” arXiv preprint arXiv:2311.07919, 2023. [7] P. K. Rubenstein, C. Asawaroengchai, D. D. Nguyen, A. Bapna, Z. Borsos, F. d. C. Quitry, P. Chen, D. E. Badawy, W. Han, E. Kharitonov et al., “Audiopalm: A large language model that can speak and listen,” arXiv preprint arXiv:2306.12925, 2023. [8] H. Wang, Y. Li, S. Ma, H. Liu, and X. Wang, “Listening between the frames: Bridging temporal gaps in large audio-language models,” arXiv preprint arXiv:2511.11039, 2025. [9] A. Mesaros, T. Heittola, T. Virtanen, and M. D. Plumbley, “Sound event detection: A tutorial,” IEEE Signal Processing Magazine, vol. 38, no. 5, pp. 67–83, 2021. [10] C.-H. H. Yang, S. Ghosh, Q. Wang, J. Kim, H. Hong, S. Kumar, G. Zhong, Z. Kong, S. Sakshi, V. Lokegaonkar et al., “Multi-domain audio question answering toward acoustic content reasoning in the DCASE 2025 challenge,” arXiv preprint arXiv:2505.07365, 2025. [11] X. Xu, H. Dinkel, M. Wu, and K. Yu, “Text-to-audio grounding: Building correspondence between captions and sound events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2021, pp. 606–610.
[17] J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing, vol. 568, p. 127063, 2024. [18] A. K. Sridhar, Y. Guo, and E. Visser, “Enhancing temporal understanding in audio question answering for large audio language models,” in Conference of the Nations of the Americas Chapter of the ACL (NAACL), 2025, pp. 1026–1035. [19] Y. Wu, X. Hu, Y. Sun, Y. Zhou, W. Zhu, F. Rao, B. Schiele, and X. Yang, “Number it: Temporal grounding videos like flipping manga,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 13 754–13 765. [20] Y. Wang, Z. Wang, B. Xu, Y. Du, K. Lin, Z. Xiao, Z. Yue, J. Ju, L. Zhang, D. Yang et al., “Time-r1: Post-training large vision language model for temporal video grounding,” arXiv preprint arXiv:2503.13377, 2025. [21] R. Paulus, C. Xiong, and R. Socher, “A deep reinforced model for abstractive summarization,” in International Conference on Learning Representations (ICLR). OpenReview.net, 2018. [22] S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Selfcritical sequence training for image captioning,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 7008–7024. [23] D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948, 2025. [24] A. Mesaros, T. Heittola, and T. Virtanen, “Metrics for polyphonic sound event detection,” Applied Sciences, vol. 6, no. 6, p. 162, 2016. [25] N. Turpault, R. Serizel, A. P. Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classification of Acoustic Scenes and Events (DCASE), 2019. [26] S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005, pp. 65–72. [27] Y. Chu, J. Xu, Q. Yang, H. Wei, X. Wei, Z. Guo, Y. Leng, Y. Lv, J. He, J. Lin et al., “Qwen2-audio technical report,” arXiv preprint arXiv:2407.10759, 2024. [28] J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y. Fan, K. Dang et al., “Qwen2.5-omni technical report,” arXiv preprint arXiv:2503.20215, 2025.
[12] D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, and M. D. Plumbley, “Detection and classification of acoustic scenes and events,” IEEE Transactions on Multimedia, vol. 17, no. 10, pp. 1733–1746, 2015.
[29] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in International Conference on Machine Learning (ICML). PMLR, 2023, pp. 28 492–28 518.
[13] Z. Xie, X. Xu, Z. Wu, and M. Wu, “Audiotime: A temporally-aligned audio-text benchmark dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5.
[30] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “Lora: Low-rank adaptation of large language models.” in International Conference on Learning Representations (ICLR), vol. 1, no. 2, 2022, p. 3.
[14] H. Wang, Z. Xu, Y. Cheng, S. Diao, Y. Zhou, Y. Cao, Q. Wang, W. Ge, and L. Huang, “Grounded-videollm: Sharpening finegrained temporal grounding in video large language models,” in Findings of the Association for Computational Linguistics (EMNLP). Association for Computational Linguistics, 2025, pp. 959–975.
[31] S. Ghosh, Z. Kong, S. Kumar, S. Sakshi, J. Kim, W. Ping, R. Valle, D. Manocha, and B. Catanzaro, “Audio flamingo 2: An audiolanguage model with long-audio understanding and expert reasoning abilities,” in International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 267. PMLR / OpenReview.net, 2025.
[15] X. Li, Z. Yan, D. Meng, L. Dong, X. Zeng, Y. He, Y. Wang, Y. Qiao, Y. Wang, and L. Wang, “Videochat-r1: Enhancing spatio-temporal perception via reinforcement fine-tuning,” arXiv preprint arXiv:2504.06958, 2025.
[32] D. Ding, Z. Ju, Y. Leng, S. Liu, T. Liu, Z. Shang, K. Shen, W. Song, X. Tan, H. Tang et al., “Kimi-audio technical report,” arXiv preprint arXiv:2504.18425, 2025.
[16] L. Dong, H. Zhang, H. Lin, Z. Yan, X. Zeng, H. Zhang, Y. Huang, Y. Wang, Z.-H. Ling, L. Wang et al., “Videotg-r1: Boosting video temporal grounding via curriculum reinforcement learning on reflected boundary annotations,” arXiv preprint arXiv:2510.23397, 2025.