TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios Hong Lyu ID , Mingru Yang ID , Qianhua He ID ∗∗ , Yanxiong Li ID ∗∗ , Jinxin Huang, Zhengyu Pei School of Electronic and Information Engineering, South China University of Technology, Guangzhou, China [email protected], [email protected], [email protected]
arXiv:2607.06179v1 [eess.AS] 7 Jul 2026
Abstract There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an Automatic Audio Annotation Pipeline–TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledgeguided subset (TriAGK ) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriAGK , TriAGK could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriAGK and the TriA Pipeline. Index Terms: Automatic audio annotation pipeline, Audio dataset, Audio classification
1. Introduction Audio Classification (AC) enables the recognition of various environmental sound events, serving as a core component in applications ranging from multimedia content analysis [1] and audio captioning [2] to bio-acoustic monitoring [3, 4]. Ultimately, the efficacy of these AC systems heavily depends on the availability of large-scale audio datasets that cover diverse acoustic scenes. Existing AC datasets can be broadly categorized into general-purpose (e.g., AudioSet [1], FSD50K [5], ESC-50 [6], FSC-89 [7]) and specialized ones (e.g., Kitchen20 [8], CHiMeHome [9], NSynth-100 [10]). The former encompass a wide variety of acoustic scenes, while the latter focus on specific domains, such as domestic environments or human non-speech sounds. Among general-purpose AC datasets, AudioSet is the most extensive, consisting of a class-balanced subset (AS-20K), a class-unbalanced subset (AS-2M), and an evaluation set, with AS-2M being widely utilized for pre-training and fine-tuning audio models [11, 12, 13, 14]. FSD50K is another large-scale audio dataset, but remains class-unbalanced, leaving certain sound classes underrepresented, such as wails, moans, wheezes, squeals, and specific domestic sounds. ESC-50, by contrast, is a widely used class-balanced dataset of 50 classes across 5 major categories, though its scale is notably limited. In summary, existing general-purpose AC datasets suffer from two primary limitations: (i) insufficient data for specific acoustic scenes, and (ii) limited dataset scale. Among specialized AC datasets, DESED [15] comprises real recordings (DESEDreal) and synthetic data, covering 10 ** indicates the corresponding author.
classes of domestic audio events, and is widely used for AC and sound event detection (SED) in domestic scenes [16, 17, 18]. However, DESEDreal contains only 5955 annotated clips. Kitchen20 is designed for kitchen AC tasks, while both HTAD [19] and CHiMe-Home target domestic activity recognition. CIRDO [20] and BiMP [21] are simulated datasets for safety monitoring of elderly individuals living alone, and Nonspeech7k [22], originally developed for paralinguistic classification, can similarly be applied to domestic safety monitoring. Nevertheless, all these specialized datasets suffer from limited scale. In general, annotated audio data remains scarce across many scenarios, with the problem becoming more pronounced in highly specific domains. To address the scarcity of annotated audio data in specific scenarios, we propose a large-scale Automatic Audio Annotation Pipeline–TriA Pipeline. TriA Pipeline consists of four stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering, which collectively convert raw audio from diverse streaming platforms and scenarios into high-quality training data with audio event annotations. The most closely related works are Emilia-Pipe [23], NVSpeech-Pipe [24], and NonVerbalSpeech-Pipe [25]. However, Emilia-Pipe relies on Automatic Speech Recognition (ASR) and supports only speech data annotation, while both NVSpeech-Pipe and NonVerbalSpeech-Pipe focus exclusively on paralinguistic speech annotation. In contrast, by integrating an AED module, TriA Pipeline supports audio event annotation across a broad range of scenarios, resolving the annotation scarcity problem in the specific domains described above. Both objective metrics and subjective listening evaluations confirm that TriA Pipeline produces high-quality and diverse audio data with high annotation reliability. Based on TriA Pipeline, the TriA dataset is further constructed, containing over 2130 hours of high-quality audio data covering 431 audio classes. To evaluate the effectiveness of TriA, we partition a prior-knowledge-guided subset, TriAGK , from TriA and setup three specific classification tasks: DESED for Audio Classification (DESEDAC ), Kitchen20, and Nonspeech7k. They represent three domestic AC tasks: general domestic AC, kitchen AC, and domestic safety monitoring. We conduct comparative experiments using TriAGK and manually annotated datasets. Experimental results show that fine-tuning models using only TriAGK achieves performance comparable to models fine-tuned using manually labeled data. Furthermore, fine-tuning the model by combining TriAGK with manually annotated data can achieve average relative improvements of 3.97% in accuracy and 3.35% in Macro-F1. These results indicate that TriAGK can help the model achieve better performance in the specified classification task, and further validate the effectiveness of the proposed TriA Pipeline.
AAD segments
Standardization
AED segments
Audio Event Detection
Audio Activity Detection
Filtering
mp3 mp3
[Speech, Music] [Speech]
Figure 1: Overview of the TriA Pipeline. AAD and AED denote Audio Activity Detection and Audio Event Detection.
2. TriA Pipeline
Overview of the TriA Pipeline
This section details the TriA Pipeline and its evaluation on minibatch data. As illustrated in Figure 1, the TriA Pipeline consists of four main stages: Standardization, Audio Activity Detection (AAD), Audio Event Detection (AED), and Filtering. 2.1. Standardization This step follows the same procedure as Emilia-Pipe and aims to standardize audio with heterogeneous formats for subsequent processing. Specifically, the original audio recordings are converted into mono-channel WAV format with a sample rate of 24 kHz and a 16-bit sample width. The target loudness is then normalized to -20 dBFS, while the signal amplitude is constrained within the range of -3 dB to 3 dB to prevent distortion. Finally, each waveform is normalized by dividing all sample points by the maximum amplitude. 2.2. Audio Activity Detection The purpose of this step is to split long audio recordings into clips of appropriate length while removing redundant segments. The auditok tool1 is used to split each sample into AAD segments based on a preset energy threshold, with the maximum segment duration constrained to 30 s. To determine the appropriate minimum duration for sound activity segments and maximum duration for silent segments, subjective listening tests can be conducted on datasets related to specific scenarios. For each audio class in each dataset, 10 instances are randomly selected for listening. During the tests, the minimum duration required to identify the event in each instance and the temporal interval between two adjacent events of the same class are recorded. Based on the test results, the average minimum time required to identify events of each class is calculated and referred to as the Event Critical Time (ECT). The average interval between adjacent events of each class is also calculated and referred to as the Silent Critical Time (SCT). During the AAD processing, retained activity segments are required to exceed the minimum ECT, as shorter segments contain insufficient information to identify. The retained silent segments are restricted to be shorter than the maximum SCT, since excessively long silent segments introduce redundant information and reduce efficiency. In our experiments, the target scenario is the domestic scene, where the minimum ECT and the maximum SCT are 1.2 s and 2.0 s, respectively. 2.3. Audio Event Detection To enable the TriA dataset to be directly applied to AC tasks, the BEATs model2 is employed to annotate audio events for AAD segments. The BEATs model achieves state-of-the-art (SOTA) performance on the AS-2M dataset under the single au1 https://github.com/amsehili/auditok 2 https://github.com/microsoft/unilm/tree/
master/beats
dio modality and can detect 527 audio classes [12]. Specifically, for each AAD segment, the AS-2M fine-tuned BEATsiter3+ model is used for event detection, dividing it into multiple segments with event annotations. To detect short events in the segment, local detection is conducted by scanning each AAD segment with a fixed detection window length and window shift, producing preliminary detection segments. Adjacent detection segments within the same AAD segment are then concatenated if their annotated Top-1 events are identical. To further detect long audio in the segment and continuous event across adjacent detection windows, global detection is performed on each concatenated detection segment. In global detection, the window length is equal to the length of the segment to be detected. If the globally detected class differs from the original class, the newly class is appended to the segment annotation. The resulting segments are called AED segments. The window length for AED local detection is required to exceed the maximum ECT (3 s). The purpose is to enable the model to process one or more target events as completely as possible in a single detection, providing sufficient information to the model and improving the reliability of the detection results. However, if the window length is too long, the audio segment detected by the model at one time may include too many events, which will bring difficulties to the model detection. The higher the event confidence threshold for AED detection, the higher the reliability of the detection results. In our experiment, the event confidence threshold is set to 0.6. The window length and window shift for local detection are set to 5 s and 3 s, respectively. 2.4. Filtering The original audio may exhibit varying quality and the BEATs model may produce detection errors. To ensure the data quality and improve the matching degree between data and event annotations, the audiobox-aesthetics3 and the CLAP model4 are used to filter AED segments. Specifically, Production Complexity (PC) and Production Quality (PQ) in aesthetics are used as filtering indicators. PC and PQ are relatively objective indicators, focusing on the complexity of the audio scene and the technical quality respectively [26]. For each AED segment, the PC and PQ, as well as the CLAP similarity between the segment and event annotation [27], are calculated. Segments with PC, PQ, or CLAP similarity below predefined thresholds are filtered. The remaining high-quality segments are stored in MP3 format, and a JSONL (JSON Lines) file containing the associated metadata is generated for efficient indexing and retrieval. 2.5. Evaluation on minibatch data To validate the effectiveness of each individual module and the overall TriA Pipeline, 284.7 hours of original audio are ran3 https://github.com/facebookresearch/ audiobox-aesthetics 4 https://github.com/microsoft/CLAP
Table 1: Statistical results of 285 hours of original data processed by the TriA Pipeline. Filtering1 uses PC, PQ, and CLAP similarity thresholds of 1.8, 5.5, and 5, respectively, while Filtering2 uses thresholds of 2.24, 5.85, and 7.43, respectively. Data
Original Audio Annotated w/o Filtering Annotated w Filtering1 Annotated w Filtering2
Aesthetics PC | PQ
Duration (s)
CLAP Similarity
min
max
avg ± std
min
avg ± std
min
avg ± std
2.04 1.20 1.20 1.20
17605.37 30.00 30.00 30.00
233.05 ± 841.70 10.87 ± 9.22 11.57 ± 9.38 11.70 ± 9.83
1.39 | 3.12 1.33 | 2.85 1.80 | 5.50 2.24 | 5.85
5.17 ± 1.66 | 7.20 ± 0.89 4.60 ± 1.71 | 6.52 ± 1.07 4.90 ± 1.62 | 6.92 ± 0.77 5.18 ± 1.45 | 7.10 ± 0.71
— -4.19 2.00 7.43
— 7.71 ± 2.52 8.04 ± 2.24 9.54 ± 1.65
domly collected from streaming platforms and processed using the pipeline. The experiment is conducted on an NVIDIA RTX 3090 GPU, with a total processing duration of about 10 hours and the RTF of 0.03. As shown in Table 1, statistics are performed across multiple objective evaluation perspectives. The original audio exhibits a broad duration distribution, with a relatively long average duration and a large variance. In contrast, the filtered data shows a more standardized and appropriate duration. Compared with the unfiltered data, the filtered data achieves much higher average PC, PQ, and CLAP Similarity, indicating that the data quality and annotation reliability have been significantly improved. However, the quality of the unfiltered data is lower than that of the original data. The reason is that the original data has longer average duration and higher average sampling rate, which favor PC and PQ measures. After increasing the filtering threshold, the quality and annotation reliability are further improved. However, it reduces the class diversity and the total duration of the data. In addition, subjective listening tests are also conducted, randomly selecting 100 samples to assess annotation accuracy, and an average accuracy of 93.67% is obtained. Overall, the results demonstrate that the TriA Pipeline is feasible and effective. It can increase the amount of data in each class and continuously expand the scale of training data through repeated large-scale audio data collection and pipeline processing.
Total Duration (hours)
Total Classes
284.70 (100.00%) 242.62 (85.22%) 182.37 (64.06%) 80.08 (28.13%)
— — 325 258
Figure 2: Class distribution of TriA. Table 2: PC, PQ, and CLAP similarity of different datasets. The metrics for TriA are calculated on a randomly sampled 300hour subset. Data
Aesthetics PC | PQ
CLAP Similarity
DESEDreal Kitchen20 Nonspeech7k
3.39 ± 0.86 | 5.85 ± 0.85 2.49 ± 0.37 | 6.38 ± 0.74 2.24 ± 0.67 | 6.34 ± 0.80
7.43 ± 4.89 12.87 ± 3.41 11.71 ± 3.14
TriA
5.13 ± 1.48 | 7.08 ± 0.67
9.62 ± 1.54
TRIAL MODE − Click here for more information
3. TriA Dataset 3.1. Statistics and Analysis Over 8706 hours of original audio data were collected from streaming platforms including Bilibili and Douyin, covering diverse topics such as daily life, entertainment media, and technology. After processing with the TriA Pipeline, the TriA dataset was constructed. The TriA dataset contains over 2130 hours of audio data, covering 431 audio classes. As illustrated in Figure 2, the class distribution of TriA is presented. Among them, the class with the largest number of samples is Music, which is attributed to the widespread presence of background music in streaming videos. Although the class distribution of the data is imbalanced, task-specific balanced subsets can be constructed for downstream applications. To analyze the data quality and annotation reliability of the TriA dataset, Table 2 summarizes statistics on PC, PQ, and CLAP similarity of different datasets. TriA achieves the best PC and PQ, indicating that the data quality of TriA is high, which benefits from the filtering stage in the TriA Pipeline. The CLAP similarity of TriA ranks third among all datasets, showing that its annotation reliability is not as good as that of Kitchen20 and Nonspeech7k, but better than that of DESEDreal . This difference can be attributed to the annotation method of the datasets. Kitchen20 and Nonspeech7k are manually annotated. TriA is automatically annotated using the pipeline, whereas DESEDreal is annotated through crowdsourcing.
3.2. Prior-knowledge-guided Subset The prior-knowledge refers to the threshold used for specific scenario, which was determined through subjective listening tests and statistics related to the specific scenario. The samples related to these prior-guided thresholds are collected to a subset, TriAGK . For convenience of subsequent experiment, TriAGK is further partitioned into three subsets: TriADESED , TriAKitchen20 , and TriANonspeech7k . The audio classes they cover are the same as the classes of DESEDreal , Kitchen20, and Nonspeech7k respectively. The class nomenclature of the TriA follows the AudioSet ontology [1]. For example, the class Electric shaver toothbrush in DESEDreal corresponds to Electric shaver, electric razor and Electric toothbrush in TriA. Table 3 reports the number of clips and total duration for each dataset. The scale of each subset of TriAGK is larger than that of the corresponding dataset.
4. Experiments This section evaluates the effectiveness of TriAGK for three domestic AC tasks: DESEDAC , Kitchen20, and Nonspeech7k. For each task, three experiments are conducted. The baseline model is trained with different datasets, the first on the manually annotated data, the second on the TriAGK , and the third on both, the TriAGK first and then the manually annotated data (re-
Table 3: Number of clips and total duration of different datasets. Unlabeled data are excluded from the statistics. Dataset
Table 4: Experimental results on three AC tasks. A + B denotes sequential fine-tuning, where the model is first fine-tuned on A and then further fine-tuned on B.
Clips
Total Durations (hours)
5955 1070 7014
16.50 1.47 6.75
Task
Dataset
Acc
F1
Manual annotate
DESEDreal Kitchen20 Nonspeech7k
DESEDAC
23517 1688 9618
42.60 2.45 11.29
0.7837 0.8255 0.8258
0.7943 0.7810 0.8256
TriAGK
TriADESED TriAKitchen20 TriANonspeech7k
DESEDreal TriADESED TriADESED + DESEDreal
Kitchen20
Kitchen20 TriAKitchen20 TriAKitchen20 + Kitchen20
0.9250 0.9375 0.9813
0.9272 0.9355 0.9812
Nonspeech7k
Nonspeech7k TriANonspeech7k TriANonspeech7k + Nonspeech7k
0.9448 0.8938 0.9490
0.9437 0.8734 0.9464
fer to Table 4), to analyze whether the pipeline-processed data can help the model learn the specified AC task. 4.1. Implementation Details 4.1.1. Baseline The BEATs model [12] is adopted. The backbone consists of a 12-layer Transformer encoder with 90M parameters, initialized with the pre-trained weights BEATsiter3+ . A linear classification head, comprising two fully connected layers, is appended to the backbone. Input audio is resampled to 16 kHz. We extract 128-dimensional Mel-filter bank features using a Povey window of 25 ms and a hop size of 10 ms. All experiments are conducted on an RTX 3090 GPU. The cross entropy loss and the AdamW optimizer are used. The maximum training epoch is set to 50, with an early stopping patience of 15. For the DESEDAC and Nonspeech7k tasks, the entire BEATs model is fully fine-tuned. For the Kitchen20 task, only the classification head is fine-tuned. The learning rates for full fine-tuning and adapter fine-tuning are set to 5e-5 and 6e-3, respectively, following a cosine annealing learning rate schedule. Accuracy and Macro-F1 are used as evaluation metrics. Accuracy measures the overall classification correctness, while Macro-F1 reflects the consistency of model performance across different classes. The validation metric is accuracy + Macro-F1. 4.1.2. Task Setup The DESEDreal training set contains 4429 clips, consisting of the real strongly labeled set and the real weakly labeled set. The validation set and test set contain 373 and 1153 clips respectively, which are derived from the real strongly labeled validation set and the real strongly labeled test set. To adapt DESEDreal to the AC setting, the strongly labeled sets are converted into weak labels by retaining only the audio paths and event annotations. TriADESED is split into training and validation sets with a 9:1 ratio. The test set for the DESEDAC task is the DESEDreal test set. Previous research evaluates Kitchen20 using 5-fold crossvalidation. To unify the test set for the Kitchen20 task, the fifth fold of Kitchen20 dataset is used as the test set, which contains 160 clips. The training and validation sets of Kitchen20 dataset contain 480 and 160 clips, corresponding to the first three folds and the fourth fold, respectively. TriAKitchen20 is divided into training and validation sets with a 4:1 ratio. The Nonspeech7k dataset provides 6289 training and 725 test clips. We further split the training clips into training and validation sets with a 9:1 ratio. TriANonspeech7k is also split into training and validation sets at a 9:1 ratio. The test set for the Nonspeech7k task is the Nonspeech7k test set.
4.2. Results Table 4 reports the experimental results on three AC tasks. The model fine-tuned on TriADESED achieves higher accuracy than that fine-tuned on DESEDreal , but obtaining a lower MacroF1. It suggests that TriADESED provides better overall data quality but exhibits slightly lower consistency across different classes. When applying sequential fine-tuning (TriADESED → DESEDreal ), the model significantly outperforms the one finetuned solely on DESEDreal , achieving relative gains of 5.37% in accuracy and 3.94% in F1. It indicates that TriADESED helps the pre-trained model acquire task-relevant knowledge, and then DESEDreal further improves the performance of the model. TriAKitchen20 outperforms Kitchen20 in both overall data quality and consistency across different audio classes. Sequential fine-tuning (TriAKitchen20 → Kitchen20) achieves relative improvements of 6.09% in accuracy and 5.82% in Macro-F1 compared to fine-tuning on Kitchen20 alone, demonstrating that TriAKitchen20 can help the model to learn the Kitchen20 task. TriANonspeech7k performs worse than Nonspeech7k in both accuracy and Macro-F1. However, sequential fine-tuning (TriANonspeech7k → Nonspeech7k) still leads to performance gains over fine-tuning only on Nonspeech7k, with relative gains of 0.44% in accuracy and 0.29% in F1. It further demonstrates that TriAGK can help the model learn the specific AC tasks. Overall, fine-tuning on pipeline-processed data achieves performance comparable to, and in some cases better than, that obtained using manually annotated data. Compared to fine-tuning on manually annotated data alone, combining TriAGK subsets with manual data through sequential finetuning achieves average relative improvements of 3.97% in accuracy and 3.35% in Macro-F1. It validates the effectiveness of TriAGK for domestic AC tasks and further confirms the feasibility and practical value of the TriA Pipeline.
5. Conclusion In this paper, we propose the TriA Pipeline, a large-scale automatic audio annotation pipeline. It efficiently converts audio collected from various streaming platforms into highquality training data with event annotations. Based on the TriA Pipeline, the TriA dataset is constructed, which contains over 2130 hours of audio data covering 431 audio classes. Through comparative experiments, we verify the effectiveness of the prior-knowledge-guided subset, TriAGK , on domestic AC tasks, and further confirm that the TriA Pipeline is effective. The TriA Pipeline code and dataset is now released5 . 5 https://github.com/huanxian/TriA
6. Acknowledgments This work was partly supported by the national natural science foundation of China (62371195, 62111530145), and the exchange project of the 10th Meeting of China-Croatia Science and Technology Cooperation Committee (10-34).
7. Generative AI Use Disclosure We used GPT-5.2 to assist in polishing the manuscript.
8. References [1] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2017, pp. 776–780. [2] K. Drossos, S. Lipping, and T. Virtanen, “Clotho: An audio captioning dataset,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 736–740. [3] R. R. Kvsn, J. Montgomery, S. Garg, and M. Charleston, “Bioacoustics data analysis–a taxonomy, survey and open challenges,” IEEE Access, vol. 8, pp. 57 684–57 708, 2020. [4] A. Terenzi, N. Ortolani, I. Nolasco, E. Benetos, and S. Cecchi, “Comparison of feature extraction methods for sound-based classification of honey bee activity,” IEEE/ACM transactions on audio, speech, and language processing, vol. 30, pp. 112–122, 2021. [5] E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 829–852, 2021. [6] K. J. Piczak, “ESC: Dataset for Environmental Sound Classification,” in Proceedings of the 23rd Annual ACM Conference on Multimedia. ACM Press, 2015, pp. 1015–1018. [Online]. Available: http://dl.acm.org/citation.cfm?doid=2733373.2806390 [7] Y. Li, W. Cao, J. Tan, Q. Li, and G. Chen, “Few-shot class-incremental audio classification using pseudo-incrementally trained embedding learner and continually updated stochastic classifier,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 3880–3895, 2025. [8] M. Moreaux, M. G. Ortiz, I. Ferrané, and F. Lerasle, “Benchmark for kitchen20, a daily life dataset for audio-based human action recognition,” in 2019 International Conference on Content-Based Multimedia Indexing (CBMI), 2019, pp. 1–6. [9] P. Foster, S. Sigtia, S. Krstulovic, J. Barker, and M. D. Plumbley, “Chime-home: A dataset for sound source recognition in a domestic environment,” in 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2015, pp. 1–5. [10] Y. Li, J. Tan, Q. Li, G. Chen, S. Huang, and T. Virtanen, “Few-shot open-set audio classification using attention information-fused prototypes,” IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 1929–1943, 2026. [11] Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “Panns: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 28, pp. 2880–2894, 2020. [12] S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “Beats: audio pre-training with acoustic tokenizers,” in Proceedings of the 40th International Conference on Machine Learning, 2023, pp. 5178–5193. [13] X. LI and X. Li, “Atst: Audio representation learning with teacher-student transformer,” in Proc. Interspeech 2022, 2022, pp. 4172–4176. [14] H. Dinkel, Z. Yan, Y. Wang, J. Zhang, Y. Wang, and B. Wang, “Streaming audio transformers for online audio tagging,” in Proc. Interspeech 2024, 2024, pp. 1145–1149.
[15] N. Turpault, R. Serizel, A. Parag Shah, and J. Salamon, “Sound event detection in domestic environments with weakly labeled data and soundscape synthesis,” in Workshop on Detection and Classification of Acoustic Scenes and Events, 2019. [16] Z. Lin, Y. Li, Z. Huang, W. Zhang, Y. Tan, Y. Chen, and Q. He, “Domestic activities clustering from audio recordings using convolutional capsule autoencoder network,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021, pp. 835–839. [17] T. Khandelwal and R. K. Das, “A Multi-Task Learning Framework for Sound Event Detection using High-level Acoustic Characteristics of Sounds,” in Interspeech 2023, 2023, pp. 1214–1218. [18] T. Khandelwal, R. K. Das, A. Koh, and E. S. Chng, “Leveraging audio-tagging assisted sound event detection using weakified strong labels and frequency dynamic convolutions,” in 2023 IEEE Statistical Signal Processing Workshop (SSP), 2023, pp. 329–333. [19] E. Garcia-Ceja, V. Thambawita, S. A. Hicks, D. Jha, P. Jakobsen, H. L. Hammer, P. Halvorsen, and M. A. Riegler, “Htad: A home-tasks activities dataset with wrist-accelerometer and audio features,” in International Conference on Multimedia Modeling. Springer, 2021, pp. 196–205. [20] M. Vacher, S. Bouakaz, M.-E. Bobillier-Chaumon, F. Aman, R. A. Khan, S. Bekkadja, F. Portet, E. Guillou, S. Rossato, and B. Lecouteux, “The cirdo corpus: comprehensive audio/video database of domestic falls of elderly people,” in 10th International Conference on Language Resources and Evaluation, 2016, pp. 1389–1396. [21] J. Dibble and M. C. Bazzocchi, “Bi-modal multiperspective percussive (bimp) dataset for visual and audio human fall detection,” IEEE Access, 2025. [22] M. M. Rashid, G. Li, and C. Du, “Nonspeech7k dataset: Classification and analysis of human non-speech sound,” IET Signal Processing, vol. 17, no. 6, p. e12233, 2023. [23] H. He, Z. Shang, C. Wang, X. Li, Y. Gu, H. Hua, L. Liu, C. Yang, J. Li, P. Shi et al., “Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,” in IEEE Spoken Language Technology Workshop (SLT), 2024, pp. 885–890. [24] H. Liao, Q. Ni, Y. Wang, Y. Lu, H. Zhan, P. Xie, Q. Zhang, and Z. Wu, “Nvspeech: An integrated and scalable pipeline for human-like speech modeling with paralinguistic vocalizations,” arXiv preprint arXiv:2508.04195, 2025. [25] R. Ye, Y. Zhou, R. Yu, Z. Lin, K. Li, X. Li, X. Liu, G. Zeng, and Z. Wu, “A scalable pipeline for enabling non-verbal speech generation and understanding,” arXiv preprint arXiv:2508.05385, 2025. [26] A. Tjandra, Y.-C. Wu, B. Guo, J. Hoffman, B. Ellis, A. Vyas, B. Shi, S. Chen, M. Le, N. Zacharov, C. Wood, A. Lee, and W.-N. Hsu, “Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound,” arXiv preprint arXiv:2502.05139, 2025. [27] B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5.