ConceptioArchivearXiv CS
arXiv CSopen access

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin Zhiheng Qian1 , Aini Li2 , Hai Hu3 , Liang Zhao4 1

3

Shanghai Jiao Tong University; 2 City University of Hong Kong; The Hong Kong Polytechnic University; 4 Beijing Foreign Studies University,

[email protected], [email protected], [email protected], [email protected]

arXiv:2607.21332v1 [cs.CL] 23 Jul 2026

Abstract Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language varieties. We address this by training text-dependent and text-independent aligners for Chengdu Mandarin using a 17-hour corpus and a custom G2P dictionary. We trained a text-dependent GMM-HMM model (Chengdu-MFA) and fine-tuned a pretrained audio encoder on frame classification with Chengdu-MFA’s pseudo label for textindependent alignment (Chengdu-FC). Evaluation on an expertannotated test set show that both methods significantly outperform Standard Mandarin baselines. Chengdu-MFA reduced average phone boundary differences by 31.8%, while ChengduFC achieved a 61.2% reduction. This work establishes a practical bootstrapping pipeline for developing accurate aligners for under-resourced varieties without labor- and time-intensive manual annotation. Index Terms: phonetic forced alignment, low-resource speech data, Chengdu Mandarin

1. Introduction The increasing availability of spoken language data has heightened the need for reliable automated methods in phonetic analysis. Phonetic forced alignment is a critical technique for synchronizing speech transcription with audio at the utterance, word, and phone levels [1, 2]. By automating time-aligned annotations, dedicated aligners such as the Penn Forced Aligner [3], the Prosodylab-Aligner [4], FAVE [5], the Montreal Forced Aligner (MFA) [6], and Charsiu Forced Aligner [7], have greatly facilitated large-scale phonetic and sociolinguistic research. However, existing models for phonetic forced alignment are predominantly trained on standardized, high-resource languages (e.g., Standard Mandarin). When applied to regional or non-standard language varieties, their performance may decline due to systematic differences in the sound systems. Although some alignment toolkits allow researchers to train custom models, doing so from scratch is particularly challenging for low-resource varieties, which typically lack both massive speech corpora and specialized phonetic dictionaries. This presents a significant methodological bottleneck: how can researchers efficiently develop forced aligners for regional varieties that achieve reasonably accurate alignment with a manageable amount of data? In particular, how can this be accomplished when manual annotations are scarce or entirely absent, and when text transcriptions are also lacking? To address these challenges, we propose a transferable pipeline for developing variety-specific aligners for both text-dependent and -independent alignment. We used Chengdu Mandarin as a case study. Chengdu

Mandarin belongs linguistically to the Southwestern Mandarin group [8] and is spoken by more than 20 million speakers. Despite having tens of millions of native speakers, Chengdu Mandarin is severely under-resourced in speech technology. The aligners we trained for this variety thus serve as valuable tools for phonetic research, and our work provides practical guidance for developing similar resources for other low-resource language varieties. Previous studies have demonstrated the effectiveness of GMM-HMM-based systems for phonetic forced alignment [6]. More recently, pretrained speech encoders have been adapted for this task, although they do not consistently outperform traditional GMM-HMM approaches [7]. In this study, we employed both approaches to develop forced aligners for Chengdu Mandarin. Specifically, we collected approximately 17 hours of Chengdu Mandarin speech and constructed an expert-annotated grapheme-to-phoneme (G2P) dictionary tailored to its sound system. Using these resources, we first trained a GMM-HMMbased acoustic model, for text-dependent alignment. We then applied Chengdu-MFA to the corpus to automatically generate phone-level pseudo-labels. These pseudo-labels served as supervision for fine-tuning a pretrained speech encoder on a frame-classification task, enabling text-independent phonetic segmentation at inference time. To assess model performance, two phoneticians manually annotated a small set of Chengdu recordings as the gold standard annotations, and compared them with the output generated by our trained models for Chengdu Mandarin as well as the baseline models pre-trained on Standard Mandarin. Alignment evaluation followed the methods in the previous studies [2, 6, 7, 9, 10, 11]. Our main contributions are threefold: (1) We release the first dedicated aligners for Chengdu Mandarin, supporting both text-dependent and text-less forced alignment, as well as a specialized G2P dictionary. (2) We provide empirical evidence demonstrating the limitations of applying standard-language models to regional varieties, highlighting the necessity of variety-specific training. (3) Most importantly, we establish and validate an end-to-end bootstrapping pipeline (G2P Dictionary → text-dependent aligner →pseudo-labels → text-independent aligner) that provides the speech research community with a reproducible workflow for developing alignment tools for other under-resourced language varieties.

2. Method 2.1. G2P Dictionary for Chengdu Mandarin While Mandarin varieties share a character-based writing system, they differ considerably in their sound inventories. The phone set used by Standard Mandarin models does not apply

to the sound inventory of Chengdu Mandarin. Therefore, we compiled a Chengdu Mandarin dictionary covering all the 2876 Chinese characters in the master dataset. The dictionary was first automatically annotated using the Pypinyin library and DeepSeek-v3 [12]. Two native speakers of Chengdu Mandarin then reviewed the model-generated annotations and corrected errors. 2.2. Chengdu-MFA We developed Chengdu-MFA, a GMM-HMM-based acoustic model for text-dependent forced alignment, using the Montreal Forced Aligner (MFA) framework [6]. In a GMM-HMM system, an utterance is represented as a sequence of hidden phonetic states, whose temporal transitions are modeled by hidden Markov models, while the acoustic distribution associated with each state is modeled using Gaussian mixture models. Given an audio recording, its orthographic transcription, and a pronunciation dictionary, the model identifies the most likely state sequence and thereby estimates word- and phone-level boundaries. GMM-HMM systems do not require manually annotated phone boundaries for training. Instead, the model jointly estimates its acoustic parameters and latent state alignments from the utterance-level transcripts and their corresponding recordings. During forced alignment, the observed audio was constrained by the phone sequence derived from the transcript, and the most likely alignment path was decoded to obtain word- and phone-level timestamps. In addition to performing text-dependent alignment, Chengdu-MFA served as the first stage of our bootstrapping pipeline. We applied the trained model to the Chengdu Mandarin training corpus to generate phone-level alignments automatically. 2.3. Chengdu-FC To exploit the acoustic representations learned by pretrained speech encoders for phonetic forced alignment, a prevalent strategy is to cast the alignment process as a frame-level classification problem [7]. Given an input audio sequence of T frames, the model predicts a phone label ŷt for each frame t. The standard training objective minimizes the average framelevel cross-entropy loss: T

Lseq (y, ŷ) =

1 X LCE (yt , ŷt ), T t=1

boundary-weighted loss:  PT  t=1 wt LCE (yt , ŷt ) Lseq (y, ŷ) = , PT wt t=1   wt = 1 + (γ − 1) Ibound (t, r),

(2)

where γ ≥ 1 is the boundary-weighting factor and r specifies the radius, measured in frames, around each reference boundary. The indicator function Ibound (t, r) equals 1 if frame t lies within r frames of any phone boundary and 0 otherwise. Thus, boundary-adjacent frames receive a weight of γ, whereas all other frames retain a weight of 1.

3. Experiments 3.1. Dataset The master dataset for this study consisted of 15 high-quality audio recordings of Chengdu Mandarin and the corresponding text transcriptions drawn from [13]. The audios were produced by 15 speakers born and raised in urban Chengdu, a southwestern city in Sichuan, China. Each audio ranges from 0.6 to 1.7 hours; the total duration is about 17.7 hours. The sampling rate of the audios is 24 kHz. We created subsets of this master dataset for different procedures involved in our model training and evaluation: A test set was first created containing 50 minutes of recordings from the master dataset, sampled from the recordings of 10 speakers, 5 minutes per speaker. The test audios were manually aligned and transcribed at the word (Chinese characters) and phone (IPA symbols) levels by linguistic experts. This manual annotation serves as the gold standard against which the alignments from the baseline and our trained MFA and FC models are evaluated. The training set for Chengdu MFA acoustic model took the full master dataset minus the test set. To prepare MFA input, utterance-level TextGrids were created based on text transcriptions with timestamps from ELAN [14]. Irrelevant information such as the punctuation marks, paralinguistic annotations (e.g., ((laugh))), and other non-lexical elements were removed from the TextGrids. We used Chengdu-MFA model to provide phonetic annotation to its training set. Of these generated phone annotations, 80% were allocated to the Chengdu FC training set and 20 % to the Chengdu FC validation set to select hyperparameters.

(1)

where y = (y1 , . . . , yT ) denotes the reference phone-label sequence and ŷ = (ŷ1 , . . . , ŷT ) denotes the corresponding model predictions. We developed Chengdu-FC model series using this learning objective with different audio encoders. The learning objective requires speech data with framelevel phone annotations. Because manually annotated phone boundaries are costly to obtain, we used Chengdu-MFA to generate phone-level pseudo-labels for the training corpus as supervision for fine-tuning the pretrained speech encoder and its frame-classification head. To improve the model’s sensitivity to phone boundaries, we adopted a curriculum-learning strategy. The model was first trained for E epochs using the standard loss in Equation 1. In subsequent epochs, frames near phone boundaries were assigned greater weights than frames in phone-internal regions. Specifically, for epochs after E, we used the following

3.2. Training Setup 3.2.1. Chengdu MFA We trained the Chengdu-MFA acoustic model using MFA toolkit version 3.3.3 [6]. Following the standard practice in G2P modeling for Chinese varieties, each syllable nucleus combined with its tone was modeled as a single unit. 3.2.2. Chengdu FC We fine-tined Wav2Vec2-base, Wav2Vec2-large, XLS-R-300m, Charsiu Madnarin-FC on Chengdu Mandarin respectively. We dub the finetuned models as Chengdu-FC-base, Chengdu-FClarge, Chengdu-FC-xlsr and Chengdu-FC-charsiu. The training was done using AdamW with a weight decay of 1 × 10−4 , and a batch size of 8 on a single NVIDIA RTX 3090 GPU. The training ran for 10 epochs in total. Following a curriculum learning strategy, the first 2 epochs used the

Figure 1: Histograms of absolute differences (on a log scale) between forced-aligned boundaries and gold-standard annotations. The first and second rows show the word and phone tiers respectively. Dashed line represents the average boundary difference.

standard cross-entropy loss. For the remaining epochs, training switched to the boundary-aware weighted loss (2). Hyperparameters were selected based on the classification accuracy on the validation set. We searched for the optimal learning rate within [1 × 10−5 , 3 × 10−4 ], boundary weight γ ∈ [5, 15], and weighting radius r ∈ [0, 2]. 3.3. Evaluation We performed forced alignment on the test set using the baseline Mandarin models and our trained Chengdu models. Textdependent alignment was generated by both the MFA and FC models, while text-independent alignment was performed by FC models only. All the alignments were compared to the goldstandard manual annotations. Note that all our systems used IPA-based phone systems, therefore to ensure fair comparison across systems (our trained models vs. baselines), we treated the syllable nucleus and coda as a single unit in evaluation. For text-dependent alignment, we measured the absolute differences between the gold standard and our alignments at the start boundaries, and presented the percentage of word and phone boundaries under different time thresholds [2, 6, 9]. For text-independent alignment, because boundary matching is not applicable, we evaluated the precision (p), recall (r), F1-score (F1), and R-value of the boundaries at different time tolerances [7, 10, 11]. 3.4. Result 3.4.1. Text-dependent alignment Figure 1 presents the distribution of the absolute differences between the gold standard and the aligned boundaries generated using our Chengdu Mandarin forced aligners. Both the Chengdu-MFA and -FC aligners outperformed the respective baselines at both word and phone tier. Specifically, the average boundary difference of Chengdu-MFA was 22.1 ms at the word tier, representing 26.6% reduction compared to the Mandarin baseline. At phone tier, the average difference was 22.3 ms, a 31.8% improvement over the baseline (Table 1). Meanwhile, all Chengdu-FC models exhibited better performance than the Mandarin-FC baseline. In particular, the Chengdu-FCxlsr achieved an average boundary difference of 32.5 ms at word tier (62.3% shorter than the baseline) and 30.2 ms at phone tier

(61.2% shorter than the Charsiu-Mandarin-FC baseline). As shown in Table 2, although the Mandarin-MFA model achieved the highest proportion of predictions within 10 ms, the Chengdu-MFA outperformed it at all cutoffs above 10 ms. The superior performance of Mandarin-MFA at 10 ms threshold suggests a high concentration of small errors, likely due to its large amount of training data and the pronunciation similarity of these two Mandarin varieties. However, it exhibited a heavier tail in the distribution, leading to a lower cumulative proportion as the tolerance threshold increases, whereas Chengdu-MFA model demonstrates more robust performance across broader tolerance ranges. Regarding FC models in the text-dependent task, the Chengdu-FC-xlsr model demonstrated comparable performance to the MFA models. Specifically, it aligned 82.2% of phone tier boundaries within 50 ms. This significantly outperforms the Charsiu-Mandarin-FC baseline (52.8%) and beats the performance of the Mandarin-MFA baseline at all the time thresholds. Table 1: Mean and median boundary differences between MFA and FC models (Mandarin vs. Chengdu) and the gold standard. Differences reported are all statistically significant (Welch’s t-test, p < .001). Bold numbers mark the best in each column; underlines indicate the best FC model. Model

Word Tier↓ mean med.

Mandarin-MFA Chengdu-MFA (ours) Charsiu-Mandarin-FC Chengdu-FC-charsiu (ours) Chengdu-FC-base (ours) Chengdu-FC-large (ours) Chengdu-FC-xlsr (ours)

30.1 22.1 80.0 68.6 46.2 35.6 32.5

11.9 13.5 47.5 34.8 23.3 17.4 15.4

Phone Tier↓ mean med. 32.7 22.3 77.9 69.8 42.0 33.3 30.2

14.1 12.9 41.0 35.4 19.6 15.4 13.8

3.4.2. Text-independent alignment Table 3 details the text-independent alignment performance across different time tolerances (τ ). Our Chengdu-FC models demonstrate a substantial advantage in the R-value. When

Table 2: Percentage of boundary differences within different time thresholds.

Boundary Difference (ms)

< 10

Mandarin-MFA (Word) Mandarin-MFA (Phone) Chengdu-MFA (Word) Chengdu-MFA (Phone) Charsiu-Mandarin-FC (Word) Charsiu-Mandarin-FC (Phone) Chengdu-FC-charsiu (Word) Chengdu-FC-charsiu (Phone) Chengdu-FC-base (Word) Chengdu-FC-base (Phone) Chengdu-FC-large (Word) Chengdu-FC-large (Phone) Chengdu-FC-xlsr (Word) Chengdu-FC-xlsr (Phone)

.463 .414 .406 .406 .258 .234 .293 .279 .313 .337 .382 .398 .405 .429

Percentage within↑ < 25 < 50 < 100 .662 .643 .691 .678 .327 .375 .439 .429 .520 .559 .603 .634 .645 .669

.825 .814 .881 .868 .528 .560 .597 .592 .715 .743 .784 .800 .813 .822

.924 .924 .978 .969 .776 .768 .793 .774 .872 .885 .913 .925 .932 .942

τ = 20, the R-value of our Chengdu-FC-large model is 21.1% larger than the Mandarin baseline. While the Mandarin FC model exhibited slightly higher recall at large time tolerances (e.g., 0.954 when τ = 100), this comes at the cost of low precision, which indicates that the model introduced severe oversegmentation when applied to Chengdu Mandarin. Regarding training strategies, all three Chengdu-FC variants significantly outperformed the Standard Mandarin baseline (Charsiu-Mandarin-FC). Among them, Chengdu-FC-large has the largest number of parameters, followed by the ChengduFC-base, and then Chengdu-FC-Charsiu. Fine-tuning from Wav2Vec2-large yielded the best performance, achieving the highest F1-scores and R-values across almost all settings. Although Chengdu-FC-Charsiu was previously fine-tuned on Mandarin data, this did not compensate for the disparity in model size, resulting in relatively reduced performance comparing to the other Chengdu-FC models.

4. Discussion Our results demonstrate the benefits of variety-specific training for phonetic alignment. Under the present experimental conditions, Chengdu-MFA provided the most accurate textdependent alignments, despite being trained on only approximately 17 hours of utterance-transcribed speech. GMM-HMMbased systems therefore remain a practical option for underresourced language varieties when transcripts and a pronunciation dictionary are available. Chengdu-MFA also provides an efficient means of generating phone-level pseudo-labels without manual boundary annotation. These pseudo-labels can be used to train frameclassification models for transcript-free phonetic segmentation. Although the Chengdu-FC models were generally less accurate than Chengdu-MFA in text-dependent alignment, they can operate without input transcripts at inference time. The MFA and FC approaches therefore address complementary application scenarios. The Chengdu-specific models consistently outperformed their Standard Mandarin counterparts, confirming the benefits of variety-specific adaptation. Nevertheless, the Standard Mandarin models achieved non-trivial performance, possibly be-

Table 3: Precision (p), recall (r), F1, and R-value at different time tolerances (τ , ms) for Charsiu-Mandarin-FC, Chengdu-FC-base, Chengdu-FC-Charsiu, and Chengdu-FC-large models. τ (ms)

Model

p

r

F1

R-val

20

Charsiu-Mandarin-FC Chengdu-FC-base Chengdu-FC-charsiu Chengdu-FC-large Chengdu-FC-xlsr

.375 .464 .459 .488 .482

.589 .616 .643 .665 .691

.457 .528 .535 .562 .567

.289 .482 .453 .500 .463

40

Charsiu-Mandarin-FC Chengdu-FC-base Chengdu-FC-charsiu Chengdu-FC-large Chengdu-FC-xlsr

.495 .593 .588 .618 .597

.775 .787 .824 .842 .857

.602 .675 .685 .712 .703

.401 .601 .566 .613 .560

60

Charsiu-Mandarin-FC Chengdu-FC-base Chengdu-FC-charsiu Chengdu-FC-large Chengdu-FC-xlsr

.563 .641 .629 .648 .630

.882 .849 .882 .882 .904

.684 .728 .733 .746 .742

.455 .639 .597 .635 .584

80

Charsiu-Mandarin-FC Chengdu-FC-base Chengdu-FC-charsiu Chengdu-FC-large Chengdu-FC-xlsr

.596 .663 .649 .667 .648

.932 .880 .910 .908 .930

.725 .754 .756 .768 .763

.478 .656 .611 .648 .595

100

Charsiu-Mandarin-FC Chengdu-FC-base Chengdu-FC-charsiu Chengdu-FC-large Chengdu-FC-xlsr

.611 .677 .663 .681 .660

.954 .898 .930 .927 .947

.743 .770 .773 .785 .777

.488 .666 .620 .657 .603

cause of the phonological similarity between Standard Mandarin and Chengdu Mandarin and their large-scale training data. Cross-variety generalization also varied across modeling frameworks, with Mandarin-MFA outperforming Charsiu-MandarinFC in text-dependent alignment. Because the current evaluation contains speakers whose other recordings were included in training, future work should assess generalization to unseen speakers and additional speech domains.

5. Conclusion In this study, we created a Chengdu Mandarin phonetic forced alignment dataset with a G2P dictionary covering all characters in the dataset. We trained a MFA acoustic model for Chengdu Mandarin for text-dependent alignment and a Chengdu Mandarin frame classification model capable of performing both text-dependent and -independent alignment. Our trained models consistently outperformed the Mandarin baselines with identical architectures but far more training data. For suggestions on applying these models, if utterance-level transcripts are available, the MFA model can provide highly accurate word and phone annotation; if not, the FC model enables effective textindependent alignment. This work contributes to the forced alignment tools for low-resource language varieties by presenting a complete training and evaluating pipeline for a specific Mandarin variety.

6. References [1] J. Yuan, W. Lai, C. Cieri, and M. Liberman, “Using forced alignment for phonetics research,” in Chinese language resources: Data collection, linguistic analysis, annotation and language processing. Springer, 2023, pp. 289–301. [2] E. Chodroff, E. P. Ahn, and H. Dolatian, “Comparing languagespecific and cross-language acoustic models for low-resource phonetic forced alignment,” 2025. [3] J. Yuan, M. Liberman et al., “Speaker identification on the scotus corpus,” Journal of the Acoustical Society of America, vol. 123, no. 5, p. 3878, 2008. [4] K. Gorman, J. Howell, and M. Wagner, “Prosodylab-aligner: A tool for forced alignment of laboratory speech,” Canadian acoustics, vol. 39, no. 3, pp. 192–193, 2011. [5] I. Rosenfelder, J. Fruehwald, K. Evanini, S. Seyfarth, K. Gorman, H. Prichard, and J. Yuan, “Fave (forced alignment and vowel extraction) suite version 1.1. 3,” 2014. [6] M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi.” in Interspeech, vol. 2017, 2017, pp. 498–502. [7] J. Zhu, C. Zhang, and D. Jurgens, “Phone-to-audio alignment without text: A semi-supervised approach,” in ICASSP 2022 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 8167–8171. [8] C. N. Li and S. A. Thompson, Mandarin Chinese: A functional reference grammar. Univ of California Press, 1989. [9] R. Rousso, E. Cohen, J. Keshet, and E. Chodroff, “Tradition or innovation: A comparison of modern asr methods for forced alignment,” arXiv preprint arXiv:2406.19363, 2024. [10] O. J. Räsänen, U. K. Laine, and T. Altosaar, “An improved speech segmentation quality measure: the r-value.” in Interspeech, 2009, pp. 1851–1854. [11] F. Kreuk, J. Keshet, and Y. Adi, “Self-supervised contrastive learning for unsupervised phoneme segmentation,” arXiv preprint arXiv:2007.13465, 2020. [12] A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan et al., “Deepseek-v3 technical report,” arXiv preprint arXiv:2412.19437, 2024. [13] A. Li, “Regional dialect leveling in mandarin chinese: The case of locative variation in the chengdu dialect,” Asia-Pacific Language Variation, vol. 8, no. 1, pp. 32–71, 2022. [14] P. Wittenburg, H. Brugman, A. Russel, A. Klassmann, and H. Sloetjes, “Elan: a professional framework for multimodality research,” in International Conference on Language Resources and Evaluation, 2006. [Online]. Available: https://api.semanticscholar.org/CorpusID:18212263

Record · ID 394468 · SHA-256 bebdf3472651beb7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.