ConceptioArchivearXiv CS
arXiv CSopen access

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Zero-Shot Morphological Discovery in Low-Resource Bantu Languages via Cross-Lingual Transfer and Unsupervised Clustering Hillary Mutisya Thiomi NLP

John Mugane Harvard University

arXiv:2604.22723v1 [cs.LG] 24 Apr 2026

Abstract We present a method for discovering morphological features in low-resource Bantu languages by combining cross-lingual transfer learning with unsupervised clustering. Applied to Giriama (nyf), a language with only 91 labeled paradigms, our pipeline discovers noun class assignments for 2,455 words and identifies two previously undocumented morphological patterns: an a- prefix variant for Class 2 (vowel coalescence—the merger of two adjacent vowels—of wa-, 95.1% consistency) and a contracted k’- prefix (98.5% consistency). External validation on 444 known Giriama verb paradigms confirms 78.2% lemmatization accuracy, while a v3 corpus expansion to 19,624 words (9,014 unique lemmas) achieves 97.3% segmentation and 86.7% lemmatization rates across all major word classes. Our ensemble of transfer learning from Swahili and unsupervised clustering, combined via weighted voting, exploits complementary strengths: transfer excels at cognate detection (leveraging ∼60% vocabulary overlap) while clustering discovers language-specific innovations invisible to transfer. We release all code and discovered lexicons to support morphological documentation for low-resource Bantu languages.

1. Introduction 1.1 Motivation Morphological analysis is fundamental to linguistic documentation and natural language processing, yet most of the world’s 7,000+ languages lack comprehensive morphological resources. This is particularly acute for the Bantu language family (500+ languages, 300M+ speakers), where noun class systems—a defining typological feature—remain undocumented for many languages. Consider Giriama (Bantu E.72b in Guthrie’s classification, ∼600,000 speakers, Kenya coast): despite being a living language with substantial speaker populations, only 91 morphological paradigms (annotated word-lemma pairs with grammatical features) exist in computational form. Standard supervised learning approaches would achieve poor coverage with such minimal data. Yet Giriama shares approximately 60% vocabulary with Swahili (a high-resource relative with 816+ paradigms), suggesting cross-lingual transfer may be viable. 1.2 Research Questions The discovery pipeline builds on BantuMorph [Mutisya and Mugane, 2026], a ByT5-small character-level model trained on 16 Bantu languages for morphological analysis. BantuMorph’s encoder maps words from any Bantu language into a shared embedding space where morphologically similar words—including cross-lingual cognates—cluster together. We exploit this property for zero-shot noun class discovery in Giriama and 15 other Bantu languages. This work addresses three key questions:

1. Zero-shot discovery: Can we discover morphological features in a low-resource language using minimal supervision (n < 100 paradigms)? 2. Language-specific innovation: Can unsupervised methods identify morphological patterns unique to the target language? 3. Method complementarity: How do cross-lingual transfer and unsupervised clustering complement each other for morphological discovery?

1.3 Contributions 1. Novel multi-method approach: We combine transfer learning (K-nearest neighbors in embedding space), unsupervised clustering (UMAP + K-means), and ensemble validation. 2. Empirical validation: On Giriama, we discover 2,455 noun class labels (27× increase) and validate the underlying model on 444 known paradigms (78.2% lemmatization accuracy). 3. Linguistic discoveries: Two previously undocumented Giriama patterns: the a- prefix variant (Class 2, 95.1% consistent) and the k’- contracted prefix (98.5% consistent). 4. Scalability: Our approach requires only a character-level pretrained model, a related high-resource language, and a small unlabeled corpus. 5. Open resources: Code, discovered lexicons, and visualizations.

2. Related Work 2.1 Morphological Analysis for Low-Resource Languages Supervised approaches [Kirov et al., 2018, Sylak-Glassman et al., 2015] require large annotated datasets. Semi-supervised methods reduce annotation burden but still require seed data [Cotterell et al., 2017, Kann et al., 2017]. Unsupervised morphology induction [Creutz and Lagus, 2007, Goldsmith, 2001, Hammarström and Borin, 2011] discovers structure without supervision but struggles with rare affixes. Cross-lingual transfer [Buys and Botha, 2016, Cotterell et al., 2018, McCarthy et al., 2019] exploits typological similarity, showing promise for related languages but missing language-specific innovations. Our work combines both approaches.

Swahili

Giriama

(816 labeled nouns)

(7,812 sentences)

Transfer Learning

Unsupervised Clustering

(KNN, K=5)

(UMAP+K-means)

Ensemble Voting (weighted, conf ≥ 0.70)

Noun Class Labels + Confidence Figure 1: Three-stage discovery pipeline. Row 1: Data sources (labeled Swahili nouns + unlabeled Giriama corpus). Row 2: Two complementary methods—KNN transfer finds cognates, clustering discovers innovations. Row 3: Weighted ensemble produces highconfidence noun class labels.

2.2 Bantu Language Morphology Bantu languages exhibit rich agglutinative morphology with noun class systems [Maho, 1999, Marten and Kula, 2012]. Each noun belongs to one of 15–20 classes marked by prefixes that trigger agreement on verbs, adjectives, and other words in the sentence. Computational work on Bantu morphology includes analyzers for Swahili [Hurskainen, 1992], Zulu [Pretorius and Bosch, 2009], and recent neural methods [Vylomova et al., 2020]. Most of the 500+ Bantu languages remain understudied computationally.

Listing 1: Transfer learning algorithm 1. Extract embeddings: For each word w in Ls: es[w] = M.encode(w) For each word w in Lt: et[w] = M.encode(w) 2. For each target word wt: a. Find K=5 nearest source neighbors b. Vote for class via majority vote c. Confidence = vote_conf x sim_conf 3. Return: {(wt, c_pred, conf)}

2.3 Embedding-Based Morphology Peters et al. [2018] and Devlin et al. [2019] show that contextualized embeddings capture morphosyntactic information. Character-level models [Kim et al., 2016, Xue et al., 2022] handle morphological variation naturally. Cross-lingual embeddings [Conneau et al., 2020] enable transfer. ByT5 [Xue et al., 2022], our base model, operates at the character level and has shown strong cross-lingual transfer for morphologically rich languages.

How it works: BantuMorph’s encoder maps words from any Bantu language into a 1,472-dimensional space where morphologically similar words cluster together. For a targetlanguage word like Giriama akimbola, the 5 nearest Swahili neighbors might all be Class 2 plural forms (e.g., wanafunzi, watu), yielding a Class 2 prediction with confidence proportional to neighbor agreement and cosine similarity. Advantages: High precision on cognates (∼60% vocabulary overlap for Giriama–Swahili); interpretable; leverages labeled source data. Limitations: Cannot detect language-specific innovations absent from the source language (see Section 6 for how clustering addresses this gap).

3. Methodology Figure 1 illustrates our three-component pipeline. 3.1 Problem Formulation Input: • M : Character-level pretrained model (ByT5) • Ls : High-resource source language with noun class labels (Swahili) • Lt : Low-resource target language (Giriama) s • Ps = {(wi , ci )}N i=1 : Labeled paradigms in Ls • Ct : Unlabeled corpus in Lt Output: Noun class assignments Ĉ = t {(wj , ĉj , confj )}N j=1 for words in Lt .

3.3 Method 2: Unsupervised Clustering Intuition: Words in the same noun class cluster together in embedding space due to shared prefix patterns and agreement contexts (where nouns trigger matching prefixes on verbs and modifiers). Listing 2: Clustering algorithm 1. Extract noun candidates from corpus Ct 2. Dimensionality reduction: - UMAP: reduce to d=50 dimensions 3. Cluster: - K-means with K=12 clusters 4. Analyze each cluster: - Extract prefix patterns (first 1-3 chars) - Map to noun class via prefix-class table

3.2 Method 1: Transfer Learning via Cross-Lingual Projection Intuition: Related languages share cognates with similar embeddings; nearest neighbors likely have the same noun class. 2

set heuristically and found insensitive to moderate variations); minimum confidence 0.70. We distinguish two quality metrics: confidence is the ensemble’s per-word prediction score, combining neighbor agreement and cosine similarity; consistency is the percentage of words in a cluster that share the dominant prefix pattern. Runtime: ∼15 minutes for 1,000 sentences.2

5. Return: {(w, c_cluster, consistency)}

How it works: UMAP projects the high-dimensional embeddings to 50 dimensions preserving local structure. Kmeans (K=12, matching the typical number of productive (actively used to form new words) Bantu noun classes) partitions words into clusters. Each cluster is mapped to a noun class by extracting the dominant prefix pattern (first 1–3 characters) and matching against a cross-linguistically compiled Bantu prefix inventory (e.g., ma- → Class 6, ki- → Class 7). Clusters with no clear prefix match are labeled “unknown.” Advantages: Discovers language-specific patterns invisible to transfer; requires no labeled data. Limitations: Lower precision; cluster-to-class mapping is heuristic and may fail for classes with ambiguous prefixes (e.g., mu-: Class 1 or 3).

We first present the Giriama case study in detail (Sections 5.1– 5.3), then the multi-language scaling results (Section 5.4).

3.4 Method 3: Ensemble Validation

5.1 Giriama: Noun Class Discovery

Intuition: Multi-method agreement indicates high-confidence predictions; disagreements reveal ambiguity or innovation.

Applied to the Giriama corpus (7,812 sentences), the ensemble pipeline discovers noun class assignments for 2,455 words (27× increase over the 91 known paradigms). Transfer learning contributes 8,698 predictions (mean confidence 0.71); unsupervised clustering contributes 18,508; the high-confidence ensemble retains 5,279.

4.3 Baselines (1) Frequency baseline: assign most common class (Class 6) to all words; (2) Random baseline; (3) Transfer-only; (4) Clustering-only. 5. Results

Listing 3: Ensemble voting For each word w predicted by multiple methods: score(class) = Sum weight(m) x confidence(m, c) weights = {transfer: 1.0, clustering: 0.8} Require minimum threshold 0.70

Cross-method agreement. Transfer–clustering agreement on Giriama is 36.7%. Agreement is highest for morphologically transparent features (non-finite forms 78.3%, present tense 61.2%) and lowest for complex or rare forms (future 29.6%, perfect 21.2%). The low overall agreement reflects the complementarity of the two methods (Section 6).

Advantages: Highest precision; conservative; identifies ambiguous cases. 4. Experimental Setup 4.1 Data Model: BantuMorph v7 (ByT5-small, 300M parameters), trained on 16 Bantu languages with 80,765 paradigms across 5 tasks (segmentation, lemmatization, inflection, feature extraction, noun class prediction). Embeddings are extracted from the encoder’s final layer with mean pooling over the byte sequence. Source Language (Swahili): 816 entries with noun class labels; 14 noun classes (1–11, 14–16). Target Language (Giriama): 91 training paradigms (verb only, from UniMorph); 7,812 sentences from the EnglishGiriama parallel dataset [Lingua-Connect, 2025]; ∼600,000 speakers (Kenya coast); Bantu E.72b (a member of the Mijikenda group of coastal Kenya Bantu languages). Giriama shares approximately 60% vocabulary with Swahili.

5.2 Novel Giriama Morphological Discoveries 5.2.1 The a- Prefix Variant (Class 2) Transfer learning from Swahili would assign Class 2 words the standard wa- prefix. Unsupervised clustering identified Cluster 1 (266 words, 95.1% consistency) using the a- prefix variant: a vowel coalescence of wa- → a- characteristic of coastal Bantu dialects: akimbola "they ran" (cf. Swahili wakimbilia) akimanywa "they were known" (cf. Swahili walijulikana) akimwamba "they told him" (cf. Swahili walimwambia)

This pattern accounts for 19.6% of all Class 2 words in the Giriama corpus and was undetectable by transfer learning.

4.2 Implementation Transfer Learning: K = 5 nearest neighbors; cosine similarity in ByT5 embedding space; confidence threshold 0.60. Clustering: UMAP (reducing to 50 dimensions); K-means (K = 12).1 Ensemble: Weights: transfer=1.0, clustering=0.8 (transfer weighted higher due to its use of labeled source data; weights 1 Full UMAP hyperparameters:

5.2.2 Giriama k’- Contraction (98.5% Consistency) Clustering identified Cluster 8 (206 words, 98.5% consistency) with k’- (apostrophe = elision). Transfer learning did not detect this pattern; no Swahili equivalent exists: 2 Python 3.10, PyTorch 2.0, Transformers 4.30, UMAP 0.5, scikit-learn

n_neighbors = 15, min_dist = 0.1;

K-means random_state=42.

1.3.

3

Task

Accuracy

Lemmatization Inflection (completion)

78.2% (347/444) 54.3% (241/444)

Table 1: BantuMorph v7 accuracy on 444 known Giriama verb paradigms (external evaluation). The model was not trained on these paradigms.

k’adzamuhala k’ahendzeze k’ululu

"he/she did not care" "he/she pleased" "freedom/liberty"

Probable interpretation: ku- → k’- infinitive contraction (fast speech) or Proto-Bantu narrative ka- → k’-. This pattern requires validation by Giriama linguists. 5.3 External Validation on Known Giriama Paradigms To address the absence of a gold standard for noun class discoveries, we evaluate the underlying BantuMorph model on 444 known Giriama verb paradigms (95 unique lemmas) from UniMorph, which were not used in the discovery pipeline. The 78.2% lemmatization accuracy demonstrates that the model has learned productive Giriama morphological patterns through cross-lingual transfer from related languages, validating the foundation on which the noun class discoveries are built. Of the 95 known Giriama lemmas, 25 (26%) appear in the transfer-based discoveries and 7 (7%) in the ensemble discoveries, confirming that the pipeline correctly identifies known vocabulary while extending coverage to previously undocumented forms.

Language

Zone

Paradigms

Consist.%

Lemmas

Swahili Luganda Shona Chichewa Giriama Lingala Kirundi Zulu Kisukuma Kinyarwanda Xhosa Kimeru Kamba Kongo Kikuyu N. Sotho

G J S N E C J S F J S E E H E S

1,276 1,036 1,009 983 854 826 749 728 685 682 629 605 526 504 454 377

42.4 29.2 31.8 28.5 18.3 14.6 19.8 43.8 15.0 17.7 17.6 12.9 10.1 38.7 7.0 12.7

919 796 782 753 653 651 523 555 571 576 447 503 416 455 382 338

Total

11,923

24.6

9,320

Table 2: Discovery results across 16 Bantu languages using the full transfer+clustering+ensemble pipeline. Consist.% = internal consistency (generated form matches corpus form).

We measure internal consistency as the percentage of discovered words for which the model can regenerate the exact surface form from the predicted morphological features—a strict metric that penalizes any character-level deviation. Table 2 shows results across all 16 languages: 11,923 validated paradigms from 9,320 unique lemmas—a 130-fold increase over combined UniMorph entries. Languages with >200 training paradigms achieve >20% internal consistency; those below achieve <15%, suggesting a ∼200-paradigm minimum for effective transfer.

Expanded corpus analysis (v3). Monolingual corpus extraction extends coverage from 444 paradigms to 19,624 words (9,014 unique lemmas), achieving 97.3% segmentation and 86.7% lemmatization rates. The part-of-speech distribution—84.9% verbs (16,665), 10.5% nouns (2,066), 3.1% adjectives (618), 1.4% possessives (273)—demonstrates that BantuMorph generalizes across all major Giriama word classes, not only the nominal system targeted by the discovery pipeline. Among the 2,066 verified nouns, 9 noun classes are attested, with BANTU7 (ki-/chi-, 559 nouns) as the most productive, followed by BANTU14 (u-/bu-, 371), BANTU9 (N-, 262), and BANTU6 (ma-, 261). The high productivity of Class 7 in Giriama—surpassing Class 6—contrasts with the Swahili source data and supports the complementary value of languagespecific corpus analysis.

Language-specific innovations. Beyond Giriama, clustering discovers productive patterns absent from Swahili (and therefore invisible to transfer): • Luganda (J-zone): oku- infinitive prefix (1,107 instances; e.g., okulindiriza “to wait”, okusasulwa “to be paid”)—distinct from Swahili ku-. • Shona (S-zone): zvi- Class 8 plural prefix (846; e.g., zvinema “cinemas”)—the S-zone reflex of Proto-Bantu *bi-, distinct from Swahili vi-. • Kisukuma (F-zone): ng’- nasal prefix with elision (402; e.g., ng’wigulu “in heaven”)—F-zone specific. • Kinyarwanda (J-zone): y’i- contracted possessive (245; e.g., y’ikiyaga “of the lake”)—elision unique to Kinyarwanda.

5.4 Multi-Language Scaling

6. Discussion

We applied the same three-method pipeline (transfer + clustering + ensemble) to all 16 Bantu languages. For each language, transfer learning uses Swahili as the primary source (highestresource language in the family); for J-zone languages, Kinyarwanda also serves as a transfer source due to higher lexical overlap. Clustering operates independently per language on the unlabeled corpus.

6.1 Linguistic Significance Our discoveries contribute to Giriama documentation: 2,455 noun class labels (vs. 91 previously), two novel patterns (a-, k’-) not in existing literature. Theoretical implications: • a- variant confirms vowel coalescence in coastal Bantu 4

• k’- pattern suggests productive contraction process • Class 6 (ma-) productivity higher than Swahili (50.6% vs. ∼30%) • Expanded corpus (19,624 words) reveals Class 7 (ki-/chi) as the most productive noun class in Giriama (559/2,066 nouns), surpassing Class 6—a divergence from Swahili that merits further typological study

54th Annual Meeting of the Association for Computational Linguistics (ACL), pages 1954–1964, 2016. A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, et al. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 8440–8451, 2020. R. Cotterell, C. Kirov, J. Sylak-Glassman, D. Yarowsky, J. Eisner, and M. Hulden. CoNLL-SIGMORPHON 2017 shared task: Universal morphological reinflection in 52 languages. In Proceedings of the CoNLL–SIGMORPHON 2017 Shared Task, pages 1–30, 2017.

6.2 Methodological Insights Why 36.7% agreement? Transfer (cognates) and clustering (innovations) have complementary strengths. Genuine ambiguity exists in the language. Class imbalance (Class 6 dominates) skews single-method predictions. Value of low agreement: Disagreements reveal languagespecific features (clustering finds a-, k’-), ambiguous cases needing context, and errors for manual correction.

R. Cotterell, C. Kirov, M. Hulden, D. Yarowsky, et al. The CoNLL–SIGMORPHON 2018 shared task: Universal morphological reinflection. In Proceedings of the CoNLL– SIGMORPHON 2018 Shared Task, pages 1–27, 2018. M. Creutz and K. Lagus. Unsupervised models for morpheme segmentation and morphology learning. ACM Transactions on Speech and Language Processing, 4(1):1–34, 2007.

6.3 Error Analysis Transfer learning errors: False cognates (loanwords), sound changes (missed th/s, k’/ku correspondences), class shifts. Clustering errors: Low-consistency clusters (mixed patterns), ambiguous prefixes (mu-: Class 1 or 3?), loanwords not following native morphology.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), pages 4171–4186, 2019.

6.4 Limitations

J. Goldsmith. Unsupervised learning of the morphology of a natural language. Computational Linguistics, 27(2):153– 198, 2001.

Coverage is limited to nouns in the corpus; rare classes are underrepresented (Class 11: 6 words, Class 16: 4 words). Quality lacks a gold standard for full validation. Generalization requires a related high-resource language.

H. Hammarström and L. Borin. Unsupervised learning of morphology. Computational Linguistics, 37(2):309–350, 2011.

7. Conclusion We presented a method for zero-shot morphological discovery combining cross-lingual transfer and unsupervised clustering. Applied to Giriama (91 training paradigms): • 2,455 noun class labels discovered (27× increase) • Two novel patterns: a- prefix variant (95.1% consistency) and k’- contraction (98.5% consistency) • External validation: 78.2% lemmatization accuracy on 444 known paradigms; v3 corpus expansion to 19,624 words confirms generalization across verbs, nouns, adjectives, and possessives (97.3% segmentation, 86.7% lemmatization) The method’s key strength is complementarity: transfer learning identifies cognates shared with Swahili while unsupervised clustering discovers Giriama-specific innovations invisible to transfer. Applied at scale to 16 Bantu languages, the pipeline discovers 11,923 paradigms. We note that the discovered labels are silver-standard (model-generated, not human-verified) and recommend linguist validation before use in language documentation.

A. Hurskainen. A two-level computer formalism for the analysis of bantu morphology. Nordic Journal of African Studies, 1(1):87–119, 1992. K. Kann, R. Cotterell, and H. Schütze. Neural multi-source morphological reinflection. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 514–524, 2017. Y. Kim, Y. Jernite, D. Sontag, and A. M. Rush. Characteraware neural language models. In Proceedings of the 30th AAAI Conference on Artificial Intelligence, pages 2741– 2749, 2016. C. Kirov, R. Cotterell, J. Sylak-Glassman, G. Walther, E. Vylomova, et al. UniMorph 2.0: Universal morphology. In Proceedings of the 11th Language Resources and Evaluation Conference (LREC), pages 1868–1873, 2018. Lingua-Connect. English-giriama parallel sentence dataset. https://huggingface.co/datasets/ English-Giriama-Dataset, 2025.

References

J. F. Maho. A Comparative Study of Bantu Noun Classes. Orientalia et Africana Gothoburgensia, Gothenburg, 1999.

J. Buys and J. A. Botha. Cross-lingual morphological tagging for low-resource languages. In Proceedings of the

L. Marten and N. C. Kula. Object marking and morphosyntactic variation in Bantu. Southern African Linguistics and

5

Applied Language Studies, 30(2):237–253, 2012. A. D. McCarthy, M. Silfverberg, R. Cotterell, M. Hulden, and D. Yarowsky. Marrying universal dependencies and universal morphology. In Proceedings of the Second Workshop on Universal Dependencies, pages 91–101, 2019. H. Mutisya and J. Mugane. Cross-lingual morphological learning with character-level transformers: Evidence from 16 Bantu languages. 2026. Under review. M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), pages 2227–2237, 2018. L. Pretorius and S. E. Bosch. Exploiting cross-linguistic similarities in Zulu and Xhosa computational morphology. In Proceedings of the EACL 2009 Workshop on Language Technologies for African Languages, pages 96–104, 2009. J. Sylak-Glassman, C. Kirov, M. Post, R. Que, and D. Yarowsky. A universal feature schema for rich morphological annotation and fine-grained cross-lingual partof-speech tagging. In International Workshop on Systems and Frameworks for Computational Morphology, pages 72–93, 2015. E. Vylomova, J. White, E. Salesky, S. J. Mielke, S. Wu, K. Gorman, et al. SIGMORPHON 2020 shared task 0: Typologically diverse morphological inflection. In Proceedings of the 17th Annual SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, pages 1–39, 2020. L. Xue, A. Barua, N. Constant, R. Al-Rfou, S. Narang, M. Kale, A. Roberts, and C. Raffel. ByT5: Towards a tokenfree future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10: 291–306, 2022.

6

Record · ID 134532 · SHA-256 bb3f97e6436a64b8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.