Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data Hillary Mutisya Thiomi NLP
John Mugane Harvard University
arXiv:2604.22730v1 [cs.LG] 24 Apr 2026
Abstract We investigate whether neural models trained exclusively on modern morphological data can recover cross-lingual lexical structure consistent with historical reconstruction. Using BantuMorph v7, a transformer over Bantu morphological paradigms, we analyze 14 Eastern and Southern Bantu languages, extract encoder embeddings for their noun and verb lemmas, and identify 728 noun and 1,525 verb cognate candidates shared across 5+ languages. Evaluating these candidates against established historical resources—the Bantu Lexical Reconstructions database (BLR3; 4,786 reconstructed Proto-Bantu forms) and the ASJP basic vocabulary—we confirm 10 of the top 11 noun candidates (90.9%) align with previously reconstructed Proto-Bantu forms, including *-ntU ‘person’ (8 languages), *gombe ‘cow’ (9 languages), and *mUn (9 languages). Extending to verbs, 12 verb cognates align with reconstructed Proto-Bantu roots, including *-bon- ‘see’ and *-jÍm- ‘stand’, each attested across wide geographic ranges. Cross-model validation using an independent translation model (NLLB-600M) confirms these patterns: both models recover cognate clusters and phylogenetic groupings consistent with established Guthrie-zone classifications (p < 0.01). Cross-lingual noun class analysis reveals that all 13 productive classes maintain >0.83 cosine similarity across languages (within-class > between-class, p < 10−9 ). Our dataset is restricted to Eastern and Southern Bantu, so we interpret these results as recovering shared Bantu lexical structure consistent with Proto-Bantu rather than definitively distinguishing Proto-Bantu retentions from later regional innovations.
1. Introduction The Bantu language family, comprising 500+ languages spoken by over 300 million people across sub-Saharan Africa, shares a common ancestor known as Proto-Bantu, spoken approximately 4,500–4,000 years ago in the Cameroon grasslands [Nurse and Philippson, 2003]. The comparative method in historical linguistics reconstructs Proto-Bantu forms by identifying regular sound correspondences across daughter languages—a labor-intensive process spanning over a century of scholarship [Bastin et al., 2002, Guthrie, 1967]. This paper asks: can a neural model trained solely on modern morphological data recover cross-lingual lexical structure consistent with historical reconstruction? Using BantuMorph v7 [Mutisya and Mugane, 2026], a character-level transformer over Bantu morphological paradigms, we analyze 14 Eastern and Southern Bantu languages and show that the model’s encoder embeddings: 1. Cluster lexemes across languages in ways consistent with cognate structure, yielding 728 noun and 1,525 verb cognate candidates, most of which align with reconstructed Proto-Bantu forms (Section 4.1). 2. Recover historically stable lexical items including numerals and core vocabulary aligned with Proto-Bantu reconstructions. 3. Encode noun class structure that generalizes across languages (p < 10−9 ), with prefix evolution patterns matching known Proto-Bantu class correspondences.
4. Capture phylogenetic relationships consistent with Guthrie-zone classifications, including fine-grained Ezone sub-branching—confirmed independently by both BantuMorph and NLLB embeddings (p < 0.01). Importantly, our dataset is restricted to Eastern and Southern Bantu languages. As a result, our findings should be interpreted as recovering shared Bantu lexical structure consistent with Proto-Bantu, rather than definitively distinguishing Proto-Bantu retentions from later regional innovations.
2. Background 2.1 Proto-Bantu Reconstruction Proto-Bantu has been reconstructed through systematic comparison of daughter languages. The Bantu Lexical Reconstructions database (BLR3; Bastin et al. 2002) contains approximately 10,000 entries, and the Automated Similarity Judgement of Languages (ASJP; Wichmann et al. 2022) provides standardized 40-item basic vocabulary wordlists (core words like ‘water,’ ‘fire,’ ‘person’ that are resistant to borrowing and thus reliable indicators of shared ancestry) for computational comparison. The noun class system is Proto-Bantu’s most distinctive typological feature: 19 classes marked by prefixes (*mU-, *ba, *kI-, etc.) with systematic singular-plural pairings [Meeussen, 1967, Nurse and Philippson, 2003].
Table 1: Top noun cognate candidates discovered by embedding similarity, validated against BLR3 (4,786 reconstructions) and ASJP (40-item vocabulary).
2.2 Scope: Eastern and Southern Bantu Our study focuses on 14 languages drawn from Eastern and Southern Bantu (Guthrie zones C, E, G, H, J, N, S). Because these languages share a relatively recent common history, lexical similarities we observe may reflect:
Lemma Langs BLR3
• Proto-Bantu retentions (inherited from the common ancestor), • Proto-East-Bantu innovations (shared within the eastern branch), • Areal diffusion (contact-driven spread across neighboring languages). Distinguishing among these requires broader sampling, particularly from Western Bantu, which we identify as future work. We therefore interpret our findings as recovering shared Bantu lexical structure consistent with Proto-Bantu, rather than reconstructing Proto-Bantu directly.
ASJP Status
ng’ombe muno moko mutwe umuntu moi ngano mpaka masoko umwe
9 *gombe BLR3 9 *mUn BLR3 9 *mok BLR3 9 head ASJP 8 *-ntU person Both 7 *moi BLR3 7 *gano BLR3 7 *-dakaBLR3 6 *oko BLR3 6 one ASJP
mali
5
—
Table 2: Proto-Bantu numeral cognates recovered across languages. These are among the most stable lexical items in the Bantu family.
2.3 Neural Historical Linguistics
Lemma Meaning Langs Proto-Bantu
Prior work has extracted phylogenetic signal from mBERT encoder representations [Bjerva and Augenstein, 2021, Chi et al., 2020], but no study has demonstrated recovery of lexical cognates and proto-form correspondences from neural morphological models.
saba sita tatu kumi kenda moja tano tisa nne nane mbili
3. Method 3.1 BantuMorph v7 Embeddings We use BantuMorph v7 (ByT5-small, 300M parameters), pretrained on Bantu morphological paradigms across 16 languages [Mutisya and Mugane, 2026]. Of these, we select 14 Eastern and Southern Bantu languages for our analysis (Guthrie zones C, E, G, H, J, N, S); the remaining two languages are omitted because they lack the cross-lingual resources required for validation. Embeddings are extracted from the final encoder layer with mean pooling. To reveal cross-lingual structure, we apply language centering: subtracting the per-language mean embedding. Each language has a characteristic direction in embedding space caused by its writing system and token distribution. By subtracting the average embedding for each language, we remove this language-specific bias, leaving only the morphological and semantic information shared across languages.
‘seven’ ‘six’ ‘three’ ‘ten’ ‘nine’ ‘one’ ‘five’ ‘nine’ ‘four’ ‘eight’ ‘two’
12 *càmbà 11 — 10 *tàtù 9 *kùmì 9 *kèndà 8 — 8 *tànò 7 — 7 *nàì 6 *nànàì 5 *bìdì
dings for all cognate candidates. NLLB is an independent translation model with no morphological training objective. Cross-lingual similarity analysis—measuring how similar a word’s embeddings are across languages—identifies candidates with strong coherence. Results of this cross-model validation are reported in Section 7. 3.4 Validation We evaluate candidates against two established historical resources: • BLR3: 4,786 reconstructed Proto-Bantu forms extracted from the Wiktionary appendix derived from Bastin et al. [2002]. • ASJP: 40-item basic vocabulary wordlists for all 14 of our study languages [Wichmann et al., 2022].
3.2 Cognate Discovery Transfer learning identifies words from different languages that are nearest neighbors in the centered embedding space. We extract lemmas shared across 5+ languages, yielding 728 noun and 1,525 verb cognate candidates. The 11 highestconfidence noun candidates are selected for detailed BLR3 validation (Section 4.1), while 12 verb cognates are validated against established Proto-Bantu verb reconstructions (Section 4.2).
4. Results: Cognate Detection 4.1 Validated Noun Cognates Table 1 presents our top noun candidates with validation status. 10 of 11 (90.9%) align with at least one historical reference resource. We also identify numerals shared across languages (Table 2), consistent with the well-established reconstructability of the Proto-Bantu counting system.
3.3 Cross-Model Validation To verify that the cognate signal is not an artifact of BantuMorph’s training, we compute NLLB-600M encoder embed2
Table 3: Verb cognate candidates matching Proto-Bantu reconstructions. Langs = number of languages attesting the cognate (of 14).
Table 4: Noun class prefix forms across selected languages, with their Proto-Bantu ancestors. Proto-Bantu forms from Meeussen [1967].
Lemma Meaning Langs Proto-Bantu Note
CL PB
ona ima bona wa enda ba koma lala soma nywa tuma andika
1 *mù- m2 *bà- wa6 *mà- ma7 *kì- ki9 *nì- n15 *kù- ku-
‘see’ ‘stand’ ‘see’ ‘fall’ ‘go’ ‘be’ ‘strike’ ‘lie down’ ‘read’ ‘drink’ ‘send’ ‘write’
14 *-bon14 *-jÍm14 *-bon14 *-gu15 *-gend15 *-bà15 *-kóm14 *-làal13 13 12 *-túm10
variant
Swh Zul
Kik Nya Lug
umu- mũ- maba- aaama- ma- maisikĩ- chiinn- nuku- kũ- ku-
omuabaamaekienoku-
E. Bantu kon
run mer
swh
E. Bantu
kin lin kik
kam
nso lug nya sna zul
The highest-confidence matches include *gombe ‘cow’ (9 languages), *mUn (9 languages), and *mok (9 languages). The Bantu root *-ntU ‘person’—the word from which the family name ‘Bantu’ derives— appears in 8 of our languages. The numeral system is particularly striking: tatu ‘three’ (from *tàtù) appears in 10 languages spanning zones C through S, and kenda ‘nine’ (from *kèndà) in 9 languages—demonstrating that the model recovers some of the most ancient and stable elements of the Bantu lexicon.
xho
Zones E C J H S G N
Figure 1: MDS projection of BantuMorph embedding distances, colored by Guthrie zone. Same-zone languages cluster together: Ezone (kik, kam, mer), J-zone (kin, run, lug), S-zone (zul, xho, sna, nso). The E-zone forms the tightest cluster (mean pairwise similarity 0.990).
4.2 Proto-Bantu Verb Reconstructions From 1,525 verb cognate candidates (5+ languages each), 12 match established Proto-Bantu reconstructions [Guthrie, 1967]. Table 3 presents these validated verb cognates, ordered by breadth of attestation. The verb ona ‘see’ (Proto-Bantu *-bon-) appears in all 14 languages across zones C, E, G, H, J, N, and S—a span of over 3,000 km from Lingala (Central Africa) to Shona (Southern Africa). Its variant bona, retaining the initial consonant of the proto-form, also appears in all 14 languages. The verbs andika ‘write’ and soma ‘read’ lack BLR3 attestations, consistent with their status as East Bantu innovations for concepts post-dating Proto-Bantu.
Kirundi, Luganda) cluster in the center-top; and the S-zone languages (Zulu, Xhosa, Shona, N. Sotho) occupy the lower left. Same-zone cosine similarity (0.988) significantly exceeds cross-zone (0.977; p = 0.028, permutation test). Figure 2 shows the Ward-linkage dendrogram. The first merge is Kamba–Kikuyu (distance 0.004, similarity 0.996), confirming the close E50–E55 relationship. Kongo–Lingala merge second (distance 0.004), reflecting their shared Central Bantu heritage, followed by Kinyarwanda–Kirundi (distance 0.005), the closely related J60 pair. The E-zone sub-branch (Kamba, Kikuyu, Kimeru) forms a pure sub-tree before joining any other zone.
5. Results: Noun Class Prefix Evolution 7. Discussion
Analysis of noun class prefixes across our 14 languages reveals systematic correspondences matching Proto-Bantu reconstructions. All 13 productive classes maintain >0.83 cosine similarity in the centered embedding space, with within-class cross-lingual similarity significantly exceeding between-class (p = 4.6 × 10−9 ).
7.1 Why Modern Data Encodes Historical Structure Morphological systems are historically conservative: noun class prefixes, subject-verb agreement patterns, and derivational templates change slowly relative to the vocabulary that instantiates them. By learning these systems from modern data, neural models indirectly encode the historical relationships that produced them. For example, BantuMorph learns that Zulu abantu and Swahili watu (‘people’) have similar morphological roles because both are Class 2 plural human nouns with regular agreement patterns. This functional similarity is a consequence of shared ancestry—the model recovers historical signal as a byproduct of learning modern morphological structure.
6. Results: Phylogenetic Tree Recovery Ward-linkage clustering on the 14-language embedding similarity matrix recovers known family structure. Figure 1 shows the languages projected into two dimensions via multidimensional scaling (MDS) on the embedding distance matrix. Languages from the same Guthrie zone cluster together: the E-zone languages (Kikuyu, Kamba, Kimeru) form a tight group in the upper right; the J-zone languages (Kinyarwanda, 3
Distance
0.05
0.03
0.01 0
kin
run
mer
kam
kik
xho
nya
zul
lug
sna
kon
lin
nso
swh
Figure 2: Ward-linkage dendrogram from BantuMorph embedding distances. Label colors indicate Guthrie zone (see Fig. 1 legend). The first merges recover known sub-groupings: kam–kik (Kamba–Kikuyu, E50–E55), kon–lin (Kongo–Lingala, H–C), kin–run (Kinyarwanda–Kirundi, J60). The E-zone (mer, kam, kik) forms a pure sub-tree before joining the J-zone.
7.2 Interpreting “Recovery”
personal names like Bernard and Christine in 9 each. An automated system that equates cross-lingual sharing with cognate status would incorrectly identify these as Proto-Bantu vocabulary. They must be identified and excluded.
Our results should be interpreted as recovering cross-lingual lexical structure consistent with Proto-Bantu reconstruction, rather than reconstructing Proto-Bantu directly. Given that our dataset is restricted to Eastern and Southern Bantu, we cannot distinguish among three possible sources of observed similarity:
Adapted loanwords. More insidious than transparent foreign names are loanwords that have been phonologically and morphologically adapted into Bantu languages, making them resemble native vocabulary. Consider the word for ‘hospital’: Swahili sipitali, Kikuyu thibitarı̃, Kamba sivitalı̃, Kimeru cibitare, Kinyarwanda mubitaro, Chichewa chipatala, Shona chipatara, Zulu isibhedlela, Xhosa isibhedlela, N. Sotho sepetlele. Each language has adapted the borrowed form using its own noun class prefixes and sound-combination rules, making the words appear structurally Bantu. Similar patterns hold for church (Swahili kanisa, Kamba kanisya, Kikuyu kanitha, Luganda ekkanisa, but Zulu isonto, Xhosa icawe—showing divergent borrowing sources) and school (Swahili shule, Kikuyu thukuru, Kinyarwanda ishuri, Zulu isikole, Xhosa isikolo). These adapted loans can cluster with genuine cognates in embedding space because they share both semantic content and Bantu morphological framing.
• Proto-Bantu retentions inherited from the common ancestor, • Proto-East-Bantu innovations shared within the eastern branch, • Areal diffusion from sustained contact among neighbors. Resolving this distinction requires the comparative method and broader language sampling (particularly Western Bantu), which we leave to future work. Of our top 11 noun candidates, only mali remains unvalidated; it may represent an Arabic borrowing common to East African trade languages, a ProtoEast-Bantu innovation absent from BLR3, or areal diffusion— a distinction our method cannot make. 7.3 Cognate Detection vs. Comparative Method Embedding similarity identifies candidates at scale but does not replace phonological reconstruction or sound correspondence analysis. Our method generates 728 noun and 1,525 verb cognate candidates across 14 languages without requiring expert linguistic knowledge, achieving 90.9% alignment with BLR3 among top noun candidates. We therefore position this approach as a screening mechanism: generating high-quality cognate candidates for expert evaluation, complementing rather than replacing traditional comparative methods.
Domain bias in training corpora. Our training data draws from sources that over-represent certain semantic domains, particularly news and religious texts. Concepts prominent in these domains—government (Swahili serikali, Kikuyu gı̃thirikari, Kinyarwanda guverinoma, Luganda leta, Zulu uhulumeni), minister, bible (Swahili bibilia, Shona bhaibheri, Kongo bibiliya), Jesus (Swahili Yesu, Zulu uJesu, Kongo Yésu, Lingala Yésu)—appear in many languages not because they are inherited vocabulary but because the training texts share topical coverage. This creates a risk of under-discovery: the method may preferentially surface domain-frequent vocabulary while missing genuine cognates that happen to be rare in the available corpora. A word like ng’ombe ‘cow’ (ProtoBantu *gombe) appears in our data because cattle terminology is common in both news and everyday language, but equally ancient terms for concepts under-represented in these domains may be missed.
7.4 Filtering Modern Artifacts Modern corpora introduce challenges that any embeddingbased approach must address. Foreign proper nouns. Modern texts contain proper nouns from other languages—place names, personal names, and organizations—that appear across multiple Bantu languages in near-identical form. For example, Madagascar appears in 8 of our languages (Swahili, Kikuyu, Kamba, Kinyarwanda, Kongo, Lingala, Kimeru, N. Sotho); Argentina in 10; and 4
Mitigation. We address these challenges through a combination of approaches: proper nouns and transparent loanwords are identified via POS classification and translation analysis; adapted loanwords are detected by cross-lingual embedding coherence patterns (loanwords from a common source show uniformly high similarity, while genuine cognates show the graded divergence expected from regular sound change); and domain bias is partially mitigated by validating against domain-independent reference resources (BLR3 and ASJP basic vocabulary). The 90.9% alignment rate of our top candidates suggests these mitigations are effective, though we acknowledge that recall—the proportion of true Proto-Bantu cognates our method recovers—remains difficult to estimate.
able computational tool that generates high-quality candidates for expert evaluation.
7.5 Cross-Model Agreement
Ethan A. Chi, John Hewitt, and Christopher D. Manning. Finding Universal Grammatical Relations in Multilingual BERT. In Proceedings of ACL 2020, pages 5564–5577, 2020.
References Yvonne Bastin, André Coupez, Evariste Mumba, and Thilo C. Schadeberg. Bantu Lexical Reconstructions 3, 2002. Royal Museum for Central Africa, Tervuren. Online: https://www.africamuseum.be/ en/research/discover/human_sciences/ culture_society/blr. Johannes Bjerva and Isabelle Augenstein. Does Typological Information Transfer Cross-Lingually? In Proceedings of NAACL-HLT 2021, 2021.
The convergence of two independent models—BantuMorph (morphological) and NLLB (translational)—on the same cognate candidates and phylogenetic structure provides strong evidence that the recovered structure reflects genuine linguistic relationships rather than artifacts of any single training objective. Both models independently recover zone clustering (p < 0.01, permutation test), and NLLB’s cross-lingual similarity analysis identifies 594 strong cognates from among our candidates—an independent confirmation from a model with no explicit morphological training.
Malcolm Guthrie. Comparative Bantu: An Introduction to the Comparative Linguistics and Prehistory of the Bantu Languages. Gregg Press, 1967. Achille E. Meeussen. Bantu Grammatical Reconstructions. Africana Linguistica, 3:79–121, 1967. Hillary Mutisya and John Mugane. Cross-Lingual Morphological Learning with Character-Level Transformers: Evidence from 16 Bantu Languages. 2026. Under review. Derek Nurse and Gérard Philippson, editors. The Bantu Languages. Routledge, 2003.
8. Limitations • Automatically generated noun class labels (not humanverified) • 14 eastern/southern languages—cannot distinguish ProtoBantu from Proto-East-Bantu without western data • Character-level, not phonological—does not capture systematic sound correspondences • BLR3 matching uses substring comparison; formal cognate coding requires expert judgment
Søren Wichmann, Cecil H. Brown, and Eric W. Holman. The ASJP Database. https://asjp.clld.org/, 2022.
9. Conclusion A character-level transformer trained on modern morphological data from 14 Bantu languages recovers cross-lingual lexical structure consistent with historical reconstruction: from 728 noun and 1,525 verb cognate candidates, 10 of the top 11 noun candidates (90.9%) and 12 verb cognates align with reconstructed Proto-Bantu forms, including widely attested verbs such as *-bon- ‘see’ and *-jÍm- ‘stand’, and Proto-Bantu numerals like *tàtù ‘three’ (10 languages) and *kèndà ‘nine’ (9 languages). An independent translation model (NLLB) confirms the signal, with both models recovering Guthrie-zone clustering (p < 0.01). The model’s noun class embeddings encode historically meaningful cross-lingual correspondences (p < 10−9 ), and its phylogenetic trees recover fine-grained zone sub-structure. Because our dataset is restricted to Eastern and Southern Bantu, we interpret these results as recovering shared Bantu lexical structure consistent with Proto-Bantu rather than reconstructing Proto-Bantu directly. This approach does not replace the comparative method; it introduces a scal5
Appendix A
Reference Phylogeny 14 Bantu Languages
E-zone
kik, kam
mer
zul, xho
G, H, C, N
J-zone
S-zone
nso, sna
kin, run
lug
swh
kon, lin
nya
Figure 3: Reference language family tree (Glottolog classification) for the 14 Bantu languages in our study, grouped by Guthrie zone. Our embedding-derived dendrogram (Figure 2) recovers this zone-level structure, including the E-branch (kik, kam, mer) as a pure sub-tree.
6
B
Cross-Lingual Cognate Candidates
Tables 5–7 list cognate candidates discovered by embedding similarity across 10 or more of our 14 languages. Each lemma appears as a shared root form; the “N” column gives the number of languages in which the lemma was found. These candidates are drawn from the loanword-filtered vocabulary (proper nouns and confirmed borrowings excluded). We include this list for expert review—not all entries may represent genuine Proto-Bantu cognates, but the breadth of attestation makes them strong candidates for further investigation.
7
Table 5: Cognate candidates attested in 12–14 languages. Glosses are approximate (NLLB machine translation). Lemma POS N Gloss ona ima koma enda goma nyama kana mana nya lala manya wanda ina gwa bona dia nama nene tama
Languages
verb 14 ‘see’ All 14 languages verb 14 ‘stand’ All 14 languages verb 14 ‘strike, beat’ All 14 languages verb 14 ‘go, walk’ All 14 languages verb 14 ‘drum, dance’ All 14 languages noun 14 ‘meat, animal’ All 14 languages verb 14 ‘refuse, deny’ All 14 languages verb 13 ‘finish, end’ All except run verb 13 ‘drink, eat’ All except nya verb 13 ‘sleep, lie down’ All except sna verb 13 ‘know, learn’ All except run verb 13 ‘increase, be many’ All except zul verb 13 ‘have, sing’ All except run verb 13 ‘fall’ All except zul verb 12 ‘see’ kam, kik, kin, kon, lin, lug, mer, nso, nya, run, swh, xho, zul verb 12 ‘eat’ kam, kik, kin, kon, lin, lug, mer, nso, nya, sna, swh, zul noun 12 ‘meat, flesh’ kam, kin, kon, lin, lug, mer, nso, nya, run, sna, swh, xho noun 12 ‘big, fat’ kam, kik, kin, kon, lin, lug, mer, nso, nya, sna, swh, xho verb 12 ‘desire, want’ kam, kik, kin, kon, lin, lug, mer, nso, nya, run, sna, swh
Table 6: Cognate candidates attested in 10–12 languages. Lemma POS N Gloss tunga ngoma saba seka menya nywa kora sola
Languages
verb 12 ‘build, compose’ kam, kik, kin, kon, lin, lug, mer, nya, run, sna, swh, xho noun 11 ‘drum, song’ kam, kik, kin, kon, lin, lug, mer, nya, sna, swh, xho verb 10 ‘seven, cross’ kik, kin, kon, lin, lug, mer, nso, run, swh, zul verb 11 ‘laugh, smile’ kam, kin, kon, lin, lug, nso, nya, sna, swh, xho, zul verb 12 ‘know’ kam, kik, kin, kon, lin, lug, mer, nya, run, sna, swh, xho verb 12 ‘drink’ kam, kik, kin, lin, lug, mer, nso, run, sna, swh, xho, zul verb 12 ‘work, do’ kam, kik, kin, kon, lin, lug, mer, nso, run, sna, swh, zul verb 12 ‘choose’ kam, kik, kin, kon, lin, lug, mer, nso, nya, sna, swh, zul
Table 7: Cognate candidates attested in 10–11 languages. Lemma POS N Gloss andika mama tuma linda nena tenda genda teka muka paka baba soma kamba
Languages
verb 11 ‘write’ kam, kik, kin, kon, lin, lug, mer, nya, run, swh, xho noun 11 ‘mother’ kam, kik, kin, kon, lin, lug, mer, sna, swh, xho, zul verb 11 ‘send’ kam, kik, kin, kon, lin, lug, mer, nso, nya, run, swh verb 11 ‘wait, guard’ kam, kik, kin, kon, lin, lug, mer, nso, swh, xho, zul verb 11 ‘speak, say’ kam, kik, kin, kon, lin, lug, mer, nya, swh, xho, zul verb 11 ‘do, act’ kam, kik, kin, kon, lin, lug, mer, nya, sna, swh, zul verb 11 ‘go, travel’ kam, kik, kin, kon, lin, lug, mer, nya, run, swh, xho verb 11 ‘take, fetch’ kam, kik, kin, kon, lin, lug, mer, nso, run, sna, swh noun 11 ‘woman, wife’ kam, kik, kin, kon, lin, lug, mer, run, sna, swh, zul verb 10 ‘cat, smear’ kin, kon, lin, lug, nso, nya, run, sna, swh, xho noun 10 ‘father, wing’ kik, kin, kon, lin, lug, mer, run, sna, swh, zul verb 10 ‘read, learn’ kam, kik, kin, lug, mer, nya, run, sna, swh, zul noun 10 ‘rope’ kam, kik, kin, kon, lin, lug, mer, nya, run, sna
8