ConceptioArchivearXiv CS
arXiv CSopen access

A Pāninian Foundation for Indic Language Processing

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

A Pāṇinian Foundation for Indic Language Processing One Metagrammar for a Billion Voices: Benchmarks and Architecture

arXiv:2606.24172v1 [cs.CL] 23 Jun 2026

RITWIK BANERJEE) ∗ , Department of Computer Science, Stony Brook University, USA LAV R. VARSHNEY† , AI Innovation Institute, Stony Brook University, USA More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity. Through more than two millennia of convergence around Sanskrit, Indic languages came to share a morphosyntactic architecture formalized in Pāṇini’s grammar, the Aṣṭādhyāyī. This cuts across genealogical lines, uniting languages through a common framework. We argue that this Pāṇinian framework supplies a unifying computational architecture the field has lacked, and that benchmarks grounded explicitly in it would make Indic language systems more accurate, more data-efficient, and more transferable, effectively merging many apparently disparate and sparse Indic language resources into a single high-resource metalanguage bedrock. We propose a four-part benchmark suite to render this shared architecture explicit, measurable, and ready to be leveraged for practical applications. Moreover, we underscore the question it raises for interpretability research: whether neural models trained on these languages come to represent Pāṇini’s categories on their own. CCS Concepts: • Computing methodologies → Machine translation; Speech recognition; Phonology / morphology; Language resources; Discourse, dialogue and pragmatics; Information extraction; Natural language generation; Lexical semantics. Additional Key Words and Phrases: Indic languages, Natural Language Processing, Pāṇinian grammar, Sanskrit, cross-lingual transfer, low-resource languages, multilingual models, morphological analysis, dependency parsing, semantic role labeling, benchmarks and evaluation, computational linguistics

1 Introduction Over a billion people across South Asia and beyond communicate daily in Indic languages — Hindi, Bengali, Tamil, Telugu, Marathi, Santali, and dozens more — yet the computational infrastructure serving these speakers/writers remains critically underdeveloped relative to its global significance. As AI-powered tools for translation, accessibility, education, and information retrieval become central to modern life, the gap between what NLP can do for English or Chinese and what it can do for Indic languages carries real costs, both human and economic. Closing that gap is an important engineering challenge facing the computing community today. Progress, however, has been slow. The primary reason is not a lack of data or talent, but a structural problem in how the field has organized itself. The longstanding emphasis on genealogical divisions in the taxonomy of Indic languages has led to a fragmentation of computational approaches: separate morphological analyzers, parsers, and annotated datasets for each language family, or even each individual language. The result is redundant engineering effort, acute resource constraints, and benchmark infrastructure that cannot transfer across languages. The scale of this duplication is concrete: the Universal Dependencies [17] collection alone, as of release 2.18 in May ∗

Also with AI Innovation Institute, Stony Brook University. Also with Department of Electrical and Computer Engineering, Stony Brook University.

Authors’ Contact Information: Ritwik Banerjee, [email protected], Department of Computer Science, Stony Brook University, Stony Brook, NY, USA; Lav R. Varshney, [email protected], AI Innovation Institute, Stony Brook University, Stony Brook, NY, USA.

2

Ritwik Banerjee and Lav R. Varshney

2026 [58], maintains 24 separately annotated treebanks across 18 Indic languages1 , each built as an independent effort, and conversions between their annotation schemata are frequently lossy [49, 50]. Every new language effectively requires building from scratch. This is not a denial of the value or the tremendous engineering achievements of massively multilingual models [41]. But their design objective is breadth, not depth: covering hundreds of languages optimizes translation fluency and coverage, not the morphosyntactic and semantic structure that distinguishes Indic languages, and no current benchmark would reveal the difference. Such models do achieve cross-lingual transfer, but the units that carry it across Indic languages — the shared and borrowed vocabulary, the parallel inflectional patterns, and the subwords that encode them — are shared precisely because these languages are organized by a common Pāṇinian architecture2 , whether through inheritance, sustained contact, or independent convergence. The transfer is therefore already running on Pāṇinian rails; the models merely exploit them implicitly and incompletely, through incidental surface signals rather than the deeper structure reflected by those signals. Empirical evidence bears the fingerprints of this shared architecture: cross-lingual transfer is stronger within Indic languages than between Indic and non-Indic languages even when script differences are controlled for [3, 39], and morphological analyzers built on Pāṇinian principles transfer across language families rather than merely within them [19, 20]. Scale alone, however, quickly reaches a real ceiling. A fixed-capacity model spread across many languages dilutes per-language quality, with the gains gravitating towards high-resource languages [15], leaving most Indic languages underserved. Furthermore, generic subword tokenizers fragment morphologically rich Indic words rather than recovering their morphemes [11]. Making the shared substrate explicit converts this transfer into an interpretable and designed mechanism, so that the meaning of a lexeme can be transparently tracked as it is inflected and compounded across languages, rather than lost to subword fragmentation determined by ad hoc statistics. Naïve data sharing across Indic languages, with no Pāṇinian grounding, already increases tagging accuracy [46], suggesting that there is ample room for empirical improvements through explicit grounding. Since the transfer behavior of multilingual models already reveals that they implicitly learn a structural prior commonly held across languages, it is likely that giving Indic languages an interpretable grammar already known to be a common foundation will boost language comprehension across the board. The fragmented, language-by-language approach and the opacity of multilingual models err in opposite directions around the same fact: the structural unity of Indic languages is real — the former ignores it and pays in redundant effort, whereas the latter gains from it without ever recognizing its central role. That unity is not a modeling artifact but the structural commons observed consistently by native multilingual users, across the boundaries of Indic languages and genealogies: remarkable parallels in morphological patterns such as verb conjugation, case marking, and agglutination; syntactic structures including postpositions and participial relatives; and shared frameworks for expressing agency, causation, aspect, and evidentiality. Far from superficial lexical borrowings, these phenomena reflect deep architectural similarities engendered by over two millennia of linguistic convergence where Sanskrit has functioned as a formal intellectual “metalanguage” across South Asia and beyond. This is not an isolated observation. 1

universaldependencies.org Pāṇini was a Sanskrit grammarian from around the fifth century BCE; his Aṣṭādhyāyī (“Eight Chapters”) specifies the language in some 4,000 ordered rules built from a formal metalanguage and explicit rule-precedence mechanisms. The parallel between this rule system and modern formal grammars was drawn early in the history of modern computing: writing in the Communications of the ACM, Ingerman [24] proposed that Backus-Naur form be renamed “Pāṇini-Backus form”. The correspondence is at best cosmetic, however — the Pāṇinian rule formalism is in fact strictly more powerful than the context-free grammars described by BNF [47]. 2

A Pāṇinian Foundation for Indic Language Processing

3

In other linguistic traditions, Latin shaped the prestige registers and grammatical self-conception of English [10, 35]. In the case of Indic languages, however, the shared architecture runs considerably deeper: Tamil grammarians, working within their own tradition, arrived at structural conclusions that converge with Pāṇinian categories — a finding that is computationally significant precisely because it transcends genealogical boundaries. Dialectal and register variations, too, operate along predictable continua rather than discrete breaks: dialects typically differ in phonological realization or lexical choice while preserving fundamental morphosyntactic architecture. The widespread diglossia3 across Indic contexts — Sanskrit-Prakrit, literary and colloquial Tamil, Hindi and Khariboli — follows systematic alternations within shared structural constraints. The computational treatment of Indic languages has long suffered from a denial of these regularities. The result is a landscape in which the structural unity underlying these languages remains invisible to the tools meant to process them. We argue that Pāṇini’s grammatical framework, formalized in the Aṣṭādhyāyī [56], provides a unifying computational architecture the field needs; and that building benchmarks explicitly grounded in this framework will unlock a new generation of more capable, resource-efficient, and transferable Indic language processing systems. 2 The Pāṇinian Framework as Unifying Architecture The unifying architecture we propose is not a modern invention. Pāṇini’s Aṣṭādhyāyī, composed around 500 BCE, is one of the most sophisticated formal grammatical systems in human history — and for over two millennia, it functioned as the shared intellectual operating system of South Asian discourse, regardless of language. Philosophy, law, science, and aesthetics across the subcontinent were all conducted within its formal categories. Sanskrit was not merely one language among many; it was the metalanguage that supplied the ontological primitives, syntactic templates, and morphophonological regularities through which thought was organized and communicated across the entire region. A computing professional might find the following analogy useful: think of Marathi and Tamil as having arisen from distinct kernels — their genealogical origins differ — but with their higher-level design patterns, discourse semantics, and formal structures built to the specifications of Pāṇinian Sanskrit grammar. Just as software components built to a common interface specification remain interoperable regardless of their underlying implementation, Indic languages built on this shared specification remain structurally compatible at the level that matters most for computation: morphology, syntax, and semantic organization. This compatibility, rather than being a mere theory, is concretely manifested in texts such as Śaiva Siddhānta (சைவ சித்தா ந்தம்) that exist simultaneously in Sanskrit and Tamil literary traditions, with the underlying semantic and argumentative architecture intact across both. The analogy understates a property of Pāṇini’s system that speaks with unusual directness to modern computational practice. Any grammar formalism powerful enough to describe a natural language is, almost inevitably, powerful enough to overgenerate, i.e., to admit strings that the language does not. Penn and Kiparsky [47] show that Pāṇini’s formalism is no exception in its raw expressive power. What is exceptional, however, is the discipline imposed upon it. Through rule precedence and a layer of meta-conventions governing how rules compete and apply, the Aṣṭādhyāyī yields a single derivation for every grammatical Sanskrit sentence — disambiguation built into the architecture rather than bolted on afterward. 3

Diglossia is the coexistence of two varieties of the same language throughout a community. Often, one form is the literary dialect (the “high” register; e.g., Katharevousa, which is heavily influenced by classical Greek, and used in official communications), and the other is a common dialect of everyday usage (the “low” register, e.g., Demotic Greek, which is the standard vernacular).

4

Ritwik Banerjee and Lav R. Varshney

No generation system in the Chomskyan tradition has this property, and it is precisely this absence, Penn and Kiparsky [47] argue, that obliges modern NLP to reach for its heavy statistical and probabilistic machinery: the numerical methods are, in large part, a means of curbing the overgeneration these formalisms cannot restrain on their own. Seen this way, the marvel of the Aṣṭādhyāyī is not how many correct analyses it produces, but how many incorrect ones it entirely avoids. An architecture grounded in this tradition would inherit determinacy as a core design principle. It does not obviate the statistical components of modern NLP, but provides a reason to expect that explicit Pāṇinian structure carries information that modern systems recover only indirectly, and at a cost. Demonstrated advantages of transfer learning within Indic languages [3, 39] and existing morphological analyzers based on Pāṇinian principles [19, 20] support this picture directly. Models trained on Pāṇinian dependency labels show improved argument detection and semantic role labeling [42]. Highly domain-specific tasks, such as translation of technical lexicon, show improvements in zero-shot setting when trained on Sanskrit tokens [38]. These are not marginal gains, but strong evidence that the shared architecture is computationally real and exploitable. Furthermore, historical records support this picture independent of the computational rewards. Tamil grammarians did not merely borrow Sanskrit categories. Rather, to a great extent, they discovered the same underlying categories independently within Tamil itself: Tolkāppiyam, the oldest extant Tamil grammar text, describes sandhi and ideas similar to vibhakti; the 11th-century grammatical treatise Vīracōl ̲iyam incorporated kāraka analysis and presented older traditional components of Tamil grammar through a comparative lens; and Ilakkaṇakkottu argued that Sanskrit grammatical features not mentioned in Tamil grammar nevertheless occur in the Tamil language, and therefore belong in its grammar [2]. These are records of recognition, not imposition [1, 48]: two sophisticated grammatical traditions, arriving independently, and through interactions, at shared structural conclusions. That convergence is precisely what enables computational leveraging of the Pāṇinian framework across language families, not merely within a single genealogical tree (e.g., Karthika et al. [29] achieved improved tokenization by clustering Punjabi together with Dravidian languages); and it is what one would expect if the Pāṇinian framework captures genuine computational primitives rather than being impositions of cultural prestige alone. The relationship between genealogical and functional perspectives is worth exploring with precision, since the distinction matters for how we design computational systems. Academic linguistics rightly focuses on genealogical descent to reconstruct historical development — and it does not deny the structural commonalities we describe. But genealogical taxonomy prioritizes inheritance, whereas computational transfer depends on functional and architectural similarity. These are different questions, and they do not always have the same answer. Recent work on cross-lingual transfer in programming languages makes this point sharply: genealogical relationships were not the most predictive features for transfer learning success; structural and corpus-specific features were far more reliable [4]. The same principle applies here. The Pāṇinian framework captures precisely the functional and architectural commonalities — in morphology, syntax, and semantics — that genealogical classification leaves implicit, and that computational systems need made explicit. What this means practically is that the shared semantic infrastructure of Pāṇinian grammar enables richer, more comprehensive knowledge representations for modern Indic languages. The common syntactic templates enable transfer learning of foundational NLP processes — semantic role labeling, tokenization, semantic composition, clause and phrase representation — in ways that are both more efficient and more interpretable, because empirical methods can be grounded in an explicit, well-understood formal system. Dialectal and register variation, rather than requiring separate handling, can be treated as parametric variation within this shared framework: dialects differ

A Pāṇinian Foundation for Indic Language Processing

5

in phonological realization or lexical choice while preserving the underlying morphosyntactic architecture. This is the key insight that the fragmented, language-by-language approach has been missing; and it points directly toward a new generation of unified, transferable benchmarks for Indic language technologies. 3 Challenges in Computational Indic Language Processing There are four distinct but interacting facets of Indic languages that challenge the development of effective Indic language technologies. 3.1

Morphological complexity

Indic languages are morphologically rich in ways that standard NLP pipelines — designed primarily for English — are poorly equipped to handle. Words are not atomic units but structured compositions of roots and affixes encoding case, number, tense, mood, and aspect. Two phenomena are particularly demanding. Sandhi — the euphonic fusion of sounds at word boundaries — means that segmenting a sentence into meaningful units is itself a non-trivial task requiring linguistic knowledge before any downstream processing can begin. Samāsa, the compounding of multiple semantic units into a single surface form, means that a single word may compress what English would express as an entire phrase. These phenomena are pervasive in high-register communication and frequent even in everyday speech. Handling them correctly is a prerequisite for almost every NLP task, yet most current tools treat them as edge cases rather than central architectural concerns. 3.2

Diglossia, register variation, and code-mixing

Across Indic language communities, speakers routinely navigate multiple registers — literary and colloquial Tamil, Sanskrit-inflected Hindi, Khariboli — and switch between them fluidly depending on context. This diglossia is not noise, but a structural feature of how these languages function socially. Code-mixing, particularly between an Indic language and English, adds a further layer of complexity that is especially pronounced on social media and in urban speech. Current NLP models, trained predominantly on written standard varieties, struggle with all of these variations. The deeper problem is that existing tools treat each register or mixed variety as a separate data distribution requiring its own handling, when in fact these variations follow systematic patterns within shared structural constraints. The Hindi–Urdu pair makes this concrete. In their everyday spoken forms the two are neardialects of a single language, yet their formal registers diverge so sharply in vocabulary (Hindi drawing on Sanskrit, Urdu on Persian and Arabic) that the high varieties become mutually unintelligible [30]. What survives that divergence is the grammar: across the full register range the morphosyntactic scaffolding stays fixed, down to shared function morphemes such as the genitive -kī. The variation is large, but it is lexical and graphemic, layered over an invariant Indo-Aryan morphology and syntax. That an entire prestige vocabulary can be swapped out while the structure holds is the sharpest demonstration that what these varieties share is architectural — precisely the regularity that a register-by-register modeling approach leaves unexploited. 3.3

The nominal semantics mismatch

Perhaps the least appreciated challenge is a fundamental mismatch between the semantic assumptions embedded in standard NLP frameworks and the actual semantic organization of Indic languages. Contemporary NLP, shaped by its development on Western languages, treats nouns as static, discrete entities whose meaning is grounded in extra-linguistic reference. Indic languages,

6

Ritwik Banerjee and Lav R. Varshney

by contrast, encode a process-oriented ontology: nominal meaning is derivationally and conceptually rooted in verbal roots. A noun, in Sanskrit — and across Indic languages — is best understood not as a static label but as denoting a role within an implicit event structure. The word dhātu, for instance, means both “root” and “metal”, but its Pāṇinian definition translates as “that which supports/holds” — an action-grounded description. Standard benchmark tasks built on Englishderived semantic assumptions systematically fail to evaluate this dimension of meaning, leaving a critical gap in how we measure language comprehension for Indic languages. 3.4

Dataset scarcity and annotation incompatibility

Even where datasets exist, they are fragmented and mutually incompatible. The annotated resources that do exist were each developed for individual languages or for a single family, and even the attempts at a shared standard were never carried across the Indic family of languages. The Indian Language Corpora Initiative (ILCI) parallel corpus [25] is POS-annotated using the common BIS tagset for Indian languages [51]; the most linguistically ambitious effort, the multilayered Hindi/Urdu treebank built at IIIT-Hyderabad [8, 43] — which covers what is essentially one Indo-Aryan language in two scripts [30] — pairs a Pāṇinian kāraka dependency layer with a PropBank4 -style predicate-argument layer adapted from English. The common BIS tagset spans many languages, but only at the surface, while the treebank’s deeper, Pāṇinian-informed annotation never reaches beyond Hindi and Urdu. No resource carries the deep Pāṇinian annotations across the family boundary, and their schemata — a POS tagset, a kāraka dependency scheme, an English-derived semantic layer — do not align with one another or with the universal scheme discussed next. Universal Dependencies5 corpora exist for several Indic languages [50], but their universal scheme does not map onto the grammatical categories these languages actually use, making crosslingual benchmark transfer difficult in practice. Sanskrit NLP, meanwhile, has concentrated on segmentation and lemmatization, with little work connecting these tasks to morpheme-level semantics: projects such as the Sanskrit Sembank [21] assign WordNet synsets to words but do not encode the semantic roles — kāraka — that are central to how meaning is organized in these languages. The cumulative result is that virtually no existing benchmark is designed around the structural unity of Indic languages; nearly all follow evaluation frameworks built for English, measuring what is easy to measure rather than what matters most. 4 The State of the Art: Promising Signals, Fragmented Progress The research community has not been idle. Across morphology, syntax, and semantics, there are genuine advances in computational Indic language processing — and a consistent pattern within them: when Pāṇinian structure is explicitly exploited, performance improves. The problem is that these advances remain isolated. No one has connected them into a coherent, unified framework. The field has the ingredients; what it lacks is the architecture. Highly multilingual large language models (MLLMs), such as NLLB [41], are tremendous engineering achievements that offer rigorous benchmarks within their scope. But their scope is translation quality — fluency, adequacy, toxicity identification, etc.; not morphological, morphosyntactic, 4

The Proposition Bank, or PropBank, is a shallow and broad foundational natural language processing resource created by Palmer et al. [44], that took a practical approach to semantic representation by adding a layer of predicate-argument information, or semantic role labels, to the syntactic structures of the Penn Treebank. Essentially, it provides sentence-level annotations for “who did what to whom”. 5 Universal Dependencies — available at universaldependencies.org — is a collaborative, open project with more than 600 contributors who have built over 200 treebanks spanning more than 150 languages. It offers a unified scheme for annotating grammar (including lexical categories, morphological attributes, and syntactic relations) across diverse human languages.

A Pāṇinian Foundation for Indic Language Processing

7

or semantic role structures, cross-register performance, morpheme-level semantics, or dialectal robustness across intra-language variations. These models do not resolve the underlying problem of fragmented corpora, and thus, the field’s design limitations are continually inherited, not overcome. The wide coverage of MLLMs is an engineering objective different from deep comprehension. Without benchmarks designed to probe the underlying structural understanding of Indic languages, we cannot know how much these models have achieved, or how much more they may achieve when leveraging the universal substrate of Pāṇinian metagrammar. 4.1

Morphological analysis: strong foundations, limited reach

The most mature body of work addresses morphological segmentation. Word segmentation benchmarks built on sandhi and samāsa exist for Sanskrit [31], and recent models such as ByT5-Sanskrit [40] achieve state-of-the-art results on segmentation and lemmatization. The MorphTok benchmark [11] grounds morphological analysis explicitly in Pāṇinian grammar and has demonstrated downstream improvements in practice. Earlier work established similar gains in machine translation [5] and named entity recognition [45]. The signal is clear and consistent: Pāṇinian morphological grounding helps. Yet these benchmarks are almost entirely confined to Sanskrit. The extension to modern Indic languages, which share the same underlying morphological architecture, has not been done systematically. A foundation exists, but it has not been built upon. 4.2

Morphological tagging: cross-lingual gains left on the table

Annotated corpora for morphological tagging exist for several modern Indic languages, but they trade breadth against depth. The ILCI parallel corpus [25] spans several Indo-Aryan as well as Dravidian languages — Hindi, Bengali, Marathi, Tamil, Telugu, and others — but stops at the shallow syntactic level of part-of-speech tagging. The multi-layered treebank at IIIT-Hyderabad [8] annotates more, including gender, case, and number and a Pāṇinian kāraka dependency layer, but reaches only the near-identical Hindi/Urdu language pair. Pratyaya-Kosh [52] is itself grounded in Pāṇinian pratyaya analysis but covers only Sanskrit noun derivation. These are valuable resources; the limitation they share is not that they ignore Pāṇinian structure — a few embrace it — but that none provides unified Pāṇinian morphological annotation across the Indic languages. Breadth comes without depth, and depth without breadth, leaving the structural unity of Indic morphology fragmented across partial, mutually incompatible resources rather than exploited by a common scheme. That even a fraction of this unity is exploitable is already demonstrated by Pawar et al. [46], who trained a single multilingual model for morphosyntactic tagging, spanning both Indo-Aryan and Dravidian languages. That this was shown even with no explicit Pāṇinian grounding at all is both encouraging and sobering: if sharing data alone yields empirical gains, the advantage from a unified Pāṇinian annotation could be substantially larger. 4.3 Syntactic parsing A Pāṇinian schema exists, but remains isolated. The AnnCorra project [7] developed a dependency annotation schema explicitly grounded in Pāṇinian kāraka relations — one of the most direct implementations of classical grammar in computational form. The downstream benefits are real: incorporating kāraka over universal dependencies has shown immediate improvements in Indic-language question-answering (QA) [57]. Yet AnnCorra remains largely an isolated effort. Universal dependency corpora exist for some modern Indic languages [50], but their annotation schemata do not map to Sanskrit grammar, limiting cross-lingual transfer. An initial effort to build Pāṇinian universal dependencies was made by Tandon et al. [54], but has not been extended across multiple languages. The infrastructure for a unified multilingual Pāṇinian parser is within reach — the conceptual work has been done — but the execution remains incomplete.

8

4.4

Ritwik Banerjee and Lav R. Varshney

Morpheme semantics: almost entirely uncharted

The deepest gap is at the level of meaning. The Sanskrit Sembank [21] assigns WordNet synsets to Sanskrit words — a valuable resource — but does not encode kāraka roles or root-level meanings. State-of-the-art segmentation models like ByT5-Sanskrit correctly identify lemmas and morphological tags, but make no connection to what those morphemes mean. There is virtually no work in the NLP literature that explicitly models the semantics of individual roots and affixes in Indic languages — despite the fact that this morpheme-level semantic transparency is a defining feature of how these languages organize meaning. This is not a minor gap. It means that current systems can parse the surface form of an Indic sentence with reasonable accuracy while remaining blind to the semantic architecture that gives that sentence its meaning. No benchmark currently exists that would even reveal this blindness, let alone measure it.

4.5

Cross-lingual transfer

The signal is strong, but the infrastructure is lacking. We underscored an important finding in recent literature, where Bafna et al. [3] and Nag et al. [39] demonstrated that transfer across Indic languages yields markedly stronger performance than transfer from Indic to non-Indic languages. This is precisely the pattern one would predict if the Pāṇinian framework captures genuine computational structure shared across these languages. That latent structure, however, is not what today’s deployed systems actually exploit. The most recent benchmarks show transfer still riding on surface signal. For instance, CorIL [9], a parallel corpus and translation evaluation spanning both Indo-Aryan and Dravidian languages, finds a performance hierarchy organized by script. IndicTrans2, state-of-the-art for Brahmi-script languages, collapses on Perso-Arabic Sindhi with a near-zero BLEU score, whereas the massively multilingual NLLB [41] and BhashaVerse [37] score several times higher. The larger models close this gap through sheer breadth of script coverage rather than any structural insight. The deciding variable is a surface property (in this case, the writing system), not the grammar shared by the languages. Neither outcome reflects structural understanding; both track orthographic and lexical overlap. The starkest case is the one that should be easiest: Urdu is, grammatically, nearly identical to Hindi [30], yet a model relying on Indic-script lexical sharing cannot bridge the script divide. The shared morphosyntactic architecture, the very thing Hindi and Urdu hold in common, is precisely what these systems fail to utilize. On the other hand, Goyal and Huet [19] and Hellwig [20] demonstrated how morphological analysis based on Pāṇinian principles can work across language families. A handful of multilingual question-answer datasets exist as well [53]. But these remain isolated data points rather than a coordinated research program. The empirical evidence for the Pāṇinian hypothesis is accumulating — yet the benchmark infrastructure needed for its systematic testing, along with deliberate extensions, does not exist. There is a clear pattern across every dimension of Indic language processing, be it morphological analysis, syntactic parsing, semantic role labeling, or cross-lingual transfer. The same story repeats: isolated advances, consistent positive signals when Pāṇinian structure is exploited, and no unified framework connecting them. What the field needs is not more isolated experiments but a coordinated benchmark suite that makes the shared Pāṇinian architecture explicit, measurable, and exploitable across all Indic languages. The benefits of such a unified approach are measurable. Chang et al. [13] find that adding moderate amounts of multilingual data improves low-resource language modeling about as much as

A Pāṇinian Foundation for Indic Language Processing

9

enlarging the low-resource corpus itself by up to a third, and that the gain is driven by the syntactic similarity of the added data, with vocabulary overlap mattering only marginally6 . The advantage, thus, lies in the structural core, not the lexical surface, consistent with evidence that large models encode grammatical organization along directions shared across languages [12] and align cross-lingually without depending on shared vocabulary [16]. This is what makes the Indic case so favorable: Indic languages converge structurally across family lines with comparable morphological richness, through long contact, despite belonging to different genealogical families [28], and cross-lingual transfer tracks exactly this morphological and structural similarity [6]. For Indian languages, exploiting this relatedness is already established practice: a substantial body of lowresource machine translation and transliteration leverages the orthographic and lexical substrate these languages share [34]. A common Pāṇinian substrate would carry this deeper by making the shared morphosyntactic structure formal and explicit, so the additive advantage to low-resource languages, as demonstrated by Chang et al. [13], can operate across the entire group of Indic languages. Thereby, separately scarce Indic language resources become, at the level that actually drives language comprehension and multilingual transfer, a single large hub. 5 A Benchmark Suite for Indic Language Processing The gap analysis above reveals a consistent pattern: when Pāṇinian structure is explicitly leveraged, performance improves. But no coordinated benchmark infrastructure exists to make this leverage systematic and reproducible across languages. We propose four thematic benchmark clusters that together would constitute a unified Pāṇinian evaluation suite for Indic language processing. Each cluster is (a) directly motivated by empirical results, and (b) designed to be multilingual from the ground up rather than extended post hoc and ad hoc to other languages. 5.1

Morphological segmentation and tagging

The first cluster addresses the most foundational level of Indic language processing: the decomposition of words into their Pāṇinian roots and affixes, and the interpretation of what those morphemes mean. It provides the common skeleton to consolidate (a) multilingual morphological segmentation and tagging, (b) tasks on etymology and morpheme semantics, and (c) lexical semantics. Existing Sanskrit segmentation benchmarks [31, 40] and the MorphTok benchmark [11] provide an important starting point, but their coverage is almost entirely confined to Sanskrit. The first task extends these benchmarks to modern Indic languages, requiring models to split sentences into Pāṇinian roots (dhātu) and affixes (pratyaya and lakāra) and label their grammatical and semantic functions. The empirical case for this extension is already made: Pawar et al. [46] demonstrated a 7% gain in morphosyntactic tagging accuracy from sharing data across Indo-Aryan and Dravidian languages — without any explicit Pāṇinian grounding. Grounding the annotation explicitly in Pāṇinian categories should yield further gains (cf. Brahma et al. [11]), since producing large groundtruth corpora is demonstrably feasible: finite-state systems for segmentation and morphological analysis [22, 32] and rule-derivation engines that simulate the Aṣṭādhyāyī’s rules [36] can generate analyses at scale, and their outputs can be validated and normalized into tagged gold corpora [33]. The second task goes deeper: given an inflected word in a modern Indic language such as Marathi or Bengali, the task is to identify its Sanskrit dhātu and interpret its core meaning. This is morpheme-level semantic grounding, the testbed for the process-oriented nominal semantics that distinguishes Indic languages from the static noun-entity model assumed by standard NLP. 6

This is precisely what separates beneficial multilinguality from the drawbacks of relentless scaling — structurally aligned data adds signal, unrelated data dilutes it. This has been called the “curse of multilinguality” [15].

10

Ritwik Banerjee and Lav R. Varshney

Cross-lingual word-sense disambiguation and semantic textual similarity tasks can be built on top of this etymological grounding by using resources like the Sanskrit Sembank [21]. 5.2

Syntactic parsing via Pāṇinian dependencies

This benchmark pursues a single, high-impact goal: a unified multilingual dependency parser grounded in kāraka relations rather than universal dependencies. The conceptual groundwork exists — the AnnCorra project [7] developed a kāraka-based dependency schema, and Tandon et al. [54] made an initial effort toward Pāṇinian universal dependencies — but neither has been extended systematically across multiple languages or language families. The benchmark task here is both to build or convert dependency treebanks for multiple Indic languages using Pāṇinian dependency labels, and to evaluate parsers trained on this unified schema against language-specific baselines. A single robust multilingual parser, once available, can serve as infrastructure for virtually every downstream task across Indic languages. 5.3

Semantic role labeling and inference across registers and languages

This cluster addresses meaning at the sentence and discourse level. It consolidates semantic role comprehension across Indic languages and language families, and tests model performance across dialectal continua and register variations across Indic languages. The core benchmark annotates sentence pairs and QA examples across multiple Indic languages (Hindi, Marathi, Bengali, etc.) with kāraka roles as semantic frames, analogous to PropBank but grounded in Pāṇinian categories. Our proposal is to build directly on the kāraka-grounded work by Verma et al. [57], extending it from QA to entailment and inference. A distinct feature of these benchmarks is their explicit engagement with registers, as most Indic languages use multiple registers in both speech and text7 . Language models must capture the same underlying semantic content across these surface variations. Benchmark tasks would require models to, for example, answer questions about a narrative given in colloquial Tamil that is derived from a Sanskrit original, or to identify entailment relations across literary and spoken register pairs. These tasks directly test whether models have learned the underlying semantic architecture or merely memorized surface patterns. Dialectal robustness and code-mixing generalization. India exhibits an extremely rich set of local dialectal variations. Tamil, for instance, exhibits significant variation across Chennai, Madurai, and Sri Lanka, and Bengali forms a comparable continuum — from the southwestern districts of West Bengal through the northern belt around Cooch Behar to the eastern varieties of Tripura and Sylhet — in phonology as well as morphology [14]. These distinctions, much like the sharp register variations, share core morphosyntactic templates. Key benchmark tasks ought to include cross-dialectal semantic role labeling, where training and test sets use different dialects or registers; register adaptation, requiring models to translate between literary and colloquial forms while 7

The coexistence of formal and colloquial registers within a single language community (i.e., diglossia, cf. footnote 3) is a defining feature of modern Indic language use. Bengali, for instance, distinguishes sādhu bhāṣā, a highly Sanskritized written style, from cālitā bhāṣā, the colloquial form used in most modern communication, except possibly in official documents. The gap between these registers is wider than, say, the difference between formal and informal English; it is closer to the difference between Latin and Italian. This pattern recurs across Indic languages: Tamil maintains a sharp divide between centamil (literary) and koṭuntamil (spoken); Hindi distinguishes a Sanskritized formal register from the colloquial Khaṛiboli. The phenomenon is not unique to India — Arabic’s fuṣḥā versus ’āmmiyya, Japanese keigo (敬語) versus plain speech, or Chinese classical/literary wényánwén versus the vernacular báihuà, are structural parallels — but in the Indic context, register variation is especially significant for NLP because the formal register draws systematically on Sanskrit morphology and vocabulary, while the colloquial register reflects centuries of phonological and grammatical simplification.

A Pāṇinian Foundation for Indic Language Processing

11

preserving semantic content; and code-mixing generalization across sociolinguistic contexts ranging from formal written text to informal spoken interaction. These tasks measure something that current benchmarks almost never measure: whether a model has learned structural invariants that generalize across surface variation, or whether it has simply memorized the patterns of a particular variety — a distinction of enormous for real-world deployment across India’s linguistic diversity. 5.4

The extended Indic “sprachbund” and cross-lingual information disorder

This final cluster pushes the unified framework in two directions simultaneously: outward to the broader Indic linguistic sphere, and toward an urgent social concern. These benchmarks address what Emeneau [18] identified as the Indic sprachbund8 and what one might think of, following the “World Englishes” framework of Kachru [26, 27], as World Indic Languages. This extension is striated, however, as one moves outward, and the benchmarks must respect that gradient. At the first level are languages that remain Indo-Aryan and carry the full morphosyntactic substrate, however dispersed or divergent: diaspora varieties such as Caribbean Hindustani and Fiji Hindi, which preserve features lost in the modern standard; Romani, which retains Indo-Aryan morphosyntax despite centuries in Europe; and Sinhala and Dhivehi, the southernmost Indo-Aryan languages, which share the Pāṇinian substrate along distinct developmental paths. Here, the benchmark question is whether models trained on standard varieties generalize to these structurally Indic but divergent forms; and, conversely, whether the conservative among them, precisely because they preserve older features, can inform models about earlier stages of the shared Pāṇinian architecture. Then, there are languages connected not by structure but by civilizational influence. Tibetan, Burmese, and Thai belong to different families altogether (Sino-Tibetan and Kra-Dai), yet they were written in Brahmi-derived scripts, inherited the Sanskrit phonological taxonomy that orders those scripts, and absorbed extensive Sanskrit vocabulary. Herein lies a harder open question for the benchmark: how far can shared orthography, phonetic taxonomy9 , and lexicon — in the absence of shared morphosyntactic substrate — support transfer at all? Together, these benchmarks establish how far the structural unity genuinely extends, and where it gives way to contact alone. The applied direction addresses information disorder. Misinformation in India spreads rapidly across linguistic boundaries, mutating as it moves between language communities — a Hindi claim reappears in Punjabi with a subtle distortion, then in Bhojpuri or Bengali with another, each step compounding the original narrative while remaining semantically traceable to it. Detecting and tracking this propagation requires models that understand the shared semantic substrate across languages, which is precisely what the Pāṇinian framework provides. Benchmark datasets for cross-lingual information disorder would require models to identify semantically equivalent claims expressed across multiple Indic languages and registers, track the variations introduced as content crosses linguistic boundaries, and flag the inconsistencies that signal deliberate manipulation. For a computing community that has watched misinformation destabilize societies worldwide, building the infrastructure to detect it across one of the world’s most linguistically complex regions is not just an academic exercise, but an engineering priority. 6 Conclusion Nearly no existing benchmarks for Indic languages are designed around the structural foundation that actually unifies them. Most follow evaluation frameworks built for English, measuring what is convenient rather than what is linguistically meaningful. This article has argued that Pāṇini’s 8

A “sprachbund” is an area of linguistic convergence, corresponding to a group of languages with similarities in syntax, morphological structure, cultural vocabulary, and sound systems [55]. 9 The Pāṇinian/Śikṣā phonetic taxonomy is embedded in writing systems from Tibet to the farthest reaches of South-East Asia, including the Philippines.

12

Ritwik Banerjee and Lav R. Varshney

grammatical framework, formalized over two millennia ago and active as an intellectual infrastructure across South and South-East Asia ever since, provides precisely the unifying computational architecture the field has been missing. The architectural unity it captures is not a matter of genealogical inheritance alone: it extends even to Austroasiatic (e.g., Munda, whose speakers span India, Bangladesh, and Nepal) and Tibeto-Burman languages, where Sanskrit shaped not just the lexicon but the very phonological taxonomy that orders their scripts, and supplied the philosophical and technical vocabulary of their high registers. This unity demonstrates that over two millennia of cultural-linguistic synthesis have created structural commonalities that run deeper than genealogical trees, and offer a framework for rapid development of practical AI-driven tools. The practical payoff is substantial. A multilingual parser trained jointly on Sanskrit, Hindi, and Marathi and outputting Pāṇinian cases would be more accurate, more data-efficient, and more transferable than three separate parsers trained in isolation. A multilingual question-answering benchmark where answers must be validated against kāraka semantic roles would reveal capabilities and failure modes that remain invisible to current English-derived benchmarks. The four benchmark clusters we have proposed would provide the evaluation infrastructure for a new generation of tokenizers, morphological analyzers, dependency parsers, and semantically grounded language models — all built on a foundation that reflects how these languages actually work. There is also a deeper scientific question at stake, one that should interest the computing community beyond its immediate engineering implications. The Pāṇinian framework’s explicit formal structure — kāraka roles, dhātu roots, morphological composition rules — provides natural targets for mechanistic interpretability research. We can directly probe whether neural models trained on Indic languages learn internal representations that align with Pāṇinian categories, or whether they discover alternative organizational principles. This raises a question analogous to the Platonic representation hypothesis in vision models [23]: Do neural language models trained on Indic languages spontaneously learn internal representations that correspond to Pāṇinian categories, even without explicit supervision? If the answer is yes, it would mean that Pāṇini’s framework does not merely offer a convenient analytical lens, but it captures genuine computational primitives of Indic linguistic structure, and it is the natural and canonical way in which morphosyntactic information is organized in (and across) these languages. A framework conceived as a formal grammar would turn out to be a discovery about the deep structure of an entire sprachbund. This prospect also reframes a methodological default the field rarely questions: if models already gravitate toward Pāṇinian structure on their own, then leaning on statistical learning alone is less a first principle than a costly recovery mechanism from data. There is a formal disambiguating structure that an explicit grammatical algebra supplies by design. The most capable Indic language systems, then, are likely to come not from statistics alone, but from statistical models disciplined by such structure. Building the benchmarks to answer this question is, in itself, a contribution to our understanding of how both artificial and human minds process language. It is an invitation to the computing community to deeply engage with a rich and consequential linguistic tradition. References [1] Amitav Acharya. 2013. Civilizations in Embrace: The Spread of Ideas and the Transformation of Power : India and Southeast Asia in the Classical Age. Institute of Southeast Asian Studies, Singapore. [2] E. Annamalai. 2024. The Sanskrit Paradigm of Tamil Grammar: Embrace and Resistance. Bhasha 3, 1 (April 2024), 1–16. doi:10.30687/bhasha/2785-5953/2024/01/002

A Pāṇinian Foundation for Indic Language Processing

13

[3] Niyati Bafna, Cristina España-Bonet, Josef Van Genabith, Benoît Sagot, and Rachel Bawden. 2023. Cross-lingual Strategies for Low-resource Language Modeling: A Study on Five Indic Dialects. In Actes de CORIA-TALN 2023. Actes de la 30e Conférence sur le Traitement Automatique des Langues Naturelles (TALN), volume 1 : travaux de recherche originaux – articles longs, Christophe Servan and Anne Vilnat (Eds.). ATALA, Paris, France, 28–42. https: //aclanthology.org/2023.jeptalnrecital-long.3 [4] Razan Baltaji, Saurabh Pujar, Martin Hirzel, Louis Mandel, Luca Buratti, and Lav R. Varshney. 2025. Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study. Transactions on Machine Learning Research 2025, June (2025), 26 pages. https://openreview.net/forum?id=1PRBHKgQVM [5] Tamali Banerjee and Pushpak Bhattacharyya. 2018. Meaningless yet meaningful: Morphology grounded subwordlevel NMT. In Proceedings of the Second Workshop on Subword/Character LEvel Models, Manaal Faruqui, Hinrich Schütze, Isabel Trancoso, Yulia Tsvetkov, and Yadollah Yaghoobzadeh (Eds.). Association for Computational Linguistics, New Orleans, 55–60. doi:10.18653/v1/W18-1207 [6] Ajitesh Bankula and Praney Bankula. 2025. Cross-Linguistic Transfer in Multilingual NLP: The Role of Language Families and Morphology. arXiv:2505.13908 [7] Akshar Bharati, Rajeev Sangal, Vineet Chaitanya, Amba Kulkarni, Dipti Misra Sharma, and K.V. Ramakrishnamacharyulu. 2002. AnnCorra: Building Tree-banks in Indian Languages. In COLING-02: The 3rd Workshop on Asian Language Resources and International Standardization. Association for Computational Linguistics, Taipei, Taiwan, 8 pages. https://aclanthology.org/W02-1202 [8] Rajesh Bhatt, Bhuvana Narasimhan, Martha Palmer, Owen Rambow, Dipti Sharma, and Fei Xia. 2009. A MultiRepresentational and Multi-Layered Treebank for Hindi/Urdu. In Proceedings of the Third Linguistic Annotation Workshop (LAW III), Manfred Stede, Chu-Ren Huang, Nancy Ide, and Adam Meyers (Eds.). Association for Computational Linguistics, Suntec, Singapore, 186–189. https://aclanthology.org/W09-3036 [9] Soham Bhattacharjee, Mukund K. Roy, Yathish Poojary, Bhargav Dave, Mihir Raj, Vandan Mujadia, Baban Gain, Pruthwik Mishra, Arafat Ahsan, Parameswari Krishnamurthy, Ashwath Rao, Gurpreet Singh Josan, Preeti Dubey, Aadil Amin Kak, Anna Rao Kulkarni, Narendra V. G., Sunita Arora, Rakesh Balbantray, Prasenjit Majumdar, Karunesh K. Arora, Asif Ekbal, and Dipti Mishra Sharma. 2025. CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems. arXiv:2509.19941 [10] Norman Blake. 1996. A History of the English Language. New York University Press, New York. [11] Maharaj Brahma, N J Karthika, Atul Singh, Devaraj Adiga, Smruti Bhate, Ganesh Ramakrishnan, Rohit Saluja, and Maunendra Sankar Desarkar. 2025. MorphTok: Morphologically Grounded Tokenization for Indian Languages. arXiv:2504.10335 [12] Jannik Brinkmann, Chris Wendler, Christian Bartelt, and Aaron Mueller. 2025. Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Luis Chiruzzo, Alan Ritter, and Lu Wang (Eds.). Association for Computational Linguistics, Albuquerque, New Mexico, 6131–6150. doi:10.18653/v1/2025.naacl-long.312 [13] Tyler A. Chang, Catherine Arnett, Zhuowen Tu, and Benjamin K. Bergen. 2024. When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 4074–4096. doi:10.18653/v1/2024.emnlp-main.236 [14] Suniti Kumar Chatterji. 1926. The Origin and Development of the Bengali Language. Calcutta University Press, Calcutta, India. [15] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised Cross-lingual Representation Learning at Scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 8440–8451. doi:10.18653/v1/2020.acl-main.747 [16] Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Emerging Cross-lingual Structure in Pretrained Language Models. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Association for Computational Linguistics, Online, 6022–6034. doi:10.18653/v1/2020.acl-main.536 [17] Marie-Catherine de Marneffe, Christopher D. Manning, Joakim Nivre, and Daniel Zeman. 2021. Universal Dependencies. Computational Linguistics 47, 2 (July 2021), 255–308. doi:10.1162/coli_a_00402 [18] Murray B. Emeneau. 1956. India as a Linguistic Area. Language 32, 1 (1956), 3–16. doi:10.2307/410649 [19] Pawan Goyal and Gerard Huet. 2016. Design and analysis of a lean interface for Sanskrit corpus annotation. Journal of Language Modelling 4, 2 (2016), 145–182. doi:10.15398/jlm.v4i2.108

14

Ritwik Banerjee and Lav R. Varshney

[20] Oliver Hellwig. 2016. Improving the Morphological Analysis of Classical Sanskrit. In Proceedings of the 6th Workshop on South and Southeast Asian Natural Language Processing (WSSANLP2016), Dekai Wu and Pushpak Bhattacharyya (Eds.). The COLING 2016 Organizing Committee, Osaka, Japan, 142–151. https://aclanthology.org/W16-3715/ [21] Oliver Hellwig and Erica Biagetti. 2025. The Sanskrit Sembank. Language Resources and Evaluation 59 (2025), 3635– 3658. doi:10.1007/s10579-025-09852-1 [22] Gérard Huet. 2005. A Functional Toolkit for Morphological and Phonological Processing, Application to a Sanskrit Tagger. Journal of Functional Programming 15, 4 (2005), 573–614. doi:10.1017/S0956796804005416 [23] Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. 2024. Position: The Platonic Representation Hypothesis. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp (Eds.). PMLR, Vienna, Austria, 20617–20642. https://proceedings.mlr.press/v235/huh24a.html [24] Peter Zilahy Ingerman. 1967. “Pānini-Backus Form” suggested. Commun. ACM 10, 3 (March 1967), 137. doi:10.1145/ 363162.363165 [25] Girish Nath Jha. 2010. The TDIL Program and the Indian Langauge Corpora Intitiative (ILCI). In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, and Daniel Tapias (Eds.). European Language Resources Association (ELRA), Valletta, Malta, 982–985. https://aclanthology.org/L10-1602 [26] Braj B. Kachru. 1992. The other tongue: English across cultures (2 ed.). University of Illinois Press, Urbana, Illinois. [27] Braj B. Kachru. 1992. World Englishes: approaches, issues and resources. Language Teaching 25, 1 (1992), 1–14. doi:10.1017/S0261444800006583 [28] Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020. IndicNLPSuite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, Online, 4948–4961. doi:10. 18653/v1/2020.findings-emnlp.445 [29] N. J. Karthika, Maharaj Brahma, Rohit Saluja, Ganesh Ramakrishnan, and Maunendra Sankar Desarkar. 2025. Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights. arXiv:2506.17789 [cs.CL] [30] Robert D. King. 2006. The Poisonous Potency of Script: Hindi and Urdu. International Journal of the Sociology of Language 2001, 150 (2006), 43–59. doi:10.1515/ijsl.2001.035 [31] Amrith Krishna, Pavan Kumar Satuluri, and Pawan Goyal. 2017. A Dataset for Sanskrit Word Segmentation. In Proceedings of the Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, Beatrice Alex, Stefania Degaetano-Ortlieb, Anna Feldman, Anna Kazantseva, Nils Reiter, and Stan Szpakowicz (Eds.). Association for Computational Linguistics, Vancouver, Canada, 105–114. doi:10.18653/v1/W17-2214 [32] Sriram Krishnan and Amba Kulkarni. 2019. Sanskrit Segmentation revisited. In Proceedings of the 16th International Conference on Natural Language Processing, Dipti Misra Sharma and Pushpak Bhattacharya (Eds.). NLP Association of India, International Institute of Information Technology, Hyderabad, India, 105–114. https://aclanthology.org/2019. icon-1.12/ [33] Sriram Krishnan, Amba Kulkarni, and Gérard Huet. 2023. Validation and Normalization of DCS corpus and Development of the Sanskrit Heritage Engine’s Segmenter. In Proceedings of the Computational Sanskrit & Digital Humanities: Selected papers presented at the 18th World Sanskrit Conference, Amba Kulkarni and Oliver Hellwig (Eds.). Association for Computational Linguistics, Canberra, Australia (Online mode), 38–58. https://aclanthology.org/2023.wsc-csdh.3/ [34] Anoop Kunchukuttan and Pushpak Bhattacharyya. 2022. Machine Translation and Transliteration involving Related and Low-Resource Languages. CRC Press, Boca Raton, USA and Abingdon, UK. [35] David B. Lurie. 2023. The Vernacular in the World of Wen: Sheldon Pollock’s Model in East Asia? In Cosmopolitan and Vernacular in the World of Wen 文, Ross King (Ed.). Language, Writing and Literary Culture in the Sinographic Cosmopolis, Vol. 5. Brill, Leiden, The Netherlands, 49–68. doi:10.1163/9789004529441_003 [36] Anand Mishra. 2009. Simulating the Pāṇinian System of Sanskrit Grammar. In Sanskrit Computational Linguistics, Gérard Huet, Amba Kulkarni, and Peter Scharf (Eds.). Lecture Notes in Computer Science, Vol. 5402. Springer, Berlin, 127–138. doi:10.1007/978-3-642-00155-0_4 [37] Vandan Mujadia and Dipti Misra Sharma. 2025. BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages. arXiv:2412.04351 [38] Karthika N J, Krishnakant Bhatt, Ganesh Ramakrishnan, and Preethi Jyothi. 2025. LEVOS: Leveraging Vocabulary Overlap with Sanskrit to Generate Technical Lexicons in Indian Languages. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), Ekaterina Kochmar, Bashar Alhafni, Marie Bexte, Jill Burstein, Andrea Horbach, Ronja Laarmann-Quante, Anaïs Tack, Victoria Yaneva, and Zheng Yuan (Eds.). Association for Computational Linguistics, Vienna, Austria, 258–265. doi:10.18653/v1/2025.bea-1.20

A Pāṇinian Foundation for Indic Language Processing

15

[39] Arijit Nag, Bidisha Samanta, Animesh Mukherjee, Niloy Ganguly, and Soumen Chakrabarti. 2023. Transfer Learning for Low-Resource Multilingual Relation Classification. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 22, 2, Article 50 (March 2023), 24 pages. doi:10.1145/3554734 [40] Sebastian Nehrdich, Oliver Hellwig, and Kurt Keutzer. 2024. One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser AlOnaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 13742–13751. doi:10.18653/v1/2024.findings-emnlp.805 [41] NLLB Team. 2024. Scaling neural machine translation to 200 languages. Nature 630 (2024), 841–846. doi:10.1038/ s41586-024-07335-x [42] Riya Pal and Dipti Sharma. 2019. Towards Automated Semantic Role Labelling of Hindi-English Code-Mixed Tweets. In Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019), Wei Xu, Alan Ritter, Tim Baldwin, and Afshin Rahimi (Eds.). Association for Computational Linguistics, Hong Kong, China, 291–296. doi:10.18653/v1/D195538 [43] Martha Palmer, Rajesh Bhatt, Bhuvana Narasimhan, Owen Rambow, Dipti Misra Sharma, and Fei Xia. 2009. Hindi Syntax: Annotating Dependency, Lexical Predicate-Argument Structure, and Phrase Structure. In Proceedings of the 7th International Conference on Natural Language Processing (ICON). Macmillan Publishers, Hyderabad, India, 259– 268. [44] Martha Palmer, Paul Kingsbury, and Daniel Gildea. 2005. The Proposition Bank: An Annotated Corpus of Semantic Roles. Computational Linguistics 31, 1 (2005), 71–106. doi:10.1162/0891201053630264 [45] Priyaranjan Pattnayak, Hitesh Patel, and Amit Agarwal. 2025. Tokenization Matters: Improving Zero-Shot NER for Indic Languages. In 2025 IEEE International Conference on Electro Information Technology (eIT). IEEE, Valparaiso, Indiana, USA, 456–462. doi:10.1109/eIT64391.2025.11103625 [46] Siddhesh Pawar, Pushpak Bhattacharyya, and Partha Talukdar. 2023. Evaluating Cross Lingual Transfer for Morphological Analysis: a Case Study of Indian Languages. In Proceedings of the 20th SIGMORPHON workshop on Computational Research in Phonetics, Phonology, and Morphology, Garrett Nicolai, Eleanor Chodroff, Frederic Mailhot, and Çağrı Çöltekin (Eds.). Association for Computational Linguistics, Toronto, Canada, 14–26. doi:10.18653/v1/2023. sigmorphon-1.3 [47] Gerald Penn and Paul Kiparsky. 2012. On Pāṇini and the Generative Capacity of Contextualized Replacement Systems. In Proceedings of COLING 2012: Posters, Martin Kay and Christian Boitet (Eds.). The COLING 2012 Organizing Committee, Mumbai, India, 943–950. https://aclanthology.org/C12-2092 [48] Sheldon I. Pollock. 2000. Cosmopolitan and Vernacular in History. Public Culture 12, 3 (2000), 591–625. Project MUSE. https://muse.jhu.edu/article/26221 [49] Pooja Rai, Ayan Das, and Sanjay Chatterji. 2025. Mapping of the Nepali Dependency Treebank to Universal Dependencies. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 24, 11, Article 132 (Nov. 2025), 22 pages. doi:10.1145/3749643 [50] Vinit Ravishankar. 2017. A Universal Dependencies Treebank for Marathi. In Proceedings of the 16th International Workshop on Treebanks and Linguistic Theories, Jan Hajič (Ed.). Association for Computational Linguistics, Prague, Czech Republic, 190–200. https://aclanthology.org/W17-7623 [51] Baskaran Sankaran, Kalika Bali, Monojit Choudhury, Tanmoy Bhattacharya, Pushpak Bhattacharyya, Girish Nath Jha, S. Rajendran, K. Saravanan, L. Sobha, and K.V. Subbarao. 2008. A Common Parts-of-Speech Tagset Framework for Indian Languages. In Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, and Daniel Tapias (Eds.). European Language Resources Association (ELRA), Marrakech, Morocco, 1331–1337. https://aclanthology. org/L08-1544 [52] Arun Kumar Singh, Sushant Dave, Prathosh A. P., Brejesh Lall, and Shresth Mehta. 2020. A Benchmark Corpus and Neural Approach for Sanskrit Derivative Nouns Analysis. arXiv:2010.12937 [53] Abhishek Kumar Singh, Vishwajeet Kumar, Rudra Murthy, Jaydeep Sen, Ashish Mittal, and Ganesh Ramakrishnan. 2025. INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages. In Findings of the Association for Computational Linguistics: NAACL 2025, Luis Chiruzzo, Alan Ritter, and Lu Wang (Eds.). Association for Computational Linguistics, Albuquerque, New Mexico, 2607–2626. doi:10.18653/ v1/2025.findings-naacl.141 [54] Juhi Tandon, Himani Chaudhary, Riyaz Ahmad Bhat, and Dipti Misra Sharma. 2016. Conversion from Pāṇinian Karakas to Universal Dependencies for Hindi Dependency Treebank. In Proceedings of the 10th Linguistic Annotation Workshop held in conjunction with ACL 2016 (LAW-X 2016), Annemarie Friedrich and Katrin Tomanek (Eds.). Association for Computational Linguistics, Berlin, Germany, 141–150. doi:10.18653/v1/W16-1716 [55] Sarah Grey Thomason. 2000. Linguistic Areas and Language History. In Languages in Contact, D. G. Gilbers, J. Nerbonne, and J. Schaeken (Eds.). Studies in Slavic and General Linguistics, Vol. 28. Brill, Leiden, The Netherlands, 311– 327. doi:10.1163/9789004488472_030

16

Ritwik Banerjee and Lav R. Varshney

[56] Srisa Chandra Vasu. 1897. The Ashtādhyāyī of Pāṇini. Sindhu Charan Bose, Benares. [57] Devika Verma, Ramprasad S. Joshi, Aiman A. Shivani, and Rohan D. Gupta. 2023. Kāraka-Based Answer Retrieval for Question Answering in Indic Languages. In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Shoumen, Bulgaria, Varna, Bulgaria, 1216–1224. https://aclanthology.org/2023.ranlp-1.129 [58] Daniel Zeman, Joakim Nivre, Rimsha Abid, Mitchell Abrams, et al. 2026. Universal Dependencies 2.18. Institute of Formal and Applied Linguistics (ÚFAL), LINDAT/CLARIAH-CZ Digital Library. http://hdl.handle.net/11234/1-6149

Record · ID 303222 · SHA-256 05f12395fdff5b81
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.