IKS/ Part 2· Linguistics Chapter 5 of 14
Part 2 · Chapter 5

Linguistics

व्याकरणम्

Around the fifth century BCE, someone wrote down a complete generative description of a natural language in 3,983 rules — with a formal metalanguage, an abbreviation scheme, ordered rule application, and recursion. Nothing comparable was attempted anywhere else for another two thousand three hundred years.

This chapter takes the Aṣṭādhyāyī apart mechanism by mechanism, deriving a word by hand, and ends with why computational linguists keep returning to it.

1Why analyse language at allभाषाविचारः

The Indian grammatical tradition did not begin as curiosity about language. It began as a maintenance problem.

Chapter 2 set out the situation: a corpus that must be reproduced as exact sound, transmitted by memory, across centuries, by people whose own everyday speech was steadily drifting away from the language of the texts. Two things follow. You need to know precisely what the sounds are, and you need to know precisely how the forms are built — otherwise a reciter cannot tell a correct archaic form from an incorrect innovation.

What began as preservation turned into something much more ambitious. By Pāṇini's time the question was no longer “what does this text say?” but “what is the complete set of rules that generates every well-formed utterance of this language, and nothing else?” That is a modern question, asked and answered before Aristotle.

Why it matters again now

Every system that processes language mechanically — search, translation, dialogue, code generation — needs a way of relating surface strings to structure. The Pāṇinian tradition is the oldest sustained attempt at exactly that problem, and its solutions were designed under a constraint modern systems do not have: the whole description had to fit in a human memory. Extreme compression forced extreme regularity, and regularity is what machines want.

2The four skillsभाषाकौशलम्

The four language skills across sound and script Listening and speaking operate on sound; reading and writing operate on script; receptive skills take language in, productive skills put it out. RECEPTIVE — taking in PRODUCTIVE — putting out SOUND SCRIPT Listening śravaṇa — the primary channel for the entire Vedic corpus Speaking the skill śikṣā exists to standardise, sound by sound Reading a comparatively late arrival in this tradition Writing a record of speech, not the thing itself The upper half is where the Indian tradition put its analytical weight. Devanāgarī and its sister scripts are phonetic transcriptions of an analysis that was already complete before any of them existed — which is why the alphabet is sorted by articulation rather than by history.

3The sound inventoryवर्णमाला

Chapter 2 introduced the five places of articulation. Here is the full grid they generate — and it is a grid, not a list.

The Sanskrit sound inventory arranged as a grid Vowels by length, then five vargas of stops arranged by place of articulation down the rows and by voicing, aspiration and nasality across the columns, followed by semivowels and sibilants. VOWELS · SVARA अ आ इ ई उ ऊ ऋ ॠ ऌ ए ऐ ओ औ short / long pairsdiphthongs — always long STOPS · SPARŚA — five places × five manners VOICELESSunaspirated VOICELESSaspirated VOICEDunaspirated VOICEDaspirated NASAL ka-vargathroat · kaṇṭhya क ख ग घ ङ ka kha ga gha ṅa ca-vargapalate · tālavya च छ ज झ ञ ca cha ja jha ña ṭa-vargahard palate · mūrdhanya ट ठ ड ढ ण ṭa ṭha ḍa ḍha ṇa — retroflex ta-vargateeth · dantya त थ द ध न ta tha da dha na pa-vargalips · oṣṭhya प फ ब भ म pa pha ba bha ma SEMIVOWELS · ANTAḤSTHA य र ल व SIBILANTS + H · ŪṢMAN श ष स ह Read the stop block as a matrix. A row is a place of articulation; a column is a manner. Every cell is predicted by its coordinates — which is what lets a phonological rule be stated once, over a region of the table, instead of once per sound.
Beyond place and manner, the tradition also classifies by duration (short, long, protracted), nasality, pitch accent (udātta, anudātta, svarita) and breath effort (alpa-prāṇa against mahā-prāṇa). Each is an independent dimension, and rules may quantify over any of them.

4The Aṣṭādhyāyīअष्टाध्यायी

3,983sūtras (rules)
8chapters — hence the name
32quarters, four to a chapter
c. 6th–5th c. BCEPāṇini, of Śālātura

The work is not a beginner's textbook and was never meant as one. It is a complete formal specification, written for readers who already speak the language and need to know exactly what its rules are. Its own name for the language it describes is the reason we call it Sanskrit: संस्कृतम्, saṃskṛtam, “refined, put together properly” — refined by the very process the grammar sets out.

The Panini tradition and its principal commentaries Panini's Astadhyayi, Katyayana's varttikas, and Patanjali's Mahabhasya, together called the three sages. Pāṇini c. 500 BCE Aṣṭādhyāyī 3,983 rules. The culmination of an already long tradition — Pāṇini cites ten predecessors by name. Kātyāyana c. 4th c. BCE Vārttikas Critical notes: where a rule overgenerates, undergenerates, or needs a condition added. Patañjali c. 2nd c. BCE Mahābhāṣya “The great commentary”: adjudicates between Pāṇini and Kātyāyana, in lively dialogue form. Together the three are the muni-traya, the “three sages”, and the settled rule is that where they differ, the later authority prevails.

Six properties that make it unusual

Complete

The rules aim to generate the whole vocabulary and inflection of the language, with a small, explicitly listed set of exceptions.

Generative

Not a description of collected usage but a machine: apply the rules correctly and a valid form comes out, including forms never previously uttered.

Derivational

Every word has a derivation history — a specific sequence of rule applications from base to surface form.

Modular

Two components only: a base (verbal or nominal root), and one or more suffixes. Everything else is rule-driven adjustment.

Open

The vocabulary is not a closed list. New words are legitimate provided no rule is violated — a productive grammar, not an inventory.

Formal

It has its own metalanguage: technical terms, abbreviation conventions, and rules about how rules interact. This is the part that startles computer scientists.

5The fourteen sūtrasमाहेश्वरसूत्राणि

Everything rests on a preliminary list of fourteen short strings, in which every sound of the language appears exactly once, in a very carefully chosen order.

The fourteen Mahesvara-sutras Fourteen strings listing all the sounds of Sanskrit, each ending in a marker consonant which is not itself part of the inventory. THE SOUND SEQUENCE · EACH LINE ENDS IN A MARKER (it) THAT IS NOT ITSELF A SOUND OF THE LANGUAGE 1अ इ उण्a i u · ṆA 2ऋ ऌक्ṛ ḷ · KA 3ए ओङ्e o · ṄA 4ऐ औच्ai au · CA 5ह य व रट्ha ya va ra · ṬA 6लण्la · ṆA 7ञ म ङ ण नम्ña ma ṅa ṇa na · MA 8झ भञ्jha bha · ÑA 9घ ढ धष्gha ḍha dha · ṢA 10ज ब ग ड दश्ja ba ga ḍa da · ŚA 11ख फ छ ठ थ च ट तव्… VA 12क पय्ka pa · YA 13श ष सर्śa ṣa sa · RA 14हल्ha · LA Three things to notice 1 · Blue is vowels, green is consonants. The vowels come first, and they are grouped so that natural classes are contiguous. 2 · The red final letter of each line is a marker — an it — not a sound of the language. It exists only to be a label. 3 · ह appears twice (lines 5 and 14) and ण twice as a marker. The redundancy is deliberate: it buys shorter abbreviations, and it is the one place the ordering is visibly engineered.
Traditionally said to have been revealed to Pāṇini by Śiva — hence Māheśvara. The ordering is not phonetically natural, and that is the point: it is optimised for the abbreviation scheme in the next section.

6Pratyāhāra — the compression trickप्रत्याहारः

Here is the mechanism the whole grammar runs on. Take any sound in the list, take any marker after it, put them together: that two-letter name denotes every sound in between.

How a pratyahara names a set of sounds Four examples: ac names all vowels, hal names all consonants, ik names i u r l, and yan names ya va ra la. THE SEQUENCE, LAID FLAT अ इ उण् ऋ ऌक् ए ओङ् ऐ औच् ह य व रट् लण् ञ म ङ ण नम् झ भञ् घ ढ धष् ज ब ग ड दश् अच् · ac from अ (first sound) to the marker च् — every vowel in the language. हल् · hal from ह (line 5) to the marker ल् (line 14) — every consonant. Continues past this strip. इक् · ik from इ to the marker क् — the set { i, u, ṛ, ḷ }. The “simple vowels except a”. यण् · yaṇ from य to the marker ण् — the set { ya, va, ra, la }. The semivowels. Now look at ik and yaṇ side by side: { i, u, ṛ, ḷ } and { ya, va, ra, la }. Same size, same order, and each vowel paired with its own semivowel. The list was arranged so that this correspondence would fall out. That is not luck.
Roughly 280 usable abbreviations come out of fourteen lines. Each one names a phonologically natural class in two syllables — and because a rule can quantify over a class, one rule now does the work of dozens.
What this is, in modern terms

A pratyāhāra is an interval over a totally ordered set, named by its endpoints. The Māheśvara list is a linearisation of the phoneme inventory chosen so that every class the grammar needs to refer to happens to be a contiguous interval. Finding such an ordering is a genuine constraint-satisfaction problem, and there is a small research literature on whether Pāṇini's solution is optimal. The current answer is that it is very close.

7One rule, read properlyइको यणचि

Take a single sūtra and unpack it completely. Aṣṭādhyāyī 6.1.77 reads, in its entirety:

इको यणचि

iko yaṇ aci

Three words, two of which are pratyāhāras and one of which is a case ending doing grammatical work. Written out: in place of ik, yaṇ, when ac follows.

Applying the rule iko yan aci The three sets in the rule, the correspondence between them, and a worked example combining prati and ekam into pratyekam. इक् · ik — what is replaced इ उ ऋ ऌ i u ṛ ḷ the sixth case (“of ik”) marks what is replaced यण् · yaṇ — the replacement य् व् र् ल् y v r l the first case marks the substitute अच् · ac — the condition any vowel following the seventh case marks what must follow Because ik and yaṇ are equal in length, the substitution is positional: the nth member of ik is replaced by the nth member of yaṇ. That convention — “corresponding substitution”, sthānivad — is itself stated as a separate rule, so it does not need repeating here. WORKED EXAMPLE प्रति + एकम् prati ends in इ (a member of ik); एकम् begins with ए (a member of ac) प्रत् + य् + एकम् → प्रत्येकम् pratyekam — “each one, one by one” इ is the first member of ik, so it becomes the first member of yaṇ, which is य्.
Four syllables of rule, covering sixteen distinct sound combinations, with the case endings carrying the grammatical roles — replaced (genitive), replacement (nominative), environment (locative). The Aṣṭādhyāyī uses Sanskrit's own case system as its metalanguage.

8The algorithmप्रक्रिया

Word formation in this system is a loop. Start from a base; scan the rules; apply whichever fires; repeat until nothing applies. What comes out is the word.

Flowchart of the Paninian derivation procedure Start, read the input, iterate over the rules, apply any applicable rule and restart the scan, otherwise continue; when the rules are exhausted print the result and stop. start read the base and the operation; set i = 0 is i > 3983 ? rules exhausted? YES print the final form no rule applies any longer stop NO read sūtra i does it apply? YES perform the operation …and restart the scan from i = 0, on the modified form NO i = i + 1
Simplified. Real rule application is governed by an intricate system of priorities — some rules block others, some are exceptions to a general rule stated earlier, and a rule can be “as if” a previous state still held. Patañjali's commentary is largely occupied with these interactions, and they are the reason a naïve implementation of this flowchart produces wrong forms.

9A word derived by handरामेण

Derive the instrumental singular of rāma — “by Rāma” — from the bare nominal stem, citing the rule that fires at each step.

Step by step derivation of ramena Five steps: the stem rama, addition of the suffix ta, its replacement by ina, guna sandhi giving rame plus na, retroflexion of n to n, and the final form ramena. STEP 1 राम rule 4.1.2 supplies the case suffix. To the stem, add ​ टा (ṭā) — the third-case singular affix. The ṭ is a marker, not a sound: it exists to let other rules refer to this affix. STEP 2 राम + टा rule 7.1.12 after a stem in a, ṭā is replaced by ​ इन (ina). A rule about a rule's output — this is where the marker ṭ earns its keep. STEP 3 राम + इन rule 6.1.87 guṇa sandhi: a followed by i becomes e. The final अ of राम plus the इ of इन give ए. The same rule handles a + u → o and a + ṛ → ar. One rule, three merges. STEP 4 रामे + न rule 8.4.2 ṇatva: न becomes ण when a र or ऋ precedes in the same word, across intervening sounds. The र at the start of राम conditions a change four sounds later — a genuinely non-local rule. STEP 5 · NO RULE APPLIES · HALT रामेण rāmeṇa — “by Rāma”
Four rule firings from stem to surface. Note step 4: the retroflexion is triggered by a sound at the other end of the word, which is why a system of purely local rewriting cannot express Sanskrit phonology and why Pāṇini's does.

10Base and suffixप्रकृतिप्रत्ययौ

Every word in the language is a base plus one or more suffixes. That is the entire inventory of parts.

One verbal root generating a family of words The root kr, to do, radiating into eight derived forms: karoti, krtam, kurvan, kartum, karta, kartavyam, krtva and karotu. कृ kṛ — to do dhātu, verbal root करोति karoti — “does” कुर्वन् kurvan — “doing” कर्ता kartā — “doer” कृत्वा kṛtvā — “having done” कृतम् kṛtam — “done” कर्तुम् kartum — “to do” कर्तव्यम् kartavyam — “must be done” करोतु karotu — “let it be done”
And the same eight suffixes applied to paṭh (to read) give paṭhati, paṭhitam, paṭhan, paṭhitum, paṭhitā, paṭhitavyam, paṭhitvā, paṭhatu; applied to gam (to go) they give gacchati, gatam, gacchan, gantum, gantā, gantavyam, gatvā, gacchatu. The pattern is the grammar; the root is the variable.
What the suffixes do
BaseRole of the suffixExampleExtent
Nominal stem
प्रातिपदिक
Marks case and number — singular, dual and plural in each of seven cases.रामः · रामौ · रामाः
rāmaḥ · rāmau · rāmāḥ
रामेण · रामाभ्याम् · रामैः
rāmeṇa · rāmābhyām · rāmaiḥ
7 cases × 3 numbers = 21 forms per stem, plus a further set for the feminine.
Verbal root
धातु
Marks person, number, tense and mood.पठति · पठतः · पठन्ति
paṭhati · paṭhataḥ · paṭhanti
पठसि · पठथः · पठथ
paṭhasi · paṭhathaḥ · paṭhatha
पठामि · पठावः · पठामः
paṭhāmi · paṭhāvaḥ · paṭhāmaḥ
3 persons × 3 numbers × 10 lakāras (six tenses and four moods), in two voices.

Note the dual. Sanskrit distinguishes one, two, and more-than-two throughout — in nouns, adjectives, pronouns and verbs alike. It is a reminder that a grammar describes the distinctions a language happens to make, and that these are not universal.

11Prefixesउपसर्गाः

Before the root can go a preverb — upasarga — and these do not merely shade the meaning; they can redirect or reverse it. More than one may be stacked.

Prefixed forms of the root kr Eleven prefixed verbs formed on the root kr, with their meanings, arranged in a ring. कृ kṛ — to do upa-karotiacts so as to bring nearer — helps, benefits apa-karotiacts so as to take away — removes, harms prati-karotiacts in return, or in the reverse direction pratyupa-karotitwo prefixes: repays a kindness received ut-karotiraises up, elevates — cf. utkarṣa, excellence pra-karotisets going, brings forth saṃs-karotiputs together properly — refines; cf. saṃskṛta vyā-karotitakes apart, analyses — cf. vyākaraṇa, grammar nirā-karotithrows out, refutes adhi-karotiplaces over — cf. adhikāra, authority anu-karotifollows after — imitates
Two words on this diagram name things elsewhere in this series: saṃskṛta, the refined language, and vyākaraṇa, the discipline of taking apart. Both are transparent compounds of a prefix and this one root.

12Compounds, recursivelyसमासः

Sanskrit forms compounds by a procedure that calls itself — which is why the language can produce single words fifteen syllables long, and why parsing them mechanically is tractable.

The procedure for combining two words: strip each of its case suffix to recover the bare stem, join the stems, and treat the result as a new stem — which can then take suffixes, or be fed back into the same procedure.

Forming a compound and the recursion that generalises it Sastre and nipunah lose their endings to give the stems sastra and nipuna, which combine to sastranipuna, a new stem; the general case iterates this over n words. शास्त्रे śāstre — “in the śāstra” PŪRVA-PADA निपुणः nipuṇaḥ — “skilled” UTTARA-PADA strip endingstrip ending शास्त्र निपुण शास्त्रनिपुण A NEW STEM शास्त्रनिपुणः “skilled in the śāstra” …or feed it back in as the pūrva-pada of the next compounding The general case, stated as a recursion Let W = { w₁, w₂, … wₙ } be the stems to be joined, and Sᵢ the stem of the compound after the ith step. S₁ = w₁ Sᵢ = Sᵢ₋₁ + wᵢ for i = 2 … n (applying the sandhi rules at each join) The result Sₙ is itself a nominal stem, indistinguishable in the grammar from a simple one. Nothing limits n.
The self-similarity is the point: because a compound is a stem, no separate machinery is needed for compounds of compounds. This is also why Sanskrit prose can build a noun phrase of arbitrary depth without a single relative clause.

13Kāraka — who did whatकारकम्

Words alone do not make a sentence. “Dosa” is not a statement. Even “comes” is not: comes from where, with whom, for what? The grammar's answer is a formal theory of the roles a participant can play in an action.

A कारक (kāraka) is a participant in an action, defined by its relation to the verb (kriyā). There are six, and each is normally realised by a particular case ending — but the kāraka is the semantic role, and the case is only its usual expression.

The six karakas illustrated in one sentence A technician removes a machine from the office by truck in the morning: agent, object, instrument, source, location and recipient shown around the verb. अपकरोति apakaroti kriyā — “removes” KARTṚ · agent · 1st case यन्त्रकारकः yantrakārakaḥ — the technician the one in whom the cause of the action resides KARMAN · object · 2nd case यन्त्रम् yantram — the machine where the effect of the action lands KARAṆA · instrument · 3rd case वाहनेन vāhanena — by truck the most effective means of accomplishing it ADHIKARAṆA · locus · 7th case प्रातःकाले prātaḥkāle — in the morning the substratum: where or when it happens APĀDĀNA · source · 5th case कार्यालयात् kāryālayāt — from the office the fixed point from which separation occurs SAMPRADĀNA · recipient · 4th case the one the karman is intended to reach not present in this sentence — nothing is given to anyone Full sentence: यन्त्रकारकः प्रातःकाले कार्यालयात् वाहनेन यन्त्रम् अपकरोति । “In the morning the technician removes the machine from the office by truck.”
The sixth case — the genitive — is deliberately absent from this list. It relates one noun to another (“the king's horse”) rather than a participant to an action, and so is not a kāraka at all. The distinction between semantic role and case marking is stated explicitly, which is more than many modern grammars manage.

14Why word order is freeपदक्रमः

Because roles are carried by endings rather than by position, the words of a Sanskrit sentence can be shuffled without changing who did what. English cannot do this, and the contrast is instructive.

Five permutations of one Sanskrit sentence against their English word-order counterparts The Sanskrit sentence keeps its meaning under permutation because the case endings fix the roles, while the corresponding English reorderings become nonsense. SANSKRIT — ALL FIVE MEAN THE SAME THING ENGLISH — SAME WORDS, SAME ORDER स्थूलः बालकः स्वादु भोजनं हस्तेन खादति । sthūlaḥ bālakaḥ svādu bhojanaṃ hastena khādati The fat boy eats the tasty food with the hand. स्थूलः हस्तेन खादति स्वादु भोजनं बालकः । sthūlaḥ hastena khādati svādu bhojanaṃ bālakaḥ The fat hand eats the tasty food with the boy. स्थूलः भोजनं खादति स्वादु हस्तेन बालकः । sthūlaḥ bhojanaṃ khādati svādu hastena bālakaḥ The fat food eats the tasty hand with the boy. स्वादु भोजनं खादति स्थूलः हस्तेन बालकः । svādu bhojanaṃ khādati sthūlaḥ hastena bālakaḥ The tasty food eats the fat hand with the boy. स्वादु बालकः खादति स्थूलः भोजनं हस्तेन । svādu bālakaḥ khādati sthūlaḥ bhojanaṃ hastena The tasty boy eats the fat food with the hand. Four of the five English lines are nonsense. All five Sanskrit lines are the first sentence.
Watch bālakaḥ and bhojanam in the Sanskrit column: wherever they move, the first-case ending on one and the second-case ending on the other keep the boy eating and the food eaten. The adjectives sthūlaḥ and svādu track their nouns by agreement, not by adjacency. Word order in Sanskrit carries emphasis and metre — not grammatical relations.

15Parsing a sentenceवाक्यविश्लेषणम्

Run the machinery backwards. Given a surface sentence, split each word into base and suffix, and read the roles off the suffixes.

बालः वृक्षस्य फलं खादति — bālaḥ vṛkṣasya phalaṃ khādati — “the boy eats the fruit of the tree”
SurfaceBaseSuffixAnalysisRole
बालःबाल bālaसु sunoun stem + 1st case singularkartṛ — the boy, who eats
वृक्षस्यवृक्ष vṛkṣaस्य syanoun stem + 6th case singularnot a kāraka — relates “tree” to “fruit”, not to the eating
फलंफल phalaअम् amnoun stem + 2nd case singularkarman — the fruit, which is eaten
खादतिखाद् khādति tiverbal root + 3rd person singular, presentkriyā — the action itself

Every cell in the fourth column was recovered mechanically from the ending. No dictionary of sentence patterns was consulted, and no statistical model was needed — the information is in the morphology, put there by the rules of section 9 and recoverable by running them in reverse.

16Sanskrit and NLPयन्त्रभाषाविज्ञानम्

Natural language processing has two halves, and they are not equally hard.

Natural language generation compared with understanding Generation goes from meaning to text and is tractable with rules; understanding goes from text to meaning and is much harder. Generation · NLG meaning rules, lexicon, syntax text Comparatively tractable where the rules are explicit — which is exactly what the Aṣṭādhyāyī provides. Understanding · NLU text ? ambiguity, context, world meaning Much harder. Ambiguity is the enemy, and natural languages are full of it by design — most of it resolved by context a machine does not have. What Sanskrit specifically offers · Roles are marked morphologically, not positionally — a parser gets the kāraka structure from the endings. · Word formation is derivational and rule-governed, so an unseen word can be analysed rather than merely looked up. · The grammar is already written as a formal system with an explicit metalanguage — it does not have to be reverse-engineered from a corpus.
A claim to state carefully

It is sometimes said that Sanskrit is “the best language for computers” or “unambiguous”. Neither is right, and repeating them costs credibility. Sanskrit prose is perfectly capable of ambiguity — compounds are notoriously so, since rāja-puruṣa could be the king's man, a man who is a king, or several other things, and the tradition wrote a great deal on how to decide. What is genuinely true is narrower and still remarkable: the language has an explicit, complete, formal grammar of great economy, in which sentence roles are marked on the words themselves. That makes it an unusually good testbed, and a source of design ideas — the kāraka scheme in particular anticipates the case-role and dependency representations that modern parsers use.

The influence is real and traceable. Dependency grammar's notion that a sentence is a verb plus labelled arguments is the kāraka scheme in other clothing. The idea of a phonological rule quantifying over natural classes defined by feature intervals is the pratyāhāra. And “Pāṇinian” is a live adjective in computational linguistics, attached to a family of grammar formalisms for Indian languages.

17Self-checkपरीक्षा

Check your reading

1 · What is a pratyāhāra?

An interval over an ordered inventory, named by its endpoints — ac = all vowels, hal = all consonants, ik = i u ṛ ḷ.

2 · In iko yaṇ aci, what does the locative aci contribute?

Genitive = what is replaced (ik), nominative = the substitute (yaṇ), locative = the environment (ac). The grammar uses Sanskrit's own cases as its metalanguage.

3 · Why does the derivation of rāmeṇa end in ṇa rather than na?

The trigger is the r at the beginning of rāma, four sounds away. Guṇa sandhi is the earlier step that gives rāme.

4 · Which of the following is not a kāraka?

The genitive relates noun to noun, not participant to action, so it falls outside the theory. That the tradition drew this line explicitly is the impressive part.

5 · Why can the words of a Sanskrit sentence be freely reordered?

The five permutations in section 14 all mean the same thing in Sanskrit; the same reorderings in English produce nonsense, because English marks roles by position.

Questions worth arguing about

Is the Aṣṭādhyāyī a description or a prescription?

Both, uncomfortably. It describes a language Pāṇini observed — he records regional variants and cites earlier grammarians who disagreed. But its adoption made it prescriptive: after it, correct Sanskrit largely means Pāṇinian Sanskrit, and the language stopped changing in the ordinary way. A grammar that succeeds completely freezes its object.

What did the extreme compression cost?

Accessibility. The Aṣṭādhyāyī cannot be read without a teacher or a commentary, and the entire subsequent tradition is occupied with explaining it. A less compressed grammar would have been more usable and less durable — it would not have survived oral transmission. The design is optimal for its constraint, not in the abstract.

Does a formal grammar of one language tell us anything about language in general?

The method transfers even where the content does not. The idea that a language has a finite rule set generating an infinite set of well-formed strings, that rules apply in order and can block each other, and that a metalanguage is needed to state them — none of that is specific to Sanskrit, and all of it is in the Aṣṭādhyāyī before it appears anywhere else.

18Glossaryशब्दकोशः

IASTDevanāgarīSense
adhikaraṇaअधिकरणThe locus kāraka: where or when the action occurs. Seventh case.
apādānaअपादानThe source kāraka: the fixed point from which separation occurs. Fifth case.
dhātuधातुA verbal root.
itइत्A marker letter attached to an element for reference purposes and then deleted.
karaṇaकरणThe instrument kāraka. Third case.
kārakaकारकA participant role in an action; there are six.
karmanकर्मन्The object kāraka: where the effect lands. Second case.
kartṛकर्तृThe agent kāraka. First case.
kriyāक्रियाThe action; the verb around which the kārakas are organised.
lakāraलकारOne of the ten tense–mood classes of the Sanskrit verb.
māheśvara-sūtraमाहेश्वरसूत्रOne of the fourteen ordered strings listing the sounds of the language.
padaपदA word: anything ending in a nominal or verbal affix.
prātipadikaप्रातिपदिकA nominal stem, before case endings.
pratyāhāraप्रत्याहारAn abbreviation naming a contiguous set of sounds in the Māheśvara list.
pratyayaप्रत्ययAn affix.
samāsaसमासA compound; also the recursive procedure that forms one.
sampradānaसम्प्रदानThe recipient kāraka. Fourth case.
sandhiसन्धिThe systematic sound changes at a junction between elements.
sūtraसूत्रA rule of the Aṣṭādhyāyī.
upasargaउपसर्गA verbal prefix.
varṇaवर्णA speech sound; the minimal unit of the phonological analysis.
vārttikaवार्त्तिकA critical note on a sūtra, as in Kātyāyana's collection.
Previous · Chapter 4Wisdom through the Ages Next · Chapter 6Number System and Units of Measurement