Sanskrit · Research
The Meter and the Molecule
Part One: two codes written in threes
Every cell of your body is, at this moment, reciting. A ribosome settles onto a strand of messenger RNA and moves along it three letters at a time, reading each triplet as one unit of meaning, assembling a protein syllable by syllable. Twenty-three centuries ago — give or take a scholarly argument — a grammarian named Piṅgala wrote down a way of reading sequences three symbols at a time, out of an alphabet of exactly two, and worked out the complete mathematics of every pattern that reading could produce. He was describing the meters of Sanskrit poetry. The cell, it turns out, keeps its library in a strikingly similar format.
This essay is the first in a series exploring what these two systems — chandas, the Vedic science of meter, and the molecular machinery of DNA and protein — might have to say to each other. I want to walk the ground carefully. Some of the connections turn out to be mathematically exact, and they are more beautiful than the loose ones. Some are structural rhymes: the same organizing principle appearing in two unrelated media. And some familiar claims in this territory are numerology wearing a lab coat, and part of respecting both traditions is saying so. By the end you will have, if nothing else, two downloadable reference documents and a working dictionary between the sixty-four codons of the genetic code and the eight gaṇas of Sanskrit prosody — tools for the more adventurous work the later parts of this series will attempt.
A science of weighing syllables
Chandas is one of the six Vedāṅgas, the auxiliary sciences that guard the Veda; the Pāṇinīya Śikṣā calls it the feet of the Veda, the part it stands on. What the science actually does is weigh syllables. Every Sanskrit syllable is either laghu — light — or guru — heavy. A syllable is heavy if its vowel is long, or if a short vowel is closed by a consonant cluster, an anusvāra, or a visarga; otherwise it is light. There is no third option. A line of verse, before it is anything else, is a string over a two-letter alphabet.
The tradition heard this as duration: a laghu lasts one mātrā, one beat; a guru lasts two. But notice what has happened the moment you write the weights down. The opening of the Bhagavad-gītā's most famous meter, the śloka, becomes something like ×××× LGG× — and any verse at all becomes a binary sequence. The prosodists took this reduction utterly seriously. Their names for meters are not descriptions but coordinates: Vasantatilakā, "the ornament of spring," is defined as the gaṇa formula ta bha ja ja ga ga — precisely the fourteen-syllable string GGLGLLLGLLGLGG, every position specified, repeated four times per stanza.1
A gaṇa is a triplet of syllable weights. Two possibilities per syllable, three syllables: eight gaṇas, and the tradition names each with a single letter — ma, ya, ra, sa, ta, ja, bha, na. Meters are spelled in gaṇas the way proteins are spelled in codons: a long binary string, parsed in threes.
Piṅgala's arithmetic
The Chandaḥśāstra of Piṅgala — traditionally placed around the third century BCE — closes with a set of procedures called pratyayas that treat the space of all possible meters as a mathematical object. The first, prastāra, is an algorithm for laying out every possible pattern of n syllables in a fixed canonical order: begin with all gurus; at each step find the leftmost guru, turn it to laghu, restore everything to its left to guru, and copy the rest. Run it for three syllables and out come the eight gaṇas, in order: ma ya ra sa ta ja bha na — which is exactly the order in which the tradition has always recited them. The gaṇa alphabet is an enumeration; if you write laghu and guru as digits, the prastāra is counting in binary, two millennia before Leibniz.2
The other pratyayas complete the toolkit. Naṣṭa, "the lost one," recovers the pattern at any row number without listing the rows before it: if the number is odd write guru and add one before halving, if even write laghu and halve — repeated n times. Its inverse, uddiṣṭa, takes a pattern and computes its row: scan right to left, doubling for each laghu, doubling-minus-one for each guru. These are, position for position, conversions between binary and decimal. They make the abstraction vivid: Vasantatilakā is not just a lovely rhythm, it is row 2,933 of the 16,384 possible fourteen-syllable pādas. Saṃkhyā computes that 16,384 — as 2¹⁴ — by a halve-and-square shortcut that is the earliest known statement of exponentiation by squaring. And lagakriyā answers "how many patterns have exactly k gurus?"; its tabular form, the meru-prastāra worked out by the tenth-century commentator Halāyudha, is Pascal's triangle, seven centuries before Pascal.24
Sit with the shape of this for a moment. The named meters of classical poetry — a few hundred of them, catalogued in treatises like the Vṛttaratnākara5 — float in a space the tradition itself computed to hold over 134 million possible pāda-patterns. The meters actually used are perhaps a millionth of what is combinatorially available. Molecular biology knows this picture intimately: functional proteins are a vanishingly sparse subset of the space of possible amino-acid sequences. Both traditions selected, from an astronomical space, the tiny fraction that works — that scans, that folds.
Counting with time itself
Alongside the syllable-counting meters runs a second system that counts pure duration: mātrā meters, where a laghu contributes one beat and a guru two, and only the total is fixed. The great Āryā stanza asks for thirty mātrās in its first half and twenty-seven in its second, built from four-beat gaṇas — and there are exactly five ways to fill four beats with lights and longs. How many ways to fill m beats? The answer — 1, 2, 3, 5, 8, 13, 21… — was derived from precisely this question by the prosodist Virahāṅka around the seventh century CE: each count is the sum of the two before it, because a pattern of m beats ends either in a laghu (leaving m−1) or a guru (leaving m−2). Europe would meet the sequence five hundred years later and name it after Fibonacci.3
The memory wheel
To hold the eight gaṇas in mind, students of prosody have long used a nonsense word: yamātārājabhānasalagāḥ — यमाताराजभानसलगाः. It looks like a mere string of the gaṇa names. It is far more. Weigh its ten syllables and you get L G G G L G L L L G; now read any three consecutive syllables, starting from each of the first eight positions, and you get all eight gaṇas, each exactly once, each beginning at the syllable of its own name. Every triplet is packed into a ten-syllable word by letting the windows overlap.
Mathematics now calls a sequence with this property a de Bruijn sequence, after the twentieth-century Dutch mathematician who studied them; Donald Knuth notes this Sanskrit mnemonic as the oldest known example.4 And here the rhyme with biology stops being decorative. When modern sequencers read a genome, they recover millions of short overlapping fragments — k-mers — and reassemble the original text by walking a de Bruijn graph, the structure in which every length-k window of a sequence becomes a node stitched to its overlapping neighbors.10 The overlapping-window trick a Sanskrit student uses to carry the gaṇa alphabet in memory is, formally, the trick an assembler uses to carry a genome across the gap between fragments and text. Same mathematics, twenty-two centuries apart.
The cell's prosody
Now the other code. DNA writes in a four-letter alphabet — A, C, G, T (U in RNA) — and the protein-making machinery reads it in threes. Four letters, three positions: sixty-four codons, mapped onto twenty amino acids and a stop sign. The first codon humanity ever deciphered was read in 1961, when Nirenberg and Matthaei fed ribosomes a synthetic message of pure uracil — UUU UUU UUU… — and got back a chain of pure phenylalanine.6 Within five years the whole table was cracked.
The table has structure, and the structure has a binary flavor. Because sixty-four codons serve only twenty-one meanings, the code is degenerate — and the degeneracy is not scattered randomly. It concentrates overwhelmingly in the third position, where Crick's "wobble" hypothesis explained how a single tRNA can read several codons that differ only in their final letter.7 In 1966 — the same year, remarkably — the Soviet physicist Yuri Rumer noticed that the sixty-four codons split into two perfect halves of thirty-two: those whose first two letters fix the amino acid no matter what the third says, and those where the third letter still matters — and that one elegant global substitution (swap U with G, A with C) maps each half exactly onto the other.8
Chemists have long known three natural ways to see the four bases as two: purine versus pyrimidine (A,G vs C,U — the two-ring bases against the one-ring), strong versus weak (G,C vs A,U — three hydrogen bonds against two), and amino versus keto. Each is a projection of the four-letter alphabet onto a two-letter one, and the first of them, "RY-coding," is a standard working tool in molecular phylogenetics, used to strip noise from deep evolutionary comparisons.12 The cell's four-letter text, in other words, casts several perfectly well-defined binary shadows.
Sixty-four codons, eight gaṇas
Here is the exact bridge. Take any of those binary shadows and look at what it does to codons. Two possibilities per position, three positions: the sixty-four codons collapse into exactly eight classes of exactly eight — and eight binary triplets is precisely the gaṇa alphabet. Under the purine/pyrimidine reduction, every codon is a gaṇa, in the same sense that every Sanskrit triplet is one. The correspondence is not a metaphor; it is an isomorphism between the reduced codon space and the space Piṅgala enumerated. What remains a choice is the polarity — which class plays guru. Read purines as heavy (they are the larger molecules) and you get one dictionary; read the strong G≡C pairs as heavy (three hydrogen bonds, the more tightly bound, the longer-burning — much closer to what guru means in a mouth) and you get another. Both dictionaries are in the downloadable data below, all sixty-four codons in each polarity, and I would rather test them than adjudicate them from an armchair.
One small omen for the road: under either polarity, UUU — the first codon ever read, poly-U's monotonous hymn — maps to na, three laghus, the lightest foot in the system. The first word the cell ever spoke to us scans.
The relaxed third position
And one rhyme I have not seen written down elsewhere. In nearly every classical meter, the final syllable of the pāda is anceps — free. The poet may set a laghu where the scheme says guru; the position at the boundary of the unit simply is not enforced. The genetic code relaxes in the same place: the third position of the codon, the boundary of its unit, is where wobble lives and mutations go quietest. Two reading systems, both strict in the interior, both loose at the seam. This is a structural analogy, not an identity — the pāda relaxes once at its edge, the codon at every third letter — but it is the kind of analogy that suggests a shared engineering logic: put your tolerance where the units join.
Weighing the parallels
Because this territory attracts enthusiasm, let me weigh my own claims on the tradition's own scale — heavy, light, and unsound.
Heavy — mathematically exact. The binary reduction of the codon alphabet collapsing sixty-four codons onto the eight gaṇas: exact, as a projection. The de Bruijn structure shared by the gaṇa mnemonic and assembly graphs: the same formal object. Piṅgala's pratyayas as binary enumeration, conversion, and fast exponentiation, and Virahāṅka's mātrā counts as the Fibonacci recurrence: settled history of mathematics, however under-celebrated.23
Light — structural analogies. Anceps and wobble as boundary tolerance. Vedic meter's conserved cadence and free interior beside conserved motifs and variable regions. The sparseness of named meters in pattern space beside the sparseness of working proteins in sequence space. Fixed-duration mātrā verse beside fixed-property constraints over variable sequences. These are real resemblances of principle, and they can guide hypotheses; they prove nothing by themselves.
Unsound — and worth naming. The Gāyatrī's twenty-four syllables matched to twenty-four of this or that; meters assigned to nucleotides by fiat; any argument whose engine is that two numbers are equal. The genetic code does contain striking regularities, but the sober literature — Crick's "frozen accident," the error-minimization and coevolution accounts reviewed by Koonin and Novozhilov — offers prosaic explanations competing to absorb them,911 and the cautionary tale of reading intentional "signals" into the code's arithmetic has already been written.13 When Part Two of this series runs statistics, the null models will be doing the heavy lifting, and I intend to let them.
Next in the series: Part Two will put the dictionary to work — transcribing real genes into gaṇa-strings under both polarities, asking whether anything in a coding sequence scans: do śloka-legal windows occur above chance? Do exons keep mātrā-time differently than introns? The data files below are exactly what you need to try it before I do.
Downloads for this series
Sanskrit Meter — A Structural Reference (PDF)
References
- Kedārabhaṭṭa, Vṛttaratnākara; and Piṅgala, Chandaḥśāstra with the Mṛtasañjīvanī commentary of Halāyudha. The gaṇa formulas used throughout are the standard definitions of these treatises.
- B. van Nooten, "Binary numbers in Indian antiquity," Journal of Indian Philosophy 21 (1993), 31–50.
- P. Singh, "The so-called Fibonacci numbers in ancient and medieval India," Historia Mathematica 12 (1985), 229–244.
- D. E. Knuth, The Art of Computer Programming, vol. 4A, §7.2.1.1 — on Piṅgala's enumeration and the gaṇa mnemonic as the oldest known de Bruijn cycle.
- H. D. Velankar, Jayadāman: A Collection of Ancient Texts on Sanskrit Prosody (Bombay, 1949) — the standard modern index of meters.
- M. W. Nirenberg & J. H. Matthaei, "The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides," PNAS 47 (1961), 1588–1602.
- F. H. C. Crick, "Codon–anticodon pairing: the wobble hypothesis," Journal of Molecular Biology 19 (1966), 548–555.
- Yu. B. Rumer, "Systematization of codons in the genetic code" (1966), English translation: Phil. Trans. R. Soc. A 374 (2016), 20150446.
- F. H. C. Crick, "The origin of the genetic code," Journal of Molecular Biology 38 (1968), 367–379.
- P. E. C. Compeau, P. A. Pevzner & G. Tesler, "How to apply de Bruijn graphs to genome assembly," Nature Biotechnology 29 (2011), 987–991.
- E. V. Koonin & A. S. Novozhilov, "Origin and evolution of the genetic code: the universal enigma," IUBMB Life 61 (2009), 99–111.
- M. P. Simmons, "Relative benefits of amino-acid, codon, degeneracy, DNA, and purine-pyrimidine character coding for phylogenetic analyses of exons," Journal of Systematics and Evolution 55 (2017), 85–109.
- V. shCherbak & M. Makukov, "The 'Wow! signal' of the terrestrial genetic code," Icarus 224 (2013) — cited here as an instructive example of how far pattern-reading in the code can be taken, and of the skepticism such readings must survive.
All metrical patterns, prastāra indices, codon classes, and combinatorial identities quoted in this essay and its companion documents were generated programmatically from the primary definitions and machine-verified (algorithm round-trips, class counts, degeneracy totals) rather than transcribed by hand. The companion reference PDF documents the verification conventions.