Sanskrit · Research

The Meter and the Molecule

Part One: two codes written in threes

Every cell of your body is, at this moment, reciting. A ribosome settles onto a strand of messenger RNA and moves along it three letters at a time, reading each triplet as one unit of meaning, assembling a protein syllable by syllable. Twenty-three centuries ago — give or take a scholarly argument — a grammarian named Piṅgala wrote down a way of reading sequences three symbols at a time, out of an alphabet of exactly two, and worked out the complete mathematics of every pattern that reading could produce. He was describing the meters of Sanskrit poetry. The cell, it turns out, keeps its library in a strikingly similar format.

This essay is the first in a series exploring what these two systems — chandas, the Vedic science of meter, and the molecular machinery of DNA and protein — might have to say to each other. I want to walk the ground carefully. Some of the connections turn out to be mathematically exact, and they are more beautiful than the loose ones. Some are structural rhymes: the same organizing principle appearing in two unrelated media. And some familiar claims in this territory are numerology wearing a lab coat, and part of respecting both traditions is saying so. By the end you will have, if nothing else, two downloadable reference documents and a working dictionary between the sixty-four codons of the genetic code and the eight gaṇas of Sanskrit prosody — tools for the more adventurous work the later parts of this series will attempt.

A science of weighing syllables

Chandas is one of the six Vedāṅgas, the auxiliary sciences that guard the Veda; the Pāṇinīya Śikṣā calls it the feet of the Veda, the part it stands on. What the science actually does is weigh syllables. Every Sanskrit syllable is either laghu — light — or guru — heavy. A syllable is heavy if its vowel is long, or if a short vowel is closed by a consonant cluster, an anusvāra, or a visarga; otherwise it is light. There is no third option. A line of verse, before it is anything else, is a string over a two-letter alphabet.

The tradition heard this as duration: a laghu lasts one mātrā, one beat; a guru lasts two. But notice what has happened the moment you write the weights down. The opening of the Bhagavad-gītā's most famous meter, the śloka, becomes something like ×××× LGG× — and any verse at all becomes a binary sequence. The prosodists took this reduction utterly seriously. Their names for meters are not descriptions but coordinates: Vasantatilakā, "the ornament of spring," is defined as the gaṇa formula ta bha ja ja ga ga — precisely the fourteen-syllable string GGLGLLLGLLGLGG, every position specified, repeated four times per stanza.1

A gaṇa is a triplet of syllable weights. Two possibilities per syllable, three syllables: eight gaṇas, and the tradition names each with a single letter — ma, ya, ra, sa, ta, ja, bha, na. Meters are spelled in gaṇas the way proteins are spelled in codons: a long binary string, parsed in threes.

rowpatterngaṇabits (G=1)mātrās1ma-gaṇa11162ya-gaṇa01153ra-gaṇa10154sa-gaṇa00145ta-gaṇa11056ja-gaṇa01047bha-gaṇa10048na-gaṇa0003The traditionalcitation order —ma ya ra sa ta jabha na — is theprastāra: eighttriplets listed by acounting algorithm.Binary, two millenniabefore Leibniz.
Figure 1. The eight gaṇas in their traditional citation order — which is Piṅgala's prastāra, an enumeration algorithm. The bits column uses this series' convention (laghu = 0, guru = 1); under Piṅgala's own reading (laghu 1, guru 0, leftmost syllable as the low digit), row k spells the number k−1, so the rows count from zero to seven. Mātrās: duration in beats.

Piṅgala's arithmetic

The Chandaḥśāstra of Piṅgala — traditionally placed around the third century BCE — closes with a set of procedures called pratyayas that treat the space of all possible meters as a mathematical object. The first, prastāra, is an algorithm for laying out every possible pattern of n syllables in a fixed canonical order: begin with all gurus; at each step find the leftmost guru, turn it to laghu, restore everything to its left to guru, and copy the rest. Run it for three syllables and out come the eight gaṇas, in order: ma ya ra sa ta ja bha na — which is exactly the order in which the tradition has always recited them. The gaṇa alphabet is an enumeration; if you write laghu and guru as digits, the prastāra is counting in binary, two millennia before Leibniz.2

The other pratyayas complete the toolkit. Naṣṭa, "the lost one," recovers the pattern at any row number without listing the rows before it: if the number is odd write guru and add one before halving, if even write laghu and halve — repeated n times. Its inverse, uddiṣṭa, takes a pattern and computes its row: scan right to left, doubling for each laghu, doubling-minus-one for each guru. These are, position for position, conversions between binary and decimal. They make the abstraction vivid: Vasantatilakā is not just a lovely rhythm, it is row 2,933 of the 16,384 possible fourteen-syllable pādas. Saṃkhyā computes that 16,384 — as 2¹⁴ — by a halve-and-square shortcut that is the earliest known statement of exponentiation by squaring. And lagakriyā answers "how many patterns have exactly k gurus?"; its tabular form, the meru-prastāra worked out by the tenth-century commentator Halāyudha, is Pascal's triangle, seven centuries before Pascal.24

Sit with the shape of this for a moment. The named meters of classical poetry — a few hundred of them, catalogued in treatises like the Vṛttaratnākara5 — float in a space the tradition itself computed to hold over 134 million possible pāda-patterns. The meters actually used are perhaps a millionth of what is combinatorially available. Molecular biology knows this picture intimately: functional proteins are a vanishingly sparse subset of the space of possible amino-acid sequences. Both traditions selected, from an astronomical space, the tiny fraction that works — that scans, that folds.

Counting with time itself

Alongside the syllable-counting meters runs a second system that counts pure duration: mātrā meters, where a laghu contributes one beat and a guru two, and only the total is fixed. The great Āryā stanza asks for thirty mātrās in its first half and twenty-seven in its second, built from four-beat gaṇas — and there are exactly five ways to fill four beats with lights and longs. How many ways to fill m beats? The answer — 1, 2, 3, 5, 8, 13, 21… — was derived from precisely this question by the prosodist Virahāṅka around the seventh century CE: each count is the sum of the two before it, because a pattern of m beats ends either in a laghu (leaving m−1) or a guru (leaving m−2). Europe would meet the sequence five hundred years later and name it after Fibonacci.3

the five ways to fill four mātrāsGGLLGLGLGLLLLLLequal width = equal duration1122335485136217348patterns per total of m mātrās — Virahāṅka (Fibonacci)m
Figure 2. Fixed time, free sequence. Left: the five ways to fill four mātrās — the Āryā's building blocks — drawn so that equal width is equal duration. Right: how many patterns total m mātrās, for m = 1…8: Virahāṅka's counts, the sequence Europe later named for Fibonacci.

The memory wheel

To hold the eight gaṇas in mind, students of prosody have long used a nonsense word: yamātārājabhānasalagāḥ — यमाताराजभानसलगाः. It looks like a mere string of the gaṇa names. It is far more. Weigh its ten syllables and you get L G G G L G L L L G; now read any three consecutive syllables, starting from each of the first eight positions, and you get all eight gaṇas, each exactly once, each beginning at the syllable of its own name. Every triplet is packed into a ten-syllable word by letting the windows overlap.

Mathematics now calls a sequence with this property a de Bruijn sequence, after the twentieth-century Dutch mathematician who studied them; Donald Knuth notes this Sanskrit mnemonic as the oldest known example.4 And here the rhyme with biology stops being decorative. When modern sequencers read a genome, they recover millions of short overlapping fragments — k-mers — and reassemble the original text by walking a de Bruijn graph, the structure in which every length-k window of a sequence becomes a node stitched to its overlapping neighbors.10 The overlapping-window trick a Sanskrit student uses to carry the gaṇa alphabet in memory is, formally, the trick an assembler uses to carry a genome across the gap between fragments and text. Same mathematics, twenty-two centuries apart.

yaमाताराjaभाbhānasalaगाःgāḥya-gaṇa · LGGma-gaṇa · GGGta-gaṇa · GGLra-gaṇa · GLGja-gaṇa · LGLbha-gaṇa · GLLna-gaṇa · LLLsa-gaṇa · LLGyamātārājabhānasalagāḥ — read any three consecutive syllables:
Figure 3. yamātārājabhānasalagāḥ: ten syllables whose weights (bars: short = laghu, long = guru) contain all eight gaṇas as overlapping three-syllable windows — each beginning at the syllable of its own name. A binary de Bruijn sequence, in continuous mnemonic use for centuries.

The cell's prosody

Now the other code. DNA writes in a four-letter alphabet — A, C, G, T (U in RNA) — and the protein-making machinery reads it in threes. Four letters, three positions: sixty-four codons, mapped onto twenty amino acids and a stop sign. The first codon humanity ever deciphered was read in 1961, when Nirenberg and Matthaei fed ribosomes a synthetic message of pure uracil — UUU UUU UUU… — and got back a chain of pure phenylalanine.6 Within five years the whole table was cracked.

The table has structure, and the structure has a binary flavor. Because sixty-four codons serve only twenty-one meanings, the code is degenerate — and the degeneracy is not scattered randomly. It concentrates overwhelmingly in the third position, where Crick's "wobble" hypothesis explained how a single tRNA can read several codons that differ only in their final letter.7 In 1966 — the same year, remarkably — the Soviet physicist Yuri Rumer noticed that the sixty-four codons split into two perfect halves of thirty-two: those whose first two letters fix the amino acid no matter what the third says, and those where the third letter still matters — and that one elegant global substitution (swap U with G, A with C) maps each half exactly onto the other.8

Chemists have long known three natural ways to see the four bases as two: purine versus pyrimidine (A,G vs C,U — the two-ring bases against the one-ring), strong versus weak (G,C vs A,U — three hydrogen bonds against two), and amino versus keto. Each is a projection of the four-letter alphabet onto a two-letter one, and the first of them, "RY-coding," is a standard working tool in molecular phylogenetics, used to strip noise from deep evolutionary comparisons.12 The cell's four-letter text, in other words, casts several perfectly well-defined binary shadows.

Sixty-four codons, eight gaṇas

Here is the exact bridge. Take any of those binary shadows and look at what it does to codons. Two possibilities per position, three positions: the sixty-four codons collapse into exactly eight classes of exactly eight — and eight binary triplets is precisely the gaṇa alphabet. Under the purine/pyrimidine reduction, every codon is a gaṇa, in the same sense that every Sanskrit triplet is one. The correspondence is not a metaphor; it is an isomorphism between the reduced codon space and the space Piṅgala enumerated. What remains a choice is the polarity — which class plays guru. Read purines as heavy (they are the larger molecules) and you get one dictionary; read the strong G≡C pairs as heavy (three hydrogen bonds, the more tightly bound, the longer-burning — much closer to what guru means in a mouth) and you get another. Both dictionaries are in the downloadable data below, all sixty-four codons in each polarity, and I would rather test them than adjudicate them from an armchair.

One small omen for the road: under either polarity, UUU — the first codon ever read, poly-U's monotonous hymn — maps to na, three laghus, the lightest foot in the system. The first word the cell ever spoke to us scans.

R — purine (A, G)Y — pyrimidine (C, U)polarity shown: R→guru, Y→laghumaRRRAAA → LysAAAKAAG → LysAAGKAGA → ArgAGARAGG → ArgAGGRGAA → GluGAAEGAG → GluGAGEGGA → GlyGGAGGGG → GlyGGGGyaYRRCAA → GlnCAAQCAG → GlnCAGQCGA → ArgCGARCGG → ArgCGGRUAA → StopUAA*UAG → StopUAG*UGA → StopUGA*UGG → TrpUGGWraRYRACA → ThrACATACG → ThrACGTAUA → IleAUAIAUG → MetAUGMGCA → AlaGCAAGCG → AlaGCGAGUA → ValGUAVGUG → ValGUGVsaYYRCCA → ProCCAPCCG → ProCCGPCUA → LeuCUALCUG → LeuCUGLUCA → SerUCASUCG → SerUCGSUUA → LeuUUALUUG → LeuUUGLtaRRYAAC → AsnAACNAAU → AsnAAUNAGC → SerAGCSAGU → SerAGUSGAC → AspGACDGAU → AspGAUDGGC → GlyGGCGGGU → GlyGGUGjaYRYCAC → HisCACHCAU → HisCAUHCGC → ArgCGCRCGU → ArgCGURUAC → TyrUACYUAU → TyrUAUYUGC → CysUGCCUGU → CysUGUCbhaRYYACC → ThrACCTACU → ThrACUTAUC → IleAUCIAUU → IleAUUIGCC → AlaGCCAGCU → AlaGCUAGUC → ValGUCVGUU → ValGUUVnaYYYCCC → ProCCCPCCU → ProCCUPCUC → LeuCUCLCUU → LeuCUULUCC → SerUCCSUCU → SerUCUSUUC → PheUUCFUUU → PheUUUF
Figure 4. The collapse: sort the 64 codons by the purine/pyrimidine pattern of their three positions and exactly eight classes of eight emerge — the gaṇa alphabet. Columns follow the prastāra order of Figure 1; the small letter after each codon is its amino acid (* = stop); hover any codon for its full name. Drawn in the R→guru polarity; the S/W dictionary is in the data appendix.

The relaxed third position

And one rhyme I have not seen written down elsewhere. In nearly every classical meter, the final syllable of the pāda is anceps — free. The poet may set a laghu where the scheme says guru; the position at the boundary of the unit simply is not enforced. The genetic code relaxes in the same place: the third position of the codon, the boundary of its unit, is where wobble lives and mutations go quietest. Two reading systems, both strict in the interior, both loose at the seam. This is a structural analogy, not an identity — the pāda relaxes once at its edge, the codon at every third letter — but it is the kind of analogy that suggests a shared engineering logic: put your tolerance where the units join.

a pāda of Vasantatilakā — every position fixed, except the last× ancepsa codon — first two positions carry the identity, the third wobblesN₁N₂N₃wobble (Crick 1966)Two codes, one habit: the constraintrelaxes at the boundary of the unit.
Figure 5. Where each code loosens its grip: the pāda's final syllable is anceps — metrically free — and the codon's third position is where wobble lives. A structural analogy (boundary tolerance in both reading systems), not an identity.

Weighing the parallels

Because this territory attracts enthusiasm, let me weigh my own claims on the tradition's own scale — heavy, light, and unsound.

Heavy — mathematically exact. The binary reduction of the codon alphabet collapsing sixty-four codons onto the eight gaṇas: exact, as a projection. The de Bruijn structure shared by the gaṇa mnemonic and assembly graphs: the same formal object. Piṅgala's pratyayas as binary enumeration, conversion, and fast exponentiation, and Virahāṅka's mātrā counts as the Fibonacci recurrence: settled history of mathematics, however under-celebrated.23

Light — structural analogies. Anceps and wobble as boundary tolerance. Vedic meter's conserved cadence and free interior beside conserved motifs and variable regions. The sparseness of named meters in pattern space beside the sparseness of working proteins in sequence space. Fixed-duration mātrā verse beside fixed-property constraints over variable sequences. These are real resemblances of principle, and they can guide hypotheses; they prove nothing by themselves.

Unsound — and worth naming. The Gāyatrī's twenty-four syllables matched to twenty-four of this or that; meters assigned to nucleotides by fiat; any argument whose engine is that two numbers are equal. The genetic code does contain striking regularities, but the sober literature — Crick's "frozen accident," the error-minimization and coevolution accounts reviewed by Koonin and Novozhilov — offers prosaic explanations competing to absorb them,911 and the cautionary tale of reading intentional "signals" into the code's arithmetic has already been written.13 When Part Two of this series runs statistics, the null models will be doing the heavy lifting, and I intend to let them.

Next in the series: Part Two will put the dictionary to work — transcribing real genes into gaṇa-strings under both polarities, asking whether anything in a coding sequence scans: do śloka-legal windows occur above chance? Do exons keep mātrā-time differently than introns? The data files below are exactly what you need to try it before I do.

Downloads for this series

Sanskrit Meter — A Structural Reference (PDF)
The metrical systems in full: syllable weights, the gaṇa code, Vedic and classical catalogs with verified binary patterns, mātrā meters, Piṅgala's algorithms, and the graded meter↔molecule mapping table.

Meter Catalog & Codon–Gaṇa Dictionary — Data Appendix (PDF)
Every major meter with its binary pattern and prastāra index; all sixty-four codons with amino acids and gaṇa assignments in both R/Y and S/W polarities.

Machine-readable data (ZIP: JSON + CSV)
meter-catalog and codon-gana-dictionary, for your own experiments. Every pattern and mapping was generated and cross-checked by program, not transcribed by hand.

References

  1. Kedārabhaṭṭa, Vṛttaratnākara; and Piṅgala, Chandaḥśāstra with the Mṛtasañjīvanī commentary of Halāyudha. The gaṇa formulas used throughout are the standard definitions of these treatises.
  2. B. van Nooten, "Binary numbers in Indian antiquity," Journal of Indian Philosophy 21 (1993), 31–50.
  3. P. Singh, "The so-called Fibonacci numbers in ancient and medieval India," Historia Mathematica 12 (1985), 229–244.
  4. D. E. Knuth, The Art of Computer Programming, vol. 4A, §7.2.1.1 — on Piṅgala's enumeration and the gaṇa mnemonic as the oldest known de Bruijn cycle.
  5. H. D. Velankar, Jayadāman: A Collection of Ancient Texts on Sanskrit Prosody (Bombay, 1949) — the standard modern index of meters.
  6. M. W. Nirenberg & J. H. Matthaei, "The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides," PNAS 47 (1961), 1588–1602.
  7. F. H. C. Crick, "Codon–anticodon pairing: the wobble hypothesis," Journal of Molecular Biology 19 (1966), 548–555.
  8. Yu. B. Rumer, "Systematization of codons in the genetic code" (1966), English translation: Phil. Trans. R. Soc. A 374 (2016), 20150446.
  9. F. H. C. Crick, "The origin of the genetic code," Journal of Molecular Biology 38 (1968), 367–379.
  10. P. E. C. Compeau, P. A. Pevzner & G. Tesler, "How to apply de Bruijn graphs to genome assembly," Nature Biotechnology 29 (2011), 987–991.
  11. E. V. Koonin & A. S. Novozhilov, "Origin and evolution of the genetic code: the universal enigma," IUBMB Life 61 (2009), 99–111.
  12. M. P. Simmons, "Relative benefits of amino-acid, codon, degeneracy, DNA, and purine-pyrimidine character coding for phylogenetic analyses of exons," Journal of Systematics and Evolution 55 (2017), 85–109.
  13. V. shCherbak & M. Makukov, "The 'Wow! signal' of the terrestrial genetic code," Icarus 224 (2013) — cited here as an instructive example of how far pattern-reading in the code can be taken, and of the skepticism such readings must survive.

All metrical patterns, prastāra indices, codon classes, and combinatorial identities quoted in this essay and its companion documents were generated programmatically from the primary definitions and machine-verified (algorithm round-trips, class counts, degeneracy totals) rather than transcribed by hand. The companion reference PDF documents the verification conventions.