Sanskrit · Research

The Meter and the Molecule

Part Six: the gaṇa and the amino acid

Part One built a dictionary and this series has been reading it in one direction ever since: codon → gaṇa, DNA turned into meter. But a dictionary has two sides, and the other side was never opened. Each gaṇa names exactly eight codons; eight codons name amino acids; so each of the eight feet of Sanskrit prosody names a family of amino acids. Turn the dictionary over and the Vedic foot stops being a rhythm imposed on DNA and becomes something stranger — a specification, a class of proteins. This essay opens that side, matches the two texts codon-for-codon in reading frame, and finds the strongest result the series has produced: the gaṇa's middle syllable — light or heavy — predicts whether its amino acids are water-fearing or water-loving, by a margin of four hydropathy units at the floor of a permutation test, while the first and third syllables say nothing whatsoever. The foot's middle syllable is its meaning-bearing one. Part Three had already shown it is the one position life refuses to vary. Those two facts turn out to be the same fact.

The dictionary, turned over

The mechanics are immediate. Under Part One's R/Y polarity — purine heavy, pyrimidine light — the gaṇa ma (G G G) names the eight codons AAA, AAG, AGA, AGG, GAA, GAG, GGA, GGG, and those encode lysine, arginine, glutamate and glycine: the charged residues and the smallest one. The gaṇa bha (G L L) names alanine, isoleucine, threonine and valine — the oily, buried ones. Every gaṇa resolves to a coherent chemical family, and the families sort themselves along a single axis when you weigh them with the standard Kyte–Doolittle hydropathy scale.3

the eight gaṇas as amino-acid families (R/Y polarity), sorted by hydropathyeach gaṇa names exactly 8 codons — and therefore a family of amino acidsgaṇapatternamino acidshydropathy (Kyte–Doolittle)middlebhaLA I T V+2.45raLA I M T V+2.12saLL P S+1.30naLF L P S+1.05laghu middle ↑ · guru middle ↓jaGC H R Y-1.62taGD G N S-2.05maGE G K R-3.08yaGQ R W+3 stop-3.38hydrophobic →← hydrophilic
Figure 1. Part One's dictionary turned over: each gaṇa's eight codons resolve to a family of amino acids. Sorted by mean hydropathy, the eight feet split without exception — every hydrophobic gaṇa carries a laghu middle syllable, every hydrophilic one a guru middle.

Sort the eight feet by hydropathy and the table splits at a stroke. The four hydrophobic gaṇas are bha, ra, sa, na; the four hydrophilic are ja, ta, ma, ya. Look at what separates the groups: every foot above the line has a laghu middle syllable, every foot below it a guru middle — eight for eight, no exceptions, no overlap. The prosodist's oldest classification, four feet against four, turns out to partition the amino acids by the property that decides whether a residue faces the water or hides from it.

Which syllable carries the meaning

That deserves a test rather than an admiring glance, so here is the test. Take the sixty-one sense codons, group them by the weight of their first syllable, and compare the mean hydropathy of the amino acids they encode: guru −0.14 against laghu −0.38, a difference of a quarter unit, p = 0.75. Do it by the third syllable: −0.49 against −0.04, p = 0.57. Nothing, and nothing. Now do it by the middle syllable: guru −2.44 against laghu +1.73 — a gap of 4.18 hydropathy units, with a two-sided permutation p of 0.00005, which is simply the floor of twenty thousand shuffles. The middle syllable does all the work; its neighbours do none.

which syllable carries the meaning? mean hydropathy of the amino acids named,split by the weight of each syllable of the gaṇa (R/Y, 61 sense codons)-3-2-10+1+2+3guru -0.14laghu -0.38syllable 1p = 0.75guru -2.44laghu +1.73syllable 2p = 0.00005Δ 4.18 unitsguru -0.49laghu -0.04syllable 3p = 0.57Kyte–Doolittleonly the middle syllable separates them — the first and third say nothing
Figure 2. The test: mean Kyte–Doolittle hydropathy of the amino acids named, split by the weight of each syllable of the gaṇa. Syllables one and three separate nothing (p = 0.75, 0.57); the middle syllable separates by 4.18 units at the permutation floor (p = 0.00005).

Let me be exact about what is and is not new here, because this series has earned its caution. The biochemistry is not new: Carl Woese showed in 1966 that the second base of a codon correlates with the amino acid's "polar requirement," and the second-position rule — U in the middle means hydrophobic, A in the middle means hydrophilic — is textbook.4 What Part Six contributes is the translation and the coincidence it exposes. In metrical terms the rule reads: a light middle syllable makes the foot hydrophobic, a heavy middle makes it hydrophilic — and in that reading it collides with something this series measured for entirely unrelated reasons. Part Three, comparing eight genomes from malaria to Streptomyces, found that the codon's third position ranges across nearly the whole scale (strong-base fraction 0.17 to 0.93) while the middle position is held in a narrow band by every organism alive. We now know why life will not let the middle position drift: it is the syllable that decides the fold. The conserved position, the meaning-bearing position, and the hydropathy-carrying position are one position — and the tradition, which had no way of knowing any of this, had already made that syllable the one a gaṇa is named for.

Codon for codon, in frame

With both sides of the dictionary open, the two texts can be matched the way translators match texts: unit for unit, in frame. Earlier parts slid Vedic patterns along DNA base by base; that is a comparison of letters. A gaṇa against a codon is a comparison of words — three syllables against three bases, both read in their own proper frame, the Veda from the start of its stanza and the gene from its start codon. Against 653 human codons (beta-globin, insulin, and TP53) the four Vedas offer 459 gaṇas, and the question is how long the two can agree.

The answer is six. The longest unbroken agreement anywhere, in either polarity, is six gaṇas — eighteen bases, six full codons of protein — and shuffled Vedas manage 5.7 to 6.1, so p = 0.53 and 0.83. Chance, as ever, and the series would be suspicious of anything else. But among the tens of thousands of possible pairings only four perfect six-gaṇa matches exist, which makes them rare enough to exhibit individually, and one of them stopped me.

the hero match — six gaṇas, eighteen bases, in frameṚgveda 10.129.7: “…he who surveys it from the highest heaven — he alone knows, or perhaps he knows not”vā na yoraACGTas yādh yakmaGAGEṣaḥ pa rabhaGTTVme vi oraGTGVman so aṅmaAGGRgha ve dajaCGCRTP53 codons 170–175 · peptide TEVVRRR175 — the most mutated residue in human cancerthe other three perfect matches:1.1.3 gaṇa 3 · ya-ja-ja-bha-ta-bha → INS CDS 94–99 · QCCTSI · CAATGCTGTACCAGCATC10.90.2 gaṇa 3 · ma-ma-na-sa-ya-ma → INS CDS 83–88 · EGSLQK · GAGGGGTCCCTGCAGAAGAVS 1.3.2 gaṇa 1 · ta-bha-ya-bha-ma-ma → TP53 CDS 352–357 · DAQAGK · GATGCCCAGGCTGGGAAG
Figure 3. The longest in-frame agreement in the corpus — six gaṇas, six codons, eighteen bases: the Nāsadīya's closing verse of unknowing laid across TP53 codons 170–175, ending on the hotspot residue 175. Chance, at the rate chance predicts; four such matches exist in all.

Ṛgveda 10.129.7 is the last verse of the Nāsadīya, the hymn of creation, and it is the most famous confession of ignorance in Sanskrit: yó asyādhyakṣaḥ paramé vyoman / só aṅga veda yádi vā ná veda — "he who is its overseer in the highest heaven, he alone knows — or perhaps he knows not." Six of its gaṇas, read as codons, match TP53 at codons 170 to 175, the peptide TEVVRR. That stretch sits in p53's DNA-binding domain and it ends on residue 175 — one of the three canonical hotspot codons (with 248 and 273) whose mutation accounts for a large share of all p53 misfires in human cancer.5 The hymn of not-knowing, laid across the guardian of the genome, ending precisely at the residue where the guardian most often fails. It is chance — four such matches are about what this much text should yield, and I would not publish it as anything else. But of all the passages and all the codons, that is where the arithmetic put its finger, and a writer who pretended not to notice would be lying about what the work is like. The other three are quieter: Ṛgveda 1.1.3 onto insulin at 94–99, Ṛgveda 10.90.2 onto insulin at 83–88, Atharvaveda 1.3.2 onto TP53 at 352–357.

The Gāyatrī, as a peptide

One more consequence follows from the inverted dictionary, and it is irresistible. The Gāyatrī mantra is twenty-four syllables. Twenty-four is eight gaṇas. Eight gaṇas is eight codons. The most recited verse in the world is, in the dictionary's other direction, exactly the length of an octapeptide — and it specifies one, in the way a degenerate motif does.

the Gāyatrī as an octapeptide — 24 syllables = 8 gaṇas = 8 codonseach gaṇa names 8 codons, so the mantra specifies a family of 76,800 possible octapeptidesbha+2.5AITVra+2.1AIMTVya-3.4QRWma-3.1EGKRja-1.6CHRYsa+1.3LPSta-2.0DGNSra+2.1AIMTVbest real human match: 4 of 8 (below chance, p = 0.998) — the mantra names no gene we holdred bar = hydrophobic family · blue = hydrophilic · highlighted middle syllable decides which
Figure 4. The Gāyatrī's twenty-four syllables are exactly eight gaṇas — eight codons — so the mantra specifies an octapeptide family: 76,800 of them under R/Y. Red = hydrophobic family, blue = hydrophilic, decided by the highlighted middle syllable. No human peptide we hold matches it above chance.

Its feet are bha-ra-ya-ma-ja-sa-ta-ra, and under R/Y that reads as the motif [AITV]-[AIMTV]-[QRW]-[EGKR]-[CHRY]-[LPS]-[DGNS]-[AIMTV]: seventy-six thousand eight hundred distinct octapeptides, with one foot (ya) admitting a stop. Its hydropathy profile — hydrophobic, hydrophobic, then a long hydrophilic middle, then a brief return — is a real shape, and one can ask the obvious question: does it look like a structured peptide? It does not. Its hydrophobic moment sits at the 77th percentile of Vedic windows, comfortably inside the crowd, no amphipathic helix hiding in the mantra. And its best match among the human peptides we hold is four positions out of eight — below the chance expectation of about five, p = 0.998. The Gāyatrī names a family of proteins that our genes decline to join. I record that with some satisfaction: the one place where this series' instrument could most easily have been made to sing, it sang flat, and said so.

What the middle syllable was for

The coda to this series argued that the meter and the molecule are united by engineering rather than authorship — that both codes put their tolerance at the boundary of the unit and their meaning in the middle. Part Six turns that argument from analogy into measurement. The pāda's free syllable is its last; the codon's free base is its third; and now the codon's meaning-bearing base has been shown to be its middle, carrying the single property that most determines what a protein becomes, with the two flanking positions carrying none of it. The eight gaṇas are not a decorative overlay on the genetic code. Read in the direction this essay opened, they are a partition of the amino acids by the property biology cares about most — a partition Piṅgala's alphabet already contained, waiting two thousand years for someone to weigh it.

To close: the coda gathers what five essays of testing left standing — one medium, one engineering, one vehicle, and no shared author any test can see.

Downloads for this essay

Codon bundle (ZIP: tables, catalog, code, results)
The gaṇa→amino-acid table for both polarities, the ranked codon-frame match catalog, the hydropathy tests, and all analysis code. Python standard library only; the Part One dictionary is re-derived and checked entry-for-entry at load.

Part One · Part Two · Part Three · Part Four · Part Five · Coda
The series.

References

  1. R. Oshop, Parts OneFive of this series, AyurAstro (2026) — the codon→gaṇa dictionary, the verified gene corpus, and the positional conservation result of Part Three.
  2. Verified sequences: HBB and INS (Part Two), TP53 NM_000546.6 CDS (Part Four); the four-Veda scanned corpus (Parts Four–Five).
  3. J. Kyte & R. F. Doolittle, "A simple method for displaying the hydropathic character of a protein," Journal of Molecular Biology 157 (1982), 105–132 — the hydropathy scale used throughout.
  4. C. R. Woese, D. H. Dugre, W. C. Saxinger & S. A. Dugre, "The molecular basis for the genetic code," PNAS 55 (1966), 966–974 — the polar-requirement correlation with the second codon position, of which this essay's middle-syllable law is a restatement in metrical terms.
  5. M. Olivier, M. Hollstein & P. Hainaut, "TP53 mutations in human cancers: origins, consequences, and clinical use," Cold Spring Harbor Perspectives in Biology 2 (2010), a001008 — codons 175, 248 and 273 as the canonical hotspots.
  6. Piṅgala, Chandaḥśāstra, with Halāyudha's Mṛtasañjīvanī — the gaṇa alphabet, whose middle syllable this essay weighs.

The Part One dictionary is re-derived from first principles at the start of the analysis and checked against the published file entry by entry (64 codons, both polarities) before anything else runs. Hydropathy tests use 20,000 label permutations over the 61 sense codons; alignment nulls shuffle the Vedic gaṇa stream (300 replicates); the Gāyatrī best-match null draws from the genes' own gaṇa frequencies (2,000 replicates). Every number in this essay was diffed against results6.json by a consistency script before publication.