Sanskrit · Research

The Meter and the Molecule

Part Three: the dialects of life

A dictionary was built in Part One, and in Part Two it was spent on a single gene, which declined to contain hidden poetry but agreed to reveal its tempo. This essay widens the lens as far as it will go. In 1980, Richard Grantham examined how differently organisms choose among synonymous codons and proposed what he called the genome hypothesis: that every genome has a consistent, characteristic way of speaking — his word for it was dialect.4 The metaphor was linguistic from the start. We are simply taking it one step further, to where it was perhaps always pointing: if each genome speaks a dialect, then under the codon→gaṇa dictionary each genome scans in one — and the dialects of life turn out to be arranged along a single metrical axis, with a malaria parasite whispering at one end and a soil bacterium tolling like a bell at the other.

Eight speakers

The data are codon-usage tables: for a given organism, how often each of the sixty-four codons appears across its sequenced protein-coding DNA. The canonical source is the Kazusa CUTG database, compiled from GenBank by Nakamura, Gojobori, and Ikemura,3 and from it I took eight speakers chosen to span the living world: three animals (human, fruit fly, nematode), a plant (thale cress), a fungus (baker's yeast), a protist (the malaria parasite Plasmodium falciparum, famously the most A·T-rich genome ever sequenced9), and two bacteria chosen as opposite poles — E. coli and the antibiotic-making soil bacterium Streptomyces coelicolor, whose genome is among the most G·C-rich known.10 Eight species for eight gaṇas; together their tables summarize 118.5 million codons — over 355 million bases — of curated coding sequence.

The verification habit of this series continues, adapted to tabular data. Each Kazusa table states both a frequency per thousand and a raw count for every codon, plus the total codons behind it; the three quantities must reproduce one another exactly, so the table carries its own checksum. All eight tables were required to pass — sixty-four codons present, counts summing to the stated totals to the last digit, every per-mille value regenerating from its count within rounding — and the human table was fetched twice in different renderings and required to agree. The complete tables, parser, and analysis are in the download box, and every number in this essay is generated by that code.

Eight dialects, eight gaṇas

Figure 1 is the census: for each species, the share of its codons falling in each gaṇa under the S/W polarity (strong G·C as guru). Read it as a table of accents. Plasmodium is an extremist: 44.0% of all its codons are na — the tribrach, three lights, the emptiest foot in the system — and its top two rows of the dictionary (na and bha) hold nearly two-thirds of its entire genome's speech. At the other end Streptomyces has all but abolished na (0.5%) and speaks two-thirds of its codons in just ma and ra, the heaviest feet. The human dialect sits in the crowded middle, led by ra.

eight dialects, eight gaṇas — share of each species’ codons falling in each gaṇa (S/W)mayarasatajabhanaP. falciparum1.13.23.79.36.311.520.844.01.24S. cerevisiae5.17.89.815.511.312.418.419.81.40C. elegans5.67.412.014.714.311.718.116.21.43A. thaliana6.18.113.015.214.212.117.513.71.45E. coli15.79.418.212.810.84.814.414.01.52H. sapiens12.210.421.614.310.79.211.210.31.52D. melanogaster13.912.221.316.99.65.810.99.31.54S. coelicolor33.713.133.312.83.61.02.10.51.72temporows sorted by tempo (mātrās per base); values are percentages of codons
Figure 1. The census of dialects: percentage of each species' codons falling in each gaṇa under S/W (strong G·C = guru), rows sorted by tempo. Malaria's speech is nearly half na; Streptomyces speaks two-thirds of its codons in ma and ra.

And the leading feet form a ladder. Sorted by overall weight, the dominant gaṇa walks through na (L L L) → bha (G L L) → ra (G L G) → ma (G G G): heaviness is added first to the codon's first syllable, then to its third, and only at the far extreme to its middle. That ordering is no accident, and its reason is the heart of this essay's third section — the middle syllable is where the genetic code keeps meaning, so it is the last place a dialect is free to redecorate.

The tempo scale of life

Collapse each dialect to a single number — mātrās per base, the tempo of Part Two — and the eight species arrange themselves along one axis (Figure 2). Under S/W the scale runs from Plasmodium's 1.24 to Streptomyces' 1.72, with humanity at 1.52. The same hundred-codon protein, spoken aloud at one mātrā per beat, would take about 371 beats in malaria's dialect and 517 in Streptomyces' — a forty-percent difference in duration for the same semantic content. Under R/Y, meanwhile, the whole living world crowds between 1.49 and 1.59: purines and pyrimidines stay near balance in coding sequence everywhere, the cross-species echo of the strand-symmetry Part Two met in a single gene.2

the tempo scale of life — mātrās per base of coding sequenceS/W polarity (guru = G·C)R/Y (guru = purine)1.21.31.41.51.61.7P. falciparum1.24S. cerevisiae1.40C. elegans1.43A. thaliana1.45E. coli1.52H. sapiens1.52D. melanogaster1.54S. coelicolor1.72lightest: nearly half of malaria’s codons are na (L L L)heaviest
Figure 2. The tempo scale of life: mātrās per base of coding sequence. Under S/W (blue) the dialects spread from 1.24 to 1.72; under R/Y (orange) all eight crowd between 1.49 and 1.59 — purine balance is conserved where G·C content is free.

Because tempo under S/W is G·C content in metrical dress, this figure is not a discovery; it is the famous compositional spectrum of genomes5 translated into the tradition's vocabulary. What the translation contributes is, again, aptness. G·C pairs are the strongly bound, slow-melting pairs: the S/W guru is heavy in the thermodynamic sense before it is heavy in any poetic one. A Streptomyces gene really is a heavier, more tightly held text than a Plasmodium gene — and the meter lens states that plainly, in beats.

The free syllable, everywhere

Now split each dialect's weight by codon position (Figure 3), and the series' oldest theme returns at its grandest scale. Across the eight species, the strong-base fraction of the codon's middle position varies over a range of just 0.29 — every genome, from malaria to Streptomyces, holds it in a narrow band. The first position wanders more (range 0.41). But the third position — wobble's seat, the anceps syllable of Part Two — spans 0.17 to 0.93: a range of 0.76, two and a half times the middle position's. Streptomyces writes essentially every third syllable heavy (93%); Plasmodium writes five of six light. They are spelling the same amino acids with opposite prosody, using precisely the freedom the code's degeneracy grants.58

the free syllable, across the tree of life — strong-base (G·C) fraction by codon positionposition 1spread 0.41S. coelicolor 0.73P. falciparum 0.32position 2spread 0.29S. coelicolor 0.51P. falciparum 0.22position 3spread 0.76S. coelicolor 0.93P. falciparum 0.170.000.250.500.751.00position 3 spans 0.17–0.93
Figure 3. Anceps at planetary scale: strong-base fraction by codon position across the eight species. The meaning-bearing middle position spans just 0.29; the degenerate third position spans 0.76 — from 0.17 in malaria to 0.93 in Streptomyces.

This is the anceps principle made planetary. In a Sanskrit verse, the pāda-final syllable is left metrically free, and so it is exactly there that different words, different recitations, different regional habits can live inside one meter. In the genetic code, the third position is left semantically free, and so it is exactly there that two billion years of divergent mutational pressure has been absorbed — while the meaning-bearing middle position stays nearly uniform across the tree of life. Same engineering, same address: put the tolerance at the boundary of the unit, and let the dialects bloom there.

Testing the primeval foot

The gaṇa lens also lets us re-ask an old and beautiful question. In 1981 J.C.W. Shepherd, following ideas of Eigen and Schuster about the origin of the code, argued that the earliest genes may have been written in codons of the form R·N·Y — purine, anything, pyrimidine — a primeval pattern whose traces might still be readable in modern genes.67 In dictionary terms, R·N·Y is a claim that life once favored two specific feet — ta and bha in the R/Y polarity — and the codon tables let us ask whether any preference survives today beyond what each position's composition already dictates: compare observed R·N·Y frequency with the product of first-position purine and third-position pyrimidine fractions.

is there a residual “primeval foot”? RNY occurrence vs composition-only expectationratio of observed R•N•Y codons to what first- and third-position composition alone predicts0.900.951.001.051.10no excessP. falciparum0.94S. cerevisiae1.01C. elegans0.97A. thaliana0.94E. coli1.05H. sapiens0.96D. melanogaster1.02S. coelicolor1.06all eight sit within ±6% of expectation: modern codon usage keeps no RNY excess beyond composition
Figure 4. No residual primeval foot: observed R·N·Y codon share divided by the product of first-position purine and third-position pyrimidine fractions. All eight species sit within ±6% of no-excess, on both sides of the line.

The answer is a firm and useful no. All eight species sit within ±6% of the no-excess line (0.94 to 1.06), scattered on both sides of it. Whatever R·N·Y architecture the first genes had — and Shepherd's evidence was about subtle periodicities, detected by more delicate instruments than this one — modern codon usage keeps no gross surplus of the primeval foot once composition is accounted for. After Part Two's Sragviṇī, this series owes its readers exactly this kind of result, reported at exactly this volume: a clean answer to a clean question, with the romance declined.

Filling the envelope

One more translation, because it closes a circle opened in Part One. Under S/W a codon lasts three, four, five, or six mātrās, and Part One's Virahāṅka combinatorics counts the weight-patterns available at each duration: one, three, three, one. That is the envelope of the possible. Figure 5 shows how three species fill it: Plasmodium crowds the short durations (44% of its codons take only three mātrās), humanity sits balanced across the middle, and Streptomyces piles into five and six — filling half its speech with five-beat feet and giving the six-beat ma a full third of everything it says. The envelope is fixed by arithmetic; where a genome lives inside it is dialect. Between malaria and the soil bacterium, life occupies essentially the entire habitable range of the codon's little duration-space.

how life fills the mātrā envelope — codon durations in mātrās (S/W)dashed outline = the 1-3-3-1 space of possible weight patterns per duration (Part One)P. falciparum44%342%413%51%6(lightest)H. sapiens10%335%443%512%6(middle)S. coelicolor0%316%450%534%6(heaviest)
Figure 5. The mātrā envelope (dashed: the 1-3-3-1 count of possible weight patterns per duration, from Part One's Virahāṅka arithmetic) and how three dialects fill it — malaria crowding the short feet, humanity balanced, Streptomyces living in the five- and six-beat feet.

What the trilogy holds

Three essays in, the ledger balances cleanly. The mathematics is exact and old: a binary prosody enumerated by Piṅgala, a triplet reading shared with the ribosome, a de Bruijn mnemonic that anticipates the assembler, Fibonacci counts born from mātrā verse. The empiricism is sober and reproducible: no ślokas hide in genes, no meter survives multiple-testing correction, no primeval foot outlasts its own composition — and yet the tempo of a gene tracks its function, the anceps position is measurably the code's free syllable within a gene, across synonymous choices, and now across the whole tree of life, and every genome speaks a dialect that a prosodist could place on one scale by ear. The dictionary never found poems in the molecule. It found something better: that where the metrical language fits biology, it fits at exactly the joints biology itself considers deepest — weight where the bonds are strong, freedom where the meaning ends.

On the series' horizon: the promised chromosome-scale tempo maps and the repeat-length spectra against Virahāṅka's counts remain open — they want whole-genome data rather than summary tables, and they will get their own essay if and when they earn it. The corpus below is, as always, yours first.

Downloads for this essay

Dialect bundle (ZIP: tables, code, results)
All eight verified codon-usage tables with provenance, the checksum-verifying parser, the full analysis, results3.json, dialect-metrics CSV, and the figure generator. Python standard library only.

Part One · Part Two · Structural Reference (PDF) · Catalog & Dictionary (PDF)
The foundations this essay stands on.

References

  1. R. Oshop, "The Meter and the Molecule — Part One," AyurAstro (2026) — the gaṇa system and the codon→gaṇa dictionary.
  2. R. Oshop, "The Meter and the Molecule — Part Two," AyurAstro (2026) — tempo, the śloka nulls, and the anceps measurement within a single gene.
  3. Y. Nakamura, T. Gojobori & T. Ikemura, "Codon usage tabulated from international DNA sequence databases: status for the year 2000," Nucleic Acids Research 28 (2000), 292; the Kazusa CUTG database. Tables used: gbpri/gbinv/gbpln/gbbct entries for the eight species, identifiers in the download bundle.
  4. R. Grantham, C. Gautier, M. Gouy, R. Mercier & A. Pavé, "Codon catalog usage and the genome hypothesis," Nucleic Acids Research 8 (1980), r49–r62 — the origin of the "dialect" framing.
  5. A. Muto & S. Osawa, "The guanine and cytosine content of genomic DNA and bacterial evolution," PNAS 84 (1987), 166–169 — positional G·C constraint, loosest at codon position three.
  6. J. C. W. Shepherd, "Method to determine the reading frame of a protein from the purine/pyrimidine genome sequence and its possible evolutionary justification," PNAS 78 (1981), 1596–1600 — the R·N·Y primeval-message hypothesis.
  7. M. Eigen & P. Schuster, The Hypercycle: A Principle of Natural Self-Organization (Springer, 1979) — the theoretical setting for an R·N·Y first code.
  8. T. Ikemura, "Codon usage and tRNA content in unicellular and multicellular organisms," Molecular Biology and Evolution 2 (1985), 13–34 — how dialects are enforced by the tRNA pool.
  9. M. J. Gardner et al., "Genome sequence of the human malaria parasite Plasmodium falciparum," Nature 419 (2002), 498–511 — the A·T-extreme genome.
  10. S. D. Bentley et al., "Complete genome sequence of the model actinomycete Streptomyces coelicolor A3(2)," Nature 417 (2002), 141–147 — the G·C-extreme genome.

Every codon-usage table passed an internal checksum (counts reproduce the stated totals and per-mille values exactly, all sixty-four codons present) before analysis; the human table was additionally re-fetched in a second rendering and matched value-for-value. All metrics derive from exact codon probabilities (raw counts over stated totals), and a consistency script diffed this essay's prose against results3.json before publication. Seed-free: Part Three's numbers are deterministic arithmetic, with no Monte Carlo required.