The Language Before Language
Dr. Gustavo Caetano-Anollés on dipeptides, dual codes, and the emergence of biological meaning
One of the deepest questions in biology remains unresolved.
Life depends on a translation system that converts nucleotide sequences into proteins through the genetic code. Yet the origin of that code remains uncertain. Why does life rely on two molecular languages rather than one? Why this particular mapping between codons and amino acids? Why the ribosome?
For decades, the dominant explanation has been the RNA World hypothesis, which proposes that RNA preceded proteins and initially carried both informational and catalytic functions. It is an elegant framework, but mounting evidence suggests the story may be more complicated.
Gustavo Caetano-Anollés has approached the problem through phylogenomics, reconstructing evolutionary history from molecular structures preserved in modern organisms. Rather than asking what the first molecule was, he asks what surviving biological systems can reveal about the order in which life's fundamental components emerged.
In a 2025 study published in the Journal of Molecular Biology, his team analyzed 4.3 billion dipeptide sequences across 1,561 proteomes spanning Archaea, Bacteria, and Eukarya. The resulting chronology aligned with independent histories derived from protein domains and transfer RNA molecules. All three systems pointed to the same sequence of events during the emergence of the genetic code.
Among the most surprising findings was a striking symmetry: dipeptides and their reversed counterparts, known as anti-dipeptides, appeared synchronously throughout evolutionary history. The pattern suggests a deep connection between early coding systems and the bidirectional architecture of molecular information.
The findings raise questions that extend far beyond proteins. They touch on how biological meaning emerged, how information became linked to function, and when chemistry crossed the threshold into biology.
The Conversation
On the Origin of the Code
Gustavo Caetano-Anollés When three independently reconstructed evolutionary timelines converge on the same chronology, such agreement is unlikely to be a methodological coincidence. Instead, it signals something fundamental about the underlying structure of the evolutionary process itself. Our research found that the evolutionary histories of protein domains, transfer RNAs (tRNAs), aminoacyl-tRNA synthetases (aaRSs), and dipeptide repertoires all point to the same sequence of events during the emergence of the genetic code. These systems represent fundamentally different biological archives of structure, function, and sequence, yet they tell a remarkably consistent story.
Such congruence is difficult to reconcile with the strongest version of Francis Crick's "frozen accident" hypothesis, in which the genetic code is viewed as one of many possible solutions that happened to arise and subsequently became fixed. If the code had emerged through a genuinely contingent event, one would expect fewer signatures of constraint across independent molecular systems, and perhaps even conflicting evolutionary signals. Instead, the convergence suggests that the emergence of the genetic code was shaped by powerful historical constraints acting simultaneously on peptide structure, tRNA recognition, and amino acid assignment.
This does not imply strict determinism, nor does it suggest that the modern genetic code was the only possible outcome. Rather, once these systems became coupled, only a limited number of evolutionary trajectories remained accessible.
From a systems biology perspective, the genetic code is therefore best understood neither as an accident nor as an inevitability, but as a constrained attractor in evolutionary space: a stable evolutionary configuration toward which interacting molecular processes repeatedly converged under the constraints imposed by the underlying architecture of biology itself.
Gustavo Caetano-Anollés The word "mysteriously" is a deliberate scientific flag rather than rhetorical flourish. Even after reconstructing the chronology of genetic code emergence from dipeptide repertoires, tRNA evolution, and protein structural history, a fundamental question remains unresolved: what is the direction of causality linking the genetic code to the proteome? Our study shows that the chronology of genetic code recruitment mirrors the chronology of dipeptide appearance, revealing a surprisingly intimate connection between the evolution of coding and the evolution of protein composition. Yet the causal arrow remains unclear. Did emerging protein structures require particular amino acid combinations and thereby create selective pressures that drove the expansion of the code? Did innovations in coding and amino acid assignment permit the exploration of new dipeptide combinations and, consequently, new protein architectures? Or did both processes emerge together in a self-reinforcing feedback loop, with neither serving as the primary driver? It is even possible that both reflect a deeper physical principle governing molecular recognition, folding, catalysis, and information transfer.
The mystery lies in the fact that a seemingly local property of proteins — the frequencies and chronology of dipeptide sequences — contains information about one of biology's most global organizational systems, the genetic code itself. The code and the proteome appear to have co-evolved as mutually dependent systems, yet no obvious external reference point explains how one could have originated without the other. In this sense, the problem resembles the classic "chicken-and-egg" paradox of biological origins. The correlation is robust and reproducible; the underlying mechanism remains incomplete. Understanding how these two languages became coupled — and why life ultimately came to rely on both — remains one of the most profound unanswered questions in evolutionary biology.
Gustavo Caetano-Anollés Our findings align poorly with a strict "RNA World" scenario in which RNA achieved replication and catalytic sophistication before proteins entered the stage. Instead, they support a growing body of evidence that primitive aaRS-like catalysts participated very early in the emergence of amino acid charging and coding specificity, placing the origin of the genetic code within a coevolving RNA–peptide system. The evolutionary signal preserved in dipeptides, tRNA molecules, and protein domains is consistent with an early "operational" code centered on tRNA acceptor stems and peptide synthesis before the modern anticodon-based coding system was fully established. In this view, the first chapter of biochemistry is better represented by RNA ↔ peptides ↔ cofactors than by a simple RNA → proteins sequence.
The "first day" of life was therefore probably not a world dominated by a single self-sufficient molecule, but rather a network of mutually reinforcing chemical partners. Short peptides may have provided catalytic activities, RNA molecules supplied recognition, memory, and information transfer, while cofactors expanded the available chemical repertoire. Particularly intriguing is evidence that ancestral aaRS "urzymes" — minimal catalytic cores related to modern tyrosyl- and seryl-tRNA synthetases — may have linked aminoacylation to primitive peptide-synthesis activities. Such multifunctional catalysts would have blurred the modern distinction between information processing and protein synthesis. Rather than proteins appearing after the genetic code was established, coding and protein evolution may have emerged together through a process of mutual scaffolding: primitive peptides helping to stabilize and organize nucleic acids, and nucleic acids providing increasingly reliable templates for peptide production.
On Duality and Coevolution
Gustavo Caetano-Anollés The synchronous appearance of dipeptides and their reversed counterparts, anti-dipeptides, is among the most provocative findings of our study because it supports an ancestral duality consistent with bidirectional coding operating at the proteome level. Such simultaneity is unlikely to be trivial. In modern information systems, symmetry, redundancy, and parity relationships are often exploited to increase reliability and facilitate error detection. One should be cautious, however, about directly importing engineering concepts into early biology. Our data demonstrate synchrony, not purpose, and they do not establish that bidirectional coding evolved as an explicit error-correction mechanism. Nor do they imply design.
Nevertheless, the information-theoretic intuition is worth taking seriously. The early molecular world may have been organized around complementary relationships that linked sense and antisense strands, sequence and complement, and peptide and anti-dipeptide pairs. If sequences were generated, interpreted, or constrained in both directions, complementary molecular products could have imposed additional constraints on acceptable sequence variation, increasing robustness without requiring dedicated error-correction machinery. More generally, duality distributes information across related structures rather than concentrating it in a single sequence, making biological information less vulnerable to corruption. The deeper implication is that the first molecular language may have possessed an intrinsic symmetry arising from the chemistry of nucleic acids and peptides themselves. In that sense, robustness may not have been engineered or selected directly; it may have emerged naturally from the bidirectional and complementary architecture of the system. The redundancy was not designed into the code — it was embedded in the physics and chemistry from which the code emerged.
Gustavo Caetano-Anollés The weight of our research over several decades suggests that the bridge and its shores emerged together through a process of molecular coevolution, although peptide function may have played a particularly important role in shaping the earliest stages of the process. Phylogenomic analyses of the RNA and protein components of the ribosome reveal that the most ancient ribosomal RNA structures interacted with the most ancient ribosomal proteins, initiating a coordinated accretion process that ultimately produced the ribosomal core. This evolutionary pattern is inconsistent with a purely RNA-based ribosome later invaded by proteins, just as it is inconsistent with a protein machine that subsequently recruited RNA. Instead, both components grew together incrementally, each new structural innovation creating opportunities and constraints for the other. The ribosome was not built from one shore toward the other; the bridge and its shores co-emerged as a single evolving architecture.
The dipeptide findings strengthen this interpretation by revealing that the emergence of the genetic code was linked to the evolving structural requirements of proteins themselves. The chronology of dipeptide recruitment suggests that an early protein code was already responding to demands for folding, interaction, and function while an operational RNA-based code was simultaneously developing through tRNA recognition and amino acid assignment. The evolutionary sequence appears to have involved peptide-producing activities, operational coding through tRNA acceptor interactions, increasing aaRS specificity, the establishment of anticodon-based coding, and finally the emergence of increasingly sophisticated translation machinery. In this view, the ribosome is best understood as the fossilized record of a prolonged negotiation between RNA and proteins.
On Meaning, Function, and Stability
Gustavo Caetano-Anollés AaRSs are often described as the guardians of the genetic code because they enforce its most fundamental correspondence: the assignment of specific amino acids to specific tRNAs. This mapping constitutes the semantic core of translation. Every act of translation depends on the faithful conversion of nucleotide information into protein structure, and aaRS enzymes are responsible for maintaining that fidelity. What they guard against is mis-acylation — the attachment of the wrong amino acid to a tRNA. If, for example, phenylalanine is loaded onto a tRNA designated for serine, the resulting protein will contain an incorrect residue at a specific position. In some proteins the consequence may be minor, but in enzymes, structural proteins, or regulatory factors such errors can compromise folding, destroy catalytic activity, disrupt signaling networks, and ultimately reduce organismal fitness.
Without the fidelity imposed by aaRS enzymes, codons would lose their predictable meaning, proteins would become unreliable, and inheritance itself would lose continuity. The code must therefore be protected from charging errors, mistranslation, ambiguity, and the evolutionary drift that could erode the correspondence between nucleic acid sequences and protein products. Biologically, meaning is not symbolic but operational: a codon has meaning only insofar as it consistently produces the same biochemical consequence. Corruption occurs when that consistency breaks down. To prevent such failures, many aaRS enzymes have evolved sophisticated editing domains that proofread aminoacylated tRNAs and hydrolyze incorrectly charged molecules before they reach the ribosome. In effect, they function as molecular editors that preserve the integrity of the genetic message.
What is especially remarkable is that the guardians themselves are products of the very code they protect. The system is therefore self-referential: proteins generated by the genetic code are responsible for maintaining the fidelity of the coding relationships that allow those proteins to exist in the first place. This circular architecture is one of the deepest and most intriguing features of life's molecular logic. The genetic code does not merely store information; it actively preserves the conditions necessary for its own interpretation.
Gustavo Caetano-Anollés Looking across my research over several decades — from metabolic networks and protein domain evolution to ribosomal architecture, tRNA phylogenies, aminoacyl-tRNA synthetases, and now dipeptides — the same fundamental question has reappeared in different forms: How did molecular information become coupled to molecular function? Or, stated more simply, how did chemistry become coding? Why does life rely on a code at all, and why does that code possess the structure it does? These questions lie at the heart of every project I have pursued, whether the focus was metabolism, proteins, RNA, or the translation apparatus itself.
What makes the dipeptide work especially important is that it allows us to examine this coupling at its most elementary level. Earlier studies approached the problem through protein folds, metabolic organization, ribosomal evolution, or tRNA structure. Dipeptides bring us closer to the point where informational and functional worlds first intersect because they represent the smallest sequence units capable of carrying both structural and evolutionary signals. By integrating dipeptide chronologies with the histories of protein domains, tRNAs, aaRSs, and the genetic code, we can explore the entry of amino acids and codons into the coding system and trace the transition from an early operational code associated with the tRNA acceptor stem to the canonical code centered on the anticodon loop. In this sense, the work links protein architecture, coding systems, and evolutionary chronology within a single explanatory framework.
Does this bring us to the center of the problem? Perhaps. The remarkable congruence among these independent evolutionary histories suggests that we are approaching something fundamental. Yet the deeper mechanism underlying that congruence remains elusive. The origin of the genetic code may be one of those scientific questions that rewards asymptotic progress rather than definitive closure: each advance reveals a deeper level of organization beneath it. Whether we have reached the center or simply gained a clearer view of a center that recedes as we approach it, the dipeptide findings bring us closer to understanding how information, structure, and function became united in a single evolutionary process. That, ultimately, is the question I have been circling all along.
Gustavo Caetano-Anollés One of the most intriguing findings from our work is that signatures associated with protein thermostability appear relatively late in evolutionary history. Tracing the determinants of thermal adaptation suggests that highly thermostable proteins were not a primordial feature of life but rather a later development, consistent with the emergence of proteins in the comparatively mild environments that characterized much of the Archaean eon. This observation carries an important implication: the earliest proteins were probably not selected primarily for durability or resistance to extreme conditions. Instead, they appear to have been selected for functionality — for their ability to fold, interact, recognize molecular partners, and participate in emerging biochemical networks.
In this framework, flexibility rather than rigidity was the primitive condition. Early proteins needed to explore structural and functional space, engage promiscuously with diverse molecular partners, and support the emergence of increasingly sophisticated biochemical processes. Thermostability, by contrast, represents a specialized optimization. A molecule that remains perfectly stable but performs no useful function contributes little to evolutionary success. The first challenges facing life were more fundamental: achieving molecular recognition, catalysis, and reliable coding relationships. Before biological systems could optimize robustness, they first had to establish a dependable correspondence between sequence and function.
Information fidelity — the reliable translation of molecular structure into biological function — was likely more urgent than structural endurance. Only after coding systems, catalytic networks, and functional proteins became established would there have been strong selective pressure to optimize stability under increasingly diverse or demanding environmental conditions. In information-theoretic terms, communication preceded optimization. Life first needed a language capable of transmitting functional information; only later did it refine the infrastructure that made that language more durable. The evolutionary sequence therefore appears to have been function first, stability second — information before endurance.
On Language, Law, and Memory
Gustavo Caetano-Anollés The genetic code possesses characteristics of both a language and a law, and much of its evolutionary significance lies in the tension between these two aspects. Like a language, the code is symbolic, historically contingent, and capable of expansion. The evolutionary record suggests that amino acids entered the code in stages, that coding relationships expanded over time, and that an early operational code associated with the tRNA acceptor stem preceded the canonical code centered on the anticodon loop. In this sense, the genetic code resembles a living language that acquired new "vocabulary" and increasing semantic complexity as biological systems evolved.
At the same time, the code possesses a profoundly law-like character. Its deep architecture is remarkably conserved across all known life, deviations are rare, and the cost of change increases dramatically as biological complexity accumulates. The emerging picture from molecular chronologies is that the code became embedded within multiple layers of molecular organization. Once these layers coevolved, the code ceased to be a freely modifiable convention and became constrained by billions of years of evolutionary history. The new evidence therefore strengthens the law-like interpretation, not because the code was fixed from the outset, but because its present structure reflects a vast accumulation of historical and structural constraints.
For synthetic biology, this distinction is crucial. If the code were merely a law, rewriting it would amount to overriding a deeply entrenched biological rule. If it were merely a language, expansion and revision would be relatively straightforward. In reality, it is both. The successful incorporation of noncanonical amino acids through engineered orthogonal tRNA–synthetase pairs demonstrates that the code can acquire new "words" and expand its expressive capacity. Yet the near universality of the twenty standard amino acids and their canonical assignments reveal a level of conservation more characteristic of natural law than linguistic convention. The lesson is not that the code cannot be rewritten, but that successful redesign must respect the evolutionary logic already built into the system. The genetic code is best viewed as a language whose grammar has become so deeply embedded in the fabric of life that it now behaves like a law.
Gustavo Caetano-Anollés If I could ask a single question of the first molecular system that functioned simultaneously as gene and catalyst — before the modern distinction between nucleic acid information and protein function emerged — it would not be: How did you replicate? How did you catalyze reactions? Or how did you encode information? Those are important questions, but they concern mechanisms rather than origins. Instead, I would ask: What property allowed you to persist while all competing chemistries disappeared? At what point did a collection of molecules cease being merely chemistry and become a historical entity capable of transmitting its organization through time?
This question lies at the heart of our research because the congruence between dipeptide chronologies, tRNA evolution, protein domain histories, aminoacyl-tRNA synthetases, and coding systems suggests that life's emergence was not simply a matter of replication or catalysis appearing in isolation. The evidence increasingly points toward a self-reinforcing network in which information, structure, and function became inseparable. The amino acid and dipeptide compositions of the earliest proteins indicate that coding emerged through coevolutionary interactions between polypeptides and nucleic acid cofactors, with selective pressures favoring molecular flexibility, folding, and functional integration. In this sense, the first molecular systems were already responding to structural constraints long before the modern genetic code was fully established.
The deeper issue is not whether those molecules "knew" what they were doing, but when and how chemistry acquired something resembling agency. At what point did molecular interactions begin to preferentially preserve certain organizations because they were better at maintaining themselves? That transition — from transient chemical processes to systems capable of retaining information, accumulating history, and shaping their own evolutionary future — is the threshold these studies ultimately seek to understand. The origin of life may therefore not have been the invention of replication alone, catalysis alone, or even coding alone, but the emergence of a coupled molecular system in which information, structure, and function became mutually reinforcing. Once that coupling appeared, evolution acquired memory, and chemistry became biology.
Conclusion
There is a moment in this conversation when the scientific and the philosophical become difficult to separate.
Asked what he would say to the first molecular system that functioned as both gene and catalyst, Caetano-Anollés declines the obvious questions. He does not ask how it replicated, catalyzed reactions, or encoded information. Instead, he asks what allowed it to persist while competing chemistries disappeared.
That question reaches beyond mechanism. It asks when chemistry became history.
Across decades of work on metabolism, protein evolution, ribosomes, transfer RNA, aminoacyl-tRNA synthetases, and now dipeptides, Caetano-Anollés has pursued a single problem: how information became coupled to function. The remarkable agreement among multiple molecular archives does not solve that problem, but it narrows the possibilities. Independent histories preserved in proteins, tRNAs, and dipeptides all point toward the same underlying narrative.
The emerging picture is neither pure accident nor strict inevitability. Life appears to have developed within a landscape of powerful constraints, where certain paths became increasingly favored as information, structure, and function grew more tightly linked.
The genetic code, as Caetano-Anollés describes it, is both language and law. It carries the marks of history, yet behaves with the stability of a physical principle. Somewhere in that tension lies one of biology's deepest mysteries.
Perhaps the most profound implication of this work is that life did not begin when molecules first replicated. It began when information, structure, and function became inseparable, allowing evolution to accumulate memory.
That was the moment chemistry learned to remember.