The Language Before Language

Professor Gustavo Caetano-Anollés in his laboratory
Prof. Gustavo Caetano-Anollés in his laboratory · © University of Illinois

Dr. Gustavo Caetano-Anollés on dipeptides, dual codes, and the emergence of biological meaning

One of the deepest questions in biology remains unresolved.

Life depends on a translation system that converts nucleotide sequences into proteins through the genetic code. Yet the origin of that code remains uncertain. Why does life rely on two molecular languages rather than one? Why this particular mapping between codons and amino acids? Why the ribosome?

For decades, the dominant explanation has been the RNA World hypothesis, which proposes that RNA preceded proteins and initially carried both informational and catalytic functions. It is an elegant framework, but mounting evidence suggests the story may be more complicated.

Gustavo Caetano-Anollés has approached the problem through phylogenomics, reconstructing evolutionary history from molecular structures preserved in modern organisms. Rather than asking what the first molecule was, he asks what surviving biological systems can reveal about the order in which life's fundamental components emerged.

In a 2025 study published in the Journal of Molecular Biology, his team analyzed 4.3 billion dipeptide sequences across 1,561 proteomes spanning Archaea, Bacteria, and Eukarya. The resulting chronology aligned with independent histories derived from protein domains and transfer RNA molecules. All three systems pointed to the same sequence of events during the emergence of the genetic code.

Among the most surprising findings was a striking symmetry: dipeptides and their reversed counterparts, known as anti-dipeptides, appeared synchronously throughout evolutionary history. The pattern suggests a deep connection between early coding systems and the bidirectional architecture of molecular information.

The findings raise questions that extend far beyond proteins. They touch on how biological meaning emerged, how information became linked to function, and when chemistry crossed the threshold into biology.

Diagram showing the evolution of transfer RNA
How transfer RNA evolved over billions of years, growing arm by arm from a simple ancestral structure into the molecule that translates genetic information in every living cell today. · © Gustavo Caetano-Anollés

The Conversation

On the Origin of the Code

Our NarrativesYour study found that dipeptide sequences, protein domains, and tRNA evolutionary histories are all congruent with each other. When three independently derived timelines align so precisely, what does that tell us about the nature of the genetic code itself? Do these results suggest the code was a biochemical inevitability or a contingent "frozen accident" that happened once and stuck?

Gustavo Caetano-Anollés When three independently reconstructed evolutionary timelines converge on the same chronology, such agreement is unlikely to be a methodological coincidence. Instead, it signals something fundamental about the underlying structure of the evolutionary process itself. Our research found that the evolutionary histories of protein domains, transfer RNAs (tRNAs), aminoacyl-tRNA synthetases (aaRSs), and dipeptide repertoires all point to the same sequence of events during the emergence of the genetic code. These systems represent fundamentally different biological archives of structure, function, and sequence, yet they tell a remarkably consistent story.

Such congruence is difficult to reconcile with the strongest version of Francis Crick's "frozen accident" hypothesis, in which the genetic code is viewed as one of many possible solutions that happened to arise and subsequently became fixed. If the code had emerged through a genuinely contingent event, one would expect fewer signatures of constraint across independent molecular systems, and perhaps even conflicting evolutionary signals. Instead, the convergence suggests that the emergence of the genetic code was shaped by powerful historical constraints acting simultaneously on peptide structure, tRNA recognition, and amino acid assignment.

This does not imply strict determinism, nor does it suggest that the modern genetic code was the only possible outcome. Rather, once these systems became coupled, only a limited number of evolutionary trajectories remained accessible.

From a systems biology perspective, the genetic code is therefore best understood neither as an accident nor as an inevitability, but as a constrained attractor in evolutionary space: a stable evolutionary configuration toward which interacting molecular processes repeatedly converged under the constraints imposed by the underlying architecture of biology itself.


Our NarrativesYou describe the origin of the genetic code as "mysteriously linked" to dipeptide composition. The word "mysteriously" is evocative and unusual in scientific writing. What remains genuinely unexplained or "mysterious" to you about this link, even after completing this extensive research?

Gustavo Caetano-Anollés The word "mysteriously" is a deliberate scientific flag rather than rhetorical flourish. Even after reconstructing the chronology of genetic code emergence from dipeptide repertoires, tRNA evolution, and protein structural history, a fundamental question remains unresolved: what is the direction of causality linking the genetic code to the proteome? Our study shows that the chronology of genetic code recruitment mirrors the chronology of dipeptide appearance, revealing a surprisingly intimate connection between the evolution of coding and the evolution of protein composition. Yet the causal arrow remains unclear. Did emerging protein structures require particular amino acid combinations and thereby create selective pressures that drove the expansion of the code? Did innovations in coding and amino acid assignment permit the exploration of new dipeptide combinations and, consequently, new protein architectures? Or did both processes emerge together in a self-reinforcing feedback loop, with neither serving as the primary driver? It is even possible that both reflect a deeper physical principle governing molecular recognition, folding, catalysis, and information transfer.

The mystery lies in the fact that a seemingly local property of proteins — the frequencies and chronology of dipeptide sequences — contains information about one of biology's most global organizational systems, the genetic code itself. The code and the proteome appear to have co-evolved as mutually dependent systems, yet no obvious external reference point explains how one could have originated without the other. In this sense, the problem resembles the classic "chicken-and-egg" paradox of biological origins. The correlation is robust and reproducible; the underlying mechanism remains incomplete. Understanding how these two languages became coupled — and why life ultimately came to rely on both — remains one of the most profound unanswered questions in evolutionary biology.


Our NarrativesYour work challenges the RNA World hypothesis — the dominant idea that RNA preceded proteins. If we move away from that framework, how do you describe the "first day" of biochemical life? Did RNA and proteins emerge as a single, co-evolving system rather than a linear sequence?

Gustavo Caetano-Anollés Our findings align poorly with a strict "RNA World" scenario in which RNA achieved replication and catalytic sophistication before proteins entered the stage. Instead, they support a growing body of evidence that primitive aaRS-like catalysts participated very early in the emergence of amino acid charging and coding specificity, placing the origin of the genetic code within a coevolving RNA–peptide system. The evolutionary signal preserved in dipeptides, tRNA molecules, and protein domains is consistent with an early "operational" code centered on tRNA acceptor stems and peptide synthesis before the modern anticodon-based coding system was fully established. In this view, the first chapter of biochemistry is better represented by RNA ↔ peptides ↔ cofactors than by a simple RNA → proteins sequence.

The "first day" of life was therefore probably not a world dominated by a single self-sufficient molecule, but rather a network of mutually reinforcing chemical partners. Short peptides may have provided catalytic activities, RNA molecules supplied recognition, memory, and information transfer, while cofactors expanded the available chemical repertoire. Particularly intriguing is evidence that ancestral aaRS "urzymes" — minimal catalytic cores related to modern tyrosyl- and seryl-tRNA synthetases — may have linked aminoacylation to primitive peptide-synthesis activities. Such multifunctional catalysts would have blurred the modern distinction between information processing and protein synthesis. Rather than proteins appearing after the genetic code was established, coding and protein evolution may have emerged together through a process of mutual scaffolding: primitive peptides helping to stabilize and organize nucleic acids, and nucleic acids providing increasingly reliable templates for peptide production.


On Duality and Coevolution

Our NarrativesThe study reveals that dipeptides and their mirror reversals, "anti-dipeptides," appeared synchronously. In information theory, symmetry is often used for error correction. Does this duality suggest that the logic of early life was designed to protect its molecular language from corruption from the very beginning?

Gustavo Caetano-Anollés The synchronous appearance of dipeptides and their reversed counterparts, anti-dipeptides, is among the most provocative findings of our study because it supports an ancestral duality consistent with bidirectional coding operating at the proteome level. Such simultaneity is unlikely to be trivial. In modern information systems, symmetry, redundancy, and parity relationships are often exploited to increase reliability and facilitate error detection. One should be cautious, however, about directly importing engineering concepts into early biology. Our data demonstrate synchrony, not purpose, and they do not establish that bidirectional coding evolved as an explicit error-correction mechanism. Nor do they imply design.

Nevertheless, the information-theoretic intuition is worth taking seriously. The early molecular world may have been organized around complementary relationships that linked sense and antisense strands, sequence and complement, and peptide and anti-dipeptide pairs. If sequences were generated, interpreted, or constrained in both directions, complementary molecular products could have imposed additional constraints on acceptable sequence variation, increasing robustness without requiring dedicated error-correction machinery. More generally, duality distributes information across related structures rather than concentrating it in a single sequence, making biological information less vulnerable to corruption. The deeper implication is that the first molecular language may have possessed an intrinsic symmetry arising from the chemistry of nucleic acids and peptides themselves. In that sense, robustness may not have been engineered or selected directly; it may have emerged naturally from the bidirectional and complementary architecture of the system. The redundancy was not designed into the code — it was embedded in the physics and chemistry from which the code emerged.


Our NarrativesYou frame the ribosome as the bridge between two codes: the genetic and the protein. If the ribosome is indeed the bridge, do your findings suggest it was built from the protein side, the RNA side, or did the bridge and the two shores emerge simultaneously as a single functional architecture?

Gustavo Caetano-Anollés The weight of our research over several decades suggests that the bridge and its shores emerged together through a process of molecular coevolution, although peptide function may have played a particularly important role in shaping the earliest stages of the process. Phylogenomic analyses of the RNA and protein components of the ribosome reveal that the most ancient ribosomal RNA structures interacted with the most ancient ribosomal proteins, initiating a coordinated accretion process that ultimately produced the ribosomal core. This evolutionary pattern is inconsistent with a purely RNA-based ribosome later invaded by proteins, just as it is inconsistent with a protein machine that subsequently recruited RNA. Instead, both components grew together incrementally, each new structural innovation creating opportunities and constraints for the other. The ribosome was not built from one shore toward the other; the bridge and its shores co-emerged as a single evolving architecture.

The dipeptide findings strengthen this interpretation by revealing that the emergence of the genetic code was linked to the evolving structural requirements of proteins themselves. The chronology of dipeptide recruitment suggests that an early protein code was already responding to demands for folding, interaction, and function while an operational RNA-based code was simultaneously developing through tRNA recognition and amino acid assignment. The evolutionary sequence appears to have involved peptide-producing activities, operational coding through tRNA acceptor interactions, increasing aaRS specificity, the establishment of anticodon-based coding, and finally the emergence of increasingly sophisticated translation machinery. In this view, the ribosome is best understood as the fossilized record of a prolonged negotiation between RNA and proteins.


On Meaning, Function, and Stability

Our NarrativesYou describe aminoacyl-tRNA synthetases as "guardians of the genetic code." That is a striking metaphor. In a biological sense, what does the code need protection from? What would it mean for the "meaning" of life's molecular language to be violated or corrupted?

Gustavo Caetano-Anollés AaRSs are often described as the guardians of the genetic code because they enforce its most fundamental correspondence: the assignment of specific amino acids to specific tRNAs. This mapping constitutes the semantic core of translation. Every act of translation depends on the faithful conversion of nucleotide information into protein structure, and aaRS enzymes are responsible for maintaining that fidelity. What they guard against is mis-acylation — the attachment of the wrong amino acid to a tRNA. If, for example, phenylalanine is loaded onto a tRNA designated for serine, the resulting protein will contain an incorrect residue at a specific position. In some proteins the consequence may be minor, but in enzymes, structural proteins, or regulatory factors such errors can compromise folding, destroy catalytic activity, disrupt signaling networks, and ultimately reduce organismal fitness.

Without the fidelity imposed by aaRS enzymes, codons would lose their predictable meaning, proteins would become unreliable, and inheritance itself would lose continuity. The code must therefore be protected from charging errors, mistranslation, ambiguity, and the evolutionary drift that could erode the correspondence between nucleic acid sequences and protein products. Biologically, meaning is not symbolic but operational: a codon has meaning only insofar as it consistently produces the same biochemical consequence. Corruption occurs when that consistency breaks down. To prevent such failures, many aaRS enzymes have evolved sophisticated editing domains that proofread aminoacylated tRNAs and hydrolyze incorrectly charged molecules before they reach the ribosome. In effect, they function as molecular editors that preserve the integrity of the genetic message.

What is especially remarkable is that the guardians themselves are products of the very code they protect. The system is therefore self-referential: proteins generated by the genetic code are responsible for maintaining the fidelity of the coding relationships that allow those proteins to exist in the first place. This circular architecture is one of the deepest and most intriguing features of life's molecular logic. The genetic code does not merely store information; it actively preserves the conditions necessary for its own interpretation.


Our NarrativesYour research has spanned decades, moving from metabolic networks to protein folds to tRNA and now to dipeptides. What is the central, singular question you have been circling all this time, and do you feel the dipeptide sequences have finally brought you to the center of it?

Gustavo Caetano-Anollés Looking across my research over several decades — from metabolic networks and protein domain evolution to ribosomal architecture, tRNA phylogenies, aminoacyl-tRNA synthetases, and now dipeptides — the same fundamental question has reappeared in different forms: How did molecular information become coupled to molecular function? Or, stated more simply, how did chemistry become coding? Why does life rely on a code at all, and why does that code possess the structure it does? These questions lie at the heart of every project I have pursued, whether the focus was metabolism, proteins, RNA, or the translation apparatus itself.

What makes the dipeptide work especially important is that it allows us to examine this coupling at its most elementary level. Earlier studies approached the problem through protein folds, metabolic organization, ribosomal evolution, or tRNA structure. Dipeptides bring us closer to the point where informational and functional worlds first intersect because they represent the smallest sequence units capable of carrying both structural and evolutionary signals. By integrating dipeptide chronologies with the histories of protein domains, tRNAs, aaRSs, and the genetic code, we can explore the entry of amino acids and codons into the coding system and trace the transition from an early operational code associated with the tRNA acceptor stem to the canonical code centered on the anticodon loop. In this sense, the work links protein architecture, coding systems, and evolutionary chronology within a single explanatory framework.

Does this bring us to the center of the problem? Perhaps. The remarkable congruence among these independent evolutionary histories suggests that we are approaching something fundamental. Yet the deeper mechanism underlying that congruence remains elusive. The origin of the genetic code may be one of those scientific questions that rewards asymptotic progress rather than definitive closure: each advance reveals a deeper level of organization beneath it. Whether we have reached the center or simply gained a clearer view of a center that recedes as we approach it, the dipeptide findings bring us closer to understanding how information, structure, and function became united in a single evolutionary process. That, ultimately, is the question I have been circling all along.


Our NarrativesProtein thermostability emerged relatively late in your evolutionary timeline, long after the genetic code was established. Why would life have waited so long to prioritize structural stability? What does this delay reveal about the priorities of early biology — was "information" more urgent than "durability"?

Gustavo Caetano-Anollés One of the most intriguing findings from our work is that signatures associated with protein thermostability appear relatively late in evolutionary history. Tracing the determinants of thermal adaptation suggests that highly thermostable proteins were not a primordial feature of life but rather a later development, consistent with the emergence of proteins in the comparatively mild environments that characterized much of the Archaean eon. This observation carries an important implication: the earliest proteins were probably not selected primarily for durability or resistance to extreme conditions. Instead, they appear to have been selected for functionality — for their ability to fold, interact, recognize molecular partners, and participate in emerging biochemical networks.

In this framework, flexibility rather than rigidity was the primitive condition. Early proteins needed to explore structural and functional space, engage promiscuously with diverse molecular partners, and support the emergence of increasingly sophisticated biochemical processes. Thermostability, by contrast, represents a specialized optimization. A molecule that remains perfectly stable but performs no useful function contributes little to evolutionary success. The first challenges facing life were more fundamental: achieving molecular recognition, catalysis, and reliable coding relationships. Before biological systems could optimize robustness, they first had to establish a dependable correspondence between sequence and function.

Information fidelity — the reliable translation of molecular structure into biological function — was likely more urgent than structural endurance. Only after coding systems, catalytic networks, and functional proteins became established would there have been strong selective pressure to optimize stability under increasingly diverse or demanding environmental conditions. In information-theoretic terms, communication preceded optimization. Life first needed a language capable of transmitting functional information; only later did it refine the infrastructure that made that language more durable. The evolutionary sequence therefore appears to have been function first, stability second — information before endurance.

Scientific illustration related to the evolution of the genetic code
The genetic code mapped as a spiral through time, showing how the system of codons and amino acids expanded outward from a single ancient origin. Orange traces the older operational code; blue traces the standard code that followed. · © Gustavo Caetano-Anollés

On Language, Law, and Memory

Our NarrativesThe genetic code is described as resistant to change, constrained by its own deep history. In your view, does this make the code more like a language, which evolves and adapts, or more like a law, which persists? What are the implications of this distinction for synthetic biology's ambitions to rewrite or "hack" the code?

Gustavo Caetano-Anollés The genetic code possesses characteristics of both a language and a law, and much of its evolutionary significance lies in the tension between these two aspects. Like a language, the code is symbolic, historically contingent, and capable of expansion. The evolutionary record suggests that amino acids entered the code in stages, that coding relationships expanded over time, and that an early operational code associated with the tRNA acceptor stem preceded the canonical code centered on the anticodon loop. In this sense, the genetic code resembles a living language that acquired new "vocabulary" and increasing semantic complexity as biological systems evolved.

At the same time, the code possesses a profoundly law-like character. Its deep architecture is remarkably conserved across all known life, deviations are rare, and the cost of change increases dramatically as biological complexity accumulates. The emerging picture from molecular chronologies is that the code became embedded within multiple layers of molecular organization. Once these layers coevolved, the code ceased to be a freely modifiable convention and became constrained by billions of years of evolutionary history. The new evidence therefore strengthens the law-like interpretation, not because the code was fixed from the outset, but because its present structure reflects a vast accumulation of historical and structural constraints.

For synthetic biology, this distinction is crucial. If the code were merely a law, rewriting it would amount to overriding a deeply entrenched biological rule. If it were merely a language, expansion and revision would be relatively straightforward. In reality, it is both. The successful incorporation of noncanonical amino acids through engineered orthogonal tRNA–synthetase pairs demonstrates that the code can acquire new "words" and expand its expressive capacity. Yet the near universality of the twenty standard amino acids and their canonical assignments reveal a level of conservation more characteristic of natural law than linguistic convention. The lesson is not that the code cannot be rewritten, but that successful redesign must respect the evolutionary logic already built into the system. The genetic code is best viewed as a language whose grammar has become so deeply embedded in the fabric of life that it now behaves like a law.


Our NarrativesYou have spent your career asking how life came to speak two languages simultaneously. If you could ask one question of the very first molecular system that acted as both gene and catalyst — before those two functions diverged — what would it be?

Gustavo Caetano-Anollés If I could ask a single question of the first molecular system that functioned simultaneously as gene and catalyst — before the modern distinction between nucleic acid information and protein function emerged — it would not be: How did you replicate? How did you catalyze reactions? Or how did you encode information? Those are important questions, but they concern mechanisms rather than origins. Instead, I would ask: What property allowed you to persist while all competing chemistries disappeared? At what point did a collection of molecules cease being merely chemistry and become a historical entity capable of transmitting its organization through time?

This question lies at the heart of our research because the congruence between dipeptide chronologies, tRNA evolution, protein domain histories, aminoacyl-tRNA synthetases, and coding systems suggests that life's emergence was not simply a matter of replication or catalysis appearing in isolation. The evidence increasingly points toward a self-reinforcing network in which information, structure, and function became inseparable. The amino acid and dipeptide compositions of the earliest proteins indicate that coding emerged through coevolutionary interactions between polypeptides and nucleic acid cofactors, with selective pressures favoring molecular flexibility, folding, and functional integration. In this sense, the first molecular systems were already responding to structural constraints long before the modern genetic code was fully established.

The deeper issue is not whether those molecules "knew" what they were doing, but when and how chemistry acquired something resembling agency. At what point did molecular interactions begin to preferentially preserve certain organizations because they were better at maintaining themselves? That transition — from transient chemical processes to systems capable of retaining information, accumulating history, and shaping their own evolutionary future — is the threshold these studies ultimately seek to understand. The origin of life may therefore not have been the invention of replication alone, catalysis alone, or even coding alone, but the emergence of a coupled molecular system in which information, structure, and function became mutually reinforcing. Once that coupling appeared, evolution acquired memory, and chemistry became biology.

Conclusion

There is a moment in this conversation when the scientific and the philosophical become difficult to separate.

Asked what he would say to the first molecular system that functioned as both gene and catalyst, Caetano-Anollés declines the obvious questions. He does not ask how it replicated, catalyzed reactions, or encoded information. Instead, he asks what allowed it to persist while competing chemistries disappeared.

That question reaches beyond mechanism. It asks when chemistry became history.

Across decades of work on metabolism, protein evolution, ribosomes, transfer RNA, aminoacyl-tRNA synthetases, and now dipeptides, Caetano-Anollés has pursued a single problem: how information became coupled to function. The remarkable agreement among multiple molecular archives does not solve that problem, but it narrows the possibilities. Independent histories preserved in proteins, tRNAs, and dipeptides all point toward the same underlying narrative.

The emerging picture is neither pure accident nor strict inevitability. Life appears to have developed within a landscape of powerful constraints, where certain paths became increasingly favored as information, structure, and function grew more tightly linked.

The genetic code, as Caetano-Anollés describes it, is both language and law. It carries the marks of history, yet behaves with the stability of a physical principle. Somewhere in that tension lies one of biology's deepest mysteries.

Perhaps the most profound implication of this work is that life did not begin when molecules first replicated. It began when information, structure, and function became inseparable, allowing evolution to accumulate memory.

That was the moment chemistry learned to remember.

Dr. Gustavo Caetano-Anollés is Professor in the Department of Crop Sciences, the Carl R. Woese Institute for Genomic Biology, and Biomedical and Translational Sciences at the Carle Illinois College of Medicine, University of Illinois Urbana-Champaign. The study Tracing the Origin of the Genetic Code and Thermostability to Dipeptide Sequences in Proteomes was published in the Journal of Molecular Biology in August 2025 and was co-authored with Minglei Wang and M. Fayez Aziz.