Nirenberg, Khorana & Holley: Cracking the Genetic Code

Nirenberg Khorana Holley — scientific infographic poster

Every living thing on Earth writes its proteins in the same language. A bacterium, an oak tree, a blue whale and you all use the same three-letter words to mean the same twenty amino acids. That fact is so basic to modern medicine that it is easy to forget somebody had to work it out — and that as recently as 1960, nobody on the planet knew a single word of the language. This page is about the people who translated it, and about what the translation means for you when you are handed a genetic test result, offered a vaccine, or sold a supplement that claims to "activate your DNA."

Table of Contents

  1. The Prize and the Three Men
  2. What Was Known and What Was Not
  3. The Poly-U Experiment, 27 May 1961
  4. The Race, and How It Was Won
  5. Khorana's Chemistry: Writing RNA to Order
  6. Nirenberg's Triplet-Binding Assay, 1964
  7. Holley's tRNA: The Adaptor That Makes a Code a Code
  8. What the Code Actually Is
  9. What the Code Made Possible
  10. Genetic Testing and What a Variant Means
  11. What Universality Does Not License
  12. Where Mainstream Science Agrees — and What Remains Debated
  13. Key Research Papers
  14. Connections
  15. Featured Videos

1. The Prize and the Three Men

The 1968 Nobel Prize in Physiology or Medicine went to three men — Robert W. Holley, Har Gobind Khorana and Marshall W. Nirenberg — "for their interpretation of the genetic code and its function in protein synthesis." Between them they answered a question that in 1960 looked as though it might take a generation: given a stretch of genetic material, which sequence of bases specifies which amino acid?

They answered it in about five years.

Marshall Nirenberg (1927–2010)

Nirenberg was born in New York City and raised in Orlando, Florida, where his family moved when he developed rheumatic fever as a child. He collected specimens in the Florida wetlands, took his bachelor's and master's degrees at the University of Florida in zoology and biology, and completed a PhD in biochemistry at the University of Michigan in 1957. He then joined the National Institutes of Health in Bethesda as a postdoctoral fellow and became an independent investigator there in 1960.

This background matters to the story more than it might seem. The molecular biology of the late 1950s was a small, tightly connected, and frankly rather grand community — the Cavendish and the MRC unit in Cambridge, Caltech, the Pasteur Institute, Cold Spring Harbor, the "phage group" around Max Delbrück. Its members knew each other, corresponded constantly, and traded unpublished results. Nirenberg belonged to none of it. He was a young government scientist, trained in biochemistry rather than in genetics or physics, working in a laboratory nobody in the field was watching. When he cracked the first codon, most of the people who should have been most interested had never heard his name. He described the experience himself, decades later, in an unusually candid memoir published in Trends in Biochemical Sciences — a piece worth reading in full for its account of how a scientific race actually feels from the inside (Nirenberg 2004).

Har Gobind Khorana (1922–2011)

Khorana was born in Raipur, a village in the Punjab of British India that now lies in Pakistan. His father was a patwari, a village agricultural clerk in the colonial revenue service — a modest post, but one that required literacy. Khorana wrote later that his family was practically the only literate one in a village of perhaps a hundred people, and that his father devoted himself, against the odds of the place and the time, to educating his children. There was no school building; the first lessons happened outdoors.

From that beginning he took a bachelor's and master's degree at Punjab University in Lahore, then a government fellowship to Liverpool, where he earned a doctorate in organic chemistry in 1948. He did postdoctoral work in Zurich with Vladimir Prelog and then in Cambridge with Alexander Todd, whose laboratory was the world centre for nucleotide chemistry. He ran his own group at the British Columbia Research Council in Vancouver from 1952, moved to the Institute for Enzyme Research at the University of Wisconsin–Madison in 1960, and to MIT in 1970, where he stayed for the rest of his career.

Khorana was not primarily a biologist. He was a synthetic chemist — and that turned out to be exactly what the problem needed.

Robert W. Holley (1922–1993)

Holley was born in Urbana, Illinois, took his degree at the University of Illinois, and completed a PhD in organic chemistry at Cornell in 1947. During the Second World War he worked on the chemical synthesis of penicillin with Vincent du Vigneaud — a formative apprenticeship in the patient, grinding chemistry of natural products. He spent most of his career at Cornell, in the United States Department of Agriculture's Plant, Soil and Nutrition Laboratory and then as professor of biochemistry, before moving to the Salk Institute in California in 1968. His contribution to the prize was different in kind from the other two: not the decoding of the message, but the structure of the molecule that reads it.

And Heinrich Matthaei, who did not share the prize

Johann Heinrich Matthaei, a young German plant physiologist, arrived at Nirenberg's NIH laboratory in 1960 on a NATO fellowship. He was Nirenberg's only collaborator on the experiment that broke the code, he performed the decisive assay himself, and his name is first on one of the two 1961 papers and second on the other. He was not included in the 1968 prize.

This site keeps a running record of the collaborators the Nobel's rules leave out, and Matthaei belongs on it prominently. The Nobel statutes cap a prize at three laureates, and the 1968 award was already full — but Matthaei's exclusion is not a rounding error in a crowded field. On 27 May 1961 there were exactly two people in the world who knew what UUU meant, and one of them was Heinrich Matthaei. He returned to Germany, worked at the Max Planck Institute for Experimental Medicine in Göttingen, and lived the rest of his professional life adjacent to a discovery that is universally attributed to his colleague. Readers who find this pattern familiar will recognise it from Rosalind Franklin, from the long years Katalin Karikó spent demoted and unfunded, and from the cell-cycle laboratories whose members were left off. The science is not diminished by naming them. It is more accurately described.

2. What Was Known and What Was Not

By 1960 a reader following molecular biology would have known the following.

DNA's structure was settled. The double helix had been published in 1953 by James Watson and Francis Crick, on the strength of Rosalind Franklin's and Maurice Wilkins's X-ray data. The two strands pair A with T and G with C, which explains how DNA copies itself — separate the strands and each one templates its partner. That was the great insight, and it was rightly celebrated.

But heredity is not the point of DNA; proteins are. DNA sits in the nucleus. Proteins — enzymes, hormones, antibodies, the collagen in your skin, the haemoglobin in your blood — are built in the cytoplasm. Something had to carry the instruction from one place to the other. In 1960 and 1961 several laboratories converged on the answer: a short-lived RNA copy of the gene, which François Jacob and Jacques Monod named messenger RNA. The message existed. Nobody could read it.

The sequence hypothesis. Crick had proposed that the order of bases along a nucleic acid specifies the order of amino acids along a protein — a one-dimensional string translated into another one-dimensional string. This is now so obvious it sounds like a definition. In 1958 it was a conjecture.

The code had to be at least three letters long. This follows from simple counting. There are four bases (A, U, G and C in RNA) and twenty amino acids in proteins. One base per amino acid gives four possibilities — far too few. Two bases give sixteen — still too few. Three bases give sixty-four, which is more than enough. In late 1961 Crick, Leslie Barnett, Sydney Brenner and Richard Watts-Tobin published elegant genetic evidence from bacteriophage mutants that the code is in fact read in non-overlapping triplets from a fixed starting point: adding or deleting one or two bases wrecked the gene, but adding or deleting three restored it. That paper is the reason we speak of "reading frames" at all.

What nobody knew was any of the actual words. Sixty-four triplets, twenty amino acids, no dictionary. And — this is the part that is hard to imagine now — no way to determine the sequence of anything. DNA sequencing did not exist; Frederick Sanger's method was fifteen years away. RNA sequencing did not exist. You could not take a gene and read it, and you could not take a natural RNA and know what was in it. The problem looked, to serious people, close to intractable. Estimates circulated that assigning the codons would take decades.

The way out, when it came, was a reversal. If you cannot read a message, write one — make an artificial RNA whose composition you control, feed it to a system that builds protein, and see what protein comes out.

3. The Poly-U Experiment, 27 May 1961

Nirenberg and Matthaei had spent months on an unglamorous piece of plumbing: a cell-free protein-synthesis system. They ground up Escherichia coli, spun out the debris, and were left with a soup containing ribosomes, transfer RNAs, enzymes and salts — everything needed to build protein, in a test tube, with no living cell involved. Then they treated it with DNase to destroy the bacterium's own DNA and let the residual messenger RNA decay, so the system fell quiet. It would now make protein only if you gave it a message. That preparation is the subject of their first 1961 paper (Matthaei & Nirenberg 1961), and it is the real technical achievement underneath the famous one.

The experiment itself is almost childishly simple to describe. Take polyuridylic acid — a synthetic RNA consisting of nothing but uracil, U-U-U-U-U-U for its whole length. Add it to the silent extract. Supply radioactively labelled amino acids, one kind at a time, and see which one gets built into protein.

The laboratory notebooks record the answer in the early hours of 27 May 1961. Poly-U produced a polypeptide made entirely of phenylalanine. Nothing else was incorporated.

If the code is read in triplets, and the message contains only U, then the only triplet present is UUU. And the product is polyphenylalanine. Therefore:

UUU = phenylalanine.

That is the first word of the genetic code, and it was obtained not by reading a sequence but by composing one. The published account appeared in the Proceedings of the National Academy of Sciences in October 1961 under the deliberately flat title "The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides."

Moscow, August 1961

Two months later Nirenberg travelled to the Fifth International Congress of Biochemistry in Moscow and presented the result in a ten-minute talk in a small session room. Almost nobody came. He was an unknown, the session was obscure, and the audience numbered in the dozens.

Francis Crick heard about it. Crick was chairing a major session of the congress, and he arranged for Nirenberg to give the talk again — this time to the full assembly, an audience of over a thousand. The response was immediate and enormous. It is a genuinely gracious act by a man who had every incentive to be territorial about the genetic code, and it is the moment the field learned what had happened at NIH.

It also started a race.

4. The Race, and How It Was Won

The obvious next step was to make other synthetic RNAs and see what they produced. The enzyme that made this possible was polynucleotide phosphorylase, which stitches nucleotides into RNA-like polymers. It had been discovered by Severo Ochoa and Marianne Grunberg-Manago, and it had already won Ochoa a share of the 1959 Nobel Prize. His laboratory at the New York University School of Medicine was large, superbly funded, staffed with experienced enzymologists, and in possession of both the enzyme and the expertise to use it. Within weeks of Moscow, Ochoa's group was working full speed on the code, publishing their first codon-composition results in PNAS before the end of 1961.

By any reasonable prediction, the race should not have been close. On one side, a Nobel laureate with a large well-resourced team and a head start in the relevant chemistry; on the other, a junior NIH investigator with one postdoctoral collaborator.

What happened next is the most attractive part of the story. Nirenberg's colleagues at NIH — people with their own laboratories, their own projects and no obligation whatsoever — volunteered to help him keep pace. Maxine Singer and Leon Heppel, both experts in nucleic acid enzymology working on the same campus, supplied him with the synthetic copolymers he needed and taught his group how to make more. Others joined in. Nirenberg's own account of the period describes this collaboration frankly and gratefully, and treats "competition versus collaboration" as one of its explicit themes.

It is worth being clear about what this was and was not. It was not altruism against interest in some abstract sense — NIH scientists had reason to want the code cracked in Bethesda. But it was a group of established researchers choosing to make an outsider's problem their own, without co-authorship on the flagship papers, at a moment when the alternative was watching a much better-equipped laboratory win. That is a real thing that happened, and it is the reason the story ends where it does.

Both laboratories used the same basic trick and both hit the same wall. Random copolymers give you composition, not sequence. If you make an RNA from a mixture of U and C, you get triplets containing various numbers of each, in random order, and the relative amounts of the amino acids incorporated tell you how many U's and C's each codon contains. That is genuinely useful — by 1963 the base composition of most codons was known. But composition is not a sequence. Knowing that a codon contains two U's and one C does not tell you whether it is UUC, UCU or CUU.

Breaking that ambiguity took two more ideas, one from Wisconsin and one from Bethesda.

5. Khorana's Chemistry: Writing RNA to Order

Khorana's contribution begins from a completely different direction. He was not trying to decode anything. He had spent a decade building the chemistry of nucleotide synthesis — how to join one nucleotide to the next, reliably, in a chosen order, with protecting groups that keep the wrong reactions from happening. It was slow, difficult, deeply unfashionable work, and it produced the one thing the code problem was missing: nucleic acids of defined, known, repeating sequence.

Khorana's group could synthesise a short DNA of, say, alternating T and C, use enzymes to copy it into a long RNA, and hand over a message reading UCUCUCUCUC… — not a random copolymer with 50% U and 50% C, but a molecule whose order was known with certainty.

Now the logic becomes sharp. Read UCUCUCUC… in non-overlapping triplets and you get UCU, CUC, UCU, CUC — only two codons, strictly alternating. Feed it to the cell-free system and you get a polypeptide with two amino acids strictly alternating: serine, leucine, serine, leucine. That establishes that one of {UCU, CUC} means serine and the other leucine, and combining it with other repeats resolves which. The 1965 paper reporting exactly this result — "the in vitro synthesis of a co-polypeptide containing two amino acids in alternating sequence dependent upon a DNA-like polymer containing two nucleotides in alternating sequence" — has one of the least catchy titles in the history of molecular biology and one of the cleanest arguments.

Repeats of three bases (UACUACUAC…) are even more informative: depending on where you start reading, the same molecule presents UAC, ACU or CUA, so a single polymer yields three homopolypeptides at once. Khorana's group worked systematically through the repeating di-, tri- and tetranucleotides, and their assignments fell into place alongside Nirenberg's. By the 1966 Cold Spring Harbor symposium, where both groups presented, the table was essentially complete (Khorana et al. 1966).

Why this is the most consequential piece of the whole story

Khorana did not stop at decoding. In 1970 his group announced the first chemical synthesis of a complete gene — the yeast alanine tRNA gene, the very molecule Holley had sequenced — and in 1976 they synthesised a bacterial tyrosine suppressor tRNA gene that actually functioned when put into a cell, promoter and all. He described the whole programme in a 1979 Science review titled, with justified simplicity, "Total synthesis of a gene."

Take that capability forward and you have, in a direct and traceable line, the entire practice of modern molecular biology:

Nirenberg found the first word. Khorana built the printing press.

6. Nirenberg's Triplet-Binding Assay, 1964

Meanwhile Nirenberg, working with Philip Leder, found a way to read codons one at a time without making protein at all.

The observation behind it is this. A ribosome holding a messenger sequence will grip the matching charged transfer RNA — the tRNA carrying the correct amino acid — even before any peptide bond forms. And, crucially, the message does not have to be long. Leder and Nirenberg found that a trinucleotide — three bases, one codon, nothing more — is sufficient to make a ribosome bind the right charged tRNA. Dinucleotides do not work; three is the minimum.

The assay that follows is beautifully crude. Mix ribosomes, a single defined trinucleotide, and a pool of radioactively labelled charged tRNAs. Pour the mixture through a nitrocellulose filter. Free tRNA passes straight through; the bulky ribosome complex sticks to the filter. Count the radioactivity on the filter. If it is high, the trinucleotide you added corresponds to the labelled amino acid.

Their 1964 Science paper reports the founding cases: pUpUpU, pApApA and pCpCpC specifically direct the binding of phenylalanine-, lysine- and proline-tRNA respectively. That is UUU, AAA and CCC assigned directly, unambiguously, in an afternoon each.

The method scales. Khorana's chemistry could supply all sixty-four trinucleotides; the filter assay could test them against every amino acid. What had been an inference problem became a lookup table. Between 1964 and 1966 Nirenberg's group, with Thomas Caskey, Richard Marshall, Richard Brimacombe and others, worked through the whole set, and the complete genetic code was presented at Cold Spring Harbor in 1966 (Nirenberg et al. 1966).

Five years and a few weeks after poly-U. The decades-long project people had feared took less time than a medical degree.

7. Holley's tRNA: The Adaptor That Makes a Code a Code

Holley's problem was different, and in some ways harder.

Consider what "the code" requires physically. A codon is three bases on an RNA strand. An amino acid is a small organic molecule with no particular affinity for bases. There is no chemistry by which UUU reaches out and grabs phenylalanine — they simply do not fit each other in any meaningful way. Something has to stand between them.

Crick had predicted such a molecule years earlier and called it an adaptor: one end recognises the codon, the other end carries the amino acid, and an enzyme that recognises both ends is what enforces the pairing. The molecule turned out to be transfer RNA, and this is the reason a genetic "code" is genuinely a code rather than a cipher.

The distinction is worth a sentence, because it is the conceptual heart of the whole subject. In a cipher, the relationship between symbol and meaning is determined — shift each letter by three and you can work it out from first principles. In a code, the relationship is assigned and has to be looked up, because there is nothing about the symbol that implies its meaning. UUU means phenylalanine because a particular tRNA carries the anticodon AAA at one end and gets loaded with phenylalanine at the other. Change the loading enzyme and UUU would mean something else. The meaning lives in the adaptor, not in the bases.

Seven years and a great deal of yeast

Holley set out to determine the complete structure of one such adaptor: the tRNA that carries alanine, from baker's yeast. Nothing like it had ever been done. No nucleic acid of any kind had ever been sequenced.

The obstacles were brutal and mostly practical. Transfer RNAs are a mixed population — a cell contains dozens of different ones, chemically similar, and Holley first had to invent a way to separate them, developing a countercurrent distribution method to pull the alanine species away from the rest. Because tRNA is present in tiny proportion, the starting material had to be enormous: commercial baker's yeast processed in industrial quantities to yield a workable amount of one purified species. Then the sequencing itself, which had to be done by controlled destruction — cut the molecule into fragments with ribonucleases, work out the composition and order of each fragment, cut it again a different way to get overlapping pieces, and reassemble the whole from the overlaps. This is the same logic later used for proteins and genomes, but Holley had to invent its application to RNA while doing it.

The work ran from the late 1950s to 1965 — roughly seven years for one molecule of just under eighty nucleotides. The announcement, in Science in March 1965, is one of the shortest and most consequential abstracts in the literature: "The complete nucleotide sequence of an alanine transfer RNA, isolated from yeast, has been determined. This is the first nucleic acid for which the structure is known." A companion paper in the Journal of Biological Chemistry laid out the fragment analysis in full.

The cloverleaf

With the sequence in hand, Holley noticed that stretches of it could pair with other stretches, the way the two strands of DNA pair. Folded to maximise that pairing, the chain forms a cloverleaf: three or four stem-and-loop arms radiating from a central junction. One loop carries the anticodon, the three bases that read the codon. The opposite end carries the site where the amino acid is attached. Every transfer RNA ever sequenced since folds the same way, and X-ray crystallography later showed that the cloverleaf folds again in three dimensions into a compact L, with the anticodon at one tip and the amino acid at the other — about eighty ångströms apart, a molecular tool holding its cargo at arm's length. A useful history of this period, written by two chemists who were in it, is RajBhandary & Köhrer (2006).

Holley's molecule is why the code table has a physical existence. Without the adaptor, the table is a list. With it, the table is machinery.

8. What the Code Actually Is

Here is the finished product, in plain language. If you read only one section of this page, read this one and the two that follow it.

The basic arrangement

Your genes are written in DNA using four chemical letters: A, C, G and T. When a gene is used, it is first copied into messenger RNA, which uses the same letters except that U replaces T. The ribosome reads that message in groups of three, without gaps and without overlaps. Each group of three is a codon.

Four letters in groups of three gives 64 possible codons. Of those:

The code is redundant but not ambiguous

Sixty-one codons for twenty amino acids means most amino acids have more than one codon. Leucine, serine and arginine have six each. Methionine and tryptophan have exactly one. This property is called redundancy (or degeneracy), and it is easy to misread, so it is worth stating both halves:

Ambiguity would be catastrophic — a cell could not build a reliable protein if a codon's meaning were a coin-flip. Redundancy, by contrast, turns out to be a feature.

Third-position wobble and built-in error tolerance

The redundancy is not scattered randomly across the table. It is concentrated almost entirely in the third position of the codon. In many cases, the first two letters determine the amino acid outright and the third can be anything: CU_ is leucine whatever fills the blank; GG_ is glycine; AC_ is threonine.

Crick explained the mechanism in 1966 with the wobble hypothesis: the pairing between the third base of the codon and the corresponding base of the anticodon is looser than standard base pairing, which lets a single tRNA read two or more codons that differ only at that position. This is why a cell needs far fewer than 61 different tRNAs.

The consequence for you is direct and important. A large fraction of single-letter mutations change nothing. A random substitution at a third position very often produces a different codon for the same amino acid, so the protein comes out identical. Beyond that, the code's layout tends to place chemically similar amino acids at codons that differ by one letter, so that even when a mutation does change the amino acid, the replacement is often chemically similar and the protein still works. Freeland and Hurst quantified this in 1998, comparing the real code against a million randomly generated alternatives that used the same amino acids and the same number of codons each. Their conclusion, in their own words, was that only about one in a million random alternative codes was better at minimising the effects of error than the natural one.

The genetic code, in other words, has shock absorbers.

It is very nearly universal — and the exceptions are real

The same 64 assignments are used by bacteria, archaea, plants, fungi, insects, fish and humans. This is the single most practically important fact on this page, and section 9 is about what it buys us.

But "nearly" is doing real work in that sentence, and this site does not round it off:

These exceptions are genuine, they are not trivia, and they are also not a refutation. The overwhelming majority of coding, in the overwhelming majority of organisms, uses the standard table — which is what makes the next section possible.

9. What the Code Made Possible

The universality of the code has a blunt practical consequence: a gene taken from a human being can be read correctly by a bacterium. The bacterium does not know or care that the sequence is foreign. It reads the codons with its own ribosomes and its own tRNAs and produces the human protein. Everything below follows from that single fact.

Recombinant human insulin

Until the late 1970s, every person with diabetes who injected insulin was injecting an animal's. Insulin was extracted from the pancreases of slaughtered cattle and pigs — a supply chain that tied the world's insulin availability to the meat industry, and that delivered a protein differing from the human one by a few amino acids. Most patients did well on it. A minority developed antibodies, injection-site reactions, or lipoatrophy at the injection site.

In 1979 a group at Genentech and the City of Hope reported the expression in E. coli of chemically synthesised genes for human insulin. Note what that title contains: genes written base by base by chemists, in the tradition Khorana founded, read by bacterial ribosomes because the code is universal, producing the authentically human protein. Recombinant human insulin reached the market in 1982 as the first approved recombinant DNA drug for humans. Banting, Best, Collip and Macleod made insulin a treatment in 1922; the genetic code made it an unlimited and human one.

Everything that followed

The same manoeuvre — put a human gene into a cell that will read it — produced an entire class of medicines:

Reading a genome and predicting a protein

Sequencing gives you letters. The genetic code turns letters into meaning. When a laboratory sequences your BRCA1 gene, the raw output is a string of A, C, G and T; software applies the code table to derive the protein sequence, compares it to the reference, and reports where they differ. Without the code the sequence is inert data. Every part of clinical genomics — carrier screening, tumour panels, newborn screening confirmation, pharmacogenetics — sits on the table Nirenberg and Khorana completed in 1966.

mRNA vaccines: a sentence written in this code

An mRNA vaccine is, at bottom, the most literal possible use of the discovery on this page. It is a written message in the genetic code, packaged in a lipid particle, delivered to your cells, and read by your own ribosomes to produce a single protein — typically a viral surface protein — which your immune system then learns to recognise. The technology's hard problems were about delivery and about stopping the immune system from destroying the message on arrival; Karikó and Weissman's nucleoside modification solved the second, and won the 2023 Nobel Prize. But the message itself is written in Nirenberg's and Khorana's alphabet, and could not be written at all without it.

One step in the design deserves specific mention, because it makes the dependency exact. Codon optimisation is a routine part of designing a therapeutic mRNA. Because the code is redundant, the same protein can be specified by many different nucleotide sequences — and those sequences are not equivalent in practice. Cells contain different amounts of the different tRNAs, and codons matched to abundant tRNAs are translated faster and are associated with more stable messages. Designers therefore rewrite the coding sequence using synonymous codons, changing the letters while leaving the protein identical, to raise the amount of protein produced per dose. Hanson and Coller's 2018 review describes this "second genetic code" of codon bias and its effects on translation efficiency and mRNA stability (Hanson & Coller 2018); Pardi, Hogan, Porter and Weissman's review of the field describes where it sits in vaccine design (Pardi et al. 2018).

You cannot choose between synonymous codons unless you know which codons are synonymous. That knowledge is the 1966 table. A modern vaccine design step is, quite literally, an application of a Cold Spring Harbor symposium paper from the year before the moon landing was planned.

10. Genetic Testing and What a Variant Means

If you have had genetic testing — a carrier screen before pregnancy, a tumour panel, a hereditary cancer test, a diagnostic exome for an undiagnosed condition, or a consumer ancestry-and-health kit — your report described changes in your DNA using vocabulary that comes straight from the genetic code. This section explains that vocabulary. It is the most immediately useful thing on this page.

Everyone carries variants. A typical human genome differs from the reference at millions of positions. The overwhelming majority mean nothing. The question a report tries to answer is always: does this particular change break the protein? And the code determines how that question is answered.

Silent (synonymous) variants

A single letter changes, but because of redundancy the new codon still specifies the same amino acid. GGA becomes GGG; both mean glycine. The protein sequence is unchanged.

These are usually harmless, and reports generally classify them as benign. "Usually" is not "always": a synonymous change can still matter if it lands on a splice site — the signal that tells the cell where to cut and join the message — or if it shifts translation speed enough to affect how the protein folds. Both are real but uncommon, and this is one of the places where codon usage, from the previous section, quietly reappears in a clinical setting.

Missense variants

A single letter changes and the new codon specifies a different amino acid. One brick in the wall is swapped for a different brick.

This is the hardest category to interpret, and the one most likely to appear on your report as uncertain. Consequences range from nothing at all to devastating, depending entirely on which amino acid, in which position, in which protein:

Nonsense variants

A single letter changes and an amino-acid codon becomes a stop codon. CAG (glutamine) becomes TAG (stop). The ribosome halts there, and the protein is produced truncated — or, very often, not produced at all, because cells have a surveillance system (nonsense-mediated decay) that detects prematurely terminated messages and destroys them.

Nonsense variants are usually damaging, and how damaging depends heavily on position: a stop codon near the beginning of a gene typically abolishes the protein, while one in the last few percent of the coding sequence may leave a nearly complete and partly functional product. Reports often describe these as "loss-of-function."

Frameshift variants

Letters are inserted or deleted in a number not divisible by three. This is the most consistently destructive category, and the reason why is exactly the reason Crick, Barnett, Brenner and Watts-Tobin ran their 1961 phage experiments.

The ribosome has no punctuation. It starts at AUG and reads three letters at a time, mechanically, to the end. Delete one letter and every codon downstream is redrawn across the wrong boundaries. Consider a message read as three-letter words:

THE BIG RED DOG ATE THE OLD RAW HAM

Delete the B and re-read in threes:

THE IGR EDD OGA TET HEO LDR AWH AM

Nothing after the deletion means anything. That is precisely what happens to the protein: every amino acid after the frameshift is wrong, and typically a stop codon is encountered at random within a few dozen codons, truncating the product as well. This is why a frameshift is usually more damaging than a missense change — a missense variant is one wrong brick, a frameshift is every brick after that point being wrong, and then the wall stopping early.

The clinical corollary is that a three-base deletion behaves completely differently from a one- or two-base deletion. Delete three letters and the reading frame survives; you lose one amino acid and the rest of the protein is intact. The commonest cystic fibrosis variant, F508del, is exactly this: a three-base deletion removing a single phenylalanine from the CFTR protein. It is serious — the protein misfolds and is degraded — but it is in-frame, and the modern CFTR modulator drugs work by helping that nearly-complete protein fold and function. A frameshift in the same gene leaves nothing for a modulator to rescue. Which is to say: the reading frame is not an academic detail; it can determine whether a drug can help you.

"Variant of uncertain significance" — why this keeps happening

Clinical laboratories in the United States classify variants on a five-tier scale set out in a 2015 joint guideline from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology: pathogenic, likely pathogenic, uncertain significance, likely benign, benign. That guideline is the reason two different laboratories usually reach the same answer, and it is worth knowing it exists.

A variant of uncertain significance (VUS) means the laboratory found a real change in your DNA and does not have enough evidence to say whether it matters. It does not mean "probably bad." It does not mean "probably fine." It means unknown, and it is the single most common source of distress in genetic counselling.

The reasons it happens are structural, not sloppiness:

What to do with a VUS, practically: it should not be used to make surgical or treatment decisions, and management should be based on your personal and family history instead. Relatives should not be tested for it as though it were meaningful. Classifications change as evidence accumulates — frequently toward benign — so it is reasonable to ask the ordering laboratory whether they will re-contact you on reclassification, and to ask for a re-review after a few years. And it is worth asking for a genetic counsellor rather than absorbing the report alone; interpreting these categories is their actual profession. Our Genetics section covers many of the specific conditions these tests look for, and the Lab Tests section covers what other results on a panel mean.

11. What Universality Does Not License

Because the genetic code is shared by all life, it gets recruited into claims it does not support. This site's practice is to state each tier plainly rather than to gesture at "misinformation," so here they are, in order of how close they come to being reasonable.

Eating an organism does not transfer its genes to you

This is the most intuitive error and the easiest to answer. Every meal you eat is loaded with DNA and RNA — a plant's, an animal's, a fungus's. All of it is written in the same code as yours. The worry, sometimes stated explicitly and often implied, is that this genetic material might somehow be adopted by your cells.

It is not, and the reason is digestive rather than genetic. Nucleic acids are polymers, and your digestive tract dismantles polymers. Pancreatic nucleases cut DNA and RNA into short pieces; intestinal enzymes reduce those to nucleosides and free bases; those are absorbed as small molecules and enter the ordinary nucleotide pool, where they are indistinguishable from the ones your own cells make. What arrives in your bloodstream from a steak is not a cow gene. It is adenine, guanine, cytosine, thymine and ribose — the same building blocks you would have made anyway. A gene needs to arrive intact, cross into a cell, reach the nucleus and integrate or be transcribed; digestion is specifically in the business of preventing the first step.

The honest complication is worth including, because there is a real scientific dispute in the vicinity. In 2012 a Chinese group reported that a plant microRNA from rice could be detected in human and mouse blood and could regulate a mammalian gene — the "cross-kingdom microRNA" hypothesis. It attracted enormous attention. It also attracted a series of failed replications: Dickinson and colleagues, feeding mice diets rich in the relevant plant material, reported in Nature Biotechnology in 2013 that they could not detect meaningful oral bioavailability of plant microRNAs, and attributed some earlier positive findings to contamination and normalisation artefacts. The original authors replied in the same issue, and the question has never fully closed, though the weight of evidence has moved against a general dietary-microRNA effect. Our page on Ambros and Ruvkun, who discovered microRNAs, covers that dispute in more detail. Two points hold regardless of how it resolves: a regulatory microRNA is not a gene, and detecting a molecule in blood is not the same as showing it does anything there.

"GMO foods rewrite your DNA"

A genetically modified plant contains a gene that a conventional plant does not. When you eat it, that gene is digested exactly like every other gene in the meal — and the meal already contained tens of thousands of plant genes you were never worried about. There is no mechanism by which one particular sequence survives digestion when the rest does not.

There are real debates about genetically modified crops — herbicide use, seed patents and grower dependence, monoculture and biodiversity, labelling and consent. Those are agricultural, economic and political arguments, and they can be had on their merits. "The inserted gene will enter your genome" is not one of them; it is a claim about molecular biology, and it is not true.

"DNA activation," "junk DNA awakening," and 12-strand DNA

At the far end sit products and services — supplements, frequencies, sound baths, coded meditations, "quantum" devices, remote sessions — that claim to activate dormant DNA, switch on junk DNA, or unlock additional strands.

These have no mechanism behind them, and the specific claims are checkable:

The plain version: the code is universal because all life inherited it from a common ancestor. That is a statement about evolutionary history. It is not a channel, a resonance, or a susceptibility.

12. Where Mainstream Science Agrees — and What Remains Debated

Settled

Genuinely unresolved: why the code has the assignments it does

Here is a question that sounds as though it must have been settled decades ago and has not been. Why is UUU phenylalanine and not something else?

Four families of explanation have been argued for half a century, and none commands consensus:

Koonin and Novozhilov reviewed the whole question in 2009 under the title "Origin and evolution of the genetic code: the universal enigma" (Koonin & Novozhilov 2009) — a title chosen with deliberate honesty. The most defensible current position is that the four explanations are not mutually exclusive, that error minimisation is real and measurable, and that the historical question of origin may not be answerable from present-day evidence at all.

It is worth ending the scientific part of this page on that note. The code was cracked completely and the dictionary is not in doubt. Why the dictionary reads as it does is still open. Both facts are true at once, and a field is healthier for saying so.

13. Key Research Papers

  1. Nirenberg MW, Matthaei JH. The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides. Proceedings of the National Academy of Sciences USA, 1961;47(10):1588–1602. The poly-U experiment. UUU = phenylalanine, and the first word of the genetic code. PMID 14479932
  2. Crick FHC, Barnett L, Brenner S, Watts-Tobin RJ. General nature of the genetic code for proteins. Nature, 1961;192:1227–1232. Genetic proof from phage mutants that the code is read in non-overlapping triplets from a fixed start — the origin of the reading frame. PMID 13882203
  3. Lengyel P, Speyer JF, Ochoa S. Synthetic polynucleotides and the amino acid code. Proceedings of the National Academy of Sciences USA, 1961;47(12):1936–1942. The rival laboratory's entry into the race, two months after Moscow. PMID 14463983
  4. Nirenberg M, Leder P. RNA codewords and protein synthesis: the effect of trinucleotides upon the binding of sRNA to ribosomes. Science, 1964;145(3639):1399–1407. The triplet-binding filter assay; pUpUpU, pApApA and pCpCpC assign phenylalanine, lysine and proline directly. PMID 14172630
  5. Nishimura S, Jones DS, Khorana HG. Studies on polynucleotides XLVIII: the in vitro synthesis of a co-polypeptide containing two amino acids in alternating sequence dependent upon a DNA-like polymer containing two nucleotides in alternating sequence. Journal of Molecular Biology, 1965;13(1):302–324. Defined repeating sequences turn codon guesswork into systematic decoding. PMID 5323614
  6. Holley RW, Apgar J, Everett GA, Madison JT, Marquisee M, Merrill SH, Penswick JR, Zamir A. Structure of a ribonucleic acid. Science, 1965;147(3664):1462–1465. The complete sequence of yeast alanine transfer RNA — the first nucleic acid of any kind whose structure was known. PMID 14263761
  7. Crick FHC. Codon–anticodon pairing: the wobble hypothesis. Journal of Molecular Biology, 1966;19(2):548–555. Why the third codon position is loose, and why one tRNA can read several codons. PMID 5969078
  8. Nirenberg M, Caskey T, Marshall R, Brimacombe R, et al. The RNA code and protein synthesis. Cold Spring Harbor Symposia on Quantitative Biology, 1966;31:11–24. The completed table, presented at the symposium that closed the problem. PMID 5237186
  9. Khorana HG, Büchi H, Ghosh H, Gupta N, et al. Polynucleotide synthesis and the genetic code. Cold Spring Harbor Symposia on Quantitative Biology, 1966;31:39–49. Khorana's parallel assignment of the code from chemically defined repeating polymers. PMID 5237635
  10. Khorana HG. Total synthesis of a gene. Science, 1979;203(4381):614–625. The chemical synthesis of a complete, functional gene — the direct ancestor of all oligonucleotide synthesis, PCR primers and synthetic gene templates. PMID 366749
  11. Barrell BG, Bankier AT, Drouin J. A different genetic code in human mitochondria. Nature, 1979;282(5735):189–194. UGA read as tryptophan rather than stop, and AUA apparently as methionine — the first documented break in universality. PMID 226894
  12. Goeddel DV, Kleid DG, Bolivar F, Heyneker HL, et al. Expression in Escherichia coli of chemically synthesized genes for human insulin. Proceedings of the National Academy of Sciences USA, 1979;76(1):106–110. Universality cashed in: a synthetic human gene, read correctly by a bacterium, ending the animal-pancreas insulin supply. PMID 85300
  13. Freeland SJ, Hurst LD. The genetic code is one in a million. Journal of Molecular Evolution, 1998;47(3):238–248. Against a million random alternative codes, only about one is better at minimising the effects of mutation and mistranslation. PMID 9732450
  14. Richards S, Aziz N, Bale S, Bick D, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine, 2015;17(5):405–424. The five-tier framework — pathogenic through benign, with "uncertain significance" in the middle — that your genetic test report is written against. PMID 25741868

Live PubMed Searches

  1. Genetic code codon assignment history
  2. Transfer RNA structure — Holley
  3. Codon optimization and mRNA
  4. Genetic code universality and its exceptions
  5. Variant of uncertain significance classification

14. Connections

Back to top