Nirenberg, Khorana & Holley: Cracking the Genetic Code
Every living thing on Earth writes its proteins in the same language. A bacterium, an oak tree, a blue whale and you all use the same three-letter words to mean the same twenty amino acids. That fact is so basic to modern medicine that it is easy to forget somebody had to work it out — and that as recently as 1960, nobody on the planet knew a single word of the language. This page is about the people who translated it, and about what the translation means for you when you are handed a genetic test result, offered a vaccine, or sold a supplement that claims to "activate your DNA."
Table of Contents
- The Prize and the Three Men
- What Was Known and What Was Not
- The Poly-U Experiment, 27 May 1961
- The Race, and How It Was Won
- Khorana's Chemistry: Writing RNA to Order
- Nirenberg's Triplet-Binding Assay, 1964
- Holley's tRNA: The Adaptor That Makes a Code a Code
- What the Code Actually Is
- What the Code Made Possible
- Genetic Testing and What a Variant Means
- What Universality Does Not License
- Where Mainstream Science Agrees — and What Remains Debated
- Key Research Papers
- Connections
- Featured Videos
1. The Prize and the Three Men
The 1968 Nobel Prize in Physiology or Medicine went to three men — Robert W. Holley, Har Gobind Khorana and Marshall W. Nirenberg — "for their interpretation of the genetic code and its function in protein synthesis." Between them they answered a question that in 1960 looked as though it might take a generation: given a stretch of genetic material, which sequence of bases specifies which amino acid?
They answered it in about five years.
Marshall Nirenberg (1927–2010)
Nirenberg was born in New York City and raised in Orlando, Florida, where his family moved when he developed rheumatic fever as a child. He collected specimens in the Florida wetlands, took his bachelor's and master's degrees at the University of Florida in zoology and biology, and completed a PhD in biochemistry at the University of Michigan in 1957. He then joined the National Institutes of Health in Bethesda as a postdoctoral fellow and became an independent investigator there in 1960.
This background matters to the story more than it might seem. The molecular biology of the late 1950s was a small, tightly connected, and frankly rather grand community — the Cavendish and the MRC unit in Cambridge, Caltech, the Pasteur Institute, Cold Spring Harbor, the "phage group" around Max Delbrück. Its members knew each other, corresponded constantly, and traded unpublished results. Nirenberg belonged to none of it. He was a young government scientist, trained in biochemistry rather than in genetics or physics, working in a laboratory nobody in the field was watching. When he cracked the first codon, most of the people who should have been most interested had never heard his name. He described the experience himself, decades later, in an unusually candid memoir published in Trends in Biochemical Sciences — a piece worth reading in full for its account of how a scientific race actually feels from the inside (Nirenberg 2004).
Har Gobind Khorana (1922–2011)
Khorana was born in Raipur, a village in the Punjab of British India that now lies in Pakistan. His father was a patwari, a village agricultural clerk in the colonial revenue service — a modest post, but one that required literacy. Khorana wrote later that his family was practically the only literate one in a village of perhaps a hundred people, and that his father devoted himself, against the odds of the place and the time, to educating his children. There was no school building; the first lessons happened outdoors.
From that beginning he took a bachelor's and master's degree at Punjab University in Lahore, then a government fellowship to Liverpool, where he earned a doctorate in organic chemistry in 1948. He did postdoctoral work in Zurich with Vladimir Prelog and then in Cambridge with Alexander Todd, whose laboratory was the world centre for nucleotide chemistry. He ran his own group at the British Columbia Research Council in Vancouver from 1952, moved to the Institute for Enzyme Research at the University of Wisconsin–Madison in 1960, and to MIT in 1970, where he stayed for the rest of his career.
Khorana was not primarily a biologist. He was a synthetic chemist — and that turned out to be exactly what the problem needed.
Robert W. Holley (1922–1993)
Holley was born in Urbana, Illinois, took his degree at the University of Illinois, and completed a PhD in organic chemistry at Cornell in 1947. During the Second World War he worked on the chemical synthesis of penicillin with Vincent du Vigneaud — a formative apprenticeship in the patient, grinding chemistry of natural products. He spent most of his career at Cornell, in the United States Department of Agriculture's Plant, Soil and Nutrition Laboratory and then as professor of biochemistry, before moving to the Salk Institute in California in 1968. His contribution to the prize was different in kind from the other two: not the decoding of the message, but the structure of the molecule that reads it.
And Heinrich Matthaei, who did not share the prize
Johann Heinrich Matthaei, a young German plant physiologist, arrived at Nirenberg's NIH laboratory in 1960 on a NATO fellowship. He was Nirenberg's only collaborator on the experiment that broke the code, he performed the decisive assay himself, and his name is first on one of the two 1961 papers and second on the other. He was not included in the 1968 prize.
This site keeps a running record of the collaborators the Nobel's rules leave out, and Matthaei belongs on it prominently. The Nobel statutes cap a prize at three laureates, and the 1968 award was already full — but Matthaei's exclusion is not a rounding error in a crowded field. On 27 May 1961 there were exactly two people in the world who knew what UUU meant, and one of them was Heinrich Matthaei. He returned to Germany, worked at the Max Planck Institute for Experimental Medicine in Göttingen, and lived the rest of his professional life adjacent to a discovery that is universally attributed to his colleague. Readers who find this pattern familiar will recognise it from Rosalind Franklin, from the long years Katalin Karikó spent demoted and unfunded, and from the cell-cycle laboratories whose members were left off. The science is not diminished by naming them. It is more accurately described.
2. What Was Known and What Was Not
By 1960 a reader following molecular biology would have known the following.
DNA's structure was settled. The double helix had been published in 1953 by James Watson and Francis Crick, on the strength of Rosalind Franklin's and Maurice Wilkins's X-ray data. The two strands pair A with T and G with C, which explains how DNA copies itself — separate the strands and each one templates its partner. That was the great insight, and it was rightly celebrated.
But heredity is not the point of DNA; proteins are. DNA sits in the nucleus. Proteins — enzymes, hormones, antibodies, the collagen in your skin, the haemoglobin in your blood — are built in the cytoplasm. Something had to carry the instruction from one place to the other. In 1960 and 1961 several laboratories converged on the answer: a short-lived RNA copy of the gene, which François Jacob and Jacques Monod named messenger RNA. The message existed. Nobody could read it.
The sequence hypothesis. Crick had proposed that the order of bases along a nucleic acid specifies the order of amino acids along a protein — a one-dimensional string translated into another one-dimensional string. This is now so obvious it sounds like a definition. In 1958 it was a conjecture.
The code had to be at least three letters long. This follows from simple counting. There are four bases (A, U, G and C in RNA) and twenty amino acids in proteins. One base per amino acid gives four possibilities — far too few. Two bases give sixteen — still too few. Three bases give sixty-four, which is more than enough. In late 1961 Crick, Leslie Barnett, Sydney Brenner and Richard Watts-Tobin published elegant genetic evidence from bacteriophage mutants that the code is in fact read in non-overlapping triplets from a fixed starting point: adding or deleting one or two bases wrecked the gene, but adding or deleting three restored it. That paper is the reason we speak of "reading frames" at all.
What nobody knew was any of the actual words. Sixty-four triplets, twenty amino acids, no dictionary. And — this is the part that is hard to imagine now — no way to determine the sequence of anything. DNA sequencing did not exist; Frederick Sanger's method was fifteen years away. RNA sequencing did not exist. You could not take a gene and read it, and you could not take a natural RNA and know what was in it. The problem looked, to serious people, close to intractable. Estimates circulated that assigning the codons would take decades.
The way out, when it came, was a reversal. If you cannot read a message, write one — make an artificial RNA whose composition you control, feed it to a system that builds protein, and see what protein comes out.
3. The Poly-U Experiment, 27 May 1961
Nirenberg and Matthaei had spent months on an unglamorous piece of plumbing: a cell-free protein-synthesis system. They ground up Escherichia coli, spun out the debris, and were left with a soup containing ribosomes, transfer RNAs, enzymes and salts — everything needed to build protein, in a test tube, with no living cell involved. Then they treated it with DNase to destroy the bacterium's own DNA and let the residual messenger RNA decay, so the system fell quiet. It would now make protein only if you gave it a message. That preparation is the subject of their first 1961 paper (Matthaei & Nirenberg 1961), and it is the real technical achievement underneath the famous one.
The experiment itself is almost childishly simple to describe. Take polyuridylic acid — a synthetic RNA consisting of nothing but uracil, U-U-U-U-U-U for its whole length. Add it to the silent extract. Supply radioactively labelled amino acids, one kind at a time, and see which one gets built into protein.
The laboratory notebooks record the answer in the early hours of 27 May 1961. Poly-U produced a polypeptide made entirely of phenylalanine. Nothing else was incorporated.
If the code is read in triplets, and the message contains only U, then the only triplet present is UUU. And the product is polyphenylalanine. Therefore:
UUU = phenylalanine.
That is the first word of the genetic code, and it was obtained not by reading a sequence but by composing one. The published account appeared in the Proceedings of the National Academy of Sciences in October 1961 under the deliberately flat title "The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides."
Moscow, August 1961
Two months later Nirenberg travelled to the Fifth International Congress of Biochemistry in Moscow and presented the result in a ten-minute talk in a small session room. Almost nobody came. He was an unknown, the session was obscure, and the audience numbered in the dozens.
Francis Crick heard about it. Crick was chairing a major session of the congress, and he arranged for Nirenberg to give the talk again — this time to the full assembly, an audience of over a thousand. The response was immediate and enormous. It is a genuinely gracious act by a man who had every incentive to be territorial about the genetic code, and it is the moment the field learned what had happened at NIH.
It also started a race.
4. The Race, and How It Was Won
The obvious next step was to make other synthetic RNAs and see what they produced. The enzyme that made this possible was polynucleotide phosphorylase, which stitches nucleotides into RNA-like polymers. It had been discovered by Severo Ochoa and Marianne Grunberg-Manago, and it had already won Ochoa a share of the 1959 Nobel Prize. His laboratory at the New York University School of Medicine was large, superbly funded, staffed with experienced enzymologists, and in possession of both the enzyme and the expertise to use it. Within weeks of Moscow, Ochoa's group was working full speed on the code, publishing their first codon-composition results in PNAS before the end of 1961.
By any reasonable prediction, the race should not have been close. On one side, a Nobel laureate with a large well-resourced team and a head start in the relevant chemistry; on the other, a junior NIH investigator with one postdoctoral collaborator.
What happened next is the most attractive part of the story. Nirenberg's colleagues at NIH — people with their own laboratories, their own projects and no obligation whatsoever — volunteered to help him keep pace. Maxine Singer and Leon Heppel, both experts in nucleic acid enzymology working on the same campus, supplied him with the synthetic copolymers he needed and taught his group how to make more. Others joined in. Nirenberg's own account of the period describes this collaboration frankly and gratefully, and treats "competition versus collaboration" as one of its explicit themes.
It is worth being clear about what this was and was not. It was not altruism against interest in some abstract sense — NIH scientists had reason to want the code cracked in Bethesda. But it was a group of established researchers choosing to make an outsider's problem their own, without co-authorship on the flagship papers, at a moment when the alternative was watching a much better-equipped laboratory win. That is a real thing that happened, and it is the reason the story ends where it does.
Both laboratories used the same basic trick and both hit the same wall. Random copolymers give you composition, not sequence. If you make an RNA from a mixture of U and C, you get triplets containing various numbers of each, in random order, and the relative amounts of the amino acids incorporated tell you how many U's and C's each codon contains. That is genuinely useful — by 1963 the base composition of most codons was known. But composition is not a sequence. Knowing that a codon contains two U's and one C does not tell you whether it is UUC, UCU or CUU.
Breaking that ambiguity took two more ideas, one from Wisconsin and one from Bethesda.
5. Khorana's Chemistry: Writing RNA to Order
Khorana's contribution begins from a completely different direction. He was not trying to decode anything. He had spent a decade building the chemistry of nucleotide synthesis — how to join one nucleotide to the next, reliably, in a chosen order, with protecting groups that keep the wrong reactions from happening. It was slow, difficult, deeply unfashionable work, and it produced the one thing the code problem was missing: nucleic acids of defined, known, repeating sequence.
Khorana's group could synthesise a short DNA of, say, alternating T and C, use enzymes to copy it into a long RNA, and hand over a message reading UCUCUCUCUC… — not a random copolymer with 50% U and 50% C, but a molecule whose order was known with certainty.
Now the logic becomes sharp. Read UCUCUCUC… in non-overlapping triplets and you get UCU, CUC, UCU, CUC — only two codons, strictly alternating. Feed it to the cell-free system and you get a polypeptide with two amino acids strictly alternating: serine, leucine, serine, leucine. That establishes that one of {UCU, CUC} means serine and the other leucine, and combining it with other repeats resolves which. The 1965 paper reporting exactly this result — "the in vitro synthesis of a co-polypeptide containing two amino acids in alternating sequence dependent upon a DNA-like polymer containing two nucleotides in alternating sequence" — has one of the least catchy titles in the history of molecular biology and one of the cleanest arguments.
Repeats of three bases (UACUACUAC…) are even more informative: depending on where you start reading, the same molecule presents UAC, ACU or CUA, so a single polymer yields three homopolypeptides at once. Khorana's group worked systematically through the repeating di-, tri- and tetranucleotides, and their assignments fell into place alongside Nirenberg's. By the 1966 Cold Spring Harbor symposium, where both groups presented, the table was essentially complete (Khorana et al. 1966).
Why this is the most consequential piece of the whole story
Khorana did not stop at decoding. In 1970 his group announced the first chemical synthesis of a complete gene — the yeast alanine tRNA gene, the very molecule Holley had sequenced — and in 1976 they synthesised a bacterial tyrosine suppressor tRNA gene that actually functioned when put into a cell, promoter and all. He described the whole programme in a 1979 Science review titled, with justified simplicity, "Total synthesis of a gene."
Take that capability forward and you have, in a direct and traceable line, the entire practice of modern molecular biology:
- Oligonucleotide synthesis — ordering a short piece of DNA of any sequence you like, delivered in a day. Every research laboratory on Earth does this routinely.
- PCR primers, which are exactly such oligonucleotides. Every PCR-based diagnostic test you have ever taken depends on someone being able to write a specific short sequence to order.
- Site-directed mutagenesis — deliberately changing one codon in a gene to see what the protein does, which is how protein function is studied.
- Synthetic gene fragments, and therefore the DNA templates from which therapeutic mRNA is transcribed. An mRNA vaccine begins its life as a synthesised DNA template. That template exists because Khorana's chemistry exists.
Nirenberg found the first word. Khorana built the printing press.
6. Nirenberg's Triplet-Binding Assay, 1964
Meanwhile Nirenberg, working with Philip Leder, found a way to read codons one at a time without making protein at all.
The observation behind it is this. A ribosome holding a messenger sequence will grip the matching charged transfer RNA — the tRNA carrying the correct amino acid — even before any peptide bond forms. And, crucially, the message does not have to be long. Leder and Nirenberg found that a trinucleotide — three bases, one codon, nothing more — is sufficient to make a ribosome bind the right charged tRNA. Dinucleotides do not work; three is the minimum.
The assay that follows is beautifully crude. Mix ribosomes, a single defined trinucleotide, and a pool of radioactively labelled charged tRNAs. Pour the mixture through a nitrocellulose filter. Free tRNA passes straight through; the bulky ribosome complex sticks to the filter. Count the radioactivity on the filter. If it is high, the trinucleotide you added corresponds to the labelled amino acid.
Their 1964 Science paper reports the founding cases: pUpUpU, pApApA and pCpCpC specifically direct the binding of phenylalanine-, lysine- and proline-tRNA respectively. That is UUU, AAA and CCC assigned directly, unambiguously, in an afternoon each.
The method scales. Khorana's chemistry could supply all sixty-four trinucleotides; the filter assay could test them against every amino acid. What had been an inference problem became a lookup table. Between 1964 and 1966 Nirenberg's group, with Thomas Caskey, Richard Marshall, Richard Brimacombe and others, worked through the whole set, and the complete genetic code was presented at Cold Spring Harbor in 1966 (Nirenberg et al. 1966).
Five years and a few weeks after poly-U. The decades-long project people had feared took less time than a medical degree.
7. Holley's tRNA: The Adaptor That Makes a Code a Code
Holley's problem was different, and in some ways harder.
Consider what "the code" requires physically. A codon is three bases on an RNA strand. An amino acid is a small organic molecule with no particular affinity for bases. There is no chemistry by which UUU reaches out and grabs phenylalanine — they simply do not fit each other in any meaningful way. Something has to stand between them.
Crick had predicted such a molecule years earlier and called it an adaptor: one end recognises the codon, the other end carries the amino acid, and an enzyme that recognises both ends is what enforces the pairing. The molecule turned out to be transfer RNA, and this is the reason a genetic "code" is genuinely a code rather than a cipher.
The distinction is worth a sentence, because it is the conceptual heart of the whole subject. In a cipher, the relationship between symbol and meaning is determined — shift each letter by three and you can work it out from first principles. In a code, the relationship is assigned and has to be looked up, because there is nothing about the symbol that implies its meaning. UUU means phenylalanine because a particular tRNA carries the anticodon AAA at one end and gets loaded with phenylalanine at the other. Change the loading enzyme and UUU would mean something else. The meaning lives in the adaptor, not in the bases.
Seven years and a great deal of yeast
Holley set out to determine the complete structure of one such adaptor: the tRNA that carries alanine, from baker's yeast. Nothing like it had ever been done. No nucleic acid of any kind had ever been sequenced.
The obstacles were brutal and mostly practical. Transfer RNAs are a mixed population — a cell contains dozens of different ones, chemically similar, and Holley first had to invent a way to separate them, developing a countercurrent distribution method to pull the alanine species away from the rest. Because tRNA is present in tiny proportion, the starting material had to be enormous: commercial baker's yeast processed in industrial quantities to yield a workable amount of one purified species. Then the sequencing itself, which had to be done by controlled destruction — cut the molecule into fragments with ribonucleases, work out the composition and order of each fragment, cut it again a different way to get overlapping pieces, and reassemble the whole from the overlaps. This is the same logic later used for proteins and genomes, but Holley had to invent its application to RNA while doing it.
The work ran from the late 1950s to 1965 — roughly seven years for one molecule of just under eighty nucleotides. The announcement, in Science in March 1965, is one of the shortest and most consequential abstracts in the literature: "The complete nucleotide sequence of an alanine transfer RNA, isolated from yeast, has been determined. This is the first nucleic acid for which the structure is known." A companion paper in the Journal of Biological Chemistry laid out the fragment analysis in full.
The cloverleaf
With the sequence in hand, Holley noticed that stretches of it could pair with other stretches, the way the two strands of DNA pair. Folded to maximise that pairing, the chain forms a cloverleaf: three or four stem-and-loop arms radiating from a central junction. One loop carries the anticodon, the three bases that read the codon. The opposite end carries the site where the amino acid is attached. Every transfer RNA ever sequenced since folds the same way, and X-ray crystallography later showed that the cloverleaf folds again in three dimensions into a compact L, with the anticodon at one tip and the amino acid at the other — about eighty ångströms apart, a molecular tool holding its cargo at arm's length. A useful history of this period, written by two chemists who were in it, is RajBhandary & Köhrer (2006).
Holley's molecule is why the code table has a physical existence. Without the adaptor, the table is a list. With it, the table is machinery.
8. What the Code Actually Is
Here is the finished product, in plain language. If you read only one section of this page, read this one and the two that follow it.
The basic arrangement
Your genes are written in DNA using four chemical letters: A, C, G and T. When a gene is used, it is first copied into messenger RNA, which uses the same letters except that U replaces T. The ribosome reads that message in groups of three, without gaps and without overlaps. Each group of three is a codon.
Four letters in groups of three gives 64 possible codons. Of those:
- 61 specify an amino acid. These are the working words of the language.
- 3 are stop signals — UAA, UAG and UGA. They do not code for anything; they tell the ribosome the protein is finished and to release it. They are the full stop at the end of the sentence.
- AUG is both a start signal and the codon for methionine. Translation of a protein begins at an AUG, which is why essentially every protein starts life with methionine at its front end (often trimmed off afterwards).
The code is redundant but not ambiguous
Sixty-one codons for twenty amino acids means most amino acids have more than one codon. Leucine, serine and arginine have six each. Methionine and tryptophan have exactly one. This property is called redundancy (or degeneracy), and it is easy to misread, so it is worth stating both halves:
- Redundant: several different codons can mean the same amino acid. UUA, UUG, CUU, CUC, CUA and CUG all mean leucine.
- Not ambiguous: no single codon means two different things. Each of the 64 has exactly one meaning. You can always translate a message; you cannot always back-translate a protein into a unique gene.
Ambiguity would be catastrophic — a cell could not build a reliable protein if a codon's meaning were a coin-flip. Redundancy, by contrast, turns out to be a feature.
Third-position wobble and built-in error tolerance
The redundancy is not scattered randomly across the table. It is concentrated almost entirely in the third position of the codon. In many cases, the first two letters determine the amino acid outright and the third can be anything: CU_ is leucine whatever fills the blank; GG_ is glycine; AC_ is threonine.
Crick explained the mechanism in 1966 with the wobble hypothesis: the pairing between the third base of the codon and the corresponding base of the anticodon is looser than standard base pairing, which lets a single tRNA read two or more codons that differ only at that position. This is why a cell needs far fewer than 61 different tRNAs.
The consequence for you is direct and important. A large fraction of single-letter mutations change nothing. A random substitution at a third position very often produces a different codon for the same amino acid, so the protein comes out identical. Beyond that, the code's layout tends to place chemically similar amino acids at codons that differ by one letter, so that even when a mutation does change the amino acid, the replacement is often chemically similar and the protein still works. Freeland and Hurst quantified this in 1998, comparing the real code against a million randomly generated alternatives that used the same amino acids and the same number of codons each. Their conclusion, in their own words, was that only about one in a million random alternative codes was better at minimising the effects of error than the natural one.
The genetic code, in other words, has shock absorbers.
It is very nearly universal — and the exceptions are real
The same 64 assignments are used by bacteria, archaea, plants, fungi, insects, fish and humans. This is the single most practically important fact on this page, and section 9 is about what it buys us.
But "nearly" is doing real work in that sentence, and this site does not round it off:
- Mitochondria use a different code. Your mitochondria carry their own small genome and translate it with their own machinery, and their dictionary differs from the one in your nucleus. Barrell, Bankier and Drouin showed in 1979, by comparing the human mitochondrial cytochrome oxidase subunit II gene against the corresponding beef heart protein, that UGA is read as tryptophan rather than as a stop signal, and that AUA appears to mean methionine rather than isoleucine. Several further mitochondrial deviations have been catalogued since, and they differ between lineages.
- Some ciliates reassign stop codons. Horowitz and Gorovsky showed in 1985 that in the nuclear genes of Tetrahymena thermophila, UAA codes for glutamine rather than stopping translation — the first demonstration that a "universal" termination codon could carry a coding function (Horowitz & Gorovsky 1985). Comparable reassignments occur in other ciliates and in some yeasts and mycoplasmas.
- Two extra amino acids get inserted by context. Selenocysteine and pyrrolysine are incorporated at particular UGA and UAG codons when specific signals elsewhere in the message instruct the ribosome to read through. Selenocysteine matters to human biology — it is how selenium is built into the glutathione peroxidase and thyroid deiodinase enzymes.
These exceptions are genuine, they are not trivia, and they are also not a refutation. The overwhelming majority of coding, in the overwhelming majority of organisms, uses the standard table — which is what makes the next section possible.
9. What the Code Made Possible
The universality of the code has a blunt practical consequence: a gene taken from a human being can be read correctly by a bacterium. The bacterium does not know or care that the sequence is foreign. It reads the codons with its own ribosomes and its own tRNAs and produces the human protein. Everything below follows from that single fact.
Recombinant human insulin
Until the late 1970s, every person with diabetes who injected insulin was injecting an animal's. Insulin was extracted from the pancreases of slaughtered cattle and pigs — a supply chain that tied the world's insulin availability to the meat industry, and that delivered a protein differing from the human one by a few amino acids. Most patients did well on it. A minority developed antibodies, injection-site reactions, or lipoatrophy at the injection site.
In 1979 a group at Genentech and the City of Hope reported the expression in E. coli of chemically synthesised genes for human insulin. Note what that title contains: genes written base by base by chemists, in the tradition Khorana founded, read by bacterial ribosomes because the code is universal, producing the authentically human protein. Recombinant human insulin reached the market in 1982 as the first approved recombinant DNA drug for humans. Banting, Best, Collip and Macleod made insulin a treatment in 1922; the genetic code made it an unlimited and human one.
Everything that followed
The same manoeuvre — put a human gene into a cell that will read it — produced an entire class of medicines:
- Human growth hormone. Previously extracted from cadaver pituitary glands, a practice that transmitted Creutzfeldt–Jakob disease to a number of recipients before it was stopped. Recombinant growth hormone ended that risk entirely.
- Clotting factors VIII and IX for haemophilia. Previously pooled from thousands of plasma donations, which is how a generation of haemophiliacs was infected with HIV and hepatitis C. Recombinant factors removed the human plasma from the supply chain.
- Erythropoietin for the anaemia of kidney disease.
- Monoclonal antibodies — the drugs whose names end in -mab, now central to the treatment of rheumatoid arthritis, inflammatory bowel disease, asthma, migraine, high cholesterol and many cancers. Each one is a designed protein, expressed from a designed gene.
- Enzyme replacement therapies for lysosomal storage diseases such as Gaucher, Pompe and Fabry disease — conditions where a single missing enzyme can now be manufactured and infused.
- Vaccines made from a single protein, such as the hepatitis B surface antigen produced in yeast — work that grew directly out of Blumberg's identification of the virus.
Reading a genome and predicting a protein
Sequencing gives you letters. The genetic code turns letters into meaning. When a laboratory sequences your BRCA1 gene, the raw output is a string of A, C, G and T; software applies the code table to derive the protein sequence, compares it to the reference, and reports where they differ. Without the code the sequence is inert data. Every part of clinical genomics — carrier screening, tumour panels, newborn screening confirmation, pharmacogenetics — sits on the table Nirenberg and Khorana completed in 1966.
mRNA vaccines: a sentence written in this code
An mRNA vaccine is, at bottom, the most literal possible use of the discovery on this page. It is a written message in the genetic code, packaged in a lipid particle, delivered to your cells, and read by your own ribosomes to produce a single protein — typically a viral surface protein — which your immune system then learns to recognise. The technology's hard problems were about delivery and about stopping the immune system from destroying the message on arrival; Karikó and Weissman's nucleoside modification solved the second, and won the 2023 Nobel Prize. But the message itself is written in Nirenberg's and Khorana's alphabet, and could not be written at all without it.
One step in the design deserves specific mention, because it makes the dependency exact. Codon optimisation is a routine part of designing a therapeutic mRNA. Because the code is redundant, the same protein can be specified by many different nucleotide sequences — and those sequences are not equivalent in practice. Cells contain different amounts of the different tRNAs, and codons matched to abundant tRNAs are translated faster and are associated with more stable messages. Designers therefore rewrite the coding sequence using synonymous codons, changing the letters while leaving the protein identical, to raise the amount of protein produced per dose. Hanson and Coller's 2018 review describes this "second genetic code" of codon bias and its effects on translation efficiency and mRNA stability (Hanson & Coller 2018); Pardi, Hogan, Porter and Weissman's review of the field describes where it sits in vaccine design (Pardi et al. 2018).
You cannot choose between synonymous codons unless you know which codons are synonymous. That knowledge is the 1966 table. A modern vaccine design step is, quite literally, an application of a Cold Spring Harbor symposium paper from the year before the moon landing was planned.
10. Genetic Testing and What a Variant Means
If you have had genetic testing — a carrier screen before pregnancy, a tumour panel, a hereditary cancer test, a diagnostic exome for an undiagnosed condition, or a consumer ancestry-and-health kit — your report described changes in your DNA using vocabulary that comes straight from the genetic code. This section explains that vocabulary. It is the most immediately useful thing on this page.
Everyone carries variants. A typical human genome differs from the reference at millions of positions. The overwhelming majority mean nothing. The question a report tries to answer is always: does this particular change break the protein? And the code determines how that question is answered.
Silent (synonymous) variants
A single letter changes, but because of redundancy the new codon still specifies the same amino acid. GGA becomes GGG; both mean glycine. The protein sequence is unchanged.
These are usually harmless, and reports generally classify them as benign. "Usually" is not "always": a synonymous change can still matter if it lands on a splice site — the signal that tells the cell where to cut and join the message — or if it shifts translation speed enough to affect how the protein folds. Both are real but uncommon, and this is one of the places where codon usage, from the previous section, quietly reappears in a clinical setting.
Missense variants
A single letter changes and the new codon specifies a different amino acid. One brick in the wall is swapped for a different brick.
This is the hardest category to interpret, and the one most likely to appear on your report as uncertain. Consequences range from nothing at all to devastating, depending entirely on which amino acid, in which position, in which protein:
- Swapping one amino acid for a chemically similar one, in a stretch of the protein that is not doing anything critical, often has no measurable effect.
- Swapping an amino acid at the active site of an enzyme, or at a position conserved unchanged across a billion years of evolution, is frequently catastrophic.
- Sickle cell disease is caused by a single missense change in the beta-globin gene: GAG becomes GTG, glutamic acid becomes valine, at position 6 of a 146-amino-acid chain. One letter, one amino acid, and haemoglobin polymerises into rigid fibres that deform red cells into crescents. It is the textbook demonstration that a point substitution is not automatically minor.
Nonsense variants
A single letter changes and an amino-acid codon becomes a stop codon. CAG (glutamine) becomes TAG (stop). The ribosome halts there, and the protein is produced truncated — or, very often, not produced at all, because cells have a surveillance system (nonsense-mediated decay) that detects prematurely terminated messages and destroys them.
Nonsense variants are usually damaging, and how damaging depends heavily on position: a stop codon near the beginning of a gene typically abolishes the protein, while one in the last few percent of the coding sequence may leave a nearly complete and partly functional product. Reports often describe these as "loss-of-function."
Frameshift variants
Letters are inserted or deleted in a number not divisible by three. This is the most consistently destructive category, and the reason why is exactly the reason Crick, Barnett, Brenner and Watts-Tobin ran their 1961 phage experiments.
The ribosome has no punctuation. It starts at AUG and reads three letters at a time, mechanically, to the end. Delete one letter and every codon downstream is redrawn across the wrong boundaries. Consider a message read as three-letter words:
THE BIG RED DOG ATE THE OLD RAW HAM
Delete the B and re-read in threes:
THE IGR EDD OGA TET HEO LDR AWH AM
Nothing after the deletion means anything. That is precisely what happens to the protein: every amino acid after the frameshift is wrong, and typically a stop codon is encountered at random within a few dozen codons, truncating the product as well. This is why a frameshift is usually more damaging than a missense change — a missense variant is one wrong brick, a frameshift is every brick after that point being wrong, and then the wall stopping early.
The clinical corollary is that a three-base deletion behaves completely differently from a one- or two-base deletion. Delete three letters and the reading frame survives; you lose one amino acid and the rest of the protein is intact. The commonest cystic fibrosis variant, F508del, is exactly this: a three-base deletion removing a single phenylalanine from the CFTR protein. It is serious — the protein misfolds and is degraded — but it is in-frame, and the modern CFTR modulator drugs work by helping that nearly-complete protein fold and function. A frameshift in the same gene leaves nothing for a modulator to rescue. Which is to say: the reading frame is not an academic detail; it can determine whether a drug can help you.
"Variant of uncertain significance" — why this keeps happening
Clinical laboratories in the United States classify variants on a five-tier scale set out in a 2015 joint guideline from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology: pathogenic, likely pathogenic, uncertain significance, likely benign, benign. That guideline is the reason two different laboratories usually reach the same answer, and it is worth knowing it exists.
A variant of uncertain significance (VUS) means the laboratory found a real change in your DNA and does not have enough evidence to say whether it matters. It does not mean "probably bad." It does not mean "probably fine." It means unknown, and it is the single most common source of distress in genetic counselling.
The reasons it happens are structural, not sloppiness:
- Most VUSs are missense. Frameshift and nonsense variants are usually classifiable on mechanism alone. A missense variant requires knowing what that specific amino acid does in that specific position, which often nobody knows.
- Rarity is the enemy. Classification leans heavily on frequency data — a variant common in healthy populations is almost certainly benign — and on seeing the variant repeatedly in affected families. A variant seen in three people worldwide supports neither argument.
- Reference databases are not evenly populated. Population genomic databases have historically over-represented people of European ancestry, so variants common and benign in other populations appear rare and therefore suspicious. Patients of non-European ancestry consequently receive more uncertain results. This is a known, documented and slowly improving inequity, and it is worth knowing about if it is your result.
What to do with a VUS, practically: it should not be used to make surgical or treatment decisions, and management should be based on your personal and family history instead. Relatives should not be tested for it as though it were meaningful. Classifications change as evidence accumulates — frequently toward benign — so it is reasonable to ask the ordering laboratory whether they will re-contact you on reclassification, and to ask for a re-review after a few years. And it is worth asking for a genetic counsellor rather than absorbing the report alone; interpreting these categories is their actual profession. Our Genetics section covers many of the specific conditions these tests look for, and the Lab Tests section covers what other results on a panel mean.
11. What Universality Does Not License
Because the genetic code is shared by all life, it gets recruited into claims it does not support. This site's practice is to state each tier plainly rather than to gesture at "misinformation," so here they are, in order of how close they come to being reasonable.
Eating an organism does not transfer its genes to you
This is the most intuitive error and the easiest to answer. Every meal you eat is loaded with DNA and RNA — a plant's, an animal's, a fungus's. All of it is written in the same code as yours. The worry, sometimes stated explicitly and often implied, is that this genetic material might somehow be adopted by your cells.
It is not, and the reason is digestive rather than genetic. Nucleic acids are polymers, and your digestive tract dismantles polymers. Pancreatic nucleases cut DNA and RNA into short pieces; intestinal enzymes reduce those to nucleosides and free bases; those are absorbed as small molecules and enter the ordinary nucleotide pool, where they are indistinguishable from the ones your own cells make. What arrives in your bloodstream from a steak is not a cow gene. It is adenine, guanine, cytosine, thymine and ribose — the same building blocks you would have made anyway. A gene needs to arrive intact, cross into a cell, reach the nucleus and integrate or be transcribed; digestion is specifically in the business of preventing the first step.
The honest complication is worth including, because there is a real scientific dispute in the vicinity. In 2012 a Chinese group reported that a plant microRNA from rice could be detected in human and mouse blood and could regulate a mammalian gene — the "cross-kingdom microRNA" hypothesis. It attracted enormous attention. It also attracted a series of failed replications: Dickinson and colleagues, feeding mice diets rich in the relevant plant material, reported in Nature Biotechnology in 2013 that they could not detect meaningful oral bioavailability of plant microRNAs, and attributed some earlier positive findings to contamination and normalisation artefacts. The original authors replied in the same issue, and the question has never fully closed, though the weight of evidence has moved against a general dietary-microRNA effect. Our page on Ambros and Ruvkun, who discovered microRNAs, covers that dispute in more detail. Two points hold regardless of how it resolves: a regulatory microRNA is not a gene, and detecting a molecule in blood is not the same as showing it does anything there.
"GMO foods rewrite your DNA"
A genetically modified plant contains a gene that a conventional plant does not. When you eat it, that gene is digested exactly like every other gene in the meal — and the meal already contained tens of thousands of plant genes you were never worried about. There is no mechanism by which one particular sequence survives digestion when the rest does not.
There are real debates about genetically modified crops — herbicide use, seed patents and grower dependence, monoculture and biodiversity, labelling and consent. Those are agricultural, economic and political arguments, and they can be had on their merits. "The inserted gene will enter your genome" is not one of them; it is a claim about molecular biology, and it is not true.
"DNA activation," "junk DNA awakening," and 12-strand DNA
At the far end sit products and services — supplements, frequencies, sound baths, coded meditations, "quantum" devices, remote sessions — that claim to activate dormant DNA, switch on junk DNA, or unlock additional strands.
These have no mechanism behind them, and the specific claims are checkable:
- DNA has two strands. Not twelve. The double helix is a physical structure that has been imaged, crystallised and measured for seventy years. There is no version of it with more strands awaiting activation.
- "Junk DNA" was always a bad name, and correcting it does not help the claim. Non-coding DNA is not inert — it contains promoters, enhancers, insulators, non-coding RNA genes and structural elements, and working out what those do is a major and active research field. But its functions are regulatory and biochemical, discovered by experiment. Finding that a region has a function does not mean it is "asleep" or that a purchased product can wake it.
- Gene expression is regulated, and that regulation is real. Diet, exercise, sleep, stress and toxin exposure genuinely change which genes are transcribed, through mechanisms including DNA methylation and histone modification — the subject matter of epigenetics. This is precisely what gives "DNA activation" marketing its foothold: it borrows the vocabulary of a real field. The difference is that epigenetic changes are specific, measurable, and demonstrated with assays, while "activation" products offer no measurement, no proposed mechanism, and no way for the claim to be wrong.
- The genetic code itself is not adjustable. The assignment of codons to amino acids is fixed by which tRNAs exist and which enzymes charge them. Nothing you eat, hear, meditate on or purchase changes what UUU means.
The plain version: the code is universal because all life inherited it from a common ancestor. That is a statement about evolutionary history. It is not a channel, a resonance, or a susceptibility.
12. Where Mainstream Science Agrees — and What Remains Debated
Settled
- The genetic code is read in non-overlapping triplets from a fixed start point. Established genetically in 1961 and biochemically thereafter; never seriously challenged since.
- The specific assignments — which codon means which amino acid — are known, complete, and verified thousands of times over. UUU is phenylalanine in every laboratory on Earth.
- Sixty-one codons specify amino acids, three terminate translation, and AUG both initiates and encodes methionine.
- Transfer RNA is the adaptor, folded into a cloverleaf in two dimensions and an L in three, with the anticodon at one end and the amino acid at the other.
- The code is redundant, principally at the third codon position, and Crick's wobble mechanism explains how one tRNA reads several codons.
- The standard code is used by the great majority of organisms and genes; the mitochondrial, ciliate and other documented variants are real, catalogued, and limited.
- Because of universality, a human gene expressed in a bacterium or a yeast yields the human protein — the basis of recombinant medicine.
Genuinely unresolved: why the code has the assignments it does
Here is a question that sounds as though it must have been settled decades ago and has not been. Why is UUU phenylalanine and not something else?
Four families of explanation have been argued for half a century, and none commands consensus:
- The frozen accident. Crick's 1968 proposal: the assignments were arbitrary at the start, and once a substantial number of proteins depended on them, any change would have altered every protein in the organism simultaneously and been lethal. The code is what it is because it became unchangeable before it became optimal. Under this view the specific assignments have no deeper reason at all.
- Stereochemistry. The proposal that certain codons or anticodons have genuine physical affinity for their amino acids, so the assignments were chemically determined from the start. Some experimental support exists — in vitro selection experiments have found RNA sequences that bind particular amino acids and are enriched in the corresponding codons — but the effect is not general enough to explain the whole table.
- Error minimisation. The proposal that the arrangement was shaped by selection to limit the damage done by mutation and mistranslation. This has the strongest quantitative support: Freeland and Hurst's finding that only about one in a million random alternative codes is more error-tolerant than the natural one is a striking result. Its weakness is that it explains the code's structure without explaining its content — it says why chemically similar amino acids sit at neighbouring codons, not why phenylalanine in particular sits at UUU.
- Coevolution with biosynthesis. The proposal that codons were originally shared among amino acids that are made from one another, and that the table records the order in which amino acids became biologically available.
Koonin and Novozhilov reviewed the whole question in 2009 under the title "Origin and evolution of the genetic code: the universal enigma" (Koonin & Novozhilov 2009) — a title chosen with deliberate honesty. The most defensible current position is that the four explanations are not mutually exclusive, that error minimisation is real and measurable, and that the historical question of origin may not be answerable from present-day evidence at all.
It is worth ending the scientific part of this page on that note. The code was cracked completely and the dictionary is not in doubt. Why the dictionary reads as it does is still open. Both facts are true at once, and a field is healthier for saying so.
13. Key Research Papers
- Nirenberg MW, Matthaei JH. The dependence of cell-free protein synthesis in E. coli upon naturally occurring or synthetic polyribonucleotides. Proceedings of the National Academy of Sciences USA, 1961;47(10):1588–1602. The poly-U experiment. UUU = phenylalanine, and the first word of the genetic code. PMID 14479932
- Crick FHC, Barnett L, Brenner S, Watts-Tobin RJ. General nature of the genetic code for proteins. Nature, 1961;192:1227–1232. Genetic proof from phage mutants that the code is read in non-overlapping triplets from a fixed start — the origin of the reading frame. PMID 13882203
- Lengyel P, Speyer JF, Ochoa S. Synthetic polynucleotides and the amino acid code. Proceedings of the National Academy of Sciences USA, 1961;47(12):1936–1942. The rival laboratory's entry into the race, two months after Moscow. PMID 14463983
- Nirenberg M, Leder P. RNA codewords and protein synthesis: the effect of trinucleotides upon the binding of sRNA to ribosomes. Science, 1964;145(3639):1399–1407. The triplet-binding filter assay; pUpUpU, pApApA and pCpCpC assign phenylalanine, lysine and proline directly. PMID 14172630
- Nishimura S, Jones DS, Khorana HG. Studies on polynucleotides XLVIII: the in vitro synthesis of a co-polypeptide containing two amino acids in alternating sequence dependent upon a DNA-like polymer containing two nucleotides in alternating sequence. Journal of Molecular Biology, 1965;13(1):302–324. Defined repeating sequences turn codon guesswork into systematic decoding. PMID 5323614
- Holley RW, Apgar J, Everett GA, Madison JT, Marquisee M, Merrill SH, Penswick JR, Zamir A. Structure of a ribonucleic acid. Science, 1965;147(3664):1462–1465. The complete sequence of yeast alanine transfer RNA — the first nucleic acid of any kind whose structure was known. PMID 14263761
- Crick FHC. Codon–anticodon pairing: the wobble hypothesis. Journal of Molecular Biology, 1966;19(2):548–555. Why the third codon position is loose, and why one tRNA can read several codons. PMID 5969078
- Nirenberg M, Caskey T, Marshall R, Brimacombe R, et al. The RNA code and protein synthesis. Cold Spring Harbor Symposia on Quantitative Biology, 1966;31:11–24. The completed table, presented at the symposium that closed the problem. PMID 5237186
- Khorana HG, Büchi H, Ghosh H, Gupta N, et al. Polynucleotide synthesis and the genetic code. Cold Spring Harbor Symposia on Quantitative Biology, 1966;31:39–49. Khorana's parallel assignment of the code from chemically defined repeating polymers. PMID 5237635
- Khorana HG. Total synthesis of a gene. Science, 1979;203(4381):614–625. The chemical synthesis of a complete, functional gene — the direct ancestor of all oligonucleotide synthesis, PCR primers and synthetic gene templates. PMID 366749
- Barrell BG, Bankier AT, Drouin J. A different genetic code in human mitochondria. Nature, 1979;282(5735):189–194. UGA read as tryptophan rather than stop, and AUA apparently as methionine — the first documented break in universality. PMID 226894
- Goeddel DV, Kleid DG, Bolivar F, Heyneker HL, et al. Expression in Escherichia coli of chemically synthesized genes for human insulin. Proceedings of the National Academy of Sciences USA, 1979;76(1):106–110. Universality cashed in: a synthetic human gene, read correctly by a bacterium, ending the animal-pancreas insulin supply. PMID 85300
- Freeland SJ, Hurst LD. The genetic code is one in a million. Journal of Molecular Evolution, 1998;47(3):238–248. Against a million random alternative codes, only about one is better at minimising the effects of mutation and mistranslation. PMID 9732450
- Richards S, Aziz N, Bale S, Bick D, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine, 2015;17(5):405–424. The five-tier framework — pathogenic through benign, with "uncertain significance" in the middle — that your genetic test report is written against. PMID 25741868
Live PubMed Searches
- Genetic code codon assignment history
- Transfer RNA structure — Holley
- Codon optimization and mRNA
- Genetic code universality and its exceptions
- Variant of uncertain significance classification
14. Connections
- All Notable Doctors — the full index of physicians and scientists profiled on this site.
- The Nobel Prize in Physiology or Medicine — every award from 1901 to the present, with the omissions noted.
- Watson, Crick & Wilkins — the double helix, and Rosalind Franklin's data, without which there is no code to crack.
- Karikó & Weissman — mRNA vaccines: a message written in this code and delivered to your ribosomes.
- Ambros & Ruvkun — microRNA, gene regulation, and the unresolved cross-kingdom dietary microRNA dispute.
- Barbara McClintock — transposable elements, and thirty years of being right before the field caught up.
- Banting & the discovery of insulin — the treatment the genetic code later made human and unlimited.
- Svante Pääbo — ancient DNA, Neanderthal genomes, and what sequencing extinct humans revealed.
- Hartwell, Hunt & Nurse — the cell cycle, and another prize with collaborators left off.
- Blackburn, Greider & Szostak — telomeres and telomerase, and what happens at the ends of the message.
- Genetics — the inherited conditions these tests look for, from achondroplasia to Williams syndrome.
- Sickle Cell Disease — one missense codon, one amino acid, and a lifetime of consequences.
- Cystic Fibrosis — F508del, an in-frame three-base deletion, and why the reading frame decides what drugs can do.
- Phenylketonuria — what happens when the enzyme that handles phenylalanine, the code's first decoded amino acid, is broken.
- Interactive: DNA to Protein Synthesis — watch transcription and translation run, and break them deliberately.
- Interactive: CRISPR Gene Editing — how a sequence is targeted and cut, built on the same code.
- Lab Tests — what the other numbers on your results mean.