Capecchi, Evans & Smithies: Knockout Mice, and How Disease Models Are Made
Almost every headline you have ever read that begins "scientists have discovered that gene X causes disease Y" rests on a technique three men worked out between 1980 and 1989. The technique is called gene targeting. Its most famous product is the knockout mouse — a mouse in which one chosen gene, and only that gene, has been deliberately switched off, so that researchers can watch what happens next.
This page is about how that became possible, what it genuinely delivered, and — the part that matters most if you are a patient reading health news — how often a result in a mouse turns out to mean nothing at all for a person. Both halves are true at once, and most coverage gives you only one of them.
Table of Contents
- The Prize and the Three Men
- The Problem: Finding a Gene Is Not Knowing What It Does
- Homologous Recombination: Aiming DNA at a Target
- Evans's Embryonic Stem Cells: Getting the Change Into a Mouse
- What Knockouts Actually Revealed
- What It Built for Medicine
- Where Mouse Models Fail
- Animal Research Ethics, Stated Squarely
- What Came After: CRISPR, Organoids and Chips
- What This Means When You Read Health News
- Where Mainstream Science Agrees — and What Remains Debated
- Key Research Papers
- Connections
- Featured Videos
1. The Prize and the Three Men
The 2007 Nobel Prize in Physiology or Medicine was awarded jointly to Mario R. Capecchi, Sir Martin J. Evans and Oliver Smithies for their discoveries of principles for introducing specific gene modifications in mice by the use of embryonic stem cells. Three separate lines of work had to converge before any of it worked, and the three men came at the problem from directions that had almost nothing in common.
Mario Capecchi
Capecchi was born in Verona, Italy, in October 1937. His mother, Lucy Ramberg, was an American-born poet living in Italy; she was not married to his father, and she was politically active against the fascist government. In 1941 she was arrested and sent to Dachau. Capecchi was about four years old.
Before her arrest she had sold what she owned and left the money with a farming family in the Tyrol to look after him. The money ran out after roughly a year. From then until he was nine, he lived without a family — on the streets of northern Italian towns and in a succession of institutions and orphanages, through the last years of the war and the hungry period that followed it. He has described being malnourished and, at one point, hospitalised for an extended period.
His mother survived Dachau. She was released in 1945 and spent about a year searching for him. She found him in 1946, in a hospital in Reggio Emilia. He was nine. Within months they had emigrated to the United States, to a Quaker community outside Philadelphia where his uncle, the physicist Edward Ramberg, lived. Capecchi entered an American school at nine years old, in a language he did not speak.
He took a degree in chemistry and physics at Antioch College, then went to Harvard for a doctorate under James Watson, working on how the genetic code is read. In 1973 he left Harvard for the University of Utah in Salt Lake City, a move that surprised colleagues and that he later said gave him room to work on a long-odds problem without having to justify it every year.
That problem was gene targeting, and the funding history is worth stating because it is unusually well documented. When Capecchi applied to the US National Institutes of Health in 1980 for support to try homologous recombination in mammalian cells, the reviewers were not persuaded — the consensus was that the chance of success was too low to fund. His overall grant was renewed for his other work; he redirected effort into the targeting project anyway. When he came back four years later with results, the review panel's response has been widely reported in the form: we are glad you did not follow our advice. It is a story frequently told as a parable about peer review. The more useful reading is narrower: the reviewers were correct that the odds were poor, and correct about why they were poor, and the project still worked.
Sir Martin Evans
Evans was born in Stroud, Gloucestershire, in 1941. He read biochemistry at Cambridge, took his doctorate at University College London, and spent his early career on teratocarcinoma cells — cells from a peculiar kind of tumour that contain a chaotic mixture of tissue types, as if a scrambled embryo were growing inside it. Those tumours were a curiosity to most people. To Evans they were evidence that a cell capable of becoming anything could be kept alive in a dish.
In 1981, working with the anatomist Matthew Kaufman at Cambridge, he took the step that mattered: rather than starting from a tumour, they took cells directly from an early mouse embryo and got them to grow in culture while keeping that all-purpose capacity. In the same year, working independently in San Francisco, Gail Martin derived comparable cells and gave them the name that stuck — embryonic stem cells. Evans later moved to Cardiff University, where he was knighted in 2004.
Oliver Smithies
Smithies was born in Halifax, in the West Riding of Yorkshire, in 1925, and trained at Oxford. He spent his career in North America — Toronto, then the University of Wisconsin, then the University of North Carolina at Chapel Hill from 1988 until his death in 2017 at the age of 91. He was still running experiments in his late eighties.
Before any of the gene-targeting work, he had already given laboratory medicine one of its workhorse tools: in the 1950s he invented starch gel electrophoresis, a method of separating proteins that revealed variation in human blood proteins nobody had known was there. He came to gene targeting as a protein chemist who wanted to fix a defective gene rather than as a mouse geneticist.
He is also known for something quieter. Smithies kept bound laboratory notebooks continuously from 1953, and kept them for more than fifty years — ordinary daily records, with the failures left in. His Nobel lecture was titled Turning pages, and it was largely a walk through those notebooks. The University of North Carolina has since digitised them. For anyone who has been told that science is a sequence of insights, the notebooks are a useful corrective: they are mostly a record of things that did not work, kept legibly enough that the eventual thing that did work could be traced back to its origin.
2. The Problem: Finding a Gene Is Not Knowing What It Does
By the late 1970s, biologists could read DNA sequences and could isolate individual genes. What they could not do was find out what a gene was for in a living mammal.
The traditional route ran backwards. You took a population of animals, exposed them to something mutagenic, and looked for offspring that were visibly abnormal — a mouse with a kinked tail, a mouse that shook, a mouse with the wrong coat colour. That gave you an interesting animal. Then began the long part: working out which gene had been damaged. This is forward genetics — phenotype first, gene second — and it worked, slowly. It also had a built-in blind spot. It could only find genes whose loss produced something a person could notice by looking, and it could not be pointed at a gene you were specifically interested in.
The obvious inversion is reverse genetics: pick the gene you care about, break it deliberately, and see what the animal is missing. Bacteria and yeast geneticists had been doing this routinely for years. In a mammal it was simply not possible, and the reason was mechanical.
You could get foreign DNA into mammalian cells. That much had been solved. The trouble was where it went. Injected or transfected DNA integrated into the chromosomes essentially at random — it landed wherever a break happened to be, in any number of copies, in any orientation. You could add a gene to a mouse this way (that is what a classical transgenic mouse is), but you could not edit one. There was no way to say: go to chromosome 11, find this particular sequence, and replace it.
Being able to say that is the whole of what the 2007 prize was for.
3. Homologous Recombination: Aiming DNA at a Target
The solution came from a process cells already had. Homologous recombination is a normal piece of cellular machinery: when two stretches of DNA share the same sequence, the cell can line them up and swap material between them. It is how chromosomes exchange segments during the formation of eggs and sperm, and it is one of the ways a cell repairs a broken chromosome — by copying the intact matching sequence.
The insight Capecchi and Smithies each arrived at independently was that this machinery could be hijacked. If you build a piece of DNA that matches a chosen chromosomal sequence, and put your intended alteration in the middle of it, the cell's own recombination machinery may recognise the match, line the two up, and swap your version in. The cell does the aiming. You only have to supply the template.
Smithies published the first demonstration in a human chromosome in 1985. Working with a plasmid carrying globin sequences, his group showed that introduced DNA could be inserted into the human chromosomal beta-globin locus by homologous recombination — and, importantly, that it worked whether or not the target gene was switched on in those cells. Two years later, Capecchi's group with Kirk Thomas, and Smithies's group with Thomas Doetschman and colleagues, both showed the same thing in mouse embryonic stem cells, using the Hprt gene as a test case because cells with and without a working copy can be separated with drugs.
Why "extremely rare" was the entire obstacle
Here is the number that made this hard. In the 1985 beta-globin work, the planned modification was achieved in about one in a thousand of the cells that had taken up the DNA at all. In the 1987 mouse ES-cell work, roughly one in a thousand of the drug-resistant colonies turned out to carry the correctly targeted change.
One in a thousand sounds workable until you notice what the other 999 are. They are not empty. They are cells in which the same construct integrated somewhere random in the genome. Those cells are alive, they carry the drug-resistance marker you used to select them, and under a microscope they are indistinguishable from the ones you want. So the practical task was not "make recombination happen" — recombination happened. The task was finding the correct events among a thousand look-alike wrong ones, in cell colonies you have to grow, pick and test one by one.
Imagine sorting a thousand identical envelopes to find the single one containing the right letter, where opening an envelope takes a week.
Positive–negative selection: the trick that made it practical
The fix, published by Suzanne Mansour, Kirk Thomas and Capecchi in 1988, is one of those ideas that looks obvious once someone has said it. It exploits a structural difference between a correct targeting event and a random insertion.
The targeting construct is built with two kinds of gene attached:
- A positive selection marker — typically a drug-resistance gene — placed inside the region that matches the target. Any cell that took up the DNA and kept it survives the drug. This selects for "something integrated".
- A negative selection marker — a gene that kills the cell in the presence of a second drug — placed at the outer edge of the construct, outside the matching region.
Now consider the geometry. When homologous recombination occurs, the cell lines up the matching regions and swaps in only what lies between them. The outer edge — the killer gene — is trimmed off and lost. When the construct instead crashes into a random site, the whole thing goes in, killer gene included.
So you apply both drugs. The first kills everything that took up nothing. The second kills everything that integrated randomly. What survives both is heavily enriched for the cells you actually want. Mansour and colleagues reported an enrichment of roughly 2,000-fold. That single change turned a search through a thousand colonies into a manageable experiment, and it is still the logic behind targeting-vector design decades later.
4. Evans's Embryonic Stem Cells: Getting the Change Into a Mouse
Targeting a gene in a cell in a dish is a fine piece of molecular biology, and by itself it gives you nothing about a living animal. The missing piece was Evans's.
Embryonic stem (ES) cells are taken from the inner cell mass of an early mouse embryo at the blastocyst stage — a hollow ball of a few dozen cells, days after fertilisation, before any tissue has committed to being anything. Evans and Kaufman showed in 1981 that these cells could be persuaded to grow and divide indefinitely in culture without losing that uncommitted state. They were still, in principle, capable of becoming any tissue in the body.
That combination is what makes the whole scheme work, because it means you can do slow, fiddly molecular work on a population of cells — targeting, drug selection, screening, growing up the rare correct clone — and at the end still have cells that can build a mouse.
From a dish to a mouse that breeds true
The sequence runs like this:
- Target the gene in ES cells. Introduce the construct, apply positive–negative selection, screen the survivors, and pick a clone in which the intended change sits in the intended place.
- Inject those cells into a blastocyst. A few of the modified cells are injected into an early embryo taken from a different mouse strain — conventionally one with a different coat colour, so the result is visible.
- Transfer the embryo to a foster mother. The pups that result are chimeras: mosaics built from two cell populations. Their fur is often patchy, which is the crude but immediate readout that the injected cells contributed.
- Breed the chimera. This is the step that decides everything. If the modified ES cells contributed to the germ line — to the cells that make eggs or sperm — then some of the chimera's offspring inherit the targeted change in every cell of their body. If the modified cells only ended up in skin and liver, the line dies with that animal.
- Breed those offspring together. The first generation carries the change on one chromosome and a normal copy on the other. Mating two of them yields, on average, one in four pups carrying the modification on both — a homozygous knockout, an animal with no working copy of that gene at all.
Germ-line transmission from cultured ES cells was demonstrated before targeting was, which is what made the rest credible. In 1987 Evans's group — Michael Kuehn, Allan Bradley, Elizabeth Robertson and Evans — reported mice carrying Hprt mutations derived from mutagenised ES cells, transmitted through the germ line, as a candidate model for Lesch-Nyhan syndrome. Those particular mutations were made by retroviral insertion and selection rather than by homologous recombination, so it was not gene targeting; but it proved that a change made to cells in a dish could become a heritable mouse strain. Once Capecchi's and Smithies's targeting chemistry was bolted onto Evans's cells, the full method existed.
By 1990 the first genuinely targeted mice were being reported and characterised. One early example is discussed in the next section, because what it showed was not what anyone expected.
5. What Knockouts Actually Revealed
The premise was simple: remove a gene, observe what breaks, and you have learned what the gene does. Thousands of knockouts later, the honest summary is that the premise was too simple, and that finding out how it was too simple taught biology more than the original plan would have.
Surprise one: a great many knockouts look fine
Researchers repeatedly deleted a gene they were confident was important and got a mouse that was, to all appearances, normal. Not subtly impaired — normal. It ate, bred, and lived a normal lifespan.
An early and instructive case came from Smithies's collaborators at Chapel Hill in 1990. They knocked out beta-2-microglobulin, a protein required for normal display of MHC class I molecules on the cell surface. MHC class I is how a cell shows its internal contents to the immune system; without it, killer T cells cannot inspect a cell for viral proteins. The homozygous knockout mice had no detectable class I molecules on their cells and were grossly deficient in CD8-positive T cells — the entire cytotoxic arm of adaptive immunity was crippled. The mice developed normally and were obtained at the expected frequencies.
That result is startling in two directions at once. It says the immune system has more redundancy than the textbook diagram suggests, and it says that "the mouse looks fine" is not evidence that a gene is unimportant — a laboratory cage is a benign environment, and a mouse that never meets a pathogen never has to prove it can fight one.
This turned out to be general. A 2007 review of the accumulated experience concluded that in a substantial proportion of knockout mice no phenotype could be detected at all, and that where phenotypes were eventually found, they often turned up in physiological systems nobody had thought to examine, or only appeared under a specific stress. The technical name for this is biological robustness: networks with backup routes, paralogous genes that can cover for one another, and pathways that compensate when a component is removed. Twenty years of failed expectations did more to establish the reality of that robustness than any positive result had.
Surprise two: some knockouts never make it to birth
The opposite failure is just as common. Delete the gene and the embryo dies — sometimes very early, sometimes days before birth. You now know the gene is essential for development, and you have learned nothing whatsoever about what it does in an adult, which is very often the question you were asking.
The scale of this has since been measured properly. The International Mouse Phenotyping Consortium set out to knock out and systematically phenotype thousands of mouse genes on a standardised platform. In its 2016 report covering the first 1,751 unique gene knockouts, 410 genes were lethal, and the consortium's working estimate is that roughly a third of all mammalian genes are essential for life. The same analysis found something else worth holding onto: incomplete penetrance and variable expressivity were common even on a defined genetic background. Genetically identical mice, in the same facility, on the same diet, did not all show the same phenotype. (A corrigendum to that paper was published in Nature in 2017; it is linked in the citation list below, and anyone citing the paper should read it alongside the original.)
Surprise three: phenotypes nobody predicted
And sometimes deleting a gene produced something entirely off the map — a skeletal abnormality from a gene thought to be immunological, a behavioural change from a gene thought to be metabolic. These were genuine discoveries, and they are the strongest argument for the whole enterprise: an unbiased experiment can return an answer no one would have thought to look for.
The fixes: conditional knockouts and knock-ins
Embryonic lethality was solved by making the deletion controllable in space and time. The Cre-lox system does this. Two ingredients:
- loxP sites — short DNA sequences, borrowed from a bacterial virus, inserted by gene targeting so that they flank the piece of the gene you eventually want removed. A gene with loxP sites around it is described as "floxed", and on its own it behaves normally.
- Cre recombinase — an enzyme that recognises loxP sites and cuts out whatever lies between them. Cre is supplied by a second mouse strain, engineered so that Cre is only made in a chosen tissue — only in liver cells, only in T cells, only in neurons.
Cross the two strains and the gene is deleted only where Cre is present. Everywhere else it is intact. The embryo develops normally because the gene is working in the tissues that need it; the deletion happens where you want to study it. The first demonstration, from Klaus Rajewsky's group in 1994, deleted a DNA polymerase gene segment specifically in T cells. A refinement the following year added a chemical switch — Cre held inactive until a drug is given — so the deletion can also be timed, letting an animal grow up normally and then lose the gene as an adult.
The other extension is the knock-in. Instead of destroying a gene, you replace it with a precisely altered version — often the exact single-letter change found in patients with an inherited disease. That matters because most human genetic disease is not caused by a gene being absent. It is caused by a gene making a protein that is subtly wrong: mis-folded, over-active, stuck in the wrong place. A knockout models absence. A knock-in models the actual mutation, which is frequently a different disease.
6. What It Built for Medicine
The inventory is large. A partial list of what targeted mice made possible:
Cystic fibrosis
In 1992, a group at Chapel Hill including Smithies and Beverly Koller disrupted the CFTR gene and produced the first cystic fibrosis mouse. The homozygous mice reproduced a recognisable slice of the human disease: failure to thrive, meconium ileus (a bowel obstruction seen in newborns with CF), and altered mucous and serous glands. Most died of intestinal obstruction before 40 days of age.
Read that description again and notice what is not in it. The lung disease — the thick airway secretions, the chronic infection, the progressive loss of lung function that is what cystic fibrosis is for most patients — was not the phenotype these mice had. The CF mouse is therefore a genuinely useful research animal and a poor model of the part of the illness that kills people. That is an unusually clean illustration of the general problem, and it comes from one of the field's own founding papers.
Sickle cell disease
Making a mouse model of sickle cell disease required something harder than a knockout, because the disease is caused by a specific human haemoglobin variant. Two groups reported the solution in the same issue of Science in 1997: knock out the mouse's own globin genes entirely and replace them with human globin genes carrying the sickle mutation, so the animal's red cells contain exclusively human sickle haemoglobin. Those mice sickle. This is the knockout and the knock-in used together, and it is the template for modelling any disease that depends on a particular human protein variant.
Atherosclerosis
The ApoE knockout mouse, from Nobuyo Maeda's group at Chapel Hill in 1992, is one of the most widely used disease models ever made. Apolipoprotein E is the tag that lets the liver clear cholesterol-rich lipoprotein remnants from the blood. Delete it and the remnants accumulate. The mice had roughly five times normal plasma cholesterol, developed foam-cell-rich deposits in the proximal aorta by three months, and by eight months had severe narrowing at the origin of the coronary arteries — atherosclerosis arising spontaneously, in a mouse, without a special diet.
Together with the LDL-receptor knockout, this gave cardiovascular research a tractable animal in which plaque forms on a timescale of months. The intellectual line runs directly back to Michael Brown and Joseph Goldstein, whose work on the LDL receptor explained why cholesterol clearance fails in the first place; the knockout mice let that mechanism be manipulated one gene at a time.
Cancer predisposition
Tumour-suppressor genes were defined by their absence, which makes them a natural fit for knockouts. Mice lacking p53, Rb, Brca1, Apc and others develop tumours at high rates and became the standard systems for testing what a tumour suppressor actually suppresses — and, later, for testing drugs against tumours with a defined genetic lesion rather than against a generic tumour.
Immunology and immunotherapy
Knockouts are arguably more central to immunology than to any other field, because the immune system is a network of interacting cell types and the only reliable way to establish what a component does is to remove it and watch. Mice lacking individual cytokines, receptors, or entire lymphocyte lineages became the routine tools of the discipline.
The line to modern cancer treatment is direct. The checkpoint proteins whose blockade won James Allison and Tasuku Honjo the 2018 Nobel Prize were characterised in large part through knockout mice: the PD-1 knockout mouse developed autoimmune disease, which is how PD-1's role as a brake on the immune system was established. Severely immunodeficient strains, in which the mouse's own immune system is disabled by targeted mutations, are what allow human tumours and human immune cells to be grown in an animal at all.
Humanised mice
That last point generalises into an entire sub-field. A humanised mouse is a severely immunodeficient animal — unable to reject foreign tissue — into which human cells are engrafted: human blood-forming stem cells that reconstitute a partial human immune system, human liver cells, or a patient's own tumour. These animals let human cells and human pathogens be studied in a living body, and they are the closest thing available to a testbed for a human immune response short of a human being. Their limits are equally real: the human cells sit in a mouse's tissue architecture, are supported by mouse growth factors that often do not cross-react properly, and lack the lymph-node structures a human immune system builds for itself.
Drug target validation
Underneath all of this sits the least visible and possibly largest use. When a pharmaceutical company considers a protein as a drug target, one of the first questions is what happens to an animal that lacks it — because a knockout mouse is a preview of what a perfectly effective drug against that target would do, including the side effects. A widely cited 2003 analysis in Nature Reviews Drug Discovery looked backwards at the targets of the 100 best-selling drugs then on the market and found that the knockout phenotypes correlated well with the drugs' known effects. Deleting the gene and blocking the protein tended to produce recognisably similar results.
That analysis is a retrospective on drugs that had already succeeded, so it does not tell you how often a promising knockout phenotype fails to become a drug — a genuinely important limitation. But it does explain why target validation in mice became a standard gate in drug development, and why a large share of the compounds now in clinical trials passed through a knockout mouse somewhere in their history.
7. Where Mouse Models Fail
This is the section that matters, and it is the one usually left out.
A mouse is not a small human. It is a separate species that has been evolving away from us for something on the order of 90 million years, weighs about 3,000 times less, has a heart rate around 600 beats per minute, lives two to three years, and has an immune system that differs from ours in ways that are not cosmetic. A frequently cited 2004 review in the Journal of Immunology catalogued the differences bluntly — different balance of white cell types in blood, differences in antibody classes, in Toll-like receptors, in the chemokines and their receptors, in how T cells develop and are regulated. Its title, Of mice and not men, was the point.
Add to biology the artificiality of the setup. Laboratory mice are inbred to near-genetic-identity, which removes exactly the genetic variability that makes human populations respond differently to the same drug. They are kept pathogen-free, so their immune systems are naive in a way no human's is. They are young — most experiments use the mouse equivalent of a person in their twenties, while most of the diseases being modelled are diseases of the old. They are housed at temperatures below their thermal comfort zone, which measurably alters metabolism and immune function. And a mouse "model" of a human disease is usually not the disease at all: it is an induced injury or a genetic lesion chosen because it produces some of the same readouts.
The stroke example
The single best-documented case of translation failure is neuroprotection in acute stroke, and it is worth going through carefully because it is where the field learned to audit itself.
The idea was straightforward. When an artery in the brain blocks, the tissue immediately downstream dies fast, but a surrounding rim — the penumbra — is damaged and not yet dead. A drug that protected those cells for a few hours could preserve a great deal of brain. In rodent models of stroke, drug after drug did exactly that. Infarcts shrank, animals recovered function. There was no shortage of positive results.
In 2006, a group led by Victoria O'Collins and David Howells published a systematic catalogue in Annals of Neurology of every experimental treatment they could identify for acute stroke. They found 1,026. Of these, 114 had been taken into clinical use or clinical testing and 912 had been tested only in animals. Their central finding was not the count — it was that there was no evidence that the drugs selected for human trials had performed better in animals than the ones left behind. In focal models the two groups' average improvements were statistically indistinguishable, and the scope of testing each drug had received varied enormously. The authors' conclusion was that the field could not tell which of its own candidates were the good ones.
Meanwhile, in human trials, neuroprotectants failed with remarkable consistency. Through this period the treatments that actually established benefit in acute ischaemic stroke were the ones that restore blood flow — thrombolysis, and later mechanical clot removal — not the ones that were supposed to protect neurons from lack of it.
Why the animal results were misleading
A companion analysis in 2010 supplied a large part of the answer, and it is uncomfortable. Using a database of systematic reviews of animal stroke studies — 16 reviews covering 525 publications — the authors found that only 10 of those 525 publications (2%) reported no significant effect on infarct volume, and only 6 (1.2%) failed to report at least one significant finding. A literature in which 98% of experiments work is not a literature reporting what happened. Statistical tests for publication bias were positive across the board. The authors estimated that publication bias accounted for roughly a third of the apparent efficacy in the published record — reported benefit falling from about 31% to about 24% after adjustment — and that some 214 experiments had been done and never reported.
Alongside publication bias sat straightforward methodological weakness. Preclinical studies were routinely small, and small studies with a real effect of modest size will mostly miss it — so the ones that reach significance overstate the effect. Randomisation of animals to treatment groups was often not done. Blinded assessment of outcome was often not done, in experiments where the outcome is a human being scoring an animal's behaviour. None of this requires anyone to have behaved dishonestly. It is what a field looks like when nobody has yet agreed on the standards.
It is not only stroke
A 2013 analysis in PNAS compared the pattern of gene activity in human patients after burns, trauma and endotoxin exposure with the pattern in the corresponding mouse models. The human conditions resembled one another closely. The mouse models correlated poorly with the human conditions and poorly with each other — for genes significantly changed in humans, the mouse counterparts matched close to randomly.
And that paper was directly contradicted. In 2015 a Japanese group re-analysed the same datasets, restricting the comparison to genes significantly changed in both species, and reported strong correlations in the same direction (Spearman coefficients of roughly 0.43 to 0.68, with 77–93% of genes moving the same way). Their title was a deliberate mirror image of the original: Genomic responses in mouse models greatly mimic human inflammatory diseases. The exchange generated several published rebuttals in both directions. The disagreement is substantially about which genes belong in the comparison, and it has not been cleanly resolved. We cite both papers below, and a reader should treat any single-sentence summary of "mice don't model human inflammation" — or of the reverse — as an argument someone is having, not a settled fact.
The reform that followed
The response to all this was to import into animal research the methodological machinery that clinical trials had adopted decades earlier. The STAIR recommendations (Stroke Therapy Academic Industry Roundtable, 1999) set standards for preclinical stroke work: dose-response data, defined time windows, blinded outcome assessment, replication in more than one species, testing in animals with the co-morbidities real patients have — old animals, hypertensive animals, diabetic animals. The ARRIVE guidelines (2010, updated as ARRIVE 2.0 in 2020) specified what an animal experiment must report: how many animals, how they were allocated, whether the assessor was blinded, what the exclusion criteria were, and how the sample size was decided. In 2012 the US National Institute of Neurological Disorders and Stroke convened a workshop that issued a similar call for transparent reporting, prompted in part by the recognition that preclinical results were failing to replicate. In the same year, two cancer-drug researchers published a widely discussed commentary in Nature reporting that of 53 "landmark" preclinical cancer papers their team had attempted to reproduce, the findings of only six could be confirmed.
What this argument is and is not
It would be easy to read the above and conclude that animal research does not work. That conclusion does not follow, and the same researchers who produced these audits do not draw it. They are, almost without exception, animal researchers.
What the evidence supports is narrower and more useful:
- Animal experiments are excellent for mechanism — for establishing what a gene or protein does in a living body. That is what the knockout technology was built for, and it delivers.
- Animal experiments are much weaker as predictors of clinical benefit. A drug that works in mice has cleared a low bar, not a high one.
- Much of the historical failure rate is attributable not to the animals but to how the experiments were run and reported — small samples, no randomisation, no blinding, and a publication record that printed the successes.
- The correct response is better weighting, not dismissal: treat a preclinical result as a reason to run a careful human study, never as evidence that a treatment works in people.
8. Animal Research Ethics, Stated Squarely
A page about knockout mice that does not address this is avoiding the obvious. What follows is the governing framework, the trade-off the technology created, and the disagreement — without an attempt to settle it.
The 3Rs
The actual ethical framework governing animal research in most of the world is the 3Rs, set out by the zoologist William Russell and the microbiologist Rex Burch in their 1959 book The Principles of Humane Experimental Technique. (The book predates PubMed's coverage and is not indexed there; it is a book, not a paper.) The three principles are:
- Replacement — use a non-animal method wherever one can answer the question: cell culture, tissue models, computational simulation, human volunteer studies.
- Reduction — use the smallest number of animals that will give a statistically sound answer. Note that this cuts both ways: an underpowered study wastes animals by producing an uninterpretable result, so proper sample-size calculation is a 3Rs obligation, not a statistical nicety.
- Refinement — minimise suffering in the animals that are used: anaesthesia and analgesia, humane endpoints so an animal is euthanised before a disease runs its full course, and housing that permits normal behaviour.
The 3Rs are written into law in the European Union (Directive 2010/63/EU), in the United Kingdom (the Animals (Scientific Procedures) Act 1986, under which every project requires a Home Office licence and a harm–benefit assessment), and in practice in the United States through the Animal Welfare Act, the Public Health Service Policy, and mandatory review by an Institutional Animal Care and Use Committee before any funded work begins. A 2015 review in the Journal of the American Association for Laboratory Animal Science argued that the three terms are used inconsistently across the field and would benefit from clearer definition — a reminder that a framework being universal is not the same as it being applied uniformly.
The trade-off gene targeting created
Here is the part that is usually skipped in both directions.
Knockout technology made experiments more precise. A knockout answers a question that could previously only be approached with crude tools — surgical ablation, poisons with multiple targets, drugs with off-target effects — and it answers it cleanly, often replacing a series of ambiguous experiments with one interpretable one. In several respects that is a Reduction and a Refinement.
It also massively increased the number of mice used, and this is not a small effect. Making and maintaining a genetically altered line consumes animals in a way that a simple experiment does not. Chimeras must be produced and bred. Offspring must be genotyped, and the ones carrying the wrong genotype — often most of a litter — are not used in any experiment. Lines must be maintained across generations, backcrossed onto standard genetic backgrounds, and cryopreserved or bred continuously. In jurisdictions that publish statistics, breeding and maintenance of genetically altered lines now accounts for a substantial fraction of all animal use, distinct from the experiments themselves.
So the honest accounting is that the technology improved the quality of individual experiments and increased the total number of animals. Both statements are true, and the ethical weight you assign to that depends on how you trade a count of animals against the information obtained.
The disagreement
Reasonable, informed people land in different places, and the positions do not reduce to caring versus not caring:
- One position holds that the medical benefit is enormous and demonstrable, that the 3Rs framework plus enforced welfare standards makes the practice defensible, and that abandoning animal research would stall the development of treatments for diseases that currently kill people.
- A second holds that the translation failure rate documented in the section above changes the calculation — that if a large proportion of animal studies do not usefully predict human outcomes, the benefit side of the harm–benefit assessment has been systematically overestimated, and far more should be replaced.
- A third holds that the question is not consequentialist at all — that causing suffering to a sentient animal for a benefit that animal cannot share is wrong regardless of how the arithmetic comes out.
These are not resolvable by citing another study, because the first two disagree about facts that further research could in principle settle, while the third disagrees about values that it could not. Our position is that a reader is better served by seeing the framework and the argument stated accurately than by being handed a conclusion.
9. What Came After: CRISPR, Organoids and Chips
CRISPR
The ES-cell route worked, and it was slow. Targeting a gene, screening colonies, injecting blastocysts, breeding chimeras and then breeding to homozygosity typically took a year or more per line, required real expertise in ES-cell culture, and worked well in the mouse and poorly in most other species — because good germ-line-competent ES cells were hard to derive elsewhere.
CRISPR-Cas9 removed most of that. Instead of relying on the low background rate of homologous recombination and then hunting for the rare success, CRISPR cuts the chromosome at a chosen site, which forces the cell to repair it — and repair is where the edit gets made. The frequency of useful events rises by orders of magnitude, and the targeting component is a short guide RNA that can be ordered rather than a construct that must be built.
The decisive practical consequence, demonstrated in 2013, is that CRISPR components can be injected directly into a one-cell embryo, producing modified mice in a single generation and modifying several genes at once. The ES-cell step drops out. The same approach works in species where ES cells were never available, which is why targeted rats, pigs, monkeys and zebrafish now exist.
ES-cell targeting has largely been displaced for making new lines. What has not been displaced is everything downstream of it: the logic of the experiment, the conditional Cre-lox design, the vast existing library of knockout lines built over three decades, and the conceptual framework of reverse genetics itself. CRISPR made Capecchi, Evans and Smithies's method faster. It did not replace the idea.
Organoids
Organoids are three-dimensional structures grown from stem cells that self-organise into something resembling a miniature organ — gut, kidney, liver, retina, or a rudimentary brain. Crucially they can be grown from human cells, including cells from a specific patient, so a drug can be tested against that individual's own tissue. In cystic fibrosis, patient-derived rectal organoids have been used to test how an individual responds to CFTR-modulating drugs, which is a direct answer to the problem the CF mouse could not solve. This work connects closely to Shinya Yamanaka's induced pluripotent stem cells, which are what make patient-specific starting material possible.
Organ-on-chip
An organ-on-a-chip is a microfluidic device in which human cells are cultured under mechanical and chemical conditions resembling the living organ — a lung chip that stretches with breathing, a gut chip with flow and peristalsis-like motion. These reproduce the physical environment that flat cell culture omits, and several can be linked to study how a drug metabolised by liver cells affects kidney cells downstream.
The honest limit of all three
None of these fully replaces a whole animal, and it is worth being precise about why. An organoid has no blood supply, no immune system, no nervous system and no endocrine signalling from other organs. A chip has a few cell types in a defined geometry. Neither ages, neither has a microbiome, and neither can tell you whether a drug given by mouth reaches its target at a useful concentration, or what it does to an organ you did not think to include.
For any question that is intrinsically about a whole organism — how a drug distributes and is cleared, whether a treatment causes harm somewhere unexpected, how a disease progresses over a lifetime, how the immune and nervous systems interact — there is currently no complete substitute. The realistic near-term trajectory is that these methods progressively take over the questions they can answer, shrinking the set of questions for which an animal is required, rather than eliminating that set. A reader is entitled to be sceptical of anyone claiming either that animal research will be replaced imminently or that it can never be.
One clarifying distinction, because the terms are constantly conflated in reporting: Andrew Fire and Craig Mello's RNA interference silences a gene's message without touching the gene — the DNA is unchanged, the effect is partial and reversible, and it is the basis of a class of drugs. Gene targeting and CRISPR edit the DNA itself — permanent, heritable, complete. "Knockdown" and "knockout" are not synonyms, and a study that used one has not shown what a study using the other would show.
10. What This Means When You Read Health News
If you take one practical thing from this page, take this.
"In mice" is the single most useful phrase to look for in a health headline. It is frequently absent from the headline and present in paragraph nine, or in the study's methods and nowhere else. Finding out which of these you are reading takes about fifteen seconds and changes the meaning of the article completely.
What a mouse result actually licenses
A positive result in mice means: in this species, in this artificial model, under these conditions, a measurable effect occurred. That is a real finding and it is worth something. It is a reason to do the next experiment.
It does not mean the treatment works in people. It does not mean it is safe in people. It does not mean the dose is achievable in people — mice are routinely given doses per kilogram that a human could not tolerate, or that would require eating implausible quantities of a food. It does not mean the mechanism is the same in people. And it does not mean a treatment is coming: the distance from a mouse result to an approved therapy is typically ten to fifteen years, most of the journey is failure, and the failures are not reported with anything like the enthusiasm of the original announcement.
The five questions
When you meet a health headline, these five questions sort most of the wheat from the chaff in under a minute:
- What species? Mice, rats, cells in a dish, or people? Cells in a dish is a weaker claim than mice. "Reduced cancer cell growth in the laboratory" means a compound was applied to cells in a plastic well, often at a concentration no human bloodstream could reach.
- Was it the disease, or a model of it? A mouse given a chemical that damages its joints is not a mouse with rheumatoid arthritis. A mouse with an artery tied off is not a person with a stroke. The gap between the model and the illness is where most translation failures live.
- How many, and was it blinded? Eight mice per group is normal and it is not many. If the write-up does not say whether the researchers assessing outcomes knew which animals got the treatment, assume they did, and discount accordingly.
- What was actually measured? A change in a blood marker, a smaller tumour on day 21, or an animal living longer are three very different claims. Marker changes are the weakest, and the most common.
- Who is saying it and when? A university press release describing work published this week is at the very beginning of the pipeline, and the press release is written to be noticed.
Calibrated interest, not dismissal
The wrong lesson here is cynicism. Mouse work is how almost everything in modern medicine started, including the treatments you or your family currently depend on. Statins, monoclonal antibodies, checkpoint inhibitors, and most of what is in a modern pharmacy passed through animal experiments on the way. Dismissing preclinical research as meaningless is as wrong as treating it as a cure announcement.
The right posture is calibrated interest. A mouse result is a promising early signal from a stage of research where most promising signals do not survive. File it. Note the mechanism, because the mechanism may be right even when the treatment is not. Do not change what you take, what you eat, or what you tell your doctor on the strength of it. And when the same idea reappears years later attached to a randomised controlled trial in humans, that is the point at which it has earned a decision.
A useful mental translation: when a headline says "scientists discover cure for X" and the study was in mice, read it as "scientists have found something worth testing properly, and we will know in about a decade." That version is almost always accurate, and it is the version the researchers themselves would recognise.
11. Where Mainstream Science Agrees — and What Remains Debated
Agreed
- Homologous recombination can be used to make a designed, specific change at a chosen site in a mammalian chromosome. This is not contested; it is the basis of a standard laboratory technique.
- Embryonic stem cells from a mouse blastocyst can be grown in culture, modified, returned to an embryo, and contribute to the germ line, producing a heritable strain.
- Positive–negative selection dramatically enriches for correctly targeted cells and made the method practical.
- A large fraction of mouse gene knockouts produce either no detectable phenotype or embryonic lethality, and this reflects genuine biological redundancy and the developmental essentiality of many genes rather than technical failure.
- Conditional (Cre-lox) targeting solves embryonic lethality by restricting deletion to a chosen tissue or time.
- Mouse models of human disease frequently fail to predict clinical outcomes, and preclinical research has historically suffered from publication bias, small sample sizes, and inadequate randomisation and blinding.
- CRISPR-based editing has largely superseded ES-cell targeting for generating new modified lines, without displacing the experimental logic those methods established.
Debated
- How badly mice model human disease overall. The 2013 and 2015 PNAS exchange on inflammatory conditions is the cleanest example: the same datasets, analysed with different gene-inclusion criteria, produced opposite headline conclusions. The general question is not settled, and it almost certainly has different answers for different diseases.
- How much of the translation failure is the species and how much is the methodology. If small, unblinded, unrandomised studies with selective publication were the main problem, then reformed preclinical research should predict clinical outcomes far better. Whether the reforms have delivered that is still being measured.
- How far non-animal methods can go. Estimates of what fraction of animal research organoids, chips and computational models could replace vary widely, and the estimates tend to correlate with the estimator's position on animal research generally.
- The ethical weight of the animal count. Whether the sharp increase in mice used for breeding and maintaining genetically altered lines is an acceptable cost of more precise experiments is a value judgement, and the disagreement is genuine rather than a misunderstanding of the facts.
- Whether standard inbred strains are the right animals. There is an active argument that genetically diverse mouse populations, older animals and animals with co-morbidities would predict human outcomes better than the near-identical young healthy mice that dominate published work.
12. Key Research Papers
- Evans MJ, Kaufman MH. Establishment in culture of pluripotential cells from mouse embryos. Nature 1981;292(5819):154-6 — the derivation of mouse embryonic stem cells. Independently that year: Martin GR. Isolation of a pluripotent cell line from early mouse embryos cultured in medium conditioned by teratocarcinoma stem cells. Proc Natl Acad Sci U S A 1981;78(12):7634-8, which is where the term "embryonic stem cell" comes from. The germ-line step followed from Evans's group: Kuehn MR, Bradley A, Robertson EJ, Evans MJ. A potential animal model for Lesch-Nyhan syndrome through introduction of HPRT mutations into mice. Nature 1987;326(6110):295-8 — mutations made in cultured ES cells transmitted through the germ line into mouse strains. Those particular mutations were made by retroviral insertion, not homologous recombination; the significance is the germ-line route.
- Smithies O, Gregg RG, Boggs SS, Koralewski MA, Kucherlapati RS. Insertion of DNA sequences into the human chromosomal beta-globin locus by homologous recombination. Nature 1985;317(6034):230-4 — the planned modification was achieved in about one per thousand transformed cells, whether or not the target gene was expressed.
- Thomas KR, Capecchi MR. Site-directed mutagenesis by gene targeting in mouse embryo-derived stem cells. Cell 1987;51(3):503-12 — gene targeting in ES cells, at a frequency of roughly 1 in 1,000 drug-resistant colonies. Reported the same year from Smithies's group: Doetschman T, Gregg RG, Maeda N, et al. Targetted correction of a mutant HPRT gene in mouse embryonic stem cells. Nature 1987;330(6148):576-8.
- Mansour SL, Thomas KR, Capecchi MR. Disruption of the proto-oncogene int-2 in mouse embryo-derived stem cells: a general strategy for targeting mutations to non-selectable genes. Nature 1988;336(6197):348-52 — positive–negative selection, reported to enrich roughly 2,000-fold for correctly targeted cells.
- Snouwaert JN, Brigman KK, Latour AM, et al. An animal model for cystic fibrosis made by gene targeting. Science 1992;257(5073):1083-8 — the first CFTR knockout mouse. The reported phenotype is intestinal (failure to thrive, meconium ileus, death from obstruction usually before 40 days), not the airway disease that dominates human cystic fibrosis. For a disease requiring a specific human protein variant rather than a simple deletion, see the paired 1997 reports: Ryan TM, Ciavatta DJ, Townes TM. Knockout-transgenic mouse model of sickle cell disease. Science 1997;278(5339):873-6, published alongside Pászty C, Brion CM, Manci E, et al. Transgenic knockout mice with exclusively human sickle hemoglobin and sickle cell disease. Science 1997;278(5339):876-8 — knockout and knock-in used together.
- Zhang SH, Reddick RL, Piedrahita JA, Maeda N. Spontaneous hypercholesterolemia and arterial lesions in mice lacking apolipoprotein E. Science 1992;258(5081):468-71 — roughly five times normal plasma cholesterol, foam-cell-rich aortic lesions by 3 months, severe coronary ostial occlusion by 8 months.
- Gu H, Marth JD, Orban PC, Mossmann H, Rajewsky K. Deletion of a DNA polymerase beta gene segment in T cells using cell type-specific gene targeting. Science 1994;265(5168):103-6 — the first tissue-restricted conditional knockout using Cre-lox. Timing control followed in Kühn R, Schwenk F, Aguet M, Rajewsky K. Inducible gene targeting in mice. Science 1995;269(5229):1427-9.
- Koller BH, Marrack P, Kappler JW, Smithies O. Normal development of mice deficient in beta 2M, MHC class I proteins, and CD8+ T cells. Science 1990;248(4960):1227-30 — an early targeted knockout in which the homozygotes appeared normal despite lacking detectable class I molecules and being grossly deficient in CD8-positive T cells. The general phenomenon is reviewed in Barbaric I, Miller G, Dear TN. Appearances can be deceiving: phenotypes of knockout mice. Brief Funct Genomic Proteomic 2007;6(2):91-103 — a review of why so many knockouts show no detectable phenotype, and of biological robustness. Read alongside Zambrowicz BP, Sands AT. Knockouts model the 100 best-selling drugs — will they model the next 100? Nat Rev Drug Discov 2003;2(1):38-51, a retrospective finding that knockout phenotypes correlated well with the known effects of drugs against the same targets.
- Dickinson ME, Flenniken AM, Ji X, et al.; International Mouse Phenotyping Consortium. High-throughput discovery of novel developmental phenotypes. Nature 2016;537(7621):508-14 — 410 lethal genes among the first 1,751 unique knockouts, an estimate that roughly one-third of mammalian genes are essential, and the finding that incomplete penetrance and variable expressivity are common even on a defined genetic background. A corrigendum was subsequently published (Nature 2017;551(7680):398) and should be read alongside the original.
- O'Collins VE, Macleod MR, Donnan GA, Horky LL, van der Worp BH, Howells DW. 1,026 experimental treatments in acute stroke. Ann Neurol 2006;59(3):467-77 — the canonical translation-failure audit: 114 drugs taken into clinical use or testing versus 912 tested only in animals, with no evidence that the ones taken forward had performed better experimentally. Its methodological companions: Sena ES, van der Worp HB, Bath PM, Howells DW, Macleod MR. Publication bias in reports of animal stroke studies leads to major overstatement of efficacy. PLoS Biol 2010;8(3):e1000344; van der Worp HB, Howells DW, Sena ES, et al. Can animal models of disease reliably inform human studies? PLoS Med 2010;7(3):e1000245; and the earlier standards document, Stroke Therapy Academic Industry Roundtable (STAIR). Recommendations for standards regarding preclinical neuroprotective and restorative drug development. Stroke 1999;30(12):2752-8.
- Kilkenny C, Browne WJ, Cuthill IC, Emerson M, Altman DG. Improving bioscience research reporting: the ARRIVE guidelines for reporting animal research. PLoS Biol 2010;8(6):e1000412 (co-published in several journals; this is the primary record), updated as Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0. PLoS Biol 2020;18(7):e3000410. Alongside them: Landis SC, Amara SG, Asadullah K, et al. A call for transparent reporting to optimize the predictive value of preclinical research. Nature 2012;490(7419):187-91, and the commentary by Begley CG, Ellis LM. Drug development: raise standards for preclinical cancer research. Nature 2012;483(7391):531-3, which reported that only six of 53 "landmark" preclinical cancer studies could be confirmed on attempted reproduction.
- Seok J, Warren HS, Cuenca AG, et al. Genomic responses in mouse models poorly mimic human inflammatory diseases. Proc Natl Acad Sci U S A 2013;110(9):3507-12 — and its direct rebuttal, Takao K, Miyakawa T. Genomic responses in mouse models greatly mimic human inflammatory diseases. Proc Natl Acad Sci U S A 2015;112(4):1167-72 (a subsequent erratum corrected details of that paper). Same datasets, different gene-inclusion criteria, opposite conclusions. Background on the underlying species differences: Mestas J, Hughes CC. Of mice and not men: differences between mouse and human immunology. J Immunol 2004;172(5):2731-8.
- Wang H, Yang H, Shivalila CS, et al. One-step generation of mice carrying mutations in multiple genes by CRISPR/Cas-mediated genome engineering. Cell 2013;153(4):910-8 — targeted mice made directly in the embryo, bypassing the ES-cell step. On the partial replacements: Shultz LD, Brehm MA, García-Martínez JV, Greiner DL. Humanized mice for immune system investigation: progress, promise and challenges. Nat Rev Immunol 2012;12(11):786-98; Clevers H. Modeling development and disease with organoids. Cell 2016;165(7):1586-97; Ingber DE. Human organs-on-chips for disease modelling, drug development and personalized medicine. Nat Rev Genet 2022;23(8):467-91.
- The three Nobel lectures, which are the primary sources for the biographical material on this page: Capecchi MR. The making of a scientist II (Nobel Lecture). Chembiochem 2008;9(10):1530-43; Evans M. Embryonic stem cells: the mouse source — vehicle for mammalian genetics and beyond (Nobel Lecture). Chembiochem 2008;9(11):1690-6; Smithies O. Turning pages (Nobel Lecture). Chembiochem 2008;9(9):1342-59. Also useful as an overview: Capecchi MR. Gene targeting in mice: functional analysis of the mammalian genome for the twenty-first century. Nat Rev Genet 2005;6(6):507-12.
A note on one source cited in section 8: Russell WMS and Burch RL, The Principles of Humane Experimental Technique (Methuen, London, 1959), is the origin of the 3Rs. It is a book published before PubMed's coverage begins and has no PubMed record, so no link is given here rather than a guessed one. The modern review discussing how its terms are used is Tannenbaum J, Bennett BT. Russell and Burch's 3Rs then and now: the need for clarity in definition and purpose. J Am Assoc Lab Anim Sci 2015;54(2):120-32.
Live PubMed Searches
- Gene targeting in embryonic stem cells in mice
- Conditional knockout with Cre-lox
- Preclinical to clinical translation failure
- ARRIVE guidelines and animal research reporting
- Humanized mouse models
13. Connections
- All Notable Doctors
- Nobel Prize in Physiology or Medicine — every prize from 1901 to the present, with what each one was actually for
- Genetics — inherited conditions, and how a mutation becomes a disease
- Shinya Yamanaka — the other stem-cell prize: turning an adult cell back into a pluripotent one, which is what makes patient-derived organoids possible
- Watson, Crick & Wilkins — the structure of DNA; Capecchi took his doctorate in Watson's laboratory
- Goldstein & Brown — the LDL receptor and cholesterol clearance, the mechanism the ApoE and LDL-receptor knockout mice were built to dissect
- Allison & Honjo — immune checkpoints and cancer immunotherapy, characterised largely through knockout mice
- Fire & Mello — RNA interference: silencing a gene's message rather than editing the gene, and why "knockdown" is not "knockout"
- Barbara McClintock — transposable elements, and the earlier demonstration that genomes are not static
- Hartwell, Hunt & Nurse — the cell cycle, worked out largely by finding and breaking individual genes
- Cystic Fibrosis — the disease the 1992 CFTR knockout mouse modelled only partly
- Sickle Cell Disease — modelled by knocking out mouse globin and knocking in the human sickle variant
- Atherosclerosis — where the ApoE knockout mouse became the standard research animal
- Stroke — the field in which 1,026 experimental treatments worked in animals and the translation record forced preclinical research to reform