Case Study 35.2 — What AlphaFold Did and Did Not Change
A claim-evaluation exercise. The method is Chapter 5's. The claim is about a technology rather than a drug, and that is the point.
Why do this at all
Chapter 5 gave you a procedure for evaluating a claim about a molecule: restate the claim with a population and an endpoint, find out what was actually measured, ask what the comparator was, check whether the result would have been reported had it come out the other way, and name what would change your mind.
Almost everyone who learns that procedure applies it only to drugs. That is a mistake, and this case exists to correct it. The procedure is not about pharmacology. It is about the general problem of deciding how much to believe a confident sentence, and technology claims are, if anything, harder — because they are usually true in some form, which makes the exaggeration harder to see.
So: take the highest-profile scientific claim of the last several years, and run the method on it.
The claim, as usually encountered
You will meet it in roughly this form, in a newspaper, a conference keynote, or an investor deck:
"AI has solved protein folding. We can now design drugs on a computer. AI will find cures for diseases that have resisted treatment for decades, and it will do it in years instead of decades."
Four sentences. They are not equally true. Pulling them apart is the whole exercise.
Step 1 — Restate the claim with a population and an endpoint
Chapter 5's first move is always the same, and it is always the move that does the most work.
The sentence "AI has solved protein folding" has an implicit population (proteins) and an implicit endpoint (structure prediction accuracy). That is a technical claim, and it is testable.
The sentence "AI will find cures... in years instead of decades" has a very different population (patients with diseases) and a very different endpoint (approved treatments, delivered faster). That is a clinical claim, and it is also testable — just not by any of the same evidence.
The four-sentence version blends the two so that evidence for the first is heard as evidence for the second. Once you have separated them, most of the analysis is done.
Step 2 — What was actually measured?
For the technical claim, the answer is unusually clean, and it deserves to be admired rather than grudgingly conceded.
CASP — the Critical Assessment of Structure Prediction — is a biennial blind assessment. Organizers take proteins whose structures have been determined experimentally but not published. Teams submit predictions without access to the answers. Predictions are scored against the withheld experimental structures.
Consider what that design rules out. The predictor cannot have seen the answer. The assessors are not the predictors. The targets were chosen by neither. The scoring metric was agreed in advance. This is a better-controlled evaluation than a great many clinical trials.
At CASP14, in 2020, AlphaFold2 produced accuracy that was, for a large fraction of targets, competitive with experimental determination and far beyond anything previously achieved. The assessment's own organizers described the problem as substantially solved for single domains.
Then came the part that turned a result into infrastructure: hundreds of millions of predicted structures were released publicly. And then the 2024 Nobel Prize in Chemistry, shared by Demis Hassabis and John Jumper for this work and David Baker for computational protein design.
Verdict on the technical claim: ✅. Blind assessment, decisive margin, independent replication in the form of extremely broad adoption across structural biology. If you were looking for an example of a strong technical result, this is close to the ideal case.
Now the clinical claim.
Step 3 — What was measured for the second claim?
Nothing yet.
That is not a rhetorical flourish. As of this writing, the evidence offered for "AI will produce cures faster" consists overwhelmingly of input metrics: numbers of candidates generated, numbers of targets modeled, numbers of molecules entering preclinical development, speed from target to lead. Those are real numbers and they are genuinely impressive.
They are also, in Chapter 16's exact sense, surrogates. And they are surrogates for an outcome that Chapters 9 and 10 already told you is governed by something else entirely.
Recall the base rates. Most compounds entering human trials never reach approval. The failures concentrate in the late, expensive phases. And the dominant reasons are lack of efficacy in humans and unacceptable toxicity — not a shortage of molecules to test.
So the second claim rests on a surrogate whose relationship to the outcome is not merely unproven but has a specific reason to be weak.
Step 4 — What did the technology not change?
Four things, and they are worth holding separately because they fail in different ways.
Structure is not function. A shape does not tell you what a protein does, when, in which tissue, or in response to what. Target selection — which is a biology question — dominates drug discovery outcomes, and no structure answers it. Chapter 22's substance P antagonists had excellent pharmacology against a target that turned out to be the wrong one.
Short peptides frequently have no single structure to predict. This is the limitation that matters most for this book. Many short peptides are largely disordered in solution and adopt a conformation only on binding. Asking for "the structure" of such a molecule is close to a malformed question, and a prediction tool will nonetheless return something.
Predicting a structure is not predicting a drug. Potency at achievable concentrations, selectivity, pharmacokinetics, toxicity, manufacturability, and clinical benefit are all downstream, and each kills programs routinely.
The predictions are predictions. Confidence scores vary, and they are systematically lowest for disordered regions — which is correct behavior for the model and inconvenient for biology, since disordered regions are where a great deal of regulatory interaction lives.
Step 5 — Name what would change your mind
This is the step people skip, and it is the step that separates evaluation from opinion.
The clinical claim is falsifiable, and here is the study.
Define a cohort: every compound entering Phase 1 within a stated window, classified by a pre-specified, auditable definition of "AI-derived." Register that definition before outcomes are known — otherwise it will drift toward the successes, which is the technology-claim equivalent of moving an endpoint. Follow the cohort to approval. Compare against the historical base rate for matched therapeutic areas and modalities.
If the AI-derived cohort's approval rate materially exceeds the matched base rate, the strong claim is vindicated and the rating moves to ✅.
If it matches or falls below the base rate while more candidates enter, then the technology increased throughput without changing the odds.
Notice that the second outcome is not a debunking. It describes a world in which the technology is real, valuable, and was sold with the wrong headline. That is the most common way a ⚠️ resolves, and recognizing it in advance is what keeps you from swinging between credulity and cynicism when the data arrives.
Duration of that study: roughly ten to fifteen years. Which is itself the chapter's argument, applied to the chapter's own claim. You cannot evaluate a claim about drug approvals faster than drugs get approved.
The verdicts, side by side
| The claim | Rating | Why |
|---|---|---|
| AlphaFold-class structure prediction is a transformative research tool | ✅ | Blind assessment at CASP; extremely broad adoption |
| De novo designed peptide binders are effective therapeutics | 🔬 | Real, published, prospectively promising; almost nothing in late-stage clinical use |
| Claim form — "AI has revolutionized drug discovery" | ⚠️ | Conflates candidate generation with approvals; attrition is governed by efficacy and toxicity in humans |
Three ratings, one technology, no contradiction. This is rating rule 6 — one molecule, many ratings — applied to something that is not a molecule at all.
Discussion questions
1. The technical claim and the clinical claim are supported by completely different kinds of evidence, yet they are routinely presented in the same sentence. Write out the version of the four-sentence claim that you would consider fully accurate. Is it something a keynote speaker could plausibly say? If not, what does that tell you about the incentives operating on technology communication?
2. Apply Chapter 5's "would this have been reported if it came out the other way?" test to CASP. Explain specifically what features of the CASP design make the answer yes — and then name a technology evaluation you have encountered where the answer would clearly be no.
3. Of the four limitations in Step 4, one is much more specific to this book's subject matter than the others. Identify it, explain why it lands hardest on peptides, and say whether it is a limitation of the technology or a statement about which problem the technology solved. Does that distinction matter?
4. Design the falsifying study from Step 5 in more detail. Where would you expect the disagreement to be fiercest — the definition of "AI-derived," the choice of comparator, or the endpoint? Whichever you pick, explain how a motivated party could exploit ambiguity in that component to guarantee a favorable answer.
5. Suppose the study runs and the AI-derived cohort clears at approximately the historical base rate, with three times as many candidates entering. Write the four-line evidence rating you would issue for the claim form at that point. Then write the press release a company with a large computational platform would issue on the same data. Compare them, and say which sentence in the press release you would challenge first.
6. This case applies a drug-evaluation method to a technology claim. Name two other kinds of claim in your own life — professional or otherwise — where the same five steps would apply. Then name one kind of claim where you think the method would break down, and say what about it resists the treatment.