Case Study 2 — The Phase 2 Graveyard
Why the most exciting result in a drug's life is the least reliable one
Type: Real, public, structural · Tier 1 process facts, Tier 2 magnitudes · Relevance: §9.6, §9.9, §9.10
Background: the shape of drug development
Chapter 5 §5.4 gave the trial phases. This case study is about a specific and reliable consequence of their structure.
Roughly nine in ten compounds that enter human trials never reach approval. That attrition is not evenly distributed. A substantial share of it happens at the Phase 2 to Phase 3 transition — the point at which a promising signal is asked to survive a properly powered test.
And the compounds that fail there mostly do not fail because they were fraudulent. They fail because the Phase 2 result was larger than the underlying truth, for structural reasons that operate on honest investigators running well-designed trials.
Why Phase 2 flatters
Five mechanisms, each independently sufficient to inflate a result, and they compound.
WHY PHASE 2 EFFECT SIZES EXCEED PHASE 3
① SMALL SAMPLES → HIGH VARIANCE
A trial of 80 people produces a much wider range of possible estimates than
one of 2,000. The unusually good estimates are the ones that generate
excitement and press coverage. The unusually bad ones quietly end programs.
② REGRESSION TO THE MEAN
Any estimate that is unusually far from the truth is likely to be followed by
one closer to it. A striking Phase 2 result is, on average, an overestimate —
BECAUSE it was striking enough to advance.
③ BEST-DOSE SELECTION
Phase 2 studies several doses. The best-performing one is quoted, taken into
Phase 3, and reported against. Selecting a maximum from several noisy
estimates produces a biased maximum.
④ FAVORABLE POPULATIONS
Phase 2 populations are often narrower, healthier, more adherent, and more
motivated than the Phase 3 population — and far more than the eventual
real-world population.
⑤ SHORTER DURATION
Effects that attenuate over time look better when measured earlier. Adverse
effects that accumulate look better too.
────────────────────────────────────────────────────────────────────────────
NONE OF THESE REQUIRES ANYONE TO DO ANYTHING WRONG. All five operate on
competent investigators running well-designed trials in good faith.
Mechanism ② deserves special attention because it is the least intuitive. The compounds that advance from Phase 2 to Phase 3 are selected for having produced a good result. That selection is appropriate — it is what Phase 2 is for. But it means the set of compounds entering Phase 3 is enriched for overestimates, and the Phase 3 result is therefore expected to be smaller on average even if the compound works exactly as well as it truly does.
This is not pessimism. It is arithmetic, and it applies to every drug in development including the ones that go on to succeed.
The historical pattern in this therapeutic area
Chapter 8's Case Study 2 covered the weight-loss drugs that failed on safety. This is the complementary list: compounds whose efficacy did not survive the transition.
The specifics vary and this book will not itemize individual programs it has not verified in detail. The pattern, however, is well documented across therapeutic areas and is not disputed:
- Compounds with striking Phase 2 weight effects that produced meaningfully smaller Phase 3 effects
- Compounds whose Phase 2 tolerability looked acceptable and whose Phase 3 discontinuation rates did not
- Compounds where a favorable Phase 2 population masked a much weaker effect in the intended one
- Combination products whose components' individual Phase 2 results did not add as expected
And, crucially, compounds that went the other way — where a modest Phase 2 signal became a substantial Phase 3 result. This is rarer, and it matters because it shows the effect is a distributional bias rather than a law.
📊 Evidence Rating — the general claim
Claim: A striking Phase 2 result predicts a comparable Phase 3 result.
Rating: ❌ Hype outpaces evidence — for the strong version of the claim.
Why: Phase 2 results are systematically larger than the Phase 3 results that follow, for five structural reasons that operate on honest investigators. Roughly nine in ten compounds entering human trials never reach approval, with substantial attrition at exactly this transition.
What this does NOT say: that Phase 2 results are worthless, or that a compound with a striking Phase 2 result will fail. Phase 2 is how drug development works, and every approved drug has one. The claim being rated is the inference, not the data.
What would change it: a demonstration that, in a defined therapeutic area with modern trial standards, Phase 2 effect sizes predict Phase 3 effect sizes without systematic inflation.
What this means for retatrutide specifically
Retatrutide's reported ~−24% at 48 weeks in Phase 2 is a genuinely striking result. Three things follow.
It is a real finding. Randomized, in humans, with a control arm. It belongs meaningfully higher on the evidence ladder than anything in Part III of this book.
It is expected to shrink. Not because anything is wrong with it, but because that is what Phase 2 figures do. A Phase 3 result somewhere below −24% would be entirely consistent with the drug working exactly as well as the Phase 2 data suggests it truly does.
And comparing it to tirzepatide's −21% is comparing incommensurable quantities. One is a Phase 2 estimate from a smaller trial at the best-performing dose; the other is a Phase 3 result in the population the drug was approved for. Ranking them is not a close call requiring judgment; it is a category error.
The honest counterweight
A chapter this insistent on Phase 2 caution owes the other side.
Phase 3 results are not truth either. They are larger, longer, better-powered estimates in more representative populations — which makes them better estimates, not perfect ones. Phase 3 populations still differ from real-world populations, trials still exclude comorbid and complicated patients, and adherence in a trial exceeds adherence outside one. Phase 4 exists because Phase 3 is incomplete.
And excessive Phase 2 skepticism has a cost. Compounds that would help people are abandoned when funders over-discount promising early data. The correct posture is not "ignore Phase 2" — it is "read Phase 2 as a signal that justifies Phase 3, and do not read it as a preview of Phase 3's number."
Which is, precisely, what ⚠️ means in this book's rating system.
Discussion questions
-
Explain regression to the mean as it applies here, and specifically why selecting compounds for having good Phase 2 results is what creates the bias.
-
Of the five mechanisms, which do you think contributes most? What evidence would settle it?
-
Some compounds go the other way — modest Phase 2, substantial Phase 3. What does the existence of those cases establish, and what does it not?
-
A funder must decide whether to invest in Phase 3 for a compound with a striking Phase 2 result. How should they discount it? What information would most improve the decision?
-
This case study rates an inference rather than a compound. Is that a legitimate use of the rating system? What other inferences in this book could be rated this way?
-
Apply it. Take any compound in your dossier whose supporting evidence is Phase 2 or earlier. Write what you would expect a properly powered Phase 3 to find, and why. Then note that you have just made a falsifiable prediction — which is more than most claims in this space manage.