Case Study 1 — The Trial That Stopped Itself
Data monitoring committees, interim analysis, and the price of stopping early
Type: Real, public, structural · Tier 1 process facts, Tier 2 magnitudes · Relevance: §10.3
Background: who decides a trial is over
A large randomized trial does not simply run to its planned end and reveal its answer. While it runs, an independent data monitoring committee — statisticians and clinicians with no stake in the outcome — periodically looks at unblinded results that nobody else, including the sponsor and the investigators, is permitted to see.
The committee exists for two reasons, and they are asymmetric.
To stop for harm. If the treatment is hurting people, continuing is indefensible. This is the uncontroversial case and it is why such committees became standard.
To stop for efficacy. If the treatment is clearly working, continuing to assign half of the participants to placebo becomes ethically difficult. This case is contested, and it is the subject of this case study.
To stop for either reason, the committee applies a pre-specified statistical boundary — a threshold that the interim result must cross. The boundary is deliberately more stringent than the final analysis would be, precisely because looking at data repeatedly increases the chance of seeing something striking by accident.
What happened with FLOW
FLOW examined semaglutide in adults with type 2 diabetes and chronic kidney disease, with a composite kidney endpoint. Its data monitoring committee recommended stopping early because the efficacy boundary had been crossed.
That is genuinely good news. A treatment that slows kidney disease progression in this population matters a great deal — kidney failure means dialysis or transplant, and the available options are limited.
And it makes the result harder to read, for a reason worth working through carefully.
Why truncated trials overestimate
THE SELECTION MECHANISM
Imagine a drug whose TRUE effect is a 20% relative risk reduction.
A trial with interim analyses looks at the data several times.
At any given look, the OBSERVED effect scatters around the truth by chance:
look 1: observed 12% → below boundary, continue
look 2: observed 31% → ABOVE boundary, STOP
look 3: would have been 18% (never happens — trial stopped)
final: would have been 21% (never happens)
THE TRIAL REPORTS 31%.
Not because anyone did anything wrong. The stopping rule SELECTED the look at
which chance happened to favor the treatment most — because that is precisely
what "crossed the boundary" means.
────────────────────────────────────────────────────────────────────────────
The effect is real (20%). The reported estimate (31%) is inflated.
The DIRECTION is trustworthy. The MAGNITUDE is not.
────────────────────────────────────────────────────────────────────────────
Three consequences.
The point estimate is biased upward. How much depends on how early the trial stopped and how many looks preceded it. Stopping very early with few events produces the largest inflation.
The confidence interval is wider. Fewer events means less precision, so the estimate is both inflated and less certain — an unfortunate combination.
And secondary endpoints are worse. A trial stopped when its primary endpoint crossed a boundary has not accrued enough events for its secondary endpoints, so any secondary findings are on much thinner ice than they would have been.
🔬 Read the Study — the meta-question
text FIGURE 10.CS1 — "Do stopped trials overestimate?" [real methodological literature] THE STUDY Not a single study — a body of methodological work comparing effect estimates from trials stopped early for benefit against estimates from trials that ran to completion, and against subsequent meta-analyses. THE QUESTION Do trials truncated for efficacy systematically report larger effects than the underlying truth? WHAT IT SHOWS Yes, consistently, and the inflation is largest when trials stop with few accrued events. This is well established methodologically and is not controversial among trialists. THE DOESN'T It does not show that any particular stopped trial's result is wrong, or that stopping was the wrong decision. The bias is a property of the STOPPING RULE, not of any individual trial's conduct. THE VERDICT Established. This is why guidance recommends caution in interpreting effect magnitudes from truncated trials. THE LESSON A statistical property of a decision rule can bias a literature without anybody making an error. Compare Chapter 9's Phase 2 selection and Chapter 6's testimonial survivorship — the same logic, three settings.
The genuine dilemma
This case study is not an argument against stopping trials early. The dilemma is real and has no clean resolution.
The case for stopping. Participants in the placebo arm have a condition that is progressing. If the treatment works, every additional month of the trial is a month in which some of them get worse avoidably. Trial participants consented to uncertainty, not to being kept in a control arm after the uncertainty resolved. And the results reach everyone else sooner.
The case for continuing. The uncertainty may not have resolved as much as the boundary suggests. Longer follow-up would produce a more precise estimate, better safety data, and usable secondary endpoints. And a treatment adopted on an inflated estimate may be used more widely than a truer estimate would justify — which harms a different, larger, invisible group of future patients.
Note the structure of the dilemma: it trades a visible benefit to identifiable current participants against an invisible cost to future patients who will be treated on the basis of a less precise estimate. Decision-making under that asymmetry reliably favors the visible party, and that is worth naming rather than resolving.
What good practice looks like: pre-specifying stringent boundaries, requiring a minimum number of accrued events before stopping is permitted, reporting the interim analysis history, and — most usefully — stating plainly in the publication that the estimate is likely inflated.
What this case teaches
A result can be trustworthy in direction and unreliable in magnitude. These come apart more often than most readers assume, and the distinction is worth carrying: does this drug help? and by how much? are different questions with different evidentiary requirements.
Selection effects appear everywhere in this book. Chapter 6's testimonials, Chapter 9's Phase 2 results, and now interim stopping — three settings, one logic: when something is reported because it crossed a threshold, the reported value is enriched for chance.
And process protections have costs. The data monitoring committee is a genuine safeguard, invented for good reasons, and one of its functions imposes a statistical price. That is normal for safeguards and it is not an argument against them.
Discussion questions
-
Work through the selection mechanism in your own words. Why does the stopping rule select the look at which chance most favored the treatment?
-
Someone argues: "if the boundary is stringent enough, the inflation is negligible." Evaluate. What determines how much inflation there is?
-
The dilemma trades visible benefit to current participants against invisible cost to future patients. Is there any mechanism that could correct for the asymmetry in how those are weighed?
-
FLOW's result is rated ✅ in this book despite the early stopping. Is that consistent? What would have to be true for early stopping to reduce a rating rather than merely qualify a magnitude?
-
A trial stopped early has weak secondary endpoint data. Suppose a secondary endpoint in such a trial looks favorable and generates enthusiasm. What should you conclude, and what would settle it?
-
Compare three selection effects. Testimonials (Ch 6), Phase 2 results (Ch 9), and interim stopping (here). Write one sentence identifying the common structure, then one sentence on what differs between them.