Instructor Notes — Chapter 10
Teaching notes
What this chapter is actually for
One deliverable: rate the claim, not the molecule — under maximum pressure.
Chapter 5 stated the rule. Chapter 10 is where students find out whether they can apply it when a single drug carries ten claims across five rating tiers, and when both available shortcuts — "it's proven, so this probably works too" and "they claim everything, so I believe nothing" — feel reasonable.
Secondary deliverables: the ⚠️/🔬 boundary (the sharpest statement of it in the book), and the observational benefit/harm asymmetry.
Common misconceptions
"A drug that works on everything is a scam." Usually true, and not here — several indications have completed randomized trials with hard endpoints. Students who over-apply the heuristic get the wrong answer confidently, which is worse than getting it wrong tentatively.
"A drug that works on everything must be fundamental." The opposite error, and equally common in the room. Both appear, usually from different students, which makes for a good discussion.
"Stopped early for efficacy means the effect is even bigger than reported." The exact reverse. This is genuinely counterintuitive and worth the time in Case Study 1.
"Observational studies with millions of patients beat small trials." Size does not fix confounding. A database of ten million people who were selected into treatment by clinicians is not better evidence about benefit than a randomized trial of two thousand — it is a different kind of evidence, strong for harms and weak for benefits.
"It's just the weight loss" is a criticism. Sometimes. For sleep apnea it is a complete explanation. §10.6 and answer 10.30 draw the line: mechanism bears on extrapolation, not on whether an established outcome occurred.
"⚠️ and 🔬 are basically the same." They are not, and §10.7 exists for this. Students who blur them will misrate every frontier compound in Parts V and VI.
The hardest point to teach
Why the Alzheimer's claim is 🔬 and not ⚠️.
Students find this distinction fussy until they see what it protects. Ask: what else would be ⚠️ if we put this there? Compounds that have actually been randomized in humans and produced ambiguous results — Semax, MK-677, thymosin alpha-1. Placing an unrandomized program alongside them erases the difference between "tested and unclear" and "not yet tested."
Then the second move: the observational data is the strongest-looking evidence and the weakest kind. Millions of patients, a clear association, and five biases all pushing the same direction. Students consistently overweight database size, and this is the cleanest available correction.
Demonstrations that work
Build the ten-row table live. Give students the ten claims with populations and have them assign ratings before reading §10.10. Then reveal. The disagreements cluster on MASH, HFpEF, and Alzheimer's, and the discussion of why is the lesson.
The interim-stopping simulation. Write a "true effect" of 20% on the board. Then generate five plausible interim observations by hand — 12%, 31%, 18%, 24%, 21% — and ask which one the stopping rule would report. Students see the selection immediately when the numbers are in front of them, and rarely before.
Two sentences, two errors. Put both compression errors on the board as quotations and ask which is worse. There is no correct answer; the point is that students argue and then notice both sentences have the same structure.
ClinicalTrials.gov, live. Search semaglutide, sort by condition, and read the status column. The number of conditions surprises them; the number of "recruiting" and "active, not recruiting" entries does the teaching.
Timing
For a 75-minute session:
| Minutes | Content |
|---|---|
| 0–8 | §10.1 — the four explanations. Frame the chapter's suspicion. |
| 8–20 | §10.3 and Case Study 1 — early stopping. The simulation. |
| 20–30 | §10.4 — the biopsy endpoint problem. |
| 30–40 | §10.5–10.6 — mechanical vs metabolic. The sleep apnea contrast. |
| 40–55 | §10.7–10.8 and Case Study 2 — the ⚠️/🔬 boundary and confounding by indication. |
| 55–68 | §10.10 — build the table live. The two compression errors. |
| 68–75 | §10.9 briefly; assign the split-by-indication dossier task. |
If short, cut §10.9 to reading. Do not cut the table exercise — it is the chapter's assessment in disguise.
Assessment notes
Discriminating items: 10.15, 10.18, 10.20, 10.23, 10.30.
10.20 (why 🔬 and not ⚠️) is the best single item; it requires students to apply a definitional distinction to a case where the evidence feels substantial.
10.18 (design a study to test weight mediation) is worth grading generously — the correct answer includes recognizing that the study may be practically infeasible, and students who reach "this question may not be cleanly answerable" have understood something most sources do not acknowledge.
10.29 (bet against one indication, dated) should be collected and kept. It is called back in Chapter 40 alongside the Chapter 1 prior-belief exercise and the Chapter 6 falsifier exercise.