Instructor Notes — Chapter 5

Teaching notes

What this chapter is actually for

This is the course. Everything before it is preparation and everything after is application.

If you have limited time and must teach a subset of this book, teach Chapter 5 and Chapter 17 and skip the rest. A student who can apply the four-tier rating system with a population, an endpoint, and a falsifier has acquired something durable. A student who has memorized which peptides are ✅ has acquired a list that expires.

Budget two sessions minimum, three if you can. Attempting this in 75 minutes produces students who recognize the vocabulary and cannot use it.

Common misconceptions

"Statistically significant means important." The most persistent and most consequential. A large trial can find a clinically meaningless difference at p < 0.001. The fix that works: give students a result with a tiny effect and a spectacular p-value and ask whether they would take the drug.

"Randomized means good." Randomization is necessary and not sufficient. A randomized trial can be underpowered, too short, unblinded with a subjective endpoint, or measuring a surrogate nobody has validated. Walk the six features.

"Anecdotes are just weak evidence." They are worse than weak for efficacy — they are close to uninformative, because improvement is common regardless of treatment. Regression to the mean is the concept that lands: anyone who starts a treatment when they feel worst will feel better afterward, whether or not it works. Students frequently report recognizing this in their own experience.

"Animal studies don't count." The overcorrection, and it appears fast in skeptical students. Animal work is essential, it is often excellent, and no drug reaches humans without it. The correct posture — a reason to run the trial, not a preview of its result — is harder to hold than either extreme, which is why §5.3 states it explicitly.

"A ❌ means it doesn't work." The single most important correction in the course and it needs repeating in every subsequent chapter. Students hear "no evidence" and translate to "no effect" automatically. The reframe: ❌ describes the literature; it does not describe the molecule.

"Relative and absolute are both fine, just different framings." They are both accurate and they are not equally honest in isolation. The three-baseline-risk table settles it.

The hardest points to teach

Estimands. Genuinely difficult and worth the time, because students will encounter the number in the wild. What works: pose the question as two different questions rather than as two analyses. "What happens if your doctor prescribes this?" versus "What happens if you take it?" Both are real questions people ask. They have different answers. Neither is the answer.

Then note the asymmetry: the efficacy estimand is always larger, so an unnamed number is systematically optimistic. Students find this more memorable than the definitions.

Rule 4 — never downgrade with distaste. Rule 3 (no upgrading with mechanism) is easy to teach because students can see the error in others. Rule 4 is invisible from inside. Nobody experiences their own skepticism as motivated.

The exercise that exposes it is 5.46 (count your ratings, notice the distribution). Do it in class if the group has enough trust, and be prepared for the discussion to become personal.

Demonstrations that work

The two claims, cold. Open the session by putting both claims on the board with no context: semaglutide produces substantial weight loss and BPC-157 heals tendon injuries. Ask students to rate their confidence in each, 1–10, on paper, privately. Collect nothing. Return to it at the end of the second session. The movement is the lesson, and it works better privately than aloud.

Live registry check. Project ClinicalTrials.gov. Search a compound with a strong online reputation and few trials. Filter PubMed for randomized controlled trials on the same compound. Do this live, not from a screenshot. Watching the count come back as zero is more persuasive than any argument, and students will do it themselves afterward.

The relative/absolute conversion, on the board. Take SELECT's 20% and derive 1.5 percentage points and an NNT in front of them. Then have them write the three headlines from exercise 5.28. The moment where they realize the honest version cannot be a headline is the moment the chapter's argument about media becomes structural rather than cynical.

The surrogate that killed people. Case Study 1 needs no embellishment. Tell it as a story, hold the ending back, and ask students to predict the result before revealing it. Most predict benefit.

Timing

Session 1 (75 min) — the ladder and the trial:

Minutes Content
0–8 The two claims, cold. Private confidence ratings.
8–20 §5.1–5.2 — "how would we know" and the ladder.
20–38 §5.3 — animal evidence. Figure 5.4 worked in full.
38–50 §5.4–5.5 — phases and the six features.
50–68 §5.6 — surrogates. Case Study 1 as a story.
68–75 Assign §5.7–5.9 as reading.

Session 2 (75 min) — the numbers and the ratings:

Minutes Content
0–15 §5.7 — p-values, CIs, HR, NNT. Fast; they read it.
15–35 §5.8 — relative vs absolute. Board work. The three headlines.
35–50 §5.9 — estimands. Two questions, not two analyses.
50–58 §5.10 — red flags. The live registry check.
58–72 §5.11–5.12 — the rating system and the two worked ratings.
72–75 Return to the cold confidence ratings. Assign Fields 5 and 6.

Assessment notes

The discriminating items are 5.12, 5.21, 5.28, 5.31, 5.40, and 5.43.

5.43 (rate a compound you personally hope works) is the one that matters. There is no key. Grade on whether the falsifier is specific enough to hand to a trialist, and on the honesty of the reflection. Many students report being unable to write a falsifier at all — flag this as the correct and valuable outcome, not a failure, or they will fabricate one to get the mark.

5.46 (count your ratings and look for drift) likewise: grade on honesty, and say so in advance, or you will receive thirty paragraphs claiming no bias.