Where Students Get Stuck

Ten failure modes, in rough order of how often they appear. Each is described the same way: what it looks like on the page or in the room, why it happens — because the cause determines the fix — and the intervention, which is a specific thing to do rather than a thing to say.

Most of these are not ignorance. Nine of the ten are a reasonable habit misapplied, which is why telling students they are wrong rarely works and giving them a case where their habit produces an absurd answer works almost immediately.


1. Rating the molecule instead of the claim

What it looks like. "BPC-157: ❌." "Semaglutide: ✅." A dossier field 6 with one row in it. In discussion: is oxytocin good or not? On an exam, a rating with no population and no endpoint attached, often with a confident justification underneath that would be perfectly adequate if the claim had been stated.

Why it happens. Because every rating system students have ever encountered rates things. Films, restaurants, appliances, hotels. The entire ambient grammar of stars-out-of-five attaches to an object, and the book is asking them to attach a rating to a proposition instead — a move with essentially no precedent in their experience. It is not carelessness; it is a strong prior.

It is also the error the whole book is built to prevent. Rule 1 and Rule 6 are two statements of the same insight, and a student who holds both has most of the method.

The intervention. Do not explain. Demonstrate the contradiction. Put a single molecule on the board and ask the class to rate it — semaglutide works best because the answer is genuinely different across claims. Take the ratings, then reveal that the same molecule carries ✅ for weight loss in adults with overweight or obesity, ✅ for glycemic control in type 2 diabetes, 🔬 for Alzheimer's disease, and ❌ for longevity in healthy adults, all simultaneously and all correctly. The room resolves it in one move: those are four different questions.

Then make the fix structural rather than conceptual. Require every rating in the course to be written in a four-line format, on exams, in dossiers, in discussion posts:

CLAIM:       [what, in whom, measured how]
EVIDENCE:    [best available, described by design]
RATING:      [✅ / ⚠️ / ❌ / 🔬]  +  [date]
WOULD CHANGE IF: [the finding that would move it]

A student cannot rate a molecule in this format. There is nowhere to put it. Formatting requirements do pedagogical work that explanation does not, and this is the clearest example of it in the course.


2. Reading ❌ as "this is bad"

What it looks like. A student describes a ❌ compound as dangerous, or as a scam, or as "debunked." The emotional temperature in discussion of ❌ compounds runs noticeably higher than the evidence supports. Occasionally the reverse: a student defends a compound they use by arguing it is not harmful — which is a real point about a claim nobody made.

Why it happens. ❌ is a red X. It reads as a mark on a test, and marks on tests are about the answer being wrong. Rule 2 exists because the symbol fights the meaning: ❌ describes the evidence, not the molecule. It says hype outpaces evidence — a statement about the gap between what is claimed and what has been shown, and silent on whether the molecule is harmful, useless, or destined to fail.

The intervention. Give them a case where the two come apart, and let them find it. A compound that is ❌ for its marketed claim and simultaneously ✅ for an approved indication makes the point instantly — the rating moved, the molecule did not. Chapter 37's table has several.

Then install a verbal habit and enforce it for a full session: students may not say "X is ❌." They must say "the claim that X does Y in Z is ❌." It sounds pedantic for twenty minutes and then stops being noticeable, which is when it has been learned. A sentence to hand them: "❌ is a statement about the literature. If the literature changes, the ❌ changes, and the molecule never moved."


3. Confusing "untested" with "tested and failed"

What it looks like. Asked which of two ❌ claims rests on stronger knowledge, students overwhelmingly choose the untested one. Asked to justify it, they say something like: at least it hasn't failed. On a dossier, the "which kind of ❌?" line is left blank or filled in wrongly. In discussion, "there's no evidence against it" is offered as a point in a compound's favor.

Why it happens. Because in ordinary life, an accusation that has not been proven is better for the accused than one that has. Students are importing a presumption-of-innocence frame, which is correct in a courtroom and backwards in an evidence base. Absence of evidence is a state of ignorance. A negative result is a state of knowledge. The second is much stronger, and it is stronger in the direction of knowing, not in the direction of the compound being good.

The intervention. Two columns on the board, and make the class populate them:

Evidence absent Evidence present and negative
What we know about the claim Nothing That it did not hold, in that population, on that endpoint
What would change it The first adequate trial Replication, or a different population/endpoint
Honest one-line summary "Nobody has looked" "Somebody looked, and the answer was no"
Which is a stronger state of knowledge? This one

Then use the pairing the book returns to: the field knows a great deal about a compound that failed a large outcome trial and very little about one that has never had one, even though both may sit at ❌. Students who feel that reversal once do not usually re-make the error.

Make the distinction mandatory on the page. Every ❌ in student work must be tagged absent or present-and-negative. It takes four words and it forces the thought.


4. Letting a good mechanism upgrade a rating

What it looks like. A dossier field 3 that runs three times the length of field 5. A rating of ⚠️ justified entirely by receptor biology. In discussion: but it makes sense that it would work. On an exam, an otherwise excellent answer that arrives at ⚠️ from a mechanism paper and no clinical data at all.

Why it happens. Mechanism is satisfying. It is a story with a causal arrow in it, it is what the course's biology component rewarded, and it is the easiest field in the dossier to fill — you can write a competent paragraph of mechanism from a review article in ten minutes, and you cannot write field 5 that way. Students drift toward mechanism the way water finds a slope. Appendix C says this outright: mechanism is where an unstructured effort always ends up.

One clarification worth making to stronger students: Rule 3 is not a claim that mechanism is uninteresting, but a claim about what it licenses — a hypothesis, not a rating.

The intervention. Chapter 22 is the whole intervention and it should be taught as a set piece, not as content. Substance P antagonists engaged their target exactly as designed and did not treat pain. The mechanism was not wrong. The mechanism was right, confirmed, and irrelevant to the outcome. Sit in that for ten minutes; it does more than any restatement of Rule 3.

Structurally: cap field 3 at one sentence in the first draft. Appendix C already asks for the plain sentence first. Enforce it. A student who cannot compress the mechanism to one sentence does not understand it well enough to be leaning on it, and a student who can has just discovered how little the sentence establishes.


5. Relative versus absolute risk

What it looks like. A student writes that a treatment "cut heart attacks by twenty percent" and believes, when asked, that this means twenty out of a hundred people were spared. Or the reverse in a skeptical direction: a student sees an absolute difference of a percentage point or two and concludes the drug does nothing, which is also an over-read.

Why it happens. Because a relative figure is the one that appears in headlines, in press releases, and in the abstract's conclusion sentence, and because the relative figure is not wrong — it is a correct number answering a question the student did not ask. There is no lie to catch, which makes this harder to teach than outright deception. Students are not being fooled; they are being handed the wrong denominator and never told there was one.

The intervention. Make them do the arithmetic by hand, once, on a case where the two numbers feel different. The book's worked example is the pattern: a relative reduction of about twenty percent corresponding to an absolute difference of roughly one and a half percentage points, with a number-needed-to-treat in the mid-sixties. Same trial, same result, three presentations, wildly different emotional weight.

Then install the standing rule: any relative figure in student work must be accompanied by the absolute figure, or by an explicit statement that the absolute figure was not reported. That second clause matters — "the abstract gave only a hazard ratio and no event rates" is a legitimate and informative finding, and students should learn that noticing the omission is itself an analytic move.

Run the conversion drill three or four times across the term. It becomes automatic quickly, and it is among the skills from this course most likely to be used outside it.


6. Surrogate endpoints

What it looks like. "It lowered the marker, so it works." A dossier field 5 that records a moved number as the result and does not flag what kind of number it is. Students treating weight, a lab value, and a mortality rate as interchangeable evidence of benefit.

Why it happens. Because surrogates are usually chosen precisely because they are plausible stand-ins, and often they are good ones. The category is not a trick but a real and frequently reasonable compromise — trials that wait for hard outcomes take years. Students who learn to sneer at surrogates will discard useful evidence.

The correct lesson is narrower: a surrogate is a bet that moving the number moves the outcome, and the bet is sometimes lost. Naming that as a bet, rather than a fallacy, keeps students calibrated.

The intervention. Have students sort a mixed list into three bins — hard outcome (death, a heart attack, a hospitalization), surrogate (a lab value, a scan finding, a weight), and ambiguous — and then defend the ambiguous ones. The defense is where the learning happens: weight is a surrogate for cardiovascular and metabolic outcomes, and it is arguably a patient-relevant outcome in itself, and the fact that a class can argue about it for fifteen minutes is exactly the sophistication you want.

On the page, require field 5's endpoint sub-line to be tagged surrogate or outcome, every time. Appendix C already asks for it. Grading it is what makes students actually do it.

The sentence to leave them with: a surrogate endpoint tells you the drug is doing something. Whether that something is a benefit is a separate question with its own evidence.


7. Ignoring the population

What it looks like. A student applies a trial result to themselves, or to a hypothetical patient, or to "people," without checking who was actually enrolled. A dossier field 5 with a design, an endpoint, and a result, and no population line. On an exam: a correct rating attached to a claim whose population is missing or silently expanded from the one studied.

Why it happens. Because the population is the least interesting sentence in an abstract and the easiest one to skim past. It is also frequently buried: the headline says reduces cardiovascular events, and the enrollment criteria that make that sentence true live in a methods section the student never reaches. And because generalization is the normal, correct behavior of a human reading a fact — students are not being lazy; they are doing what reading usually requires.

The intervention. Use the book's cleanest case, and use it early: the same drug at the same dose, in two populations differing by essentially one enrollment criterion, produced materially different results — roughly fifteen percent weight loss in adults with overweight or obesity without diabetes, roughly ten percent in adults with type 2 diabetes. One criterion. A third of the effect, gone.

Present it as a puzzle first: give students both results without the populations attached and ask them to explain the discrepancy. They will propose dose, duration, and trial quality before anyone proposes population, and the reveal does more work than the fact stated flatly.

Then the standing requirement: a claim without a population is not a claim, and receives no credit. Not a deduction — no credit, the same as a blank. Students calibrate to that in one assignment.


8. Treating regulatory status as evidence

What it looks like. In one direction: it's FDA-approved, so it works — approval treated as a ✅ for whatever the student happens to be asking about, including indications the approval never covered. In the other: it's not approved, so it's junk — non-approval read as a negative finding. Both are the same error, and students who avoid one usually commit the other.

Why it happens. Because regulatory status is the only piece of information in this whole domain that is unambiguous, public, and free. Everything else requires reading a trial. Approval status is a single bit that can be looked up in a minute, and under uncertainty people reach for the cheapest available signal.

The error is compounded because approval genuinely does correlate with evidence — it is just not the same variable, and it answers a narrower question than students think: approval is a judgment about a specific claim in a specific population in a specific jurisdiction, not a certificate attached to a molecule.

The intervention. Attack the second direction first, because it is more tractable. Chapter 38 establishes that "not approved" has five distinct meanings — never submitted, submitted and rejected, approved elsewhere only, used off-label, or not regulated as a drug at all. Have students assign compounds to those five bins from their own dossiers. The exercise collapses the category, and "never submitted" and "submitted and rejected" sitting in different bins is the moment the point lands.

For the first direction, use off-label prescribing: an approved drug being used, legally and routinely, for an indication its approval never covered. The approval is real. It is also silent about the use in front of them.

Structurally: field 6 and field 10 must never adjust each other. Appendix C separates them for this reason and says so. In grading, an entry whose rating moved because of regulatory status is a rule violation, not a judgment call, and marking it that way is fair and fast.


9. Treating a personal anecdote as data

What it looks like. My cousin lost forty pounds on it. I've been taking it for six months and my shoulder is fine. I know three people it did nothing for. Sometimes offered as a challenge to a rating, sometimes offered vulnerably, occasionally offered by a student who is disclosing their own use in the act of making the argument.

Why it happens. Because a personal outcome is the most vivid evidence a human being can have, and because it is not worthless — an anecdote is a real observation of a real event. What it lacks is a comparator, a denominator, and any protection against the reasons people report what they report. Chapter 6's treatment of the testimonial dynamic is the relevant reading, and it is worth assigning again at the moment this comes up rather than waiting for it in sequence.

This is where the classroom gets tense, and the intervention has to handle that first. A student who offers their own experience and gets a lesson in epistemology in return has been publicly corrected about their body in front of their peers. They will not speak again, and neither will anyone watching. The overview's classroom-management section applies at full strength here.

The intervention, in order:

  1. Accept the observation as an observation. "That happened. Nobody is disputing that it happened." This costs nothing and it is true.
  2. Move immediately to the claim, not the person. "So the claim on the table is that it produces that effect in people like your cousin. What would we need to know to evaluate that?"
  3. Ask the comparator question in general form, never as a challenge to their story: what would have happened otherwise, and how many people who tried it are not in the room to tell us? Selection is the concept, and it is much easier to teach with an invented example than with the one a student just offered.
  4. Never make the student's anecdote the worked example. Take the structure, drop the specifics, and work a [constructed teaching example] with the same shape. This is precisely what the book's tier-3 constructed examples are for.
  5. Name the general principle without applying it to them: an anecdote is excellent at establishing that something can happen and nearly useless at establishing how often, in whom, or compared to what.

If a student's own use is on the table, the additional rule is the one from Chapter 39: the discussion stays about the claim and never becomes a discussion of their choice. Have the pivot sentence ready before the term starts, because you will need it and improvising it under pressure goes badly.


10. Rating drift

What it looks like. A student's ratings creep upward across the term. Compounds move ⚠️ → ✅ and 🔬 → ⚠️. Almost nothing ever moves down. By week twelve, a dossier that started appropriately skeptical reads like a product catalog, and each individual upgrade was justified with something the student genuinely found.

Why it happens. This is the subtlest struggle on the list and not an error of reasoning at all. Every upgrade may be locally defensible; the problem is upstream, in the sampling:

News about a compound you are watching is selected for being newsworthy, and negative results are much less newsworthy than positive ones. A student who chose a compound in week two and paid attention to it for fourteen weeks received a filtered stream. A trial that failed to recruit, an analysis that came out flat, a quiet program discontinuation — these do not arrive in anyone's feed. So the student's information intake was biased upward, their reasoning was fine, and the output drifted anyway.

There is a second cause worth naming: students become invested in their dossier compounds. Fourteen weeks of defending a choice in seminar produces exactly the commitment that makes an honest downgrade feel like a personal loss.

The intervention. This is why Chapter 6 asks for field 12 written before investment, and the intervention is to cash that in explicitly:

  • In week two or three, have students write field 12 for every compound, and collect it. Not as a draft — as a fixed record you hold.
  • In the final third of the term, hand it back and require each student to check their current rating against what they said in week three would move it. Not against the latest news. Against their own prior specification.
  • Require an explicit drift statement in the capstone: which ratings moved, in which direction, and triggered by what. Appendix C's maintenance section frames this as measuring your own bias, and it is the single most transferable exercise in the book.
  • Grade a documented downgrade generously. Students will not produce one unless they believe it is safe, and one honest downgrade in a dossier is worth more than five upgrades.

The finding students should leave with is not "I drifted." It is "I drift in a consistent direction, and now I know which one." Some readers are systematically too generous with compounds they want to work; others are systematically too harsh with anything that sounds like marketing. Both are correctable and neither is correctable unmeasured.


Quick reference: struggle → rule → fix

# Struggle Rule violated One-line fix
1 Rating the molecule 1, 6 Mandatory four-line rating format
2 ❌ read as "bad" 2 Ban "X is ❌"; require the full claim
3 Untested vs. failed 5 Tag every ❌ absent or present-and-negative
4 Mechanism upgrades 3 One-sentence cap on field 3; teach Chapter 22 as a set piece
5 Relative vs. absolute Require both figures, or a note that one was not reported
6 Surrogate endpoints Tag every endpoint surrogate or outcome
7 Ignoring population 1 A claim without a population gets no credit
8 Regulatory status as evidence Fields 6 and 10 never adjust each other; use the five meanings of "not approved"
9 Anecdote as data 4 Accept the observation, move to the claim, work a constructed analog
10 Rating drift 5 Collect field 12 early; return it late; grade honest downgrades well

Related: Chapter 5 (the method) · Chapter 6 (hype cycle, testimonials, field 12 early) · Chapter 22 (mechanism without outcome) · Chapter 37 (master table) · Chapter 38 (regulation) · Chapter 39 (the clinical conversation) · Appendix C (dossier workbook) · Appendix D (reading a trial) · Appendix F (red flags) · Appendix H (worked evaluations) · How to Teach This Book