Chapter 5 — Quiz
Twenty-eight questions — the longest quiz in the book, because this is the chapter worth over-testing.
Multiple choice
1. The question that comes before all others when meeting a health claim is:
a) is it safe? b) how would we know? c) who funded it? d) is it approved?
2. Which sits highest on the evidence ladder?
a) a large randomized controlled trial b) a cohort study c) expert opinion d) a mechanism argument
3. Randomization's central benefit is that it makes groups comparable in:
a) age and sex b) disease severity c) all respects including unmeasured ones d) treatment adherence
4. Roughly what proportion of compounds entering human trials never reach approval?
a) one in ten b) three in ten c) half d) nine in ten
5. A Phase II trial is designed primarily to detect:
a) rare harms b) an efficacy signal and dose-response c) long-term outcomes d) acute toxicity only
6. Phase II effect sizes are, compared with subsequent Phase III results:
a) systematically smaller b) systematically larger c) identical on average d) unrelated
7. A p-value of 0.03 means:
a) the effect is large b) there is a 3% chance the drug doesn't work c) a difference this big would
occur by chance about 3% of the time if there were no real effect d) the result will replicate 97% of
the time
8. A hazard ratio of 0.80 indicates:
a) an 80% reduction in events b) events occurred at 80% of the control rate — a 20% relative
reduction c) a 20% absolute reduction d) no significant difference
9. Number needed to treat (NNT) tells you:
a) the sample size required b) how many must be treated for one additional person to benefit
c) the number of doses needed d) the trial duration
10. A surrogate endpoint is:
a) a placebo b) a secondary analysis c) a measurement believed to predict a clinically important
outcome d) an outcome measured in a subgroup
11. SELECT's cardiovascular result was approximately:
a) 20% relative reduction, ~1.5 percentage points absolute b) 20% absolute reduction
c) 1.5% relative reduction d) 50% relative reduction
12. The treatment-regimen estimand answers the question:
a) what happens if you take the drug as prescribed b) what happens if you are assigned the drug,
counting everyone including those who stopped c) what happens in the highest-dose group
d) what happens in responders
13. Which estimand is systematically larger?
a) treatment-regimen b) efficacy c) they are equal d) it varies randomly
14. SURMOUNT-1's tirzepatide 15 mg results at 72 weeks were approximately:
a) −10% and −12% b) −21% and −22.5% c) −15% and −30% d) a single figure of −22.5%
15. A trial that measures twenty outcomes and reports the one that reached significance has:
a) found a real effect b) conducted a valid subgroup analysis c) most likely found noise
d) improved statistical power
16. An unblinded trial with a subjective endpoint is problematic mainly because:
a) it is more expensive b) expectations of both participants and assessors influence the measurement
c) randomization fails d) it cannot be published
17. A case report is most valuable as evidence of:
a) efficacy b) a rare harm c) dose-response d) mechanism
18. A meta-analysis of ten poor-quality trials is:
a) automatically top-rung evidence b) an elaborate average of poor studies c) equivalent to one large
RCT d) invalid by definition
19. "Mechanism presented as evidence of effect" is a red flag because:
a) mechanism is usually wrong b) mechanism is necessary but not sufficient, and roughly nine in ten
compounds with believed mechanisms fail c) mechanisms cannot be measured d) it is always marketing
20. A ❌ rating in this book means:
a) the compound does not work b) the compound is dangerous c) the confident version of the claim is
not supported by human evidence d) the compound was rejected by regulators
Short answer
21. Explain why "the compound has no human trials" and "the compound's human trials failed" are
different situations, and why both are called "unproven."
22. Give the five reasons animal results fail to transfer, and identify which is least discussed.
23. Why must a relative risk figure always be accompanied by an absolute one? Give the reason, not
just the rule.
24. Explain how the same trial can honestly report two different weight-loss figures.
25. Explain the intensive-lifestyle-support issue in obesity trials, and state what the trial result
actually describes.
Applying the rating discipline
26. Write a full four-line rating (claim, rating, reason, falsifier) for BPC-157 for tendon healing
in humans. Then state one thing your rating does not say.
27. Rule 3 forbids upgrading a rating with mechanism, and rule 4 forbids downgrading with
distaste. Explain what each rule is protecting against, and why the second is harder to follow.
28. A source gives a single overall rating to a molecule. Explain what information is lost, using
semaglutide as your example.
Answer key
**1.** b. **2.** a. **3.** c. **4.** d. **5.** b. **6.** b. **7.** c. **8.** b. **9.** b. **10.** c.
**11.** a. **12.** b. **13.** b. **14.** b. **15.** c. **16.** b. **17.** b. **18.** b. **19.** b.
**20.** c.
**21.** A compound with **no** human trials has no human evidence in either direction — the question is
open, and the absence describes the literature rather than the molecule. A compound whose trials
**failed** has evidence *against* it: somebody looked, properly, and did not find the effect. The
second is a much stronger epistemic position, and it is a *worse* position for the compound. Both get
called "unproven" because the word describes the absence of positive evidence and does not
distinguish "not yet examined" from "examined and found wanting."
**22.** (1) different biology — species differences in metabolism, immunity, and receptors; (2)
different disease — a model captures a piece of a condition, not the condition; (3) different dose —
animal doses often do not scale safely to humans; (4) different endpoint — animals cannot report pain,
function, or quality of life; (5) **different publication pressure** — the least discussed. Animal
studies are much less likely to be pre-registered, so the published animal literature is a filtered
sample of the experiments actually run, in a way the human trial literature increasingly is not.
**23.** Because relative reduction is a *ratio* and absolute reduction is a *difference*, and when
baseline risk is low, a large ratio is a small difference. A 20% relative reduction is 10 percentage
points at 50% baseline risk and 0.2 percentage points at 1% baseline risk — a fiftyfold difference in
what it means for an individual, with an identical relative figure. Quoting only the relative number is
therefore not a communication; it is a decision made on the reader's behalf about what impression to
leave.
**24.** Because two different pre-specified analyses answer two different questions. The
**treatment-regimen estimand** counts everyone as randomized, including participants who discontinued
— answering "what happens if you are prescribed this?" The **efficacy estimand** restricts to those who
remained on treatment — answering "what happens if you take it?" Both are correct and both are
reported. The efficacy figure is systematically larger, which is why quoting a figure without naming
its estimand is systematically optimistic rather than randomly imprecise.
**25.** Nearly all obesity pharmacotherapy trials give **both** arms structured lifestyle support —
dietary counseling, activity guidance, and frequent study contact. This makes the drug effect look
**smaller** than it would against no intervention (the placebo arm is genuinely being treated) and
potentially **larger** than it will be in practice (a patient receiving a prescription and no support
is not receiving what the trial delivered). The result therefore describes **drug-plus-support versus
support alone**, which is not the comparison most readers have in mind.
**26.** **Claim:** BPC-157 accelerates healing of tendon and soft-tissue injuries in humans.
**Rating:** ❌ Hype outpaces evidence. **Why:** As of this writing there is no completed,
peer-reviewed, randomized controlled human trial of BPC-157 for any indication in the published
literature; the supporting evidence is rodent work, which sits on the animal rung. **What would change
it:** a completed, adequately powered, randomized controlled human trial with a pre-specified
functional endpoint — or even a well-conducted Phase I safety study, since human safety data does not
currently exist either.
**What it does not say:** that BPC-157 does nothing. That is not established either. ❌ describes the
state of the evidence, not the state of the molecule.
**27.** **Rule 3 (no upgrading with mechanism)** protects against the field's most common error:
treating a plausible pathway as a substitute for a trial. Between receptor activation and patient
benefit sit at least six steps, and roughly nine in ten compounds fail at one of them.
**Rule 4 (no downgrading with distaste)** protects against the opposite bias, and it is harder to
follow because it operates on you invisibly. Nobody experiences their own skepticism as motivated. A
reader who has decided the peptide space is mostly nonsense will apply a stricter standard to its
compounds without noticing, and will feel appropriately rigorous while doing it. The check is
procedural: apply the same six questions to every claim and see whether the answers, not the
conclusions, differ.
**28.** It destroys the distinction between claims with completely different evidence bases.
Semaglutide for weight loss in adults with obesity is ✅ — multiple large RCTs, regulatory approval,
characterized safety. Semaglutide for Alzheimer's disease is 🔬 — a hypothesis being properly tested,
with trials ongoing and no result. A single verdict on "semaglutide" would have to be wrong about one
of them, and whichever way it erred it would mislead: an overall ✅ implies established benefit for
indications that have none, and an overall ⚠️ understates one of the best-evidenced results in modern
metabolic medicine.