Chapter 5 — Quiz

Twenty-eight questions — the longest quiz in the book, because this is the chapter worth over-testing.


Multiple choice

1. The question that comes before all others when meeting a health claim is: a) is it safe? b) how would we know? c) who funded it? d) is it approved?

2. Which sits highest on the evidence ladder? a) a large randomized controlled trial b) a cohort study c) expert opinion d) a mechanism argument

3. Randomization's central benefit is that it makes groups comparable in: a) age and sex b) disease severity c) all respects including unmeasured ones d) treatment adherence

4. Roughly what proportion of compounds entering human trials never reach approval? a) one in ten b) three in ten c) half d) nine in ten

5. A Phase II trial is designed primarily to detect: a) rare harms b) an efficacy signal and dose-response c) long-term outcomes d) acute toxicity only

6. Phase II effect sizes are, compared with subsequent Phase III results: a) systematically smaller b) systematically larger c) identical on average d) unrelated

7. A p-value of 0.03 means: a) the effect is large b) there is a 3% chance the drug doesn't work c) a difference this big would occur by chance about 3% of the time if there were no real effect d) the result will replicate 97% of the time

8. A hazard ratio of 0.80 indicates: a) an 80% reduction in events b) events occurred at 80% of the control rate — a 20% relative reduction c) a 20% absolute reduction d) no significant difference

9. Number needed to treat (NNT) tells you: a) the sample size required b) how many must be treated for one additional person to benefit c) the number of doses needed d) the trial duration

10. A surrogate endpoint is: a) a placebo b) a secondary analysis c) a measurement believed to predict a clinically important outcome d) an outcome measured in a subgroup

11. SELECT's cardiovascular result was approximately: a) 20% relative reduction, ~1.5 percentage points absolute b) 20% absolute reduction c) 1.5% relative reduction d) 50% relative reduction

12. The treatment-regimen estimand answers the question: a) what happens if you take the drug as prescribed b) what happens if you are assigned the drug, counting everyone including those who stopped c) what happens in the highest-dose group d) what happens in responders

13. Which estimand is systematically larger? a) treatment-regimen b) efficacy c) they are equal d) it varies randomly

14. SURMOUNT-1's tirzepatide 15 mg results at 72 weeks were approximately: a) −10% and −12% b) −21% and −22.5% c) −15% and −30% d) a single figure of −22.5%

15. A trial that measures twenty outcomes and reports the one that reached significance has: a) found a real effect b) conducted a valid subgroup analysis c) most likely found noise d) improved statistical power

16. An unblinded trial with a subjective endpoint is problematic mainly because: a) it is more expensive b) expectations of both participants and assessors influence the measurement c) randomization fails d) it cannot be published

17. A case report is most valuable as evidence of: a) efficacy b) a rare harm c) dose-response d) mechanism

18. A meta-analysis of ten poor-quality trials is: a) automatically top-rung evidence b) an elaborate average of poor studies c) equivalent to one large RCT d) invalid by definition

19. "Mechanism presented as evidence of effect" is a red flag because: a) mechanism is usually wrong b) mechanism is necessary but not sufficient, and roughly nine in ten compounds with believed mechanisms fail c) mechanisms cannot be measured d) it is always marketing

20. A ❌ rating in this book means: a) the compound does not work b) the compound is dangerous c) the confident version of the claim is not supported by human evidence d) the compound was rejected by regulators


Short answer

21. Explain why "the compound has no human trials" and "the compound's human trials failed" are different situations, and why both are called "unproven."

22. Give the five reasons animal results fail to transfer, and identify which is least discussed.

23. Why must a relative risk figure always be accompanied by an absolute one? Give the reason, not just the rule.

24. Explain how the same trial can honestly report two different weight-loss figures.

25. Explain the intensive-lifestyle-support issue in obesity trials, and state what the trial result actually describes.


Applying the rating discipline

26. Write a full four-line rating (claim, rating, reason, falsifier) for BPC-157 for tendon healing in humans. Then state one thing your rating does not say.

27. Rule 3 forbids upgrading a rating with mechanism, and rule 4 forbids downgrading with distaste. Explain what each rule is protecting against, and why the second is harder to follow.

28. A source gives a single overall rating to a molecule. Explain what information is lost, using semaglutide as your example.


Answer key **1.** b. **2.** a. **3.** c. **4.** d. **5.** b. **6.** b. **7.** c. **8.** b. **9.** b. **10.** c. **11.** a. **12.** b. **13.** b. **14.** b. **15.** c. **16.** b. **17.** b. **18.** b. **19.** b. **20.** c. **21.** A compound with **no** human trials has no human evidence in either direction — the question is open, and the absence describes the literature rather than the molecule. A compound whose trials **failed** has evidence *against* it: somebody looked, properly, and did not find the effect. The second is a much stronger epistemic position, and it is a *worse* position for the compound. Both get called "unproven" because the word describes the absence of positive evidence and does not distinguish "not yet examined" from "examined and found wanting." **22.** (1) different biology — species differences in metabolism, immunity, and receptors; (2) different disease — a model captures a piece of a condition, not the condition; (3) different dose — animal doses often do not scale safely to humans; (4) different endpoint — animals cannot report pain, function, or quality of life; (5) **different publication pressure** — the least discussed. Animal studies are much less likely to be pre-registered, so the published animal literature is a filtered sample of the experiments actually run, in a way the human trial literature increasingly is not. **23.** Because relative reduction is a *ratio* and absolute reduction is a *difference*, and when baseline risk is low, a large ratio is a small difference. A 20% relative reduction is 10 percentage points at 50% baseline risk and 0.2 percentage points at 1% baseline risk — a fiftyfold difference in what it means for an individual, with an identical relative figure. Quoting only the relative number is therefore not a communication; it is a decision made on the reader's behalf about what impression to leave. **24.** Because two different pre-specified analyses answer two different questions. The **treatment-regimen estimand** counts everyone as randomized, including participants who discontinued — answering "what happens if you are prescribed this?" The **efficacy estimand** restricts to those who remained on treatment — answering "what happens if you take it?" Both are correct and both are reported. The efficacy figure is systematically larger, which is why quoting a figure without naming its estimand is systematically optimistic rather than randomly imprecise. **25.** Nearly all obesity pharmacotherapy trials give **both** arms structured lifestyle support — dietary counseling, activity guidance, and frequent study contact. This makes the drug effect look **smaller** than it would against no intervention (the placebo arm is genuinely being treated) and potentially **larger** than it will be in practice (a patient receiving a prescription and no support is not receiving what the trial delivered). The result therefore describes **drug-plus-support versus support alone**, which is not the comparison most readers have in mind. **26.** **Claim:** BPC-157 accelerates healing of tendon and soft-tissue injuries in humans. **Rating:** ❌ Hype outpaces evidence. **Why:** As of this writing there is no completed, peer-reviewed, randomized controlled human trial of BPC-157 for any indication in the published literature; the supporting evidence is rodent work, which sits on the animal rung. **What would change it:** a completed, adequately powered, randomized controlled human trial with a pre-specified functional endpoint — or even a well-conducted Phase I safety study, since human safety data does not currently exist either. **What it does not say:** that BPC-157 does nothing. That is not established either. ❌ describes the state of the evidence, not the state of the molecule. **27.** **Rule 3 (no upgrading with mechanism)** protects against the field's most common error: treating a plausible pathway as a substitute for a trial. Between receptor activation and patient benefit sit at least six steps, and roughly nine in ten compounds fail at one of them. **Rule 4 (no downgrading with distaste)** protects against the opposite bias, and it is harder to follow because it operates on you invisibly. Nobody experiences their own skepticism as motivated. A reader who has decided the peptide space is mostly nonsense will apply a stricter standard to its compounds without noticing, and will feel appropriately rigorous while doing it. The check is procedural: apply the same six questions to every claim and see whether the answers, not the conclusions, differ. **28.** It destroys the distinction between claims with completely different evidence bases. Semaglutide for weight loss in adults with obesity is ✅ — multiple large RCTs, regulatory approval, characterized safety. Semaglutide for Alzheimer's disease is 🔬 — a hypothesis being properly tested, with trials ongoing and no result. A single verdict on "semaglutide" would have to be wrong about one of them, and whichever way it erred it would mislead: an overall ✅ implies established benefit for indications that have none, and an overall ⚠️ understates one of the best-evidenced results in modern metabolic medicine.