36 min read

> "The first principle is that you must not fool yourself — and you are the easiest person to fool."

Prerequisites

  • 1
  • 2
  • 4

Learning Objectives

  • Place any piece of evidence on the evidence ladder and state what that rung licenses
  • Explain what an animal study can and cannot establish about humans
  • Describe what each clinical trial phase is designed to detect
  • Identify the six structural features that determine a trial's quality
  • Explain how a surrogate endpoint can make an ineffective drug look effective
  • Interpret p-values, confidence intervals, hazard ratios, and number needed to treat in plain language
  • Convert between relative and absolute risk reduction and explain why both must be stated
  • Explain why one trial can honestly report two different effect sizes
  • Apply the four-tier rating system to a claim, with a population, an endpoint, and a falsifier

Chapter 5: How to Evaluate Peptide Evidence: Clinical Trials, Animal Studies, and Why Reddit Is Not a Peer-Reviewed Journal

"The first principle is that you must not fool yourself — and you are the easiest person to fool." — Richard Feynman, Caltech commencement address (1974)

Overview

This is the most important chapter in this book, and it contains almost no peptides.

Every other chapter tells you what is known about some molecule. This one teaches you how anyone knows anything at all — and that skill outlasts every specific verdict in these pages. Ratings age. Trials read out, approvals happen, a ❌ becomes a ⚠️ the moment somebody finally runs the study. The method does not age.

Here is the situation the method has to handle.

You encounter two claims on the same afternoon. The first: semaglutide produces substantial and sustained weight loss. The second: BPC-157 heals tendon injuries. Both are stated with equal confidence by people who sound equally informed. Both have supporting evidence you can go and look at. Both have detractors.

One of those claims is supported by multiple randomized controlled trials involving tens of thousands of people, followed for years, with pre-specified endpoints, published in major journals, reviewed by regulators in multiple countries. The other is supported by a substantial body of rodent research and, as of this writing, no completed peer-reviewed randomized human trial at all.

The difference between those evidence bases is enormous, and it is completely invisible from the confidence of the claim. Confidence is free. Anybody can have it. What costs something — money, years, a real chance of being proven wrong — is evidence, and learning to see the difference is the entire content of this chapter.

By the end you will be able to read a study abstract and say what it does and does not establish, spot the four or five ways a technically valid trial routinely misleads, convert a scary-sounding relative risk into an honest absolute one, explain why one trial can report two different weight-loss figures without anyone lying, and assign a defensible rating to any claim — with the population, the endpoint, the reason, and the thing that would change your mind.

In this chapter, you will learn to:

  • Place evidence on the ladder and state exactly what each rung licenses you to believe
  • Explain why "works in rats" is necessary and radically insufficient
  • Say what each trial phase is powered to detect — and what it cannot detect
  • Identify the six structural features that make a trial trustworthy or not
  • Recognize a surrogate endpoint and explain how it misleads
  • Read p-values, confidence intervals, hazard ratios, and NNT in plain language
  • State a result in both relative and absolute terms, always
  • Explain estimands, and why the same trial reports two numbers
  • Apply the four-tier rating system properly

Learning Paths

All five paths read this chapter in full. It is the only chapter for which that is true, and if you read one chapter of this book, read this one.

💊 GLP-1 — §5.8 and §5.9 use the exact trials you came here for. You will finish able to read the SELECT and SURMOUNT results better than most journalists who covered them. 🏋️ Performance — §5.3 and §5.10 are where Part III gets decided. Everything this book will later say about BPC-157, TB-500, and the secretagogues is an application of these two sections. 🔬 Science — full read, twice. §5.5 through §5.9 are the technical core. 💄 Cosmetic — §5.5, §5.6, and §5.10. Cosmetic studies are a specific genre with a specific set of weaknesses, and §5.10 names all of them. 🏥 Clinical — §5.8 and §5.9 are what your patients are getting wrong when they quote a number at you, and being able to explain both distinctions in thirty seconds is a genuinely useful clinical skill.


5.1 The question every claim must answer

When you meet a health claim, there is one question that comes before all others.

How would we know?

Not "is it true." Not "does it make sense." How would anyone find out — what would you have to observe, in whom, compared with what, to distinguish a world where the claim is true from a world where it isn't?

That question does enormous work, and it does it fast.

Take: "This peptide accelerates recovery." How would we know? You would need people with a defined injury, randomly assigned to the peptide or to something indistinguishable from it, with recovery measured by something that does not depend on either party's expectations, followed long enough for recovery to occur, in enough people that a real difference would be visible above noise.

Now ask whether that has been done. For a great many claims, the answer is no — and you have reached that conclusion in under a minute, without any expertise in the compound.

The question also exposes claims that cannot be tested at all. "Optimizes cellular function." How would we know? What would you measure? In whom? Compared to what? A claim with no possible observation that would refute it is not a weak claim. It is not a claim. Chapter 2's "modulates cellular signaling pathways" and Chapter 3's "supports healthy hormone balance" are both of this type, and they are the most common form of health marketing in existence precisely because they cannot lose.

Three follow-ups make the question sharper:

"In whom?" A result in mice, in healthy young men, in people with severe disease, and in the general population are four different findings. Trial populations are narrow by design, and the distance between the studied population and the person in front of you is one of the largest sources of real-world disappointment.

"Compared with what?" Almost everything improves over time. Injuries heal, symptoms fluctuate, people who start a supplement often start sleeping better too. Without a comparison group, you cannot distinguish the treatment from the passage of time.

"Measured how?" "Felt better" is not the same as a validated instrument, which is not the same as a hard clinical event. §5.6 is entirely about this.

🔍 Check Your Understanding

  1. Apply "how would we know?" to: "This peptide reduces inflammation." What would need to be observed?
  2. Why is a claim that cannot possibly be refuted worse than a claim that is simply wrong?
  3. Someone reports that their injury healed after starting a compound. What is the "compared with what?" problem here?

5.2 The evidence ladder

Evidence is not all the same kind of thing. It sorts, roughly, by how well it controls the ways you can fool yourself.

THE EVIDENCE LADDER — what each rung can and cannot tell you
  ┌──────────────────────────────────────────────────────────────────────────────┐
  │  SYSTEMATIC REVIEW / META-ANALYSIS   many trials, pooled and appraised       │  strongest
  │  LARGE RANDOMIZED CONTROLLED TRIAL   randomized, blinded, powered, outcome   │     ▲
  │  SMALL / SHORT RCT                   randomized but underpowered or brief    │     │
  │  COHORT STUDY                        followed over time, not randomized      │     │
  │  CASE-CONTROL STUDY                  looks backward from an outcome          │     │
  │  CASE SERIES / CASE REPORT           what happened to a few people           │     │
  │  ANIMAL STUDY                        a different organism, a clean model     │     │
  │  IN VITRO / CELL STUDY               cells in a dish, no organism at all     │     │
  │  MECHANISM / PLAUSIBILITY            "it should work, because the pathway"   │     ▼
  │  EXPERT OPINION                      a credentialed person's belief          │
  │  ANECDOTE / TESTIMONIAL              one person's uncontrolled experience    │  weakest
  └──────────────────────────────────────────────────────────────────────────────┘
  The rungs are not a ranking of *worth* — every drug starts at the bottom and climbs.
  They are a ranking of *what a result licenses you to believe about the next person.*

Working up from the bottom, because the bottom is where most peptide claims live:

Anecdote. One person's experience. Genuinely worthless as evidence about anyone else, and genuinely powerful psychologically, which is a dangerous combination. The problem is not that people lie. It is that a single uncontrolled experience cannot distinguish the treatment from natural recovery, from expectation, from the four other things that changed at the same time, or from regression to the mean — the statistical near-certainty that anyone who starts a treatment when they feel worst will feel better afterward whether or not the treatment does anything.

Expert opinion. Better, because an expert has seen many cases. Still weak, because it is subject to the same biases plus a few of its own — experts remember their successes, see selected patients, and have careers invested in positions. "A doctor recommended it" is a reason to take a claim seriously and not a reason to believe it.

Mechanism and plausibility. Chapter 2's subject, and Chapter 2's warning. Necessary, not sufficient, and strong evidence against, weak evidence for.

In vitro studies. Cells in a dish. Enormously useful for working out mechanism and completely disconnected from whether a whole organism benefits. Cell studies routinely use concentrations that could never be achieved in a person, in cell lines that behave unlike the tissue they came from, with no immune system, no blood supply, no liver, and no clearance.

Animal studies. §5.3, in full.

Case reports and series. What happened to one person or a handful, described in detail. Their value is different from their strength: they are excellent at generating hypotheses and at detecting rare harms — a drug causing a distinctive rare event will show up in case reports before any trial has the power to detect it. As evidence of efficacy, they remain nearly worthless, because they have no comparison group.

Case-control studies. Start with people who have an outcome, look backward for exposures. Efficient for rare outcomes. Vulnerable to recall bias and to the difficulty of choosing appropriate controls.

Cohort studies. Follow a group over time and see what happens. Real human data, real duration, and the central problem: the people who chose the exposure differ from those who didn't, in ways that are frequently the actual cause of the difference in outcome. Statistical adjustment helps and cannot fully fix it, because you can only adjust for things you thought to measure.

Randomized controlled trials. The jump in strength here is not incremental; it is categorical. Randomization means the two groups differ only by chance at the start — including in all the ways you never thought to measure. That is the one thing no amount of statistical sophistication can retrofit onto an observational study.

Systematic reviews and meta-analyses. Find all the studies on a question, appraise their quality, and pool the results. At their best, the strongest evidence available. At their worst, an elaborate average of poor studies. A meta-analysis of ten bad trials is not better than one good trial, and the review's quality depends entirely on how rigorously it searched and appraised.


5.3 What an animal study can and cannot tell you

This section decides Part III, so it is worth doing carefully.

Animal studies are real science and they are essential. No drug reaches humans without them, and for good reason: you can control everything, induce disease deliberately, sample tissue at endpoint, use genetically identical subjects, and randomize without anyone's consent being an obstacle. A clean animal experiment can establish mechanism, dose-response, and preliminary safety in ways no human study can.

And most drugs that work in animals fail in humans. Roughly nine out of ten compounds that enter human trials never reach approval, and the attrition from promising-animal-result to approval is steeper still.

Both statements are true. The tension between them is the whole issue.

WHY ANIMAL RESULTS DON'T TRANSFER — five independent reasons

  ① DIFFERENT BIOLOGY
     Rodent metabolism runs faster. Immune systems differ. Some receptors differ in
     sequence and in distribution. A drug can be a potent agonist in one species and
     nearly inactive in another.

  ② DIFFERENT DISEASE
     Animal "models" are not the disease. A surgically transected rat tendon is not a
     human tendinopathy that developed over two years of overuse. Induced obesity in a
     mouse fed a specific diet is not human obesity. The model captures a piece.

  ③ DIFFERENT DOSE
     Doses that work in animals are frequently, when scaled by body surface area,
     far above what is achievable or safe in humans.

  ④ DIFFERENT ENDPOINT
     Animals do not report pain, function, fatigue, or quality of life. Animal studies
     measure tissue, histology, and behavior in a cage. Humans want to know whether
     they can climb stairs.

  ⑤ DIFFERENT PUBLICATION PRESSURE
     Positive animal results are far more publishable than negative ones, and animal
     studies are much less likely to be pre-registered than human trials. The published
     animal literature on a compound is a filtered sample of the experiments that were run.

Reason ⑤ deserves emphasis because it is the least discussed. When you read that "over a hundred studies show" a compound does something in animals, you are reading the published subset. Human trials are increasingly required to be registered before they begin, so a trial that runs and fails leaves a trace. Animal experiments generally leave no such trace. The negative ones may simply not exist in the record.

🔬 Read the Study — a clean animal result, read properly

text FIGURE 5.4 — "The rat tendon that launched a thousand orders" [composite — real pattern, constructed specifics] THE STUDY Randomized controlled animal study. 40 rats, surgically transected Achilles tendon, randomized to peptide or saline, 14 days, histology and load-to-failure at endpoint. Funded by the investigators' institution. THE QUESTION Does the peptide accelerate tendon healing in a standardized rat injury model? WHAT IT SHOWS Treated tendons showed greater tensile strength and more organized collagen at 14 days than saline controls. The model is clean, the randomization is real, and the effect is unlikely to be chance. WHAT IT DOESN'T It does not show a dose for humans, a safety profile in humans, an effect on pain or function (rats do not report either), durability past 14 days, or any result in a tendon that was injured by overuse rather than a scalpel — which is how nearly every human tendon injury actually happens. THE VERDICT ❌ for the human claim, and a legitimate ⚠️-in-waiting for the research program. The animal result is real; the human claim built on it is unsupported. THE LESSON A clean effect in a clean model is a reason to run a human trial. It is not a preview of the human trial's result. Roughly nine in ten compounds that look this good in animals fail somewhere in human development.

The correct posture toward a strong animal result is neither dismissal nor extrapolation. It is: this is a good reason to run the human trial, and until someone does, the human question is open.

That posture is uncomfortable because it refuses both available satisfactions. You do not get to believe, and you do not get to sneer. Chapter 17 asks you to hold it for an entire chapter.


5.4 Trial phases: what each one can detect

Human drug development is staged, and each stage is designed to answer a different question. Knowing which stage a claim rests on tells you a great deal.

Phase Question Typical size What it CAN detect What it CANNOT
Preclinical Is it worth testing in people? animals, cells mechanism, gross toxicity, dose range anything about humans
Phase I Is it safe enough to continue, and what does the body do to it? tens acute toxicity, pharmacokinetics, tolerable dose range efficacy; anything but common, immediate harms
Phase II Does it do anything, and at what dose? dozens to a few hundred a signal of efficacy, dose-response, common side effects reliable effect size; uncommon harms; long-term anything
Phase III Does it work, in the population it is for, better than the alternative? hundreds to tens of thousands efficacy on the primary endpoint, common and moderately uncommon harms rare harms; effects beyond trial duration
Phase IV What happens in the real world, at scale, over time? millions rare harms, long-term effects, real-world effectiveness clean causal attribution — it is observational

Four things about this that are constantly misunderstood.

Phase II results are systematically optimistic. A Phase II trial is small, often shorter, sometimes in a favorable population, and frequently reports the best of several doses. The effect seen in Phase II is on average larger than what Phase III finds. This is not fraud; it is regression to the mean plus selection. When you read that a compound produced a spectacular result, check the phase. Chapter 9's retatrutide discussion is exactly this situation, and the book flags it every time.

Phase III is not the end of learning about safety. A Phase III trial of 5,000 people cannot detect a harm that occurs in one person per 10,000. Post-marketing surveillance exists because approval happens with genuinely incomplete safety information. That is not a scandal — it is arithmetic — but it is a reason that a drug approved last year is less well characterized than one approved twenty years ago.

Approval is a regulatory decision, not a scientific verdict. Chapter 38 covers what an approval actually certifies, which is narrower than most people assume: that in a specified population, for a specified use, the evidence submitted showed benefit exceeding risk to the regulator's satisfaction. It is not a statement that the drug is safe, that it is better than alternatives, or that it works for anything else.

And a compound with no completed trials is not at Phase 0. It is off the ladder. There is a difference between a drug that failed Phase II and a compound that has never been tested — the first has evidence against it, the second has no evidence either way — and both are commonly described as "unproven."


5.5 Anatomy of a trial: the six features that decide quality

When you read a trial, six structural features determine what it can tell you. Read them in this order.

1. Population — who was studied? Age, sex, disease severity, comorbidities, prior treatment, and country. Trial populations are narrow by design, and the distance between them and you is the most common reason a real drug disappoints in practice. Most of the weight-loss numbers in this book come from trials in people with obesity or overweight-plus-comorbidity, usually with intensive lifestyle support alongside. Quoting them for anyone else is a population swap.

2. Comparator — compared with what? Placebo tells you whether the drug beats nothing. An active comparator tells you whether it beats what you would otherwise use — a much more useful and much less common question. "Superior to placebo" and "superior to standard care" are very different claims.

3. Randomization — how were people assigned? Proper randomization is what makes the groups comparable in every respect, including the ones nobody measured. Look for concealed allocation: whoever enrolls a participant must not be able to predict which arm they will get.

4. Blinding — who knew? Single-blind (participant), double-blind (participant and investigator), or open-label (everyone knows). Blinding matters most where the endpoint involves judgment — pain, symptom scores, investigator assessment. It matters less for hard endpoints like death. For a drug with obvious side effects, blinding may be imperfect in practice even when it is nominally maintained, because participants can guess. This is a genuine and underappreciated limitation in GLP-1 trials, where gastrointestinal effects are common and noticeable.

5. Endpoint — what was measured, and was it decided in advance? The primary endpoint must be pre-specified. A trial that measures twenty things and reports the one that reached significance has not found an effect; it has found noise, and the practice has a name (outcome switching) and a literature. Registration on a public registry before the trial begins is what makes this checkable.

6. Duration and size — for how long, in how many? A trial must be powered — large enough that a real effect of the expected size would be detected. An underpowered trial that finds nothing has told you almost nothing. And duration must match the question: a twelve-week trial cannot answer a question about two years.

🔍 Check Your Understanding — population, comparator, endpoint

  1. A trial enrolled adults with obesity and a weight-related comorbidity, and someone cites its result as a reason for a lean person to take the drug. Name the error in this section's vocabulary, and say why it is not a small one.
  2. A trial reports that a drug is "superior to placebo." What question does that answer — and which more useful question does it leave entirely untouched?
  3. A trial measured twenty outcomes and reports the one that reached statistical significance. Which of the six features has been violated, and what public record lets you check it?

💊 In the Clinic — the intensive-lifestyle-support footnote

Nearly every obesity pharmacotherapy trial provides both arms with structured lifestyle intervention: dietary counseling, physical activity guidance, and regular contact with the study team, often monthly or more.

Two consequences follow, and they run in opposite directions.

It makes the drug effect look smaller than it is in isolation, because the placebo arm is also receiving a real intervention. The placebo group in a major semaglutide trial lost meaningfully more weight than an untreated population would.

And it makes the drug effect potentially larger than it will be in practice, because a patient receiving a prescription and no structured support is not receiving what the trial delivered.

Which effect dominates is not knowable from the trial. What is knowable is that the trial result describes drug-plus-support versus support alone — and that is not the comparison most people have in mind when they read the number.

This is not a criticism of the trials. Withholding lifestyle support would be unethical and would also be a worse experiment. It is a reason to read "15% weight loss" as a statement about a specific protocol rather than about a molecule.


5.6 Surrogate endpoints: the most common way a good trial misleads

A hard endpoint is something that matters directly: death, heart attack, stroke, hospitalization, fracture, a person's ability to walk.

A surrogate endpoint is a measurement believed to predict a hard endpoint: blood pressure, cholesterol, blood glucose, bone density, tumor shrinkage, a biomarker level.

Surrogates are used because they are faster and cheaper. A trial measuring cardiovascular death needs thousands of people and years. A trial measuring blood pressure needs hundreds and weeks.

And surrogates fail regularly. Medicine has a long, humbling list of drugs that improved a surrogate beautifully and either did nothing for the hard endpoint or made it worse. The canonical example is not from peptide science at all: a class of antiarrhythmic drugs successfully suppressed the abnormal heartbeats believed to cause sudden death after heart attack, and a randomized trial found higher mortality in the treated group. The surrogate improved. The patients did worse.

WHY SURROGATES FAIL

  THE ASSUMPTION                          WHAT CAN GO WRONG
  ┌────────────┐                          The drug may affect the surrogate by a route
  │   DRUG     │                          that does not connect to the outcome ────┐
  └─────┬──────┘                                                                    │
        ▼                                 The surrogate may be a MARKER of the      │
  ┌────────────┐                          disease rather than a CAUSE of it ────┐   │
  │ SURROGATE  │ ◀── improves                                                    │   │
  └─────┬──────┘                          The drug may have a separate harmful    │   │
        ▼  (assumed)                      effect that outweighs the benefit ──┐   │   │
  ┌────────────┐                                                              │   │   │
  │  OUTCOME   │ ◀── assumed to follow ... but the arrow is a HYPOTHESIS ◀────┴───┴───┘
  └────────────┘

  A surrogate is only as good as the evidence that changing IT changes the OUTCOME —
  and that evidence is a separate research program, not an assumption.

For peptides specifically, watch for these surrogates:

  • "Raises growth hormone" — a surrogate for body composition and function, and Chapter 15 shows the arrow is not established
  • "Increases IGF-1" — same
  • "Increases lean mass" — a surrogate for strength and function, and Chapter 16 shows these dissociate
  • "Improves skin elasticity on an instrument" — a surrogate for looking better, Chapter 30
  • "Reduces an inflammatory marker" — a surrogate for almost everything, and one of the least reliable
  • "Shrinks a tumor" — a surrogate for survival, and oncology has learned the hard way that these can diverge

The question to ask is always the same: has anyone shown that changing this surrogate changes the thing I actually care about? Sometimes yes — LDL cholesterol and cardiovascular events have a large supporting literature. Often no.


5.7 Reading the numbers, in plain language

Four statistical concepts, each explained by what it means rather than how it is computed.

The p-value. If the treatment did nothing at all, how surprising would a result this large be? A p-value of 0.03 means: a difference this big would occur by chance about 3% of the time if there were no real effect. Conventionally, below 0.05 is called "statistically significant."

What it does not mean: that the effect is large, important, or reproducible. A tiny, clinically meaningless difference can be highly statistically significant if the trial is large enough. Significance is about confidence that something is there; it says nothing about whether that something matters.

The confidence interval. A range of values consistent with the data. "A 15% reduction (95% CI: 8% to 22%)" means the data are consistent with anywhere between 8% and 22%.

This is more informative than a p-value and is under-reported. Two things to check: does the interval include "no effect"? (If so, the result is not statistically significant.) And is the whole interval clinically meaningful? A result of "12% (95% CI: 1% to 23%)" is technically significant and consistent with a benefit too small to care about.

The hazard ratio. The rate of events in the treated group relative to the control group over the trial. 0.80 means events occurred at 80% of the control rate — a 20% relative reduction. 1.0 means no difference. Above 1.0 means more events with treatment.

Hazard ratios are relative, always, and that is where §5.8 comes in.

Number needed to treat (NNT). How many people must receive the treatment for one additional person to benefit? It is the most honest single number in medicine because it forces the absolute scale into view.

An NNT of 10 is excellent. An NNT of 100 may still be worthwhile for a serious outcome and a safe, cheap drug. An NNT of 500 for a minor benefit with meaningful side effects is a different proposition entirely — and the relative risk reduction may look identical in all three cases.


5.8 Relative versus absolute: the SELECT worked example

This is the single most consequential statistical skill in this book.

The same result can be stated in two ways, both accurate, that leave readers with wildly different impressions.

Work it through with a real trial.

🔬 Read the Study — SELECT

text FIGURE 5.8 — "Twenty percent, or one and a half points" [real published trial] THE STUDY Randomized, double-blind, placebo-controlled cardiovascular outcomes trial. Semaglutide 2.4 mg weekly versus placebo in adults with established cardiovascular disease and overweight or obesity but WITHOUT diabetes. Roughly 17,000 participants, followed for about three years. Sponsored by the manufacturer. THE QUESTION Does semaglutide reduce major adverse cardiovascular events (MACE: cardiovascular death, non-fatal myocardial infarction, non-fatal stroke) in this population? WHAT IT SHOWS Yes. A 20% RELATIVE reduction in MACE — hazard ratio about 0.80. In ABSOLUTE terms: the event rate fell from roughly 8% to roughly 6.5% over about three years — about 1.5 percentage points. A hard endpoint, a large population, a long follow-up, pre-specified. WHAT IT DOESN'T It does not establish benefit in people WITHOUT established cardiovascular disease, in people with diabetes (a different trial), or beyond three years. It does not isolate the mechanism — whether the benefit comes from weight loss, glucose effects, blood pressure, inflammation, or a direct vascular action is NOT established by this trial. THE VERDICT ✅ for this claim, in this population. One of the strongest cardiovascular outcome results for any metabolic drug. THE LESSON "20% reduction" and "1.5 percentage points" describe the identical result. A reader given only the first overestimates the personal benefit by roughly a factor of ten. A reader given only the second underestimates the population importance. STATE BOTH, ALWAYS.

Why the two numbers diverge so much. Relative reduction is a ratio; absolute reduction is a difference. When the baseline risk is low, a large relative reduction is a small absolute one.

THE SAME 20% RELATIVE REDUCTION AT THREE BASELINE RISKS    [constructed teaching example]

  baseline risk   →   treated risk    absolute reduction   NNT over the period
  ─────────────────────────────────────────────────────────────────────────────
      50%                40%              10 points              10
       8%                6.4%            1.6 points              ~63
       1%                0.8%            0.2 points             ~500

  IDENTICAL relative reduction. Wildly different meaning for an individual.
  This is why a relative risk reduction quoted alone is not a communication —
  it is a decision made on the reader's behalf about what impression to leave.

The rule, and it is not negotiable in this book: whenever a relative figure is given, the absolute figure must be given too, in the same sentence. A source that gives only the relative number is either careless or is choosing the more impressive framing, and you cannot tell which.

Note that this cuts both ways. A drug's harms are also usually reported relatively. "Doubles the risk" of something that occurs in one person per hundred thousand is a very different sentence from "doubles the risk" of something common. Marketing quotes relative benefits and absolute harms. Critics quote absolute benefits and relative harms. Both are doing the same thing.


5.9 Estimands: why one trial reports two numbers

A subtler problem, and one that produces genuine confusion about a drug people actually take.

Here is the situation. You run a weight-loss trial for 72 weeks. Some participants stop taking the drug — side effects, life circumstances, moving away. Some stop and stay in the study. Some leave entirely.

What is "the result"?

There are two defensible answers, and they answer two different questions.

The treatment-regimen estimand (roughly, intention-to-treat) asks: "What happens if you assign people to this drug?" It counts everyone as randomized, including those who stopped. This reflects the real world, where some people will not tolerate a drug, and it is generally the primary analysis for regulatory purposes.

The efficacy estimand (roughly, per-protocol or on-treatment) asks: "What happens if you actually take it, as prescribed, for the full period?" It restricts to participants who stayed on treatment. This reflects the drug's biological effect more directly and is systematically more favorable.

TWO QUESTIONS, ONE TRIAL — SURMOUNT-1, tirzepatide 15 mg, 72 weeks

  ┌─────────────────────────────────────────────────────────────────────────┐
  │  TREATMENT-REGIMEN ESTIMAND          "if you're PRESCRIBED it"          │
  │  counts everyone as randomized, including those who discontinued        │
  │  → about −21% from baseline                                            │
  ├─────────────────────────────────────────────────────────────────────────┤
  │  EFFICACY ESTIMAND                   "if you TAKE it"                   │
  │  restricted to participants who remained on treatment                   │
  │  → about −22.5% from baseline                                          │
  └─────────────────────────────────────────────────────────────────────────┘

  SAME TRIAL. SAME DOSE. SAME DURATION. BOTH FIGURES CORRECT.
  Popular coverage quotes whichever is larger, almost always without naming which.

This is not a trick and nobody is lying. Both analyses are pre-specified, both are reported, and the difference between them is genuinely informative — it tells you something about discontinuation and adherence.

The failure is in the reporting. A number quoted without its estimand is incomplete, and because the efficacy estimand is always the larger one, the incomplete version is systematically optimistic.

The rule for reading: whenever you see a weight-loss figure from a trial, ask which analysis. If the source cannot tell you, treat the number as an upper bound. And note that the same issue arises in every long trial of any drug with meaningful side effects — it is not specific to this class, it is just unusually visible here because the numbers are large and widely quoted.


5.10 Red flags

A checklist. None is individually disqualifying; several together should stop you.

In the study itself:

  • No control group. Everyone got the treatment and improved. Compared with what?
  • Tiny sample. Ten people. Any result is compatible with chance.
  • Short duration relative to the claim. Twelve weeks does not answer a two-year question.
  • Surrogate endpoint only, with no evidence the surrogate predicts the outcome (§5.6).
  • Unblinded, with a subjective endpoint. The two together are much worse than either alone.
  • The primary endpoint was not pre-specified, or the reported endpoint differs from the registered one. Checkable on a trial registry, and worth checking.
  • Composite endpoints that combine a serious outcome with a minor one, where the minor one drives the result.
  • Subgroup findings presented as the main result. Slice enough ways and something will be significant.

In the source:

  • Funded by whoever sells it, with no independent replication. Not disqualifying — most drug trials are industry-funded — but it changes the weight, and it should be disclosed.
  • Published in a predatory journal. Journals that publish anything for a fee exist in large numbers. Check whether anyone cites it and whether the journal has an editorial board that exists.
  • A conference abstract that never became a paper. A meaningful fraction of conference abstracts never publish in full, and the ones that do not are disproportionately negative or flawed.
  • A preprint presented as peer-reviewed. Preprints are legitimate and useful and have not been reviewed. Say which you have.
  • A citation that does not support the claim. Following citations is tedious and it is the single highest-yield check available. A surprising number of confident claims cite papers that say something else.

In the framing:

  • "Studies show" without saying which, in whom, or how many.
  • Mechanism presented as evidence of effect (Chapter 2).
  • Anecdote stacked to look like data. A hundred testimonials is one testimonial, a hundred times.
  • Relative risk quoted alone (§5.8).
  • A weight-loss or effect figure with no estimand named (§5.9).
  • The claim cannot be falsified. Ask what result would have counted as failure. If there is none, stop.

⚠️ Hype Check — "studies show"

Two words that appear on nearly every product page in this market, and they are doing something specific.

"Studies show this peptide accelerates recovery and reduces inflammation."

What's true: there are almost certainly studies. That is precisely the point. The sentence is close to unfalsifiable, because for any compound anyone has ever put in a dish, some study exists.

What it conceals is the entire content of this chapter. Three questions collapse it:

  • Which rung? (§5.2) "Studies" spans everything from a cell line to a seventeen-thousand-person outcome trial, and the phrasing is chosen so that you cannot tell which.
  • In whom? (§5.1, §5.5) Cells, rats, healthy volunteers, or people with the condition being claimed?
  • All of them, or the ones that agreed? A count of supportive studies is not an evidence base. The evidence base includes the ones that found nothing — and in the animal literature, those may never have been published at all (§5.3).

The tell is a plural with no referent. A source that has the evidence names it: this trial, this many people, this endpoint, this result. A source that says "studies show" has chosen the impression over the citation, and the check takes about thirty seconds, because all you have to ask is which ones.


5.11 The four-tier rating system

Everything above assembles into this. It is the book's spine and it runs for the remaining thirty-five chapters.

Rating Name What it requires What it licenses
Strong clinical evidence Multiple adequately powered RCTs in humans, consistent direction, regulatory approval for the indication or equivalent, known safety profile "This works in people like those studied, and we know roughly how well and at what cost."
⚠️ Promising but preliminary Real human data that does not settle the question: Phase I/II, small or short RCTs, mixed results, surrogate endpoints only, or a narrow approval being extrapolated broadly "There is something here. We don't know how big, how durable, in whom, or at what risk."
Hype outpaces evidence Animal or in-vitro data only with no completed human trials; OR human trials that failed or contradicted the popular claim; OR a claim outrunning the data that exists "The confident version of this claim is not supported. That is not the same as 'it doesn't work.'"
🔬 Frontier Early-stage science proceeding properly; mechanism established, clinical translation unproven; too soon to rate "Genuinely interesting, genuinely unsettled. Check back."

Six rules for using it. These are frozen, and every chapter follows them.

  1. A rating attaches to a claim, not a molecule. Population and endpoint, always. Never "BPC-157: ❌"; always "BPC-157 for tendon healing in humans: ❌."
  2. ❌ is a statement about evidence, not about a molecule's worth. Several ❌ compounds here will likely be ⚠️ or ✅ in a decade. Some will be shown not to work. Both futures are live.
  3. Never upgrade a rating with mechanism. Chapter 2. "It makes biological sense" is a reason to run the trial.
  4. Never downgrade a rating with distaste. If the evidence is strong for a compound with a bad reputation, say so. The willingness to rate against expectation in both directions is what makes any of the ratings worth reading.
  5. Every rating is date-stamped and falsifiable. "As of this writing," plus what would change it. A rating you cannot imagine revising is a belief.
  6. One molecule, many ratings. Semaglutide for weight loss in obesity is ✅. Semaglutide for Alzheimer's disease is 🔬. Compressing those into one verdict destroys the information.

5.12 Practice: rating two claims properly

The two claims from the Overview, worked in full. These are the book's first formal ratings, and they are the template for every one that follows.

📊 Evidence Rating — semaglutide for weight loss in adults with obesity

Claim: Semaglutide 2.4 mg weekly produces substantial, sustained weight loss in adults with obesity or overweight with a weight-related comorbidity, alongside lifestyle support.

Rating:Strong clinical evidence.

Why: Multiple large randomized, double-blind, placebo-controlled trials (the STEP program). In adults with overweight or obesity without diabetes, mean weight change of about −15% from baseline at 68 weeks versus roughly −2.4% on placebo. Consistent direction across trials, regulatory approval in multiple jurisdictions, and a characterized adverse effect profile. Cardiovascular outcome data in a related population (SELECT, §5.8) adds a hard endpoint.

What would change it: Long-term data showing that the effect does not persist, or that a serious harm emerges with extended use, would move this. The finding that weight returns after discontinuation (Chapter 8) does not change this rating — it changes what the drug is for (chronic therapy rather than a course of treatment), which is a different claim.

What this rating does NOT cover: weight loss in people who are not overweight; use without lifestyle support; use beyond the studied durations; and any of semaglutide's other indications, which carry their own separate ratings (Chapter 10).

📊 Evidence Rating — BPC-157 for tendon healing in humans

Claim: BPC-157 accelerates healing of tendon and soft-tissue injuries in humans.

Rating:Hype outpaces evidence.

Why: As of this writing, there is no completed, peer-reviewed, randomized controlled human trial of BPC-157 for this or any indication in the published literature. The supporting evidence is a substantial body of rodent work — which is genuine, is often well conducted, and sits on the animal rung of the ladder (§5.2, §5.3). The confident human claim rests on an extrapolation that roughly nine out of ten compounds fail to survive.

What would change it: A completed, adequately powered, randomized, controlled human trial with a pre-specified functional endpoint would move this to ⚠️ or ✅ depending on the result — or would move it further toward ❌ if it failed. Even a well-conducted Phase I safety study would be a meaningful addition, since human safety data does not currently exist either.

What this rating does NOT say: that BPC-157 does nothing. That is not established either. The animal data is real and the human question is genuinely open. ❌ describes the state of the evidence, not the state of the molecule.

Read those two side by side and notice what is being compared. Not "a good drug and a bad compound." Two claims, held to the same standard, with the same questions asked of each. One has answers. The other has an unanswered question and a lot of confidence.

That is the whole method.

📊 Evidence Rating — semaglutide for cardiovascular events in established cardiovascular disease

One more, in the compressed four-line form the rest of the book uses whenever a claim needs a rating rather than a full discussion.

Claim: In adults with established cardiovascular disease and overweight or obesity but without diabetes, semaglutide 2.4 mg weekly reduces major adverse cardiovascular events.

Rating:Strong clinical evidence, as of this writing.

Why: A large randomized, double-blind, placebo-controlled trial with a pre-specified hard composite endpoint found roughly a 20% relative reduction — about 1.5 percentage points in absolute terms — over about three years of follow-up (§5.8).

What would change it: A comparably powered trial in this population failing to reproduce the direction of effect, or longer follow-up showing the benefit does not persist. And note what the rating deliberately does not reach: people without established cardiovascular disease, people with diabetes, and any period beyond the trial's follow-up are each a separate claim needing separate evidence.


📋 Your Evidence Dossier

This chapter fills Fields 5 and 6 — Evidence and Rating. These are the core of the project.

FIELD 5 — EVIDENCE
  Human RCTs           how many, what size, what duration, what population, what result
  Other human data     cohort, case series, uncontrolled reports
  Animal data          what models, what endpoints, how consistent
  In vitro / mechanism  what is established at the cell level
  What does NOT exist  the most important line. Name the missing study explicitly
  Funding              who paid for the studies that exist

FIELD 6 — RATING
  The claim            with population and endpoint attached
  The rating           ✅ / ⚠️ / ❌ / 🔬
  The one reason       a single sentence
  What would change it a SPECIFIC finding, not "more research"

The two hardest lines, and why they are the point

"What does NOT exist." Most sources tell you what has been found. Very few tell you what has been looked for and not found, or never looked for at all. Writing this line forces you to notice the difference between a compound with negative evidence and a compound with no evidence — a distinction that Chapter 5 has spent its length insisting on and that almost no product page will make for you.

"What would change it." Not "more research." A specific finding: a randomized trial in this population, with this endpoint, of this duration, showing this magnitude. If you cannot write it, you do not yet understand your own position well enough to defend it — and you certainly cannot recognize the evidence when it arrives.

Your task

Complete Fields 5 and 6 for every peptide in your dossier.

Expect this to take longer than every previous field combined. That is correct and it is diagnostic: if Field 5 is easy to complete for one of your peptides and nearly empty for another, you have learned something before you have written a single rating.

Two warnings.

First, do not rate a molecule. Rate a claim. If your dossier entry says "BPC-157: ❌," go back and add the population and the endpoint. Every entry should have at least one claim; several will have three or four with different ratings.

Second, write the rating before you look at Chapter 37's master table. You will compare in Chapter 37, and the comparison is only informative if your rating was independent. Where you differ, the interesting question is not who is right — it is whether you were consistently more generous toward the compounds you were hoping would work.

Everyone drifts in a direction. Finding out which direction is yours is worth more than any individual verdict.


Conclusion

Every claim has to answer one question: how would we know? Ask it first, and a great many claims resolve before any expertise is required.

Evidence sorts onto a ladder by how well it controls self-deception — from anecdote and mechanism at the bottom, through cell and animal work, through observational human studies, to randomized trials and systematic reviews at the top. The rungs are not a ranking of worth; every drug starts at the bottom and climbs. They rank what a result licenses you to believe about the next person.

Animal studies are essential and do not transfer reliably — five independent reasons, and roughly nine in ten compounds entering human trials never reach approval. The correct posture toward a strong animal result is this is a reason to run the human trial, and until someone does, the question is open.

Trial phases each detect something different, and Phase II results are systematically optimistic. Six structural features decide a trial's quality: population, comparator, randomization, blinding, endpoint, and duration-and-size. Surrogate endpoints are the most common way a technically valid trial misleads, and the arrow from surrogate to outcome is a hypothesis rather than an assumption.

Two statistical skills matter more than the rest. Relative and absolute risk describe the same result and leave opposite impressions — SELECT's 20% relative reduction is about 1.5 percentage points absolute, and both must be stated. And estimands mean one honest trial reports two different effect sizes — SURMOUNT-1's roughly −21% and roughly −22.5% are the same trial answering two different questions.

Then the ratings: ✅ ⚠️ ❌ 🔬, attached to claims and never to molecules, never upgraded by mechanism, never downgraded by distaste, always date-stamped, always falsifiable.

You now have the method. The remaining thirty-five chapters are applications of it.

Chapter 6 examines the information environment this method has to operate in — how a real preclinical result becomes a sales page in five predictable stages, and why intelligent people with good intentions are the primary vector.


Key Terms

Randomization — assigning participants to trial arms by chance, making the groups comparable in every respect including unmeasured ones.

Blinding — concealing treatment assignment from participants (single), and from investigators (double); open-label means nobody is blinded.

Placebo — an inactive treatment indistinguishable from the active one, used as a comparator.

Control arm — the group receiving placebo, standard care, or nothing, against which the treatment group is compared.

Endpoint — the outcome a trial measures.

Primary endpoint — the pre-specified main outcome a trial is designed and powered to detect.

Surrogate endpoint — a measurement believed to predict a clinically important outcome, used because it is faster or cheaper to measure.

Hard endpoint — an outcome that matters directly: death, heart attack, stroke, hospitalization.

MACE — major adverse cardiovascular events; a composite endpoint typically comprising cardiovascular death, non-fatal myocardial infarction, and non-fatal stroke.

Phase I–IV — the stages of human drug development: safety and pharmacokinetics; efficacy signal and dose; definitive efficacy; post-marketing surveillance.

Powered — describing a trial large enough that a real effect of the expected size would be detected.

Intention-to-treat — analyzing participants in the group they were randomized to, regardless of what they actually received.

Per-protocol — analyzing only participants who completed the treatment as assigned.

Estimand — the precise quantity a trial analysis is estimating; determines whether discontinuing participants are counted, and therefore which of two correct numbers a trial reports.

P-value — the probability of observing a result at least this extreme if the treatment had no effect. Below 0.05 is conventionally "statistically significant."

Confidence interval — a range of values consistent with the observed data; more informative than a p-value and less often reported.

Hazard ratio — the rate of events in the treated group relative to the control group; 0.80 means a 20% relative reduction.

Relative risk reduction — the proportional reduction in risk; a ratio.

Absolute risk reduction — the difference in risk in percentage points; the figure that determines what an individual can expect.

Number needed to treat (NNT) — how many people must be treated for one additional person to benefit; the most honest single number in medicine.

Meta-analysis — a statistical pooling of results from multiple studies.

Systematic review — a structured search and appraisal of all evidence on a question.

Publication bias — the tendency for positive results to be published and negative ones not to be, distorting the apparent evidence base.

Predatory journal — a publication that will print essentially anything for a fee, without genuine peer review.

Preprint — a manuscript posted publicly before peer review; legitimate, useful, and not reviewed.

Effect size — the magnitude of a difference, as distinct from whether it is statistically significant.


Spaced Review

  1. (Ch 1) A source states that a compound is "the same peptide used in clinical research." Using Chapter 1 and §5.10, list what this claim does and does not establish, and name the specific check you would run.

  2. (Ch 4) A product claims oral efficacy for a peptide normally injected, citing a single uncontrolled study of twelve people who reported feeling better. Identify every problem, in order of how decisive each is. Which comes from Chapter 4 and which from Chapter 5?

  3. A trial reports a 40% relative reduction in an outcome that occurs in 0.5% of the control group per year. Express the result in absolute terms and estimate the NNT. Then write the one sentence you would use to communicate it honestly to a patient.

  4. (Ch 2) A compound reliably raises IGF-1 in humans. Explain, using §5.6 and Chapter 2's seven-step gap, why this does not establish that it improves body composition or function — and name the study that would.

  5. Explain to a friend, in under five sentences, why "there are over a hundred studies on this compound" can be true while the honest verdict is still ❌. Do not use the words ladder, surrogate, or estimand.