Chapter 5 — Key Takeaways
How to Evaluate Peptide Evidence
This is the chapter to review before every other chapter in the book.
The core claims
One question comes first: how would we know? Followed by in whom, compared with what, and measured how. A claim with no possible refuting observation is not a weak claim — it is not a claim.
Evidence sorts onto a ladder by how well it controls self-deception. Anecdote and mechanism at the bottom; cell and animal work above them; observational human studies above those; randomized trials and good systematic reviews at the top. The rungs rank what a result licenses you to believe about the next person — not the worth of the research.
Animal studies are essential and do not transfer reliably. Five reasons: different biology, different disease, different dose, different endpoint, and — least discussed — different publication pressure, since animal work is largely unregistered and the published literature is a filtered sample of unknown size. Roughly nine in ten compounds entering human trials never reach approval.
Trial phases detect different things. Phase I: safety and pharmacokinetics. Phase II: an efficacy signal. Phase III: efficacy in the target population. Phase IV: rare harms and real-world effects. Phase II results are systematically optimistic, and approval is a regulatory decision rather than a scientific verdict.
Six features decide a trial's quality: population, comparator, randomization, blinding, endpoint (pre-specified), and duration-and-size.
Surrogate endpoints are the most common way a valid trial misleads. The arrow from surrogate to outcome is a hypothesis, not an assumption — and it has failed in ways that killed people.
The two statistical skills that matter most
RELATIVE vs. ABSOLUTE. SELECT: a 20% relative reduction in cardiovascular events is about 1.5 percentage points absolute over about three years. Identical result, opposite impressions. State both, always, in the same sentence. Marketing quotes relative benefits and absolute harms; critics do the reverse. Both are doing the same thing.
ESTIMANDS. SURMOUNT-1, tirzepatide 15 mg, 72 weeks: about −21% (treatment-regimen — "if you're prescribed it") and about −22.5% (efficacy — "if you take it"). Same trial, same dose, both correct. The efficacy figure is systematically larger, so a number quoted without its estimand is systematically optimistic.
The four-tier rating system
| ✅ | Strong clinical evidence | multiple powered human RCTs, consistent, approved, safety known |
| ⚠️ | Promising but preliminary | real human data that doesn't settle it |
| ❌ | Hype outpaces evidence | animal/in-vitro only, or human trials that contradicted the claim |
| 🔬 | Frontier | early, proceeding properly, too soon |
The six frozen rules:
- A rating attaches to a claim, with a population and an endpoint — never to a molecule
- ❌ is about the evidence, not the molecule. Not "it doesn't work"
- Never upgrade with mechanism
- Never downgrade with distaste — the harder rule, because your own skepticism never feels motivated
- Every rating is date-stamped and falsifiable
- One molecule, many ratings
Evidence ratings issued in this chapter
| Claim | Rating | Why | What would change it |
|---|---|---|---|
| Semaglutide 2.4 mg weekly for weight loss in adults with obesity or overweight-plus-comorbidity, alongside lifestyle support | ✅ | Multiple large RCTs (STEP); about −15% from baseline at 68 weeks vs ~−2.4% placebo in adults without diabetes; approved in multiple jurisdictions; characterized safety profile | Long-term data showing loss of effect, or emergence of a serious harm with extended use |
| BPC-157 for tendon healing in humans | ❌ | No completed peer-reviewed randomized human trial for any indication as of this writing; supporting evidence is rodent work, on the animal rung | A completed, powered, randomized human trial with a pre-specified functional endpoint — or even a Phase I safety study, since human safety data does not exist either |
Neither rating says what people assume. The ✅ does not cover people who are not overweight, use without lifestyle support, or semaglutide's other indications. The ❌ does not say BPC-157 does nothing — that is not established either.
The red-flag checklist
In the study: no control · tiny sample · too short for the claim · surrogate only · unblinded with a subjective endpoint · endpoint not pre-specified or switched · composite driven by its softest component · a subgroup reported as the main result
In the source: funded by the seller with no replication · predatory journal · a conference abstract that never published · a preprint presented as reviewed · a citation that does not support the claim
In the framing: "studies show" · mechanism as evidence · stacked anecdotes · relative risk alone · an effect size with no estimand · a claim that cannot be falsified
Key terms
randomization · blinding · placebo · control arm · endpoint · primary endpoint · surrogate endpoint · hard endpoint · MACE · Phase I–IV · powered · intention-to-treat · per-protocol · estimand · p-value · confidence interval · hazard ratio · relative risk reduction · absolute risk reduction · number needed to treat · meta-analysis · systematic review · publication bias · predatory journal · preprint · effect size
What you can now evaluate
- ✅ any peptide claim in this book, and any claim published after it
- ✅ what a study abstract does and does not establish
- ✅ whether a surrogate endpoint is carrying an untested assumption
- ✅ a relative risk figure, converted to absolute, with an NNT
- ✅ why one trial reports two effect sizes
- ✅ whether a source's evidence base is what it appears to be — including by checking a registry yourself, free, in two minutes
The one-sentence version
Confidence is free; evidence costs money, years, and a real chance of being wrong — and the whole skill is learning to see which one you are being offered.
Next: Chapter 6 — the information environment this method has to survive in, and why intelligent people with good intentions are the primary vector.