Appendix H — Twenty Worked Evidence Evaluations

Chapter 5 gave you the method and Chapter 37 gave you the ratings. This appendix does the thing neither of them could do on its own: it runs the method, start to finish, twenty times.

That is the entire design. You do not learn an evaluation method by reading a description of it any more than you learn to sight-read by reading about music theory. You learn it by watching it applied to case after case until the shape of the work becomes familiar — until, on encountering claim twenty-one somewhere out in the world, the first thing that happens in your head is not "is that true?" but "true of whom, measured how, compared against what?" That reflex is the transferable skill. Everything else in this book is scaffolding for it.

So the repetition here is not padding. The repetition is the content. Each of the twenty evaluations below is worked through the same seven steps in the same order, including the steps that feel unnecessary for that particular claim. Especially those, in fact. The step that feels unnecessary is usually the one you would skip on a claim where it mattered, and you will not know in advance which claim that is.

A word on the selection. The twenty were chosen to span all four ratings, and that constraint did real work — it forced the inclusion of claims this book rates well alongside claims it rates badly. This matters more than it might appear. A method that only ever produces ❌ is not a method. It is a disposition wearing a method's clothing. Reflexive skepticism is not more rigorous than reflexive credulity; it is the same failure with better manners, and it is just as useless for telling one claim from another. If you work through these twenty and find that the evaluations you enjoyed most were the demolitions, notice that, and go back and reread the first five. Five of the twenty are ✅. One of them — insulin for type 1 diabetes — is about as well supported as any claim in medicine. The method has to be able to say so.

The twenty also cluster: several are the same molecule evaluated for different purposes, or the same purpose evaluated in different populations. That is deliberate too. Some of the most instructive comparisons in the appendix are between adjacent entries, and a few of them are placed side by side specifically so the contrast is unavoidable.


H.1 The template

Every one of the twenty evaluations below uses these seven headings, in this order, with no exceptions and no omissions.

THE CLAIM AS USUALLY ENCOUNTERED — how a reader actually meets it, in the wild: the phrasing on the label, in the video, in the thread, at the dinner table. Not a tidied-up version.

THE CLAIM, RESTATED PRECISELY — population, endpoint, comparator, timeframe. The four questions from Appendix D §D.2, applied to the claim rather than to a paper.

THE EVIDENCE — what actually exists, described by design rather than by count. Three randomized placebo-controlled trials is a different statement from three studies, and "a large animal literature" is a different statement from "a large literature."

WHAT IT SHOWS — the strongest honest reading. Steelman it. If the claim has a defensible core, say what it is, in its best form.

WHAT IT DOESN'T SHOW — the limits, stated flatly and without hedging them back into insignificance. This is the section people soften. Don't.

THE RATING — ✅ ⚠️ ❌ 🔬, date-stamped, and where the rating is ❌, which kind of ❌.

WHAT WOULD CHANGE IT — a specific finding that would move the rating. Not "more research."

The last step deserves a note now, because it is the one most often left out and the one that does the most to keep a rating honest. "More research is needed" is not an answer to "what would change your mind." It is a way of appearing open while committing to nothing. A real answer names a study design, a population, an endpoint, and a direction: a randomized placebo-controlled trial in adults with this injury, reporting return-to-function at this interval, showing a difference of about this size. If you cannot produce that sentence for a claim you hold, you are not holding a claim. You are holding a preference.


H.2 The two kinds of ❌, because several entries turn on it

This book's ❌ covers two situations that feel similar and are not remotely similar, and the distinction is worth fixing in mind before the evaluations start.

   ❌  EVIDENCE ABSENT              ❌  EVIDENCE PRESENT AND NEGATIVE
   ------------------------          --------------------------------
   Adequate trials have not          Adequate trials were run,
   been run.                         and they answered no.

   The claim is unsupported          The claim is refuted.
   AND UNTESTED.

   We do not know.                   We know.

Both earn ❌ because in both cases the claim, as made, should not be relied on. But they describe opposite states of knowledge. Evidence absent means nobody has looked properly. Evidence present and negative means the field looked hard, with good tools, and came back with an answer — the answer just wasn't the one anyone hoped for.

The second is a far stronger state of knowledge than the first, and readers routinely treat it as the weaker one. This inversion is so common it is almost a reflex. A claim with failed trials behind it sounds discredited, damaged, embarrassing — while a claim with no trials behind it sounds merely early, unproven, promising, not yet. So the untested claim gets the benefit of the doubt and the tested one gets written off, when the truth is the reverse: about the tested claim we have real information, and about the untested one we have none at all.

Entries 14 and 16 are placed in the same group specifically to make this collision unavoidable. Entry 18 adds a third situation that is neither of these, and is stranger than both.


H.3 Claims that hold up

Five claims that survive the method. They are worth working through carefully for two reasons. The first is calibration: if you have spent time in the corners of this subject where the claims are worst, you can lose the ability to recognize a good one, and the good ones are why any of this matters. The second is that a ✅ arrived at by this method is always narrower than the claim that prompted it — and watching that narrowing happen five times is the best preparation for the groups that follow. None of the five below licenses the sentence people usually want it to license.


1. Semaglutide causes substantial weight loss

THE CLAIM AS USUALLY ENCOUNTERED. "This drug makes people lose fifteen percent of their body weight." It arrives as a magazine cover, a before-and-after, a colleague's offhand remark, a chart with a very steep line. Sometimes it arrives with the number attached and sometimes just as a general atmosphere of this one actually works, which is unusual enough in the weight-loss category to be newsworthy on its own.

THE CLAIM, RESTATED PRECISELY. In adults with overweight or obesity without type 2 diabetes, does weekly subcutaneous semaglutide, at the dose studied, produce greater reduction in body weight at 68 weeks than placebo?

Note what the restatement has already done. It has excluded people with diabetes — which will turn out to matter enormously, see entry 2. It has fixed a timeframe, so that "loses weight" becomes "weighs less at a specific moment." And it has named the comparator as placebo, which means the answer will be better than nothing, not better than the alternatives.

THE EVIDENCE. The STEP program: randomized, placebo-controlled trials designed and powered for this question. In STEP 1, participants received 2.4 mg weekly, and mean body weight at 68 weeks was approximately −15% from baseline, against approximately −2.4% in the placebo group. This is the design you want for a question about benefit: randomization to balance the confounders nobody thought of, a placebo arm to establish what happens without the drug, a prespecified endpoint, and enough participants and duration to make the answer mean something.

WHAT IT SHOWS. That in this population, this drug at this dose produces weight loss of a magnitude that had not previously been achievable with pharmacotherapy, and that the effect is attributable to the drug rather than to enrollment, attention, or the structured behavioral support that trial participants receive — because the placebo group received those too and lost about 2.4%. The gap between the arms is the finding, and it is large.

WHAT IT DOESN'T SHOW. Three things, each of which is routinely assumed and none of which this trial addressed.

It does not show that the weight loss persists after stopping. A 68-week on-treatment result tells you what happens while the drug is being taken. What happens afterward is a separate question with a separate evidence base, and treating a 68-week figure as a permanent outcome is the single most common misreading of this literature.

It does not show long-term outcomes. Weight is a measurement; living longer or better is an outcome. That the two are related in populations does not mean this specific weight change produces those specific benefits in an individual, and that inference requires its own trial — which, for a related question, is entry 3.

And it does not show that the result applies unchanged to people with type 2 diabetes. It explicitly does not, and entry 2 is the demonstration.

THE RATING. ✅ as of this writing, in 2026 — for this population, this endpoint, this timeframe, against placebo.

WHAT WOULD CHANGE IT. Very little would overturn the core result; it has been replicated across a program. What would meaningfully change the practical rating is longer-term randomized follow-up showing that the difference between arms substantially narrows while treatment continues, or independent replication of the discontinuation trajectory showing regain sufficient to make the 68-week figure misleading as a description of what the drug does over years.


2. Semaglutide causes substantial weight loss in people with type 2 diabetes

THE CLAIM AS USUALLY ENCOUNTERED. Usually not encountered as a separate claim at all — and that is the point of including it. The result from entry 1 gets quoted, someone with type 2 diabetes hears it, and the number travels intact across a population boundary that nobody mentioned was there.

THE CLAIM, RESTATED PRECISELY. In adults with overweight or obesity and type 2 diabetes, does weekly semaglutide at the same dose produce greater reduction in body weight at 68 weeks than placebo, and of what magnitude?

The restatement here differs from entry 1 by exactly one entry criterion. Same molecule. Same weekly dose. Same endpoint. Same duration. Same comparator. One changed line in the inclusion table.

THE EVIDENCE. STEP 2, randomized and placebo-controlled, in participants with type 2 diabetes. Mean weight change at 68 weeks was approximately −10%.

WHAT IT SHOWS. That the drug works in this population as well — meaningfully, reliably, and by an amount that remains clinically substantial. This is a ✅. It should be read as one.

It also shows something the first trial could not: that the effect size is population-dependent. Approximately −15% and approximately −10%, same molecule, same dose, same duration, differing by one enrollment criterion. Whatever explanations exist for the difference, the practical lesson needs none of them. The population is not context for the result. The population is part of the result.

WHAT IT DOESN'T SHOW. That a person with type 2 diabetes should expect the figure from entry 1. More generally, it does not show anything about populations neither trial enrolled. If one changed criterion moved the number by roughly a third, the honest default for an unstudied population is that the number is unknown, not that it is approximately the same.

It also does not show why. The mechanism of the difference is a separate question, and the evaluation does not depend on answering it — which is worth noticing, because the temptation to demand a mechanistic story before accepting a result is strong and usually misplaced.

THE RATING. ✅ as of this writing, in 2026 — for this population, at this magnitude.

That phrasing is the entry's reason for existing. A ✅ is not a property of a molecule. It is a property of a claim, and a claim includes the population. Transferring a ✅ across populations without checking is not a small liberty; it is the same error as transferring it across endpoints, and here we have the arithmetic to prove it.

WHAT WOULD CHANGE IT. Direct randomized comparison across these populations showing the difference is an artifact of trial conduct rather than of population would change how the pair is taught, though not the rating of either claim. Rating changes for this entry would require the same kinds of findings described in entry 1.


3. Semaglutide reduces cardiovascular risk

THE CLAIM AS USUALLY ENCOUNTERED. "It cuts heart attacks and strokes by twenty percent." This one travels well because the number is round, large, and about something that frightens people. It is also, unusually for a claim this widely repeated, true — provided every word of the restatement below is carried along with it, which it almost never is.

THE CLAIM, RESTATED PRECISELY. In adults with established cardiovascular disease and overweight or obesity, without diabetes, does semaglutide reduce major adverse cardiovascular events over approximately three years, compared with placebo?

Three restrictions arrived in that sentence. Established cardiovascular disease — this is a secondary prevention population, people who already have the disease. Without diabetes. And a hard composite endpoint over a defined horizon.

THE EVIDENCE. SELECT: randomized, placebo-controlled, approximately 17,000 participants, followed for approximately three years. The result was a relative risk reduction of approximately 20%, hazard ratio approximately 0.80.

Now the conversion, worked in full, because this is the entry where the appendix does that job:

Way of stating the same finding Figure
Relative risk reduction ~20% (hazard ratio ~0.80)
Event rate, placebo → drug ~8% → ~6.5%
Absolute risk reduction ~1.5 percentage points
Number needed to treat, over ~3 years ~65–70

Read down that column slowly. Every row describes the identical result in the identical trial. The first row is the one that becomes a headline. The last row is the one a person deciding whether to take a drug actually needs: treat roughly sixty-five to seventy people like these trial participants for about three years, and one of them avoids an event they would otherwise have had.

Neither figure is dishonest. The 20% is arithmetically correct; 1.5 is 20% of 8. But the two sentences land in a reader's chest with completely different force, and which one a source reaches for, consistently, across many claims, tells you what that source is for. Note also that the NNT is meaningless without the duration attached — treat for longer and you prevent more events — which is why "~65–70" never appears in this book without "over about three years" following it.

WHAT IT SHOWS. A hard outcome. Not a biomarker, not a scan, not a scale reading, not a surrogate that stands in for something that matters — the events themselves, counted, in a large randomized population over a meaningful duration. Chapter 16 spends a chapter on why this distinction is the most consequential one in evidence evaluation, and entries 11 and 15 below show what its absence looks like. This is the strongest single result in this book, and it deserves to be said plainly rather than buried under the qualifications that follow.

WHAT IT DOESN'T SHOW. It does not show benefit in people without established cardiovascular disease. The entry criteria were what they were; a primary-prevention population is a different population with a different baseline risk, and since the absolute benefit depends on the baseline event rate, a lower-risk population would see a smaller absolute benefit even if the relative effect held exactly.

It does not show that weight loss is the mechanism. The trial tested a drug against placebo, not weight loss against no weight loss. The drug does many things; the trial does not partition credit among them. "Losing weight prevents heart attacks" is a different claim with a different evidence base, and SELECT is not it.

THE RATING. ✅ as of this writing, in 2026.

WHAT WOULD CHANGE IT. A completed randomized trial in a primary-prevention population failing to show benefit would not change this rating — it would establish the boundary this rating already declares. What would change this rating is a second large outcome trial in a comparable secondary prevention population failing to replicate, or a component breakdown showing the composite was driven by its softest element with no effect on the harder ones.


4. Insulin treats type 1 diabetes

THE CLAIM AS USUALLY ENCOUNTERED. Not as a claim. As a fact so settled that stating it feels strange — which is exactly why it belongs in an appendix about evaluation, because a method that cannot certify the obvious is not producing knowledge, it is producing doubt.

THE CLAIM, RESTATED PRECISELY. In people with type 1 diabetes, does administered insulin improve survival and glycemic control, compared with no insulin?

THE EVIDENCE. More than a century of clinical practice beginning in 1922, plus a substantial modern randomized literature on the effects of intensive glycemic control. And here is the honest part, which is the reason this entry exists: the foundational evidence predates modern trial methodology entirely. There was no randomized placebo-controlled trial of insulin versus no insulin in type 1 diabetes. There never will be. Withholding insulin is not an ethical comparator, and no committee anywhere would approve the trial that a strict methodological checklist would demand.

So a mechanical application of the evidence hierarchy — randomized beats observational, placebo- controlled beats uncontrolled, modern beats historical — would rate one of the best-established claims in medicine as poorly supported. That result should worry you about the checklist, not about insulin.

WHAT IT SHOWS. About as much as medical evidence ever shows. The effect is enormous, immediate, mechanistically understood, reversible on withdrawal, reproducible in every patient, and consistent across a century and every population on earth. Effect sizes this large do not require randomization to be believed, because there is no confounder capable of producing them. This is the same logic that makes a parachute trial unnecessary, and it is a legitimate part of evidence evaluation rather than an exception to it.

WHAT IT DOESN'T SHOW. That any specific insulin analog is superior to any other. That is a separate claim, it is a comparative one, and comparative claims about modest differences absolutely do require randomized trials — because there, the effects are small enough that confounding can manufacture them. The strength of "insulin treats type 1 diabetes" transfers to none of the downstream claims about formulation, delivery, or brand.

It also does not show anything about type 2 diabetes, where the disease, the role of insulin, and the comparators are all different.

THE RATING. ✅ as of this writing, in 2026 — and this entry exists to make a point about what the glyph means. Evidentiary strength and trial-design purity are not the same thing. Design quality is how you get strength when the effect is small and the confounders are plausible. When the effect is overwhelming and the mechanism is understood, other things can supply it. A reader who has learned only "randomized good, uncontrolled bad" has learned a useful heuristic and mistaken it for the principle underneath, which is: how likely is it that something other than the treatment produced this?

WHAT WOULD CHANGE IT. Essentially nothing available. This is worth stating rather than dodging, because a claim that nothing could change is normally a warning sign — and the correct response is not to pretend otherwise but to explain why this case is different: the disconfirming observation would be people with untreated type 1 diabetes doing well, which is observable, has been observable for a century, and does not occur.


5. Tesamorelin reduces visceral fat in HIV-associated lipodystrophy

THE CLAIM AS USUALLY ENCOUNTERED. Rarely in its own right. It is usually encountered as a premise in someone else's argument: a growth-hormone-axis peptide has an approved indication for reducing visceral fat, therefore growth-hormone-axis peptides reduce visceral fat, therefore this category of compound is a body-composition tool. The first step of that chain is true. The evaluation is about the second.

THE CLAIM, RESTATED PRECISELY. In adults with HIV-associated lipodystrophy, does tesamorelin reduce visceral adipose tissue compared with placebo?

THE EVIDENCE. Randomized, placebo-controlled trials — the design that supports the approved indication. This is a specific population with a specific pathophysiology, studied properly, with a prespecified endpoint measured by imaging rather than by impression.

WHAT IT SHOWS. A real effect on a real endpoint in a defined clinical population, sufficient for regulatory approval. This is a ✅ and it is not a grudging one. The trials were done, the effect was found, the indication exists.

WHAT IT DOESN'T SHOW. Anything about growth-hormone-axis peptides for body composition in healthy adults. This is the most commonly made leap from this result and it fails on every axis at once.

The population differs — HIV-associated lipodystrophy is a specific condition with a specific metabolic derangement, not "having some visceral fat." A treatment that corrects a disordered state does not necessarily do anything comparable to a physiologically normal one; that is close to a general rule in endocrinology and it has exceptions, but the exceptions have to be demonstrated rather than assumed.

The endpoint differs — visceral adipose tissue on imaging is not "getting lean," is not body composition broadly, and is not physical function or appearance.

And the risk-benefit calculus differs completely. In a patient with a disease, a given adverse effect profile buys a treated disease. In a healthy person, it buys a changed measurement. Those are not the same trade, and evaluating them as though they were is the structural error underneath a large fraction of the claims in this subject area.

THE RATING. ✅ as of this writing, in 2026 — for the approved indication only. The broader claim about healthy adults is a different claim, not rated here, and it is not entitled to borrow this rating.

WHAT WOULD CHANGE IT. For the rated claim: failure to replicate the visceral adipose tissue reduction in a comparable population, or evidence that the imaging endpoint does not track anything patients care about. For the broader claim to earn any rating at all, it would need what it does not have: randomized placebo-controlled trials in healthy adults, with body composition and function as prespecified endpoints, of a duration adequate to see harms as well as changes.


H.4 Claims that hold up narrowly

Four claims that are supported by real randomized evidence for a real indication, and that are routinely encountered in a form the evidence does not support. The failure mode in this group is not falsehood. It is scope.

These are the hardest claims to argue about, precisely because there is nothing wrong with the underlying result and the person repeating it is holding a genuine fact. The work is entirely in the restatement step — in noticing where the studied claim ends and the repeated claim continues. Notice how often, in the four entries below, the evaluation is essentially finished as soon as the population and the endpoint are written down.


6. NK1 receptor antagonists prevent chemotherapy-induced nausea and vomiting

THE CLAIM AS USUALLY ENCOUNTERED. Almost never in popular writing at all, which is itself informative — this is a solved problem, and solved problems do not generate content. Clinically, it is encountered as a standard component of antiemetic regimens.

THE CLAIM, RESTATED PRECISELY. In patients receiving emetogenic chemotherapy, do NK1 receptor antagonists reduce nausea and vomiting compared with regimens not containing them?

THE EVIDENCE. Randomized trials sufficient to support approval and routine guideline-directed use. Approved, effective, unglamorous, and in wide clinical practice.

WHAT IT SHOWS. That blocking this receptor does something specific and useful for a specific clinical problem, reliably enough that it became standard care. On the evidence, this is straightforwardly a ✅.

WHAT IT DOESN'T SHOW. Anything about the indications this drug class was actually developed for. And this is why the entry is here.

The point of this evaluation is that the antiemetic indication was not the indication the program set out to develop. The substance P / NK1 system was pursued with enormous investment and genuine scientific excitement for pain and for depression. Those programs are entry 16, and they failed — not quietly, not for want of effort, and not because the science was bad. What survived was an indication that was, in the original framing, a secondary consideration.

So the same molecular class, in the same era, from the same laboratories, produces a ✅ here and a ❌ there. A drug class does not have a rating. Claims have ratings, and the distance between the rating for "NK1 antagonists work" and the rating for "NK1 antagonists work for what they were built for" is the entire width of this appendix.

There is a second lesson available, quieter and worth taking. The antiemetic indication is a genuine contribution to the care of people undergoing chemotherapy. It is not a consolation prize. Fields routinely describe an outcome like this as a program's failure because it was not the intended outcome, which is a strange accounting when the result is a real drug helping real patients with a real problem.

THE RATING. ✅ as of this writing, in 2026 — for the antiemetic indication.

WHAT WOULD CHANGE IT. Little for the rated claim; it is embedded in routine practice with randomized support behind it. What is worth watching is the comparative literature — whether newer agents in the class outperform older ones is a separate comparative claim requiring separate head-to-head evidence, exactly as in entry 4.


7. Bremelanotide increases sexual desire

THE CLAIM AS USUALLY ENCOUNTERED. "There's a peptide for libido." Usually stated without a population, occasionally with a nod toward regulatory approval as the credential, and almost always in a context where the person being addressed does not match the trial population on any dimension.

THE CLAIM, RESTATED PRECISELY. In premenopausal women with hypoactive sexual desire disorder, does bremelanotide improve desire and reduce associated distress compared with placebo?

Write that sentence out and most of the evaluation is done. Three restrictions arrived: premenopausal, women, and diagnosed with a specific disorder. The claim as usually encountered has none of them.

THE EVIDENCE. Randomized trial evidence adequate to support approval for that narrow indication. The trials were done properly for the question they asked.

WHAT IT SHOWS. That in this population, this drug produced a measurable improvement over placebo on the prespecified endpoints. A real result in a real population with a real unmet need — an area where the therapeutic options have historically been thin and where a positive randomized result is worth taking seriously.

It is worth adding, in fairness to the evidence, that the endpoints here are subjective by necessity — desire and distress are reported, not measured with an instrument — which is exactly the situation where Appendix D §D.3 says blinding matters most. The trials were placebo-controlled, which is the right protection. But this is a domain where the honest reading includes noting that the placebo response in sexual function research is substantial and that the drug's effect is a difference from that, not a difference from nothing.

WHAT IT DOESN'T SHOW. Applicability to men. Applicability to postmenopausal women. Applicability to people without the diagnosis. And, most importantly, it does not show anything about general libido enhancement, which is a different concept from treating a diagnosed disorder in the same way that treating depression is different from improving mood in people who are not depressed.

Each of those is not a quibble — each is a distinct population or a distinct construct that would require its own trial, and for each the honest answer is that the trial has not been done or has not established the claim.

THE RATING. ✅ as of this writing, in 2026 — for the approved indication. Anything broader is a separate claim, and this appendix declines to rate it, which is itself the finding: an unrated claim is not a rated-as-fine claim.

WHAT WOULD CHANGE IT. For the rated claim: post-approval evidence that the effect does not hold up in wider clinical use, or that the effect size shrinks toward the placebo response under real-world conditions. For the broader claims: randomized placebo-controlled trials in the specific populations being invoked, with prespecified endpoints, which do not currently exist in a form that would support them.


8. Somatostatin analogs are useful in neuroendocrine tumors

THE CLAIM AS USUALLY ENCOUNTERED. Usually as a single undifferentiated statement — "they're used for neuroendocrine tumors" — collapsing two genuinely different claims into one. Splitting them is the whole exercise.

THE CLAIM, RESTATED PRECISELY. Two claims, stated separately because they are separate.

Claim 8a: In patients with carcinoid syndrome, do somatostatin analogs control the syndrome's symptoms — flushing, diarrhea — compared with control?

Claim 8b: In patients with neuroendocrine tumors, do somatostatin analogs exert an antiproliferative effect, delaying disease progression, compared with control?

One is about how a patient feels. The other is about what the tumor does. They are supported by different trials, they would be believed for different reasons, and they could easily have come apart — a drug that controls symptoms beautifully while doing nothing to the tumor is entirely conceivable, and would be a perfectly good drug for one claim and worthless for the other.

THE EVIDENCE. Randomized evidence supports both. That is a genuinely fortunate outcome and not a foregone one.

WHAT IT SHOWS. For 8a, symptom control that materially changes daily life for patients with a syndrome that is miserable to live with. For 8b, a delay in progression demonstrated in randomized trials — a tumor-directed effect, established as such rather than inferred from the symptomatic benefit.

WHAT IT DOESN'T SHOW. That either result implies the other. Had only 8a been established, 8b would remain untested; had only 8b been established, patients would still need a separate answer about symptoms. The evidence for each stands on its own trials.

More broadly, neither claim says anything about somatostatin analogs outside this disease area, and neither speaks to comparative questions among agents in the class or to sequencing against other therapies.

THE RATING. ✅ for 8a and ✅ for 8b, stated separately, as of this writing, in 2026.

This entry is the demonstration of Chapter 37's rating rule 6 — one molecule, many ratings — operating not across wildly different uses but within a single indication area. If a molecule can require two separate ratings for two claims about the same tumor in the same patient, the idea that a molecule has a rating should be finished. It does not. The unit of evaluation is the claim, always, and "is this drug good?" is not a question the method can accept as input.

WHAT WOULD CHANGE IT. For 8b specifically: longer-term outcome data failing to show that delayed progression translates into anything patients experience would not falsify the progression finding but would narrow how it should be described — the perennial question of whether a progression endpoint is an outcome or a surrogate, which Chapter 16 treats at length and which is worth holding open here.


9. Botulinum toxin reduces glabellar lines

THE CLAIM AS USUALLY ENCOUNTERED. As common knowledge, and as the premise of a comparison. The first part is fine. The comparison is entry 17.

THE CLAIM, RESTATED PRECISELY. In healthy adults, does botulinum toxin injected into the muscles that produce glabellar lines reduce the appearance of those lines, compared with placebo injection?

Every word of "injected into the muscles" is load-bearing and will be doing work again shortly.

THE EVIDENCE. Randomized trial evidence supporting approval; the cosmetic indication was approved in 2002. More than two decades of subsequent clinical use at very large scale.

WHAT IT SHOWS. A robust, reproducible, visible effect. The mechanism is understood, the effect is temporary and dose-related, and it is one of the more thoroughly characterized cosmetic interventions in existence. ✅ without hesitation.

WHAT IT DOESN'T SHOW. Anything whatsoever about topical peptide products marketed as alternatives to it.

Not "less than people think." Not "only partially." Nothing. This is worth being blunt about because the marketing structure it enables is so common: an approved injectable establishes that a mechanism works, a topical product invokes the same mechanism, and the credibility of the first silently underwrites the second. It cannot. A result about a molecule delivered by needle into muscle is not evidence about a molecule applied to the surface of intact skin, because — as Chapter 30 established at length — the intervening barrier is not incidental to the comparison. It is specifically excellent at doing the one thing that would have to fail for the comparison to work.

Entry 17 works that claim in full. It is placed in a different group for a reason.

THE RATING. ✅ as of this writing, in 2026 — for the injected product, for this indication.

WHAT WOULD CHANGE IT. Very little; this is a mature, heavily used, well-characterized intervention. The interesting adjacent questions — duration of effect across repeated use, comparative performance among products — are separate claims requiring their own comparative evidence.


H.5 Claims where the evidence is genuinely mixed

The ⚠️ group, and the hardest one to write, because ⚠️ is not a rating people find satisfying. It does not resolve anything. It is a statement that the evidence exists, is not nothing, and is not enough — and it obliges you to keep holding the question open rather than filing it.

Resist the pull toward the poles here. The claims below are neither vindicated nor debunked, and a reader who converts a ⚠️ into a ✅ or a ❌ for the comfort of it has lost information rather than gained it. Two of these four also demonstrate that a single product can carry more than one rating depending on how loudly the claim is made.


10. Thymosin alpha-1 modulates immune function

THE CLAIM AS USUALLY ENCOUNTERED. In two mutually exclusive forms, from people who are both partly right. The promotional version: "it's an approved medicine, used in dozens of countries." The dismissive version: "it's not FDA approved." Both statements can be true simultaneously, and this entry exists because that situation is common and badly handled.

THE CLAIM, RESTATED PRECISELY. In defined patient populations, does thymosin alpha-1 improve clinical outcomes compared with control?

Note that the restatement had to leave the population as a variable, because the claim as encountered does not specify one — and that itself is a finding worth flagging before any evidence is consulted.

THE EVIDENCE. It exists, it spans multiple indications and jurisdictions, and it is of variable quality and variable accessibility. Some of it is not easy to locate, evaluate, or read in the original. That combination — real evidence, unevenly distributed and unevenly reported — is precisely the situation ⚠️ is designed to describe.

The registration status is the instructive part. The compound is registered as a medicine in some jurisdictions and not in others.

WHAT IT SHOWS. That at least one national regulator, reviewing a dossier, concluded the evidence met its standard for at least one indication. That is not nothing. Regulators are not perfect, they differ in evidentiary thresholds and in what they are asked to review, but a registration is a document produced by people who read the data.

WHAT IT DOESN'T SHOW. That the evidence would meet every regulator's standard, or that it establishes the broad immune-modulation claims typically made in its name. Chapter 38 §38.4 lays out the five distinct things "not FDA approved" can mean, and the answer here is the second one: approved elsewhere, not here.

That answer is genuinely informative, and its information content is limited in a specific way: it is neither a credential nor a refutation. A US approval was not sought, or was sought on a dossier that did not meet the threshold, or was not pursued for commercial reasons — and from the outside these look identical. What you can say is that some regulator said yes and the FDA has not. What you cannot say is which of those facts should dominate. People reliably pick the one that suits them and present it as the whole picture; the disciplined move is to report both and let the ⚠️ stand.

THE RATING. ⚠️ as of this writing, in 2026.

WHAT WOULD CHANGE IT. Toward ✅: a large, well-conducted, transparently reported randomized trial in a defined population with a clinical rather than immunological endpoint, published where it can be read and appraised. Toward ❌: the same trial, reporting no difference. The reason this sits at ⚠️ is not that the answer is unknowable — it is that the study which would settle it has not been done in the form that would settle it.


11. MK-677 improves body composition

THE CLAIM AS USUALLY ENCOUNTERED. "It raises growth hormone and IGF-1" — stated as though that settled something. Sometimes with a chart of hormone levels, which is a real measurement of a real change, presented as the conclusion of an argument it has only started.

THE CLAIM, RESTATED PRECISELY. In adults, does an orally active growth hormone secretagogue improve body composition and physical function compared with placebo, over a duration long enough for the change to matter?

The restatement quietly performed the entire evaluation, and it did so by refusing to accept the hormone level as the endpoint.

THE EVIDENCE. That it raises growth hormone and IGF-1 is well established and measurable — the compound does what it says on the pharmacological tin. The evidence on body composition and function is thinner, which is the gap the whole entry is about.

WHAT IT SHOWS. Target engagement. The molecule reaches its target, the target responds, and the response is quantifiable in a blood draw. That is a genuine finding and a necessary precondition for any clinical effect. A compound that failed here would be finished.

WHAT IT DOESN'T SHOW. That any of it produces the outcomes people want. Raising a hormone is a surrogate. Body composition and function are the endpoints that matter, and evidence about the first is not evidence about the second — it is a reason to go look for evidence about the second.

This is the cleanest small example in the book of Chapter 16's surrogate problem, and it is worth sitting with because the trap here is so comfortable. The logic feels airtight: growth hormone is anabolic, this raises growth hormone, therefore this is anabolic. Every step is plausible and the conclusion may even be partly correct. But "plausible mechanism plus confirmed target engagement" is precisely the state of knowledge that has produced the largest number of failed drugs in the history of the field — entry 16 is the same argument, made with more money and better science, and it did not survive contact with a clinical trial. See also entry 15, where a real hormonal and compositional change coexists with an absence of demonstrated benefit.

THE RATING. ⚠️ as of this writing, in 2026. The mechanism claim is solid; the outcome claim is under-evidenced. That combination is what ⚠️ is for, and collapsing it in either direction discards the more interesting half.

WHAT WOULD CHANGE IT. Toward ✅: randomized placebo-controlled trials in a defined adult population, with prespecified body composition and functional endpoints — strength, mobility, something a person would notice — sustained over months rather than weeks, with adverse effects reported at the same resolution as benefits. Toward ❌: the same trials, showing composition changes without functional benefit, which would place this exactly where entry 15 already sits.


12. Topical GHK-Cu improves the appearance of aging skin

THE CLAIM AS USUALLY ENCOUNTERED. At two very different volumes, and the volume determines the rating. Quietly: "there's some evidence it helps skin appearance." Loudly, on a product page: "reverses aging," "rebuilds collagen like an injectable."

THE CLAIM, RESTATED PRECISELY. Two claims, because there are two.

Claim 12a, the modest version: In adults, does topical GHK-Cu produce measurable improvement in skin parameters compared with vehicle?

Claim 12b, the strong version: Does topical GHK-Cu reverse skin aging, or rebuild dermal collagen comparably to an injected intervention?

THE EVIDENCE. For 12a: studies reporting instrument-measured effects, generally modest in magnitude. For 12b: nothing that supports it. The strong claim is not a bigger version of the modest claim's evidence. It is a different claim with no evidence at all, and the modest claim's data cannot be stretched to reach it.

It is also worth noting that copper peptides have a longer and more respectable research history than most ingredients in this category — this is not an invented molecule with a marketing story attached. Which is precisely what makes it useful here: the strong claim's problem is not that the underlying science is fake. It is that the science supports a much smaller sentence than the one on the label.

WHAT IT SHOWS. For 12a: probably something real and probably small. Instrument-measured skin parameters are legitimate endpoints and vehicle-controlled comparisons are the right design; the honest reading is that modest effects have been reported and the literature is not strong enough to call the magnitude with precision.

WHAT IT DOESN'T SHOW. For 12b: anything. There is no bridge from "measurable change in a skin parameter" to "reverses aging," a phrase which does not name an endpoint at all. And "rebuilds collagen like an injectable" adds a comparison that no study made, against a delivery route this product does not use — the same structural error as entry 17, which is entry 9's shadow.

THE RATING. Split, explicitly, as of this writing, in 2026:

  • ⚠️ for 12a — modest instrument-measured effects.
  • ❌ for 12b — the strong marketing version.

One product, two ratings, determined entirely by how loudly the claim is made. This is Chapter 37's rating rule 1 doing real work rather than sitting in a list: the rating attaches to the claim, not to the bottle. The practical consequence is that you cannot answer "is this product any good?" Somebody has to tell you what they think it does first, and their answer determines yours.

WHAT WOULD CHANGE IT. For 12a, toward ✅: larger vehicle-controlled trials with blinded assessment of appearance — the thing the buyer cares about — rather than instrument readouts alone, with effect sizes reported in terms a person could recognize in a mirror. For 12b: a randomized comparison against an injected comparator, which is not a study anyone is likely to run, and whose result is not hard to anticipate.


13. Myostatin and activin pathway inhibitors build muscle in muscle disease

THE CLAIM AS USUALLY ENCOUNTERED. Via the mechanism, which is genuinely striking. Blocking myostatin removes a brake on muscle growth; the phenotypes associated with loss of that signal are dramatic and photogenic. From there the claim assembles itself: block the brake, treat the disease.

THE CLAIM, RESTATED PRECISELY. In patients with a defined muscle-wasting disease, do inhibitors of the myostatin/activin pathway improve function — strength, mobility, the activities patients notice — compared with placebo, over a clinically meaningful period?

The restatement's choice of endpoint is the whole entry. It could have said "muscle mass." It doesn't, and that is not a rhetorical trick: mass is what the mechanism predicts, function is what patients came for.

THE EVIDENCE. Trials exist across several agents and several diseases. The results on muscle mass have been more encouraging than the results on function. Mass has moved. Strength, mobility, and the endpoints that patients and regulators care most about have moved less, or less consistently.

WHAT IT SHOWS. That the pathway is druggable and that the predicted biological effect occurs in humans. That is a real achievement and it validates a substantial body of preclinical work. In a field where most mechanisms do not translate at all, reaching the point where the predicted change is measurable in patients is meaningful.

WHAT IT DOESN'T SHOW. That bigger muscle is better function. This turns out to be a genuinely open empirical question rather than a definitional one, and it is the reason this whole class sits at ⚠️ rather than ✅ despite considerable investment and real biological effects. Muscle mass is behaving, here, exactly as a surrogate endpoint behaves when the surrogate and the outcome come apart — the same pattern as entries 11 and 15, in a much more sophisticated setting.

A second point, technical but consequential: several of the leading agents in this space are antibodies rather than peptides. That distinction is easy to wave away — both are made of amino acids, both are biologics — and it should not be. Chapter 38 §38.2 sets out why: the two differ in size, in manufacturing complexity, in immunogenicity profile, in half-life, and in the regulatory pathway they travel. A reader collecting evidence about "peptides" and sweeping in antibody trials is building a case out of results from a different class of molecule.

THE RATING. ⚠️ as of this writing, in 2026 — mechanism confirmed, mass endpoint moved, functional benefit not established.

WHAT WOULD CHANGE IT. Toward ✅: a randomized placebo-controlled trial in a defined muscle disease with a prespecified functional primary endpoint — a walk test, a timed task, a validated functional scale — showing a clinically meaningful difference sustained over time. Toward ❌: the accumulation of adequately powered trials showing the mass-function dissociation is robust, which would move this class from "not yet demonstrated" into the same category as entry 16 — and would be a more valuable finding than the ambiguity it replaced.


H.6 Claims that do not hold up

Five ❌ ratings, and they are ❌ for four genuinely different reasons. That variety is the point of the group. If you leave this appendix able to tell these five apart, the appendix has done its job, because the failure to distinguish them is the single most common analytical error in this subject.

Two of these entries are placed together deliberately, and the contrast between them is, I think, the most useful two pages in the book.


14. BPC-157 heals tendons and soft tissue

THE CLAIM AS USUALLY ENCOUNTERED. With unusual confidence and unusual specificity. "It heals tendons." Often accompanied by a personal account with a timeline, sometimes with citations to studies that turn out to be real studies, and frequently with the observation that it is banned in sport — offered, remarkably, as evidence that it works.

THE CLAIM, RESTATED PRECISELY. In adults with a defined soft-tissue or tendon injury, does BPC-157 reduce healing time or improve function, compared with placebo or standard care?

THE EVIDENCE. A substantial rodent literature, spanning many models and many years, reporting effects on healing across a range of tissues. And, as of this writing, not one completed, peer-reviewed, randomized human trial in the published literature.

Not a small trial. Not a flawed trial. Not a trial with an equivocal result. Zero. The claim as usually encountered has never been tested in the population it is made about.

WHAT IT SHOWS. That in rodents, in laboratory injury models, effects have been reported consistently enough by enough groups to make this a reasonable thing to want to test in people. That is a real thing to say. It is the state of knowledge from which drug development begins — and Chapter 9 supplies the sobering base rate for what happens between there and a human result.

WHAT IT DOESN'T SHOW. Whether it works in humans. And here is where the wording has to be exact, because the two available errors point in opposite directions and both are common.

This is not a finding that it does not work. Nothing above says that. It might work. The rodent data are not nothing, and dismissing them as meaningless would be its own kind of overclaim.

It is a finding that nobody has looked properly. Those are different claims, and the difference is not rhetorical. If someone tells you the evidence shows BPC-157 does not heal tendons in humans, they are wrong in the same way as the person claiming it does. There is no evidence pointing either direction, because the study that would produce it has not been run.

THE RATING. ❌ as of this writing, in 2026 — evidence ABSENT.

A regulatory note that is not part of the rating. BPC-157 was added to the WADA Prohibited List under S0 — the category for substances not approved for human therapeutic use by any government health authority — effective from the 2022 List. This is worth stating precisely because it is so often deployed as an argument. It is a regulatory fact entirely independent of the evidence rating. S0 is not a finding of efficacy; it is close to the opposite, a designation applied to substances with no approved human therapeutic use anywhere. "It's banned, so it must work" reverses the meaning of the listing. The prohibition is a fact about sport governance, it has real consequences for anyone subject to testing, and it says nothing about tendons.

WHAT WOULD CHANGE IT. This is the easiest "what would change it" in the appendix, which is exactly what makes the current state so unsatisfying: one adequately powered, randomized, placebo-controlled trial in humans with a defined injury, reporting a prespecified functional or imaging endpoint at a clinically relevant interval. One. A positive result moves this to ✅ or ⚠️ depending on magnitude and replication. A negative result moves it to the other ❌ — and, as entry 16 argues, that would be a substantial gain in knowledge rather than a loss.


15. Growth hormone is an anti-aging treatment

THE CLAIM AS USUALLY ENCOUNTERED. As a story about decline and restoration: growth hormone falls with age, aging looks like growth hormone deficiency, therefore replacing it restores youth. The story is tidy, it is decades old, and it has supported a sizable commercial ecosystem.

THE CLAIM, RESTATED PRECISELY. In healthy older adults, does growth hormone administration improve function, healthspan, or mortality, compared with placebo?

Note what the restatement excluded: people with diagnosed growth hormone deficiency, for whom replacement is a different claim with different evidence. And note what it demanded: function, healthspan, or mortality — not measurements.

THE EVIDENCE. Body-composition changes are real. Lean mass rises, fat mass falls; these are measurable and reproducible and nobody disputes them. Benefits on function and outcomes are not established. And unlike most entries in this group, there is something on the other side of the ledger: adverse effects at supraphysiological exposure are established — the consequences of excess growth hormone are visible in the clinical literature on conditions of endogenous excess and in the trial record.

WHAT IT SHOWS. That administering the hormone changes the numbers it would be expected to change. The pharmacology is not in question.

WHAT IT DOESN'T SHOW. That any of it constitutes a benefit. And this is the entry's contribution to the appendix: a real, measurable, reproducible change is not a benefit. Those are separate propositions and the gap between them is where an enormous amount of this subject lives.

Lean mass on a scan is a number. Whether the person carrying it is stronger, functions better, gets sick less, or lives longer are four further questions, each requiring its own measurement, and none of them is answered by the first. The claim being sold is about the four. The evidence is about the one.

This is entry 11's argument with a longer history and a much larger literature — and it is worth seeing them next to each other, because entry 11 sits at ⚠️ while this sits at ❌. The difference is not the strength of the mechanism; both mechanisms are sound. The difference is that here the outcome trials in the relevant population have been pursued long enough, and the harms characterized well enough, that "we don't know yet" is no longer an honest summary.

THE RATING. ❌ as of this writing, in 2026 — for the anti-aging claim in healthy adults. This rating says nothing about growth hormone for diagnosed deficiency, which is a different claim in a different population, and which is not evaluated here.

WHAT WOULD CHANGE IT. A randomized placebo-controlled trial in healthy older adults with a prespecified functional or survival primary endpoint — not body composition — of a duration long enough to observe both benefit and harm, showing benefit that outweighs the adverse effect profile already documented. The bar is high because the harms are not hypothetical, and a benefit claim has to clear what is already known about the cost.


16. NK1 receptor antagonists treat pain and depression

THE CLAIM AS USUALLY ENCOUNTERED. Historically, as one of the most exciting stories in neuroscience: substance P as a pain and mood signal, a receptor to block, and a mechanism so compelling that it drew sustained investment from multiple major programs. Today, mostly as a cautionary tale — and mostly among people who work in drug development rather than among people reading about drugs.

THE CLAIM, RESTATED PRECISELY. In adults with major depressive disorder, or in adults with a defined pain condition, do NK1 receptor antagonists improve symptoms compared with placebo?

THE EVIDENCE. Human trials were run. Repeatedly. They failed.

That is the entire evidence summary and it is the most informative sentence in this group. Not "trials were inconclusive." Not "the program was deprioritized." The compounds went into properly designed human trials against the intended indications and did not beat placebo, more than once, across programs.

WHAT IT SHOWS. That this mechanism, in these populations, on these endpoints, does not produce the clinical effect the biology predicted. That is real knowledge, expensively acquired, and it is knowledge in a way that most negative outcomes are not.

Two further things it shows, both of which are Chapter 22's central lesson:

The mechanism was excellent. This was not a fishing expedition or a molecule in search of a rationale. The preclinical case was strong, coherent, and built on solid neuroscience.

Target engagement was real. The drugs reached the brain and blocked the receptor. This was verified. The compounds did exactly what they were designed to do at the molecular level, and the patients did not get better. A mechanism can be correct, a drug can do precisely what the mechanism requires, and the clinical effect can simply not be there. There is no way to know this in advance. The trial is not a formality that confirms what the biology already established; the trial is the only place the question is answered.

WHAT IT DOESN'T SHOW. That the class is useless — see entry 6, where the same class is ✅ for a different indication. And it does not show that the substance P system is irrelevant to pain or mood biology, only that blocking this receptor with these compounds did not treat these conditions.

THE RATING. ❌ as of this writing, in 2026 — evidence PRESENT AND NEGATIVE.

Now put this next to entry 14, which is why it is here. Both carry ❌. The glyph is identical. The underlying states of knowledge are not remotely comparable.

The field knows vastly more about NK1 antagonists for depression than it knows about BPC-157 for tendons. For the first, the question was asked properly and answered: no. For the second, the question has never been asked in humans at all. If you had to bet on which claim will look worse in ten years, the honest answer is that you cannot, because one of them has been tested and the other has not — and untested is not a milder form of refuted. It is a different thing entirely, and it is the one about which you know nothing.

Notice, too, which of the two you were more inclined to dismiss on reading. Most people find the failed program more discrediting than the empty one. That instinct is exactly backward, and noticing it in yourself is worth more than the rest of this entry.

WHAT WOULD CHANGE IT. A trial in a differently defined population — a biomarker-selected subgroup, say — with a prespecified hypothesis about why this group would respond where previous populations did not, showing benefit. Absent a specific and stated reason to expect a different answer, repeating the same trial is not a plan.


17. Topical acetyl hexapeptide-8 works like injected botulinum toxin

THE CLAIM AS USUALLY ENCOUNTERED. As a phrase engineered to imply a comparison without making one. "Botox in a bottle." "Injection-free alternative." The claim is delivered by juxtaposition — the product's name near the injectable's name — so that the reader performs the comparison themselves and the marketer never has to defend it.

THE CLAIM, RESTATED PRECISELY. In healthy adults, does topically applied acetyl hexapeptide-8 reduce wrinkle depth or appearance, compared with vehicle — and separately, compared with injected botulinum toxin?

Two comparators, and the difference between them is everything. Against vehicle, the question is whether the product does anything. Against injection, the question is whether it does what it is being sold as doing.

THE EVIDENCE. Small studies, frequently vehicle-controlled at best, on instrument-measured endpoints. Against the injectable comparator: essentially nothing, because that trial is not one the category has any incentive to run.

WHAT IT SHOWS. Possibly some small measured effect against vehicle. Cosmetic formulations generally do something to skin measurements — hydration alone moves several instrument readouts — and disentangling an active ingredient's contribution from the vehicle's requires more careful work than this literature typically contains.

WHAT IT DOESN'T SHOW. That it is comparable to the injected product in any respect. And here the decisive problem is not trial quality, sample size, or endpoint choice. The decisive problem is delivery.

Chapter 30 established the point in detail, and it is not a technicality: intact skin is specifically, evolutionarily excellent at keeping molecules out. That is close to its primary function. The stratum corneum is a barrier optimized over a very long time to exclude exactly the kind of large, water-loving, charged molecules that peptides are. Getting a peptide through it in biologically meaningful quantity is a genuinely hard problem that the delivery field treats as a hard problem.

So the comparison being invited is between a molecule injected directly into a target muscle and a molecule applied to the outside of a barrier built to prevent it from reaching anything. That is not a comparison of two treatments. It is not a comparison at all. Even granting the ingredient full mechanistic plausibility — granting that it would work beautifully if it arrived — the argument still fails, because arrival is the step in dispute and no evidence addresses it.

THE RATING. ❌ as of this writing, in 2026, for the comparative claim. The much weaker claim — "produces a small measured change versus vehicle" — is a different claim, and would sit near entry 12a rather than here.

WHAT WOULD CHANGE IT. Direct evidence of meaningful delivery to the relevant tissue at biologically active concentrations, followed by a randomized trial with blinded assessment of appearance — human raters, not instruments — using the injected product as an active comparator. In that order. The delivery question comes first, because without it the trial is testing whether a molecule that never arrived did something after it got there.


18. "Nobody has reported problems with this source, so it's safe"

THE CLAIM AS USUALLY ENCOUNTERED. In forums, group chats, and comment sections, usually in response to someone asking a sensible safety question. "People have been using this one for years and nobody's had issues." It is offered in good faith almost every time, by people who genuinely believe they are citing evidence.

THE CLAIM, RESTATED PRECISELY. Does the absence of publicly reported adverse events associated with a particular unregulated supplier constitute evidence that its products are safe?

Notice that this is a claim about a market rather than about a molecule, which is why it is included. The other nineteen entries evaluate whether a substance does something. This one evaluates a form of reasoning — and that form of reasoning does more practical damage than most of the molecular claims in this appendix.

THE EVIDENCE. There is no adverse-event reporting pathway for unregulated products.

Sit with that. For approved medicines there is machinery: mandatory reporting obligations, regulatory surveillance systems, published safety communications, batch traceability, recall authority, and professionals whose job includes filing reports. Every part of that apparatus exists because detecting rare harms requires infrastructure, not attention.

For an unregulated product bought outside that system, none of it exists. There is no body collecting reports, no requirement to file one, no denominator to compare against, no way to link an adverse event back to a batch, and frequently no way to determine what a product actually contained. A person harmed may not connect the harm to the product; if they do, they may have no venue to report it; if they report it somewhere, the venue may be a forum with a strong interest in the topic and an incentive to explain it away.

WHAT IT SHOWS. Nothing. Not "a little." Nothing at all.

WHAT IT DOESN'T SHOW. That the products are safe, that they are unsafe, or that the harms are rare. Chapter 19 §19.8 makes the general form of the point: an observation carries information only if it could have come out differently. The absence of reports here could not have come out differently. In a system with no reporting pathway, a product that harmed people regularly and a product that harmed nobody generate the same public record: silence.

THE RATING. ❌ as of this writing, in 2026 — and it is a third situation, distinct from both of the ❌ types described in §H.2.

Not evidence absent, in entry 14's sense — that describes a study nobody has run yet but which could be run tomorrow. Not evidence present and negative, in entry 16's sense. This is a claim whose supporting observation could not exist even if the underlying proposition were false. The evidence is not missing; it is structurally unobtainable. No amount of waiting produces it, because the mechanism that would generate it does not exist.

That is a stranger and, in a way, worse position than either. With entry 14, more time might bring an answer. Here, more time brings more silence, and the silence will be read by more people as reassurance, and the confidence will grow while the information stays at zero. A belief that grows more confident with time while receiving no new information is not being updated. It is being repeated.

WHAT WOULD CHANGE IT. A functioning surveillance system for these products — mandatory reporting, an accessible database, batch-level traceability, independent product testing — under which a period of no reports would begin to carry real information. Absent that infrastructure, the reasoning fails regardless of how long the silence lasts or how many people cite it.


H.7 Claims at the frontier

Two claims that are neither supported nor refuted, and that are not ⚠️ either, because ⚠️ describes evidence that exists and conflicts. 🔬 describes something else: a serious question, under active investigation, whose answer is not yet in.

The rating is easy to misread as a mild ✅ — a claim on its way to being confirmed, granted provisional approval in the meantime. It is not that. Read the two entries below with the base rate from Chapter 9 held firmly in mind, and then read the last one twice, because it is where the appendix ends on purpose.


19. Semaglutide may treat Alzheimer's disease

THE CLAIM AS USUALLY ENCOUNTERED. As a headline with a hedge that nobody reads: "could protect against dementia." Sometimes accompanied by a mechanistic story about neuroinflammation or brain insulin signaling, sometimes by an observational finding, occasionally by the fact that trials are running — presented as though the running of a trial were itself a result.

THE CLAIM, RESTATED PRECISELY. In adults with early Alzheimer's disease, does semaglutide slow decline on cognitive and functional endpoints compared with placebo, over a multi-year period?

THE EVIDENCE. Three things, and it matters that they are named separately.

A mechanistic rationale, which is real and reasoned rather than invented. Observational signals, which are hypothesis-generating and subject to every confounder that attends comparing people who take a drug with people who do not. And trials underway as of this writing, which are the only one of the three capable of answering the question.

WHAT IT SHOWS. That serious people, having weighed the biology and the observational data, committed the resources to run the definitive study. That is a meaningful signal about the seriousness of the hypothesis. It is not a signal about the answer.

WHAT IT DOESN'T SHOW. Whether it works. Not a little bit, not provisionally, not on balance. The trial has not read out. The question is open.

And now the part the rating exists to say: 🔬 is not a promise, and it carries no probability of success. It is not a ✅ waiting for paperwork. It is a marker meaning this is being properly tested and we should look at the answer when it arrives.

Most 🔬 becomes ❌. That is not pessimism; it is the base rate, and Chapter 9 gives it in detail. The overwhelming majority of well-motivated, well-funded, mechanistically sensible hypotheses that reach large clinical trials do not succeed. Alzheimer's disease specifically has one of the least forgiving records in medicine — a long series of biologically compelling hypotheses that did not survive their trials. A drug with an outstanding result in one disease has no claim on a second one, which is the essential lesson: entry 3 is real and this entry is open, and the first does not lend anything to the second.

THE RATING. 🔬 as of this writing, in 2026.

WHAT WOULD CHANGE IT. The trial readout. Prespecified cognitive and functional primary endpoints, in the enrolled population, reported in full. Not a mechanism paper. Not a press release. Not a subgroup. Not an interim characterized by a sponsor as encouraging. Appendix D §D.6 covers why each of those is a different object from a result — and this entry is the one where that distinction will be tested in public, probably repeatedly, before the actual answer exists.


20. Antimicrobial peptides will be the answer to antibiotic resistance

THE CLAIM AS USUALLY ENCOUNTERED. As a hopeful conclusion to an article about a genuine crisis. The setup is accurate and serious — resistance is rising, the pipeline is thin — and then the pivot: nature has been solving this problem for hundreds of millions of years with antimicrobial peptides, they kill bacteria by physically disrupting membranes, resistance to them should be harder to develop, and a new class of antibiotics is therefore in reach.

Every clause of that is defensible. The conclusion still does not follow, and the reason it does not is the note this appendix ends on.

THE CLAIM, RESTATED PRECISELY. In patients with systemic bacterial infection, do antimicrobial peptides, given systemically, cure infection at least as well as existing antibiotics, with an acceptable safety profile?

THE EVIDENCE. Decades of research. Broad in vitro activity, thoroughly documented across an enormous number of peptides and organisms. And very few approved systemic agents.

The reasons for that gap are known and are not mysterious: toxicity — membrane-disrupting mechanisms are often insufficiently selective between bacterial and human membranes; stability — peptides are degraded by the proteases that exist to degrade peptides; and manufacturing cost — producing peptides at the scale and price point antibiotics require is a genuine constraint, and antibiotics are among the least commercially forgiving drug classes there are.

WHAT IT SHOWS. That the mechanism is real, that the molecules exist in vast diversity, that they kill bacteria in a dish, and that the underlying idea has enough merit to have sustained decades of serious work by serious scientists.

WHAT IT DOESN'T SHOW. That any of it produces a systemic antibiotic. The killing-in-a-dish step was never the hard part. Selectivity, stability in circulation, distribution to the site of infection, and a manufacturable cost structure are the hard parts, and they have remained hard for a long time.

THE RATING. 🔬 as of this writing, in 2026. The research is legitimate and worth doing. The claim that it will deliver an answer to resistance is not something the evidence currently supports.

And this is where the appendix ends deliberately. The structural point underneath entry 20 is the most important one in the book, and it is the hardest to hold on to, because it requires resisting a story that is genuinely appealing and largely true:

A compelling mechanism and a long research history are not evidence of eventual success.

They feel like it. A field with thirty years of work behind it has accumulated papers, conferences, review articles, funded laboratories, and a shared confidence that the breakthrough is near — and all of that is easy to mistake for progress toward an outcome, when much of it is progress in characterizing why the outcome is hard. A field that has been promising for thirty years is telling you something, and it is not that the payoff is overdue. It is that the obstacles are real, that they have been engaged by capable people for a long time, and that they have not yielded.

This is not an argument for abandoning the field. It is an argument against treating duration of effort as though it were evidence of imminent success — the reasoning that would have you invest more confidence in a hypothesis precisely because it has resisted confirmation the longest.

Hold entry 20 next to entry 14 and the pair covers most of the ground. In one, a long research history in the wrong species is mistaken for evidence in humans. In the other, a long research history against a hard problem is mistaken for evidence that the problem is nearly solved. Both errors substitute the volume of research for its bearing on the claim, and volume is the easiest thing in science to accumulate.

WHAT WOULD CHANGE IT. Randomized controlled trials showing a systemically administered antimicrobial peptide curing serious infection with acceptable toxicity, at a manufacturing cost compatible with actual use. Not a new peptide with impressive in vitro potency — the literature has those in abundance and they are not the bottleneck. The specific findings that would move this rating are about selectivity, stability, and cost, because those are the specific things that have stopped it.


H.8 What twenty evaluations teach

1. The restatement step does most of the work. Look back at how many of these twenty were substantially settled by the second heading, before a single piece of evidence was consulted. Writing "premenopausal women with hypoactive sexual desire disorder" in place of "a peptide for libido" finished entry 7. Writing "injected into the muscles" finished entry 9 and predetermined entry 17. Writing "function, healthspan, or mortality" rather than "body composition" decided entry 15. Writing "without type 2 diabetes" made entry 2 a separate question rather than a footnote. This is the step people skip, because it feels like clerical work before the real analysis begins. It is the real analysis. Vague claims survive by being vague; a large share of them do not survive contact with a population, an endpoint, a comparator, and a timeframe. If you adopt one habit from this appendix, adopt this one — it costs a sentence and it resolves a startling fraction of disagreements before they start.

2. The ✅ ratings are narrower than they sound, and the ❌ ratings are less final than they sound. Every ✅ above came with a boundary attached, and the boundaries were not decorative: ✅ for this population (entry 2), for the approved indication only (entries 5 and 7), for the injected product (entry 9), for the antiemetic use rather than the intended one (entry 6). Strip the qualifier and you have a different and unsupported claim wearing a supported claim's credential — which is how nearly all of the overclaiming in this subject actually happens. It is rarely fabrication; it is a true sentence with its boundaries filed off. And in the other direction, ❌ mostly means not demonstrated, which is a statement about the current state of a literature and not a permanent property of a molecule. Entry 14 could become ✅ next year on the strength of a single well-run trial. Ratings are dated for a reason. Every one above is stamped "as of this writing, in 2026," and that stamp is a commitment to revisit, not a formality.

3. Population is the most frequently omitted element, and it is frequently decisive. Entries 1 and 2 are the demonstration and they are almost unfair in their clarity: the same molecule, at the same dose, for the same duration, measured on the same endpoint against the same comparator, produced approximately −15% and approximately −10% in populations differing by one entry criterion. No subtlety, no confound, no argument — one line in an inclusion table moved the headline number by a third. And population is the element that drops out first when a result travels, because it is the least quotable part of any finding and the part a headline has no room for. Entry 3's benefit belongs to people with established cardiovascular disease. Entry 5's belongs to people with HIV-associated lipodystrophy. Entry 15 turns entirely on healthy adults versus people with a diagnosed deficiency. When you encounter a claim with no population attached, the population has not been established as "everyone." It has been dropped.

4. "What would change it" separates a rating from an opinion. Every one of the twenty entries above carries one, and writing them was the most disciplining part of assembling this appendix — several ratings changed while I was drafting that section, because I could not state what would overturn them. The test is simple and it is worth applying to yourself rather than to others: name the finding that would move you. Not a topic, not "more research," not "better studies" — a population, an endpoint, a design, a direction, and roughly a magnitude. If you can produce that sentence, you hold a claim, and you are in a position to update when the evidence arrives. If you cannot, what you have is a position, and evidence will not touch it, because you have not specified any evidence that could. A claim you cannot imagine disconfirming is not being held as a claim. That test applies as much to the ✅ entries as to the ❌ ones — entry 4 was the hardest to write for exactly this reason, and it needed an explicit account of why the disconfirming observation is available in principle and simply does not occur.

5. The method is not about peptides. Nothing in the seven steps is specific to this subject. Population, endpoint, comparator, timeframe; what exists, described by design rather than by count; the strongest honest reading and the limits stated flatly; a rating with a date; and the finding that would change it. That template works on a nutrition claim, an economic forecast, an educational intervention, a policy evaluation, a management consultant's deck, a supplement, or a friend's confident theory about why something happened. The distinction between evidence absent and evidence present and negative is not a peptide concept — it is the difference between we have not looked and we looked and the answer was no, and it applies everywhere those two states get confused, which is everywhere. Entry 18 is not about peptides at all; it is about what silence means in a system with no mechanism for producing speech, and that reasoning is the same wherever the record is generated by people with an interest in what it says. Peptides are simply an unusually good training ground: the science is genuinely interesting, the commercial pressure is intense, the regulatory picture is genuinely complicated rather than merely confusing, and the claims range from among the best-supported in medicine to among the worst-supported anywhere. You are unlikely to find a better twenty claims to practice on. But the practice was never the point. The point was the twenty-first claim, in some other field entirely, that you will now take apart without being asked to.


Related: Chapter 5 (the method) · Chapter 9 (base rates for clinical success) · Chapter 10 (SELECT) · Chapter 16 (surrogate endpoints) · Chapter 19 (unregulated markets) · Chapter 22 (NK1 and the limits of mechanism) · Chapter 30 (the skin barrier) · Chapter 37 (the rating system and its rules) · Chapter 38 (regulatory status and molecule class) · Appendix C (the dossier) · Appendix D (reading a clinical trial) · Appendix F (red flags)