25 min read

> *"When a drug appears to work on everything, one of two things is true: you have found something

Prerequisites

  • 5
  • 7
  • 8

Learning Objectives

  • Explain why one drug appears to work on several unrelated diseases
  • State the cardiovascular, renal, and hepatic evidence with its populations and endpoints
  • Explain why a trial stopped early for efficacy is both good news and harder to interpret
  • Describe the biopsy endpoint problem in liver disease trials
  • Distinguish a mechanical from a metabolic explanation in sleep apnea
  • Explain why the Alzheimer's and addiction programs are 🔬 rather than ⚠️
  • Apply one rating per indication to a single molecule

Chapter 10: Beyond Weight Loss: GLP-1 for the Heart, Kidney, Liver, Brain, and Addiction

"When a drug appears to work on everything, one of two things is true: you have found something fundamental, or you have stopped asking hard questions. Telling which takes about a decade."

Overview

Something unusual is happening with this drug class, and it is worth being precise about what.

A compound developed to lower blood glucose reduced cardiovascular events. Then a kidney outcomes trial was stopped early because the benefit was clear. Then liver disease. Then heart failure with preserved ejection fraction. Then sleep apnea. There are trials running in Alzheimer's disease and in alcohol and substance use disorders.

That list should make you suspicious, and this chapter is built around that suspicion.

Drugs that appear to work on everything have a poor historical record. The pattern is familiar from supplement marketing, where a compound "supports" cardiovascular health and cognitive health and immune health and joint health, and the breadth of the claim is inversely related to the evidence for any part of it.

And yet in this case some of the claims have completed randomized trials with hard endpoints. That is a genuinely different situation, and the discipline this chapter requires is to hold the suspicion while evaluating each claim on its own evidence — because the answers turn out to be very different. Cardiovascular is ✅. Kidney is ✅ in a defined population. Liver is a real result with a genuinely problematic endpoint. Alzheimer's and addiction are 🔬, and the gap between "trials are running" and "it works" is the entire content of those sections.

One molecule. Six or more distinct claims. Ratings ranging from ✅ to 🔬.

This is Chapter 5's rule 6 — a rating attaches to a claim, not a molecule — applied at full stretch, and it is the reason that rule exists.

In this chapter, you will learn to:

  • Explain the candidate reasons one drug might genuinely affect several diseases
  • State each indication's evidence with its population, endpoint, and status
  • Explain why "stopped early for efficacy" complicates interpretation
  • Describe the biopsy endpoint problem and why liver trials are hard
  • Distinguish mechanical from metabolic explanations
  • Say precisely why the brain and addiction programs are 🔬
  • Rate one molecule six different ways without contradicting yourself

Learning Paths

💊 GLP-1 — the whole chapter. §10.10 is the section that will most change how you read news about this drug. 🔬 Science — full read; §10.1 and §10.9 are the mechanistic core and the most genuinely unresolved material in Part II. 🏥 Clinical — §10.2 through §10.6 are the indications your patients will ask about, in the order they became relevant. §10.7 and §10.8 are the ones they will have read about and that are not yet established.


10.1 Why one drug keeps working on unrelated diseases

Four explanations, not mutually exclusive, and distinguishing them is the chapter's underlying question.

Explanation 1: weight loss is upstream of many diseases.

Excess adiposity is associated with cardiovascular disease, chronic kidney disease, fatty liver disease, sleep apnea, osteoarthritis, several cancers, and more. A drug that produces substantial sustained weight loss should improve conditions that adiposity drives.

This is the least exciting explanation and possibly the correct one. It predicts that any intervention producing comparable weight loss should produce comparable benefit — diet, exercise, surgery — which is a testable and partially tested prediction.

Explanation 2: GLP-1 receptors are where nobody expected.

Chapter 7 established receptors in pancreas, stomach, and brain. They are also present in heart, blood vessels, kidney, and immune cells. A drug reaching all of them may be acting directly in each tissue, independently of weight.

Explanation 3: inflammation is a common thread.

Chronic low-grade inflammation is implicated in cardiovascular disease, kidney disease, liver disease, and neurodegeneration. GLP-1 receptor agonists reduce inflammatory markers. If they are anti-inflammatory in a meaningful sense, one mechanism could plausibly touch several diseases.

This is the most attractive explanation and it should be held loosely. Chapter 5 §5.6 warned that inflammatory markers are among the least reliable surrogates in medicine. "Reduces CRP" is not "treats inflammation," and "treats inflammation" is not "improves outcomes."

Explanation 4: some of these will not replicate.

Not every promising finding survives. Some of the indications in this chapter will look considerably less impressive in a decade. Chapter 9's Phase 2 discipline applies here too, and the correct prior for an early-stage indication in a fashionable drug class is caution.

FOUR EXPLANATIONS — and what would distinguish them

  ① WEIGHT LOSS IS UPSTREAM
     TEST: does an equivalent amount of weight loss by another method produce
     the same benefit? Partially testable; partially tested; not settled.

  ② DIRECT TISSUE EFFECTS
     TEST: does benefit appear before meaningful weight loss, or in people who
     lose little weight? Some evidence points this way in cardiovascular data.

  ③ ANTI-INFLAMMATORY ACTION
     TEST: does the benefit track inflammatory markers better than it tracks
     weight? Requires mediation analysis, which is observational and weak.

  ④ SOME WON'T REPLICATE
     TEST: time. And larger, independent, longer trials.

  These are not mutually exclusive, and the honest position as of this writing
  is that all four are probably operating to some degree, in proportions nobody
  can specify.

10.2 Cardiovascular: from SUSTAIN 6 to SELECT

Covered in Chapter 8 §8.5 and summarized here because it is the anchor for everything else.

SUSTAIN 6 — semaglutide in type 2 diabetes at high cardiovascular risk, roughly two years: a statistically significant reduction in major adverse cardiovascular events, in a trial designed to rule out harm.

SELECT — semaglutide 2.4 mg in adults with established cardiovascular disease and overweight or obesity, without diabetes, roughly 17,000 participants over about three years: a 20% relative reduction in MACE, hazard ratio about 0.80; absolute roughly 8% to 6.5%, about 1.5 percentage points.

Rating: ✅ for that population. Not rated for primary prevention.

What makes this the anchor: it is the only indication in this chapter with a completed, large, hard-endpoint outcomes trial in a well-defined population. Every other claim in the chapter should be read against it — and most of them fall well short.


10.3 Kidney: FLOW, and what "stopped early" means

Chronic kidney disease is a major complication of type 2 diabetes and a leading cause of kidney failure. It progresses slowly, which makes it expensive to study.

FLOW examined semaglutide in adults with type 2 diabetes and chronic kidney disease, with a composite kidney endpoint — progression of kidney disease, kidney failure, and death from kidney or cardiovascular causes.

The trial was stopped early for efficacy.

🔬 Read the Study — FLOW and the early-stopping problem

text FIGURE 10.3 — "Good news that is harder to read" [real published trial] THE STUDY Randomized, double-blind, placebo-controlled. Semaglutide vs placebo in adults with type 2 diabetes and chronic kidney disease. Composite kidney endpoint. Manufacturer-sponsored. STOPPED EARLY on the recommendation of its independent data monitoring committee, for efficacy. THE QUESTION Does semaglutide slow kidney disease progression in this population? WHAT IT SHOWS Yes. A significant reduction in the composite kidney endpoint, sufficient for the monitoring committee to conclude that continuing was no longer justified. WHAT IT DOESN'T It does not tell you about people WITHOUT diabetes, about kidney disease from other causes, or about the effect over the full planned duration. And because it stopped early, the effect estimate is LESS PRECISE and MORE LIKELY TO BE OVERSTATED than it would have been at full duration. THE VERDICT ✅ for this claim, in this population. THE LESSON "Stopped early for efficacy" is genuinely good news AND a statistical complication. Trials are stopped when an interim analysis crosses a threshold — and interim analyses that cross thresholds are, on average, the ones where random variation happened to favor the treatment. Truncated trials tend to overestimate effects.

Why early stopping matters, stated carefully.

A trial with interim analyses stops when a pre-specified statistical boundary is crossed. That boundary exists for an ethical reason: if a treatment is clearly working, continuing to give placebo to half your participants becomes hard to justify.

And the trials that cross the boundary are enriched for chance-favorable results. This is the same selection logic as Chapter 9's Phase 2 problem. It does not mean the effect is not real. It means the point estimate is likely larger than the truth, and the confidence interval is wider than a completed trial's would have been.

The right reading: the direction is trustworthy, the magnitude should be held loosely, and this is a case where the ethics of the trial and the precision of its estimate are genuinely in tension.

📊 Evidence Rating — semaglutide for chronic kidney disease in type 2 diabetes

Claim: Semaglutide slows progression of chronic kidney disease in adults with type 2 diabetes and established kidney disease.

Rating:Strong clinical evidence for this population.

Why: A dedicated randomized outcomes trial with a composite kidney endpoint, stopped early for efficacy.

What would change it: a failure to replicate, or long-term follow-up showing the benefit does not persist. Note that the effect magnitude should be held loosely because of early stopping.

NOT covered: kidney disease without diabetes; kidney disease from other causes; primary prevention of kidney disease.


10.4 Liver: MASH and the biopsy endpoint problem

MASLD — metabolic dysfunction-associated steatotic liver disease — is fat accumulation in the liver associated with metabolic dysfunction. Its inflammatory, progressive form is MASH. (Both were until recently called NAFLD and NASH; the terminology changed to remove the alcohol-defined negative and to name the metabolic association directly. You will see both sets of terms.)

MASH can progress to fibrosis, cirrhosis, and liver failure. It is common, it is increasing, and until recently there was no approved pharmacotherapy.

GLP-1 receptor agonists improve liver fat and inflammatory markers. Trials have reported improvements in histological measures. This is a real and important result.

And the endpoint is genuinely problematic, in a way that deserves a section.

THE BIOPSY ENDPOINT PROBLEM

  WHY BIOPSY IS USED             WHY IT IS A BAD ENDPOINT
  ──────────────────             ────────────────────────
  It looks directly at the       · INVASIVE. A needle into the liver, with real
  tissue, rather than at a         if small risks. You cannot do it often.
  blood marker.                  · SAMPLING ERROR. A biopsy takes a tiny fraction
                                   of the liver. Disease is patchy. Two biopsies
  It is the historical             from the same liver can disagree.
  reference standard.            · READER VARIABILITY. Scoring is a pathologist's
                                   judgment, and pathologists disagree.
                                 · IT IS STILL A SURROGATE. Histology stands in for
                                   cirrhosis, liver failure, and death — which take
                                   many years to accrue.

  CONSEQUENCE: liver trials are expensive, slow, and noisy, and their primary
  endpoint is both invasive AND a surrogate. This is why the field has so few
  approved drugs, and why "improved histology" is a weaker result than it sounds.

What follows. A demonstrated histological improvement is a genuine finding and is not the same as demonstrated prevention of cirrhosis or death. The trials that would establish the latter take many years, and as of this writing they are not complete for this drug class.

📊 Evidence Rating — semaglutide for MASH

Claim: Semaglutide improves metabolic dysfunction-associated steatohepatitis.

Rating: ⚠️ Promising but preliminary, trending favorable.

Why: Randomized trials report improvement in histological measures — a real result in a disease with few options. But the endpoint is a biopsy-based surrogate with sampling error and reader variability, and hard outcomes (cirrhosis, liver failure, death) accrue over many years and have not been demonstrated.

What would change it: long-term outcome data showing reduced progression to cirrhosis or reduced liver-related mortality. That would move this to ✅. A failure of histological improvement to translate would move it toward ❌.


10.5 Heart failure with preserved ejection fraction

HFpEF is heart failure in which the heart's pumping fraction is normal but it fills poorly. It accounts for roughly half of heart failure cases, it is strongly associated with obesity, and it has historically been much harder to treat than heart failure with reduced ejection fraction.

Trials of semaglutide in people with HFpEF and obesity have reported improvements in symptoms, physical limitation, and exercise capacity, alongside weight loss.

Two readings, and they are not easy to separate.

The metabolic reading: the drug is treating an obesity-related cardiac phenotype — reducing adiposity, inflammation, and cardiac loading, and thereby improving the underlying condition.

The mechanical reading: the participants lost a great deal of weight, and people who weigh less can walk further and report less breathlessness. This would be a real benefit and a different claim.

Why it matters: if the improvement is largely mechanical, then the benefit should be reproducible by any means of weight loss and should not be described as a cardiac treatment. If it is metabolic, the drug is doing something to the heart.

📊 Evidence Rating — semaglutide for HFpEF with obesity

Claim: Semaglutide improves symptoms and physical function in adults with heart failure with preserved ejection fraction and obesity.

Rating: ⚠️ Promising but preliminary.

Why: Randomized trials report improvement in symptom and function scores — genuine, and in a condition with few effective options. But the primary endpoints are symptom and function measures rather than hard outcomes (hospitalization, death), and the mechanical contribution of weight loss is not separated from any direct cardiac effect.

What would change it: a trial with hospitalization or mortality as a pre-specified primary endpoint, and ideally a design that separates weight-mediated from direct effects.


10.6 Sleep apnea: the clearest case of the mechanical question

Obstructive sleep apnea is repeated collapse of the upper airway during sleep. It is strongly associated with obesity, and excess soft tissue around the airway is a well-established contributor.

Tirzepatide has been studied in adults with moderate-to-severe obstructive sleep apnea and obesity, with the apnea-hypopnea index — events per hour of sleep — as the endpoint. Substantial reductions were reported, and this supported an approval for the indication.

And here the mechanical explanation is not a confound; it is the leading hypothesis.

Less soft tissue around a collapsible airway means less collapse. That is a straightforward mechanism and it does not require any direct drug effect on the airway. The drug works by making the person smaller, and there is nothing wrong with that.

Why this case is useful. It is the cleanest example in the chapter of a benefit that is almost certainly weight-mediated. Compare it with the cardiovascular result, where the benefit appeared earlier and larger than weight loss alone comfortably explains. Two indications, two different answers to the "is it just the weight?" question, and being able to tell them apart is the skill.

📊 Evidence Rating — tirzepatide for obstructive sleep apnea in obesity

Claim: Tirzepatide reduces apnea-hypopnea index in adults with moderate-to-severe obstructive sleep apnea and obesity.

Rating:Strong clinical evidence for this claim and population.

Why: Randomized trials with a standard objective endpoint, supporting an approval.

What would change it: little for this claim. Note that the mechanism is probably mechanical (weight-mediated), which is a complete explanation rather than a criticism — and which predicts that comparable weight loss by other means should produce comparable benefit.


10.7 Brain: why Alzheimer's is 🔬 and not ⚠️

Trials of semaglutide in early Alzheimer's disease have been conducted. This has generated a great deal of coverage.

The rationale is genuinely reasonable:

  • Metabolic dysfunction and insulin resistance are associated with Alzheimer's risk, to the point that the phrase "type 3 diabetes" has circulated
  • GLP-1 receptors are present in the brain
  • Preclinical work reports neuroprotective effects in animal models
  • Epidemiological analyses have reported lower dementia incidence among people taking these drugs for diabetes

And every one of those is a reason to run a trial rather than a result.

The epidemiology is observational and subject to confounding by indication — people who receive newer drugs differ systematically from those who do not, in ways that predict outcomes. The animal work is animal work (Chapter 5 §5.3). The receptor presence is mechanism (Chapter 2 §2.9). And Alzheimer's is a field with an exceptionally long record of promising mechanisms that did not produce clinical benefit.

📊 Evidence Rating — semaglutide for Alzheimer's disease

Claim: Semaglutide slows cognitive decline in Alzheimer's disease.

Rating: 🔬 Frontier — too soon to rate.

Why: Trials have been conducted and the field is proceeding properly. As of this writing the supporting case rests on mechanism, animal work, and observational epidemiology subject to confounding by indication — none of which establishes clinical benefit. Alzheimer's has an unusually poor record of mechanistic promise translating.

What would change it: a completed, adequately powered randomized trial with a pre-specified cognitive or functional endpoint. A clear positive result would move this to ⚠️ or ✅ depending on magnitude and replication; a null result would move it to ❌.

Why 🔬 and not ⚠️: ⚠️ requires real human data that does not yet settle the question. 🔬 is for early-stage science proceeding properly where clinical translation is unproven. A program whose supporting evidence is preclinical and observational, however well-motivated, is 🔬.


10.8 Addiction, and the observation that started it

This one began with patients.

People taking GLP-1 receptor agonists reported, unprompted, that they were drinking less. Smoking less. Less interested in behaviors they had previously found difficult to moderate. Chapter 7 §7.7's "food noise" reports are the same phenomenon in the domain the drug was designed for.

The mechanism is plausible. GLP-1 receptors are present in brain regions involved in reward processing. A compound that quiets an intrusive appetitive loop for food is an obvious candidate for other appetitive loops.

And the evidence is early. Randomized trials in alcohol use disorder and in smoking cessation are ongoing as of this writing. Some smaller studies have reported signals; some have not.

What makes this case interesting is where the hypothesis came from. It did not come from a pharmaceutical target hypothesis. It came from patients noticing something and saying so, consistently, across many independent reports. That is a legitimate and valuable source of hypotheses — and Chapter 5's ladder places unblinded self-report near the bottom, which is also true.

Both things at once, exactly as with food noise.

📊 Evidence Rating — GLP-1 receptor agonists for addiction

Claim: GLP-1 receptor agonists reduce alcohol consumption, smoking, or other addictive behaviors.

Rating: 🔬 Frontier.

Why: A consistent, unprompted, mechanistically plausible patient-derived observation, with randomized trials ongoing as of this writing. The supporting evidence is currently observational and preclinical, and blinding is difficult because the drug produces noticeable effects.

What would change it: completed randomized trials with pre-specified consumption or abstinence endpoints, ideally with an active comparator producing similar side effects to preserve blinding.

Note the asymmetry in how this should be held: the observation is real and worth studying, and nothing about it currently supports prescribing for these indications.


10.9 Inflammation as the candidate common thread

If one mechanism explains several of these, inflammation is the leading candidate. It is worth stating both why and why to be careful.

The case for. Chronic low-grade inflammation is implicated in atherosclerosis, kidney disease progression, liver fibrosis, and neurodegeneration. GLP-1 receptors are present on immune cells. Inflammatory markers fall on treatment. And a separate line of cardiovascular research has established that targeting inflammation directly can reduce cardiovascular events — so the general idea that inflammation is causal rather than merely associated has independent support.

The case for caution.

Inflammatory markers are surrogates, and poor ones. Chapter 5 §5.6. Many things lower CRP without improving anything.

Weight loss itself reduces inflammation. So an anti-inflammatory effect does not distinguish explanation ③ from explanation ①; it may be downstream of it.

And "anti-inflammatory" is the vaguest mechanism claim in medicine. Chapter 18 will make this point about "immune modulation," and it applies here. A mechanism that can explain benefit in any disease explains benefit in none, until someone specifies which inflammatory process, in which tissue, measured how.

The honest position: inflammation is a plausible partial explanation, it is under active investigation, and treating it as established would be exactly the error this book exists to prevent.

A test that would help. If inflammation mediates the cardiovascular benefit, then the benefit should be larger in people whose baseline inflammatory markers are higher — because they have more of the thing being treated. That is a pre-specifiable subgroup analysis, it is checkable in existing trial data, and it would not be conclusive but would be genuinely informative. Notice that this is a better test than "do markers fall on treatment?", which everything that reduces weight does. The useful question is never whether a proposed mediator moves; it is whether the benefit tracks it.


10.9a The indications you have not heard about

A chapter listing successes owes an account of the rest, because the list of indications being tested is much longer than the list of indications that have worked — and the difference is invisible in coverage.

Search a trial registry for any widely used drug and you will find trials in conditions nobody associates with it. Most will be small. Many will be investigator-initiated rather than company-sponsored. A substantial fraction will complete and never publish, and a further fraction will publish null results that no one covers.

For this drug class, trials have been registered or reported across a range including polycystic ovary syndrome, osteoarthritis, psoriasis, certain cancers, infertility, and several psychiatric conditions. Some of these will produce real results. Most will not. That is the normal distribution of outcomes for a fashionable drug class, and it has been the normal distribution for every fashionable drug class.

Three things follow.

The visible indication list is filtered. You hear about the ones that worked and the ones that generate compelling headlines. This is Chapter 6 §6.3's survivorship problem operating on research programs rather than testimonials — and it means the apparent hit rate of this drug class is substantially higher than its actual hit rate.

Absence of coverage is not absence of result. A trial that completed two years ago with a null finding and no publication is invisible in exactly the way Chapter 5's Case Study 2 described. Checking a registry for completed trials without posted results is the cheapest available correction, and it is a check almost nobody performs.

And the base rate should inform your priors on the frontier claims. Alzheimer's and addiction are 🔬 partly because of what their own evidence currently is, and partly because the historical base rate for "fashionable drug class turns out to also treat a hard neurological or psychiatric condition" is poor. Base-rate reasoning is legitimate here and it is not the same as dismissal. It sets a prior; a completed trial updates it.

🔍 Check Your Understanding

  1. Why is the visible list of indications for a drug class systematically more favorable than the full list of what has been tried?
  2. What single registry check would partly correct for this, and why does almost nobody do it?
  3. When is base-rate reasoning legitimate, and when does it become an excuse not to look at evidence?

10.10 One molecule, several ratings — the discipline demonstrated

Assemble the chapter.

SEMAGLUTIDE AND TIRZEPATIDE — ONE CLASS, SEVEN CLAIMS, FIVE RATINGS

  CLAIM                                          POPULATION                 RATING
  ─────────────────────────────────────────────────────────────────────────────────
  Weight loss                                    obesity, no diabetes         ✅
  Glycemic control                               type 2 diabetes              ✅
  Cardiovascular event reduction                 established CVD +            ✅
                                                 overweight, no diabetes
  Cardiovascular — PRIMARY PREVENTION            no established CVD       NOT RATED
  Chronic kidney disease progression             T2D + CKD                    ✅
  Obstructive sleep apnea (tirzepatide)          moderate-severe OSA +        ✅
                                                 obesity
  MASH                                           MASH                         ⚠️
  HFpEF symptoms and function                    HFpEF + obesity              ⚠️
  Alzheimer's disease                            early AD                     🔬
  Addiction / substance use                      various                      🔬
  ─────────────────────────────────────────────────────────────────────────────────

  A SINGLE OVERALL RATING FOR THIS CLASS WOULD HAVE TO BE WRONG ABOUT AT LEAST
  FOUR OF THESE. That is why rule 6 exists.

And notice the structure of the errors this prevents.

Upward compression — treating the ✅ indications as licensing the 🔬 ones. "It's proven to reduce heart attacks, so the Alzheimer's thing is probably right too." No: those claims rest on completely different evidence.

Downward compression — treating the 🔬 indications as discrediting the ✅ ones. "They're claiming it cures everything, so I don't believe the heart result either." Also no.

Both errors are common, they run in opposite directions, and they have the same cause, which is attaching a rating to a molecule rather than to a claim.

⚠️ Hype Check — "it might be a wonder drug"

"They're finding it helps with heart disease, kidney disease, liver disease, sleep apnea, Alzheimer's, and addiction. It might be one of those once-in-a-generation drugs."

What's true: several of those are supported by completed randomized trials with real endpoints. The cardiovascular, kidney, and sleep apnea results are genuine, and it is entirely reasonable to be impressed.

What the sentence does: it puts all six in one list, which implies they share an evidentiary status. They do not. Two have completed hard-endpoint outcome trials; two have surrogate or symptom endpoints; two have no completed randomized efficacy trials at all.

The diagnostic question: for which of those, in which population, on what endpoint, and at what stage? A source that can answer for each is worth reading. A source that lists them as a set has told you it is not distinguishing.

And the historical note: "works on everything" is the shape of a claim that usually turns out to be wrong. In this case some of it is right, which is unusual — and the unusualness is a reason to be careful rather than a reason to relax.

What a reader should actually do with this

A ten-row table is the correct level of granularity for a book and the wrong level for a conversation. So here is the compression that is legitimate.

If you or someone you know is considering this class of drug, the ten rows collapse to three questions, and only one of them is about the table:

Which claim applies to you? Not all ten. Usually one, occasionally two. The rating for that row is the one that matters, and the other nine are context rather than evidence about your situation.

Is your situation the studied population? This is the question that most often changes the answer. Every ✅ in the table carries a population — established cardiovascular disease, type 2 diabetes with kidney disease, moderate-to-severe sleep apnea with obesity. A person outside the studied population is not covered by the rating, and no amount of enthusiasm about the drug's breadth changes that.

And what would your clinician add? Baseline assessment, interaction checking, monitoring, and the judgment about whether an alternative fits better. Chapter 39 is entirely about making that conversation productive, and it is worth reading before the appointment rather than after.

The compression that is not legitimate is the one this section exists to prevent: taking the ✅ rows as evidence about the ⚠️ and 🔬 rows, or taking the 🔬 rows as evidence against the ✅ rows. Those are the two errors named above, and they are the reason the table has ten rows instead of one.

A useful habit for anything you read about this class: when a source mentions a new indication, ask which row it belongs in. If the source cannot tell you the population, the endpoint, and whether the evidence is randomized, it has not distinguished the rows either — and you are reading a summary of enthusiasm rather than a summary of evidence.


📋 Your Evidence Dossier

This chapter teaches Field 6's most important discipline: split the rating by indication.

FIELD 6, SPLIT BY INDICATION

  For each peptide, list EVERY claim you have encountered about it, then rate each
  separately:

    CLAIM                     POPULATION        ENDPOINT       RATING    FALSIFIER
    ......................    ..............    ...........    ......    .........
    ......................    ..............    ...........    ......    .........
    ......................    ..............    ...........    ......    .........

  THE TEST: if every row has the same rating, you have probably not looked hard
  enough — or the compound genuinely has only one claim, which is worth noting.

Worked demonstration — semaglutide, six ways

See §10.10's table. The point of reproducing it as a dossier entry is that each row has its own falsifier, and they are different:

Claim Falsifier
Weight loss loss of effect during continued treatment
Cardiovascular (secondary prevention) failure to replicate; reversal on longer follow-up
Kidney failure to replicate
MASH histological improvement failing to translate to hard outcomes
Alzheimer's a completed adequately powered trial reporting null
Addiction completed randomized trials reporting null

Notice that two of these falsifiers are "a trial reports null" and four are "a completed result fails to hold." That difference — between claims awaiting a first test and claims that have passed one — is what separates 🔬 from ✅, and writing the falsifiers out makes it visible.

Your task

For every peptide in your dossier, list every distinct claim you have encountered and rate each separately.

Expect this to be uncomfortable for compounds you are invested in, because the exercise usually reveals that a compound you think of as "well-supported" has one supported claim and four unsupported ones travelling under its reputation.

And check for upward compression in your own entries. If you rated a compound ✅ for one indication and then wrote favorably about another indication without a separate rating, you have done the thing this chapter is about.


Conclusion

A drug developed to lower blood glucose has produced completed randomized evidence in cardiovascular disease, chronic kidney disease, and obstructive sleep apnea; real but surrogate-endpoint evidence in liver disease and heart failure with preserved ejection fraction; and ongoing early-stage programs in Alzheimer's disease and addiction.

Four candidate explanations: weight loss is upstream of many diseases; GLP-1 receptors are present in unexpected tissues; inflammation is a common thread; and some of these will not replicate. All four are probably operating, in proportions nobody can specify.

FLOW demonstrates that "stopped early for efficacy" is good news and a statistical complication: trials that cross an interim boundary are enriched for chance-favorable results, so the direction is trustworthy and the magnitude should be held loosely.

MASH demonstrates the biopsy endpoint problem — invasive, subject to sampling error and reader variability, and still a surrogate for outcomes that take years.

Sleep apnea is the chapter's clearest weight-mediated benefit, and that is a complete explanation rather than a criticism. Cardiovascular is the clearest case where weight loss alone does not comfortably explain the result.

Alzheimer's and addiction are 🔬, and the reason matters: ⚠️ requires real human data that does not settle the question, while 🔬 is for early science proceeding properly. Mechanism, animal work, and observational epidemiology — however well-motivated — do not clear that bar.

And the discipline is one rating per claim. A single verdict on this drug class would have to be wrong about at least four of the ten claims in §10.10's table. Upward compression treats the ✅s as licensing the 🔬s; downward compression treats the 🔬s as discrediting the ✅s. Both are the same error.

And one closing observation about how this chapter will age. Several of these ratings will move within a few years, in both directions — a 🔬 will become a ⚠️ or a ❌ when a trial reports, and a ⚠️ may become a ✅ when an outcome study completes. That is the system working, not failing. A rating that could not move would not have been a rating.

What will not age is the structure: one claim per row, a population attached, an endpoint named, a falsifier written down, and the two compression errors held off in both directions.

Chapter 11 goes back a century, to the drug against which every claim in this chapter should be measured.


Key Terms

Pleiotropic effect — an effect of a drug or gene on multiple apparently unrelated systems.

Indication expansion — the process by which a drug approved for one use accumulates evidence and approvals for others.

FLOW — the randomized outcomes trial of semaglutide in type 2 diabetes with chronic kidney disease, stopped early for efficacy.

MASLD / MASH — metabolic dysfunction-associated steatotic liver disease and steatohepatitis; formerly NAFLD and NASH.

Biopsy endpoint — an endpoint assessed by tissue sampling; invasive, subject to sampling error and reader variability, and generally still a surrogate.

Albuminuria — protein in the urine; a marker of kidney damage.

eGFR — estimated glomerular filtration rate; the standard measure of kidney function.

HFpEF — heart failure with preserved ejection fraction.

Obstructive sleep apnea — repeated upper airway collapse during sleep.

Apnea-hypopnea index — events of apnea or hypopnea per hour of sleep; the standard objective measure of sleep apnea severity.

Stopped early for efficacy — termination of a trial when an interim analysis crosses a pre-specified benefit boundary; ethically motivated, and associated with overestimated effect sizes.

Confounding by indication — the bias arising when people who receive a treatment differ systematically from those who do not, in ways related to the outcome.


Spaced Review

  1. (Ch 5) FLOW was stopped early for efficacy. Explain why this is good news and a statistical complication, and state how it should change your reading of the effect size.

  2. (Ch 5) The MASH trials used a biopsy endpoint. List four distinct problems with that endpoint and state which one makes it a surrogate.

  3. Sleep apnea benefit is probably weight-mediated; cardiovascular benefit probably is not entirely. Explain what evidence distinguishes them, and why "it's just the weight loss" is a complete explanation rather than a dismissal in one case.

  4. (Ch 2) Explain why "anti-inflammatory" is a weak mechanistic claim, and what would have to be specified to make it a strong one.

  5. Explain to a friend why a drug being proven for heart disease tells you nothing about whether it works for Alzheimer's — without implying the Alzheimer's research is illegitimate.