Chapter 36 — Case Study 2: The Muscle Question, and the Trial That Would Actually Settle It

What this case study is, and is not. It is not a report of a trial. As of this writing, in 2026, §36.6 rates pharmacological prevention of lean mass loss 🔬 precisely because no completed evidence on functional endpoints exists. So this is a design exercise instead: given the question, what would have to be measured, in whom, against what, before an answer meant anything? Where a trial result appears below it is labeled [constructed teaching example] and describes no real study. Nothing here is medical advice, and nothing here describes how to obtain or use any compound.


1. The question, stated precisely enough to be gradeable

Most public versions of the muscle question are not questions. "Are these drugs making people lose muscle?" has no population, no endpoint, no comparator, and no way to be wrong — which is exactly the defect §36.2's rating of a claim form was about. Run the six questions on it and every one comes back empty.

Here is a version that survives:

Does adding a pharmacological agent that preserves lean mass during substantial weight loss on GLP-1-based therapy produce better function — strength, mobility, falls, independence — than weight loss without it, in the people in whom the concern actually concentrates?

That sentence names a population, an endpoint, a comparator, and a direction. It can be wrong. It is the claim §36.6 rates 🔬 — and almost none of the coverage you will encounter is about it. Most of it is about a different claim, which is cheaper to answer and does not mean what it appears to mean.

2. What is already established, and should not be re-argued

Start where the argument is over, because a great deal of coverage spends its energy re-establishing this and then stops.

Substantial weight loss from any cause includes loss of lean mass, not only fat mass. Dietary restriction does it. Bariatric surgery does it. Illness does it. Pharmacological weight loss does it. This has been documented across every weight-loss modality studied, for decades. It is not a novel finding, it is not specific to GLP-1 drugs, and coverage presenting it as a newly discovered hazard of this drug class is getting the history wrong.

The proportion of lost mass that is lean varies by study, by measurement method, by rate of loss, and by what the person was doing with their body during it. That variation is part of why the question is hard, and it is also why any single quoted percentage should be read as one study's number rather than a fact about the drug class.

So the open question is not whether lean mass declines. It is what that decline is made of, whether it costs anyone anything, and whether a second drug should be given to prevent it.

3. Why a body-composition scan is not the measurement that answers it

A DEXA reading is a surrogate (Chapter 16). This is the single most important sentence in the section, and it survives every technical objection anyone raises to it.

Two separate problems sit underneath it, and they compound.

The first is what "lean mass" contains. Lean mass as reported by body composition methods is not contractile skeletal muscle. It is a compartment defined by what it is not — not fat, not bone mineral — and it therefore includes fluid, connective tissue, organ mass, and glycogen with the water that travels with it. Someone shedding a large amount of weight also sheds part of the structural and fluid load that supported a larger body. It is not obvious that losing that is harmful; some of it may be appropriate remodeling. Distinguishing "lost muscle you needed" from "lost supporting tissue you no longer need" is technically difficult, and a standard scan does not attempt it.

The second is the inferential chain. Even granting that the scan is measuring muscle, getting from there to something a person would notice requires several assumptions in series:

THE CHAIN A DEXA RESULT IS ASKED TO CARRY

   preserved lean mass    →   preserved strength    →   preserved mobility
   on a scan                  and power                 (gait speed, stairs,
                                                         rising from a chair)
                                            ↓
                          fewer falls, fewer fractures, retained independence

   Each arrow is an assumption until it is measured. A trial that measures only
   the leftmost box has established the leftmost box.

Each arrow can fail. Strength does not track mass linearly, particularly in older adults, where the neural and qualitative components of force production matter as much as cross-sectional area. Mobility depends on balance, joint health, and cardiopulmonary reserve as well as strength. Falls are multifactorial in a way that makes them notoriously resistant to single-mechanism interventions.

A trial can demonstrate convincingly that a myostatin-pathway agent preserves lean mass during weight loss, that result can be entirely real, and it can still fail to mean anything for the people taking it.

4. The wrinkle that makes this surrogate worse than most

There is a further problem here that does not apply to most surrogates, and it deserves to be stated on its own.

An agent that adds lean mass to the scan has, mechanically, improved the thing being measured.

Most surrogates are indirect: a drug lowers a marker, and you hope the marker's movement reflects something downstream. Here the intervention acts directly on the yardstick. Blocking signaling through the activin type II receptors increases lean mass — that is the mechanism, it is not in dispute, and it is why these agents are candidates in the first place. So a positive scan result is close to guaranteed by the pharmacology. It confirms that the drug does what it was designed to do. It carries almost no information about whether doing that helps.

A result that the mechanism all but promises in advance is a weak test of the mechanism. A drug that raises the number you are using to judge it demands more care than usual, not less — and in practice attracts less, because a large, highly significant effect on a familiar-sounding measure reads as strong evidence to almost every audience.

5. What the trial would have to measure

So: function, not composition. Concretely, and in rough order of how hard each is to collect and how much each is worth:

Measured strength. Grip strength and knee extensor strength are cheap, standardized, and reproducible. They are not the outcome anyone cares about, but they are one link closer than a scan and they are objective.

Mobility performance. Gait speed over a fixed distance, chair-rise (sit-to-stand) time, a timed up-and-go, stair ascent. These measure the composite thing that strength contributes to. Gait speed in particular is a well-established prognostic measure in older adults, which is why it keeps appearing in the chapter's list.

Falls, ascertained prospectively. Not recalled at the end of the study — prospectively, with a diary or scheduled contact, because retrospective fall reporting is unreliable. Falls are a hard outcome. They are also rare enough that counting them adequately drives the trial's size and duration up sharply, which is exactly why they are the endpoint most often dropped.

Independence. Ability to live alone, activities of daily living, disability-free survival, institutionalization. This is the endpoint that would settle the question and the one least likely to be attempted, because it requires years.

Note what this list does to the trial. Every step rightward makes the study longer, larger, more expensive, and slower to report — the same asymmetry §36.5 describes for weight loss versus cardiovascular events. The composition trial can be run quickly and the function trial cannot, which guarantees that the composition result arrives first and dominates coverage for years. That is a structural feature of the field, not a failure of any particular sponsor, and knowing it in advance is most of what protects you from the resulting headlines.

6. In whom, and against what

Two design choices decide whether the trial answers the question or merely produces data.

The population has to include the people the concern is about. §36.6 locates the concern most sharply in older adults: less reserve, sarcopenia already a clinical consideration, and — the part that makes this genuinely hard — a group in whom substantial weight loss also has documented benefits. Both things are true at once. A trial enrolling healthy middle-aged adults with good baseline function can show a preserved scan reading and cannot show a preserved capability, because there was no functional decline available to prevent. Question 3 in its most consequential form: were older adults enrolled in meaningful numbers, or excluded by the entry criteria that make a trial easier to run?

The comparator has to be the thing that is already available. This is the design decision most likely to be made badly, and the one to look for first. Against placebo, an agent that preserves lean mass will very likely win on the scan. But the uncontroversial intervention already exists: resistance training and adequate protein intake are recommended during weight loss by essentially every clinical source, carry no pharmacological risk, and cost nothing but effort. Chapter 28's sacubitril/valsartan result carries the weight it does because the comparator was an active treatment already known to work. The equivalent here is a training-plus-protein arm. A trial without one has answered "is this better than nothing?" — which is not the question anyone with the option of doing the training is asking.

There is an honest complication, and §36.6 flags it: whether training and protein work as well during pharmacologically driven weight loss, when appetite suppression may make adequate protein intake harder, is less established than the confident version of the advice suggests. That is an argument for putting the arm in the trial, not for leaving it out.

7. Adding a second drug is a new intervention, not a correction

The framing that will accompany these agents is that they fix a side effect. Resist it.

Adding a second drug to manage the effects of a first is a claim requiring its own outcome evidence. It has its own risk profile, its own tolerability burden, its own cost, its own interactions, and its own potential to harm someone who would have been fine. That it was added to address a problem caused by something else earns it exactly zero credit. It clears the same bar as any other intervention or it does not clear a bar at all.

Two questions follow that a composition trial will not answer. How are adverse effects in the people who experience them weighed against a benefit measured on a scan? And what is the exposure profile if the agent is long-acting — §36.4's counterweight applies here too, because you cannot titrate or stop what has already been given.

8. The detail that belongs in a peptide book

Several of the leading candidate agents are antibodies rather than peptides.

Bimagrumab — an antibody against the activin type II receptors — is among the compounds studied in this context. Read the name with Chapter 1 §1.8 in hand: the -mab stem means monoclonal antibody. Roughly thirty times the mass of a peptide. Produced in living cell culture rather than by chemical synthesis, with the cost structure that follows. And a long half-life, which is not a neutral fact in a trial where reversibility matters.

Two things follow. The first is practical: none of §36.9's peptide manufacturing arithmetic applies to these candidates, and none of Chapter 32's synthesis constraints govern their supply.

The second is a caution about this book. The interesting frontier is not always peptide-shaped. §36.3's orforglipron was the other example — a small molecule winning a race that the peptide framing would have predicted a peptide to win. A text organized around peptides will systematically under-notice the antibody and small-molecule answers to the same problems, and a reader who has learned to decode drug names can catch that from six letters.

9. Where this leaves a reader today

The rating does not move, and the honest position is narrower than either the alarmed or the reassuring version.

Established: substantial weight loss of any kind includes lean mass; this is true of every modality studied; and the benefits of weight loss in people with obesity-related disease are supported by hard outcome data (Chapter 10).

Not established: that lean mass lost on these drugs is functionally worse than lean mass lost by other means, that it produces meaningful impairment in most people, or that any drug should be added to prevent it.

Available now and uncontroversial: resistance training and adequate protein intake during weight loss.

Worth raising with a clinician, particularly for older adults or anyone with existing mobility concerns: whether functional status should be tracked, and what tracking it would look like. Grip strength and gait speed are cheap, quick, and measure the thing that matters. A scan measures the surrogate. Chapter 39 is about having that conversation well.

And the caveat the chapter keeps repeating, because this is the section where readers are most likely to forget it: 🔬 means a serious question is being asked seriously. It does not mean coming soon, and it carries no probability of success. Most 🔬 becomes ❌.


Discussion Questions

  1. Write the muscle question in a form that could be false, then write it in the form you have most often encountered it. Identify precisely which of the six questions the second version fails to answer, and say what the failure lets a writer get away with.

  2. A [constructed teaching example] trial reports that adding a myostatin-pathway agent to a GLP-1 drug preserved lean mass on DEXA, with a large and highly significant effect versus placebo. State what has been established. Then list everything a reader might reasonably but wrongly conclude, and name the assumption behind each wrong conclusion.

  3. This case study argues that an agent acting directly on the measurement used to judge it is a worse-than-usual surrogate. Construct the strongest counterargument — that acting on the mechanism is exactly what you would want — and then say where it fails.

  4. Rank the four functional endpoint families (strength, mobility performance, falls, independence) by how much each would tell you and, separately, by how hard each is to collect. Explain what the disagreement between the two rankings predicts about which results will be published first and how they will be covered.

  5. Suppose a sponsor argues that a resistance-training-plus-protein comparator arm is impractical: adherence cannot be verified, and the arm would delay the trial by years. Evaluate the argument. Is it a reason to omit the arm, a reason to design it differently, or a reason to discount the trial's conclusion? Defend your answer.

  6. Several leading candidates in this area are antibodies rather than peptides. Write your own four-line 🔬 dossier entry for this question — rating, the readout that would move it, the date, and the "do not count" line — and make sure the readout you name is one an antibody trial could actually produce. Then say which kinds of news you already expect to see that must not move it.