Appendix D — Reading a Clinical Trial: A Field Guide

Chapter 5 gave you the method. This appendix gives you the mechanics: how to open an actual paper and get an actual answer out of it, in a reasonable amount of time, without a statistics degree.

Two things to say before starting. First, you do not need to understand everything in a trial report to evaluate it. A great deal of a modern paper is regulatory and technical apparatus. The things that determine whether the result means what the headline says are few, and they are almost always findable in under fifteen minutes. Second, you will not be able to evaluate everything. Some questions genuinely require expertise you do not have, and the correct response is to notice that and hold the conclusion loosely rather than to guess. Knowing which is which is most of the skill.


D.1 Read it in the wrong order

The single most useful habit in this appendix.

A paper is written in the order Abstract → Introduction → Methods → Results → Discussion. Read it in the order Methods → Results → Abstract → Discussion.

The reason is straightforward. The abstract and the discussion are where the authors tell you what they think their result means. They are interpretation, they are written to be quoted, and they are where the framing lives. The Methods and Results are the part that constrains what the result can mean. If you read the abstract first, you will read the methods looking for confirmation of a conclusion you have already absorbed — and you will find it, because that is how reading works.

Read the methods cold. Decide what result would be meaningful before you know what the result was. Then look. Then read what the authors say about it and notice any distance between their summary and yours.

That distance is the finding, more often than people expect.


D.2 The four questions, on the first pass

Everything in Chapter 5 collapses to four questions, and all four live in the Methods section. Write the answers down before going further.

1. WHO — the population. Not "adults with obesity." The actual entry criteria. Age range, BMI range, whether diabetes was an inclusion or an exclusion, prior treatment, comorbidities, geography. Then ask the question that matters: is this population you, or the person you are asking about? Chapter 8's STEP 1 and STEP 2 differ in exactly one criterion — the presence of type 2 diabetes — and produced ~−15% and ~−10% weight loss respectively at the same dose. The entry criteria are not fine print. They are frequently the whole result.

2. WHAT — the intervention. Dose, route, schedule, duration, and what else participants received. A trial where everyone also got structured lifestyle counseling is answering a different question from one where they did not.

3. AGAINST WHAT — the comparator. Placebo, an active drug, standard care, or nothing. A placebo-controlled result tells you the drug beats nothing. It does not tell you the drug beats what people are already taking. Chapter 28's PARADIGM-HF is the model of the stronger design: the comparator was an active drug already proven to reduce mortality, which is why the result changed guidelines.

4. MEASURED HOW — the endpoint. The primary endpoint, singular, and whether it is a surrogate or an outcome. Chapter 16 is entirely about this distinction and Chapter 28 supplied the demonstration: nesiritide improved the hemodynamic measures it was approved on and did not reduce death or rehospitalization when tested.

If you stop after these four, you will already read trials better than most coverage of them does.


D.3 Design: what protects the result from being an artifact

Randomization assigns participants to groups by chance, which is the only method that balances unknown confounders as well as known ones. This is why Chapter 5 rates randomized evidence above observational evidence for questions about benefit. Look for how randomization was done — a computer-generated sequence is standard; anything resembling alternate assignment or date-of-birth is not randomization.

Allocation concealment means the person enrolling a participant could not know which group they would land in. It is distinct from blinding and it is easy to overlook. Without it, randomization can be subverted without anyone intending to — an investigator who suspects the next slot is placebo may unconsciously delay enrolling a sicker patient.

Blinding. Who did not know the assignment: participants, treating clinicians, outcome assessors, statisticians. Blinding matters most for subjective endpoints and least for hard ones. Death is hard to misclassify; pain scores, global impression scales, and symptom questionnaires are extremely sensitive to expectation.

A specific and important caveat for this book's subject matter: the GLP-1 drugs are difficult to blind in practice, because their gastrointestinal effects are noticeable enough that many participants can guess their assignment. This does not invalidate the trials — the hard endpoints in SELECT are not the kind of thing expectation produces — but it is a real limitation, and a source that never mentions it is not reading carefully.

Prespecification. Was the primary endpoint declared before the data were seen? This is what trial registration exists to establish, and it is checkable: registry entries carry timestamps and revision histories. An endpoint chosen after looking at the data is not a test; it is a description.


D.4 Analysis: where results quietly change size

Intention-to-treat versus per-protocol. ITT analyzes participants in the group they were assigned to, regardless of what they actually did — including those who stopped the drug. Per-protocol analyzes only those who completed as intended. ITT is the conservative and generally correct choice, because dropping non-completers destroys the randomization: people who stop taking a drug are systematically different from people who continue, usually because it was not working or was causing side effects.

The estimand question, which is the sophisticated version of the same issue. Modern trials state explicitly what question they are answering, and two common choices give different numbers from identical data:

  • A treatment-regimen estimand asks: what happened to people assigned this treatment, including those who discontinued? This is closer to what happens in the real world.
  • An efficacy estimand asks: what happened to people who actually took the treatment as intended? This is closer to what the drug does when taken.

SURMOUNT-1 is the worked example this book uses. At the 15 mg dose over 72 weeks, the reported weight change was approximately −21% on the treatment-regimen estimand and approximately −22.5% on the efficacy estimand. Both figures are correct. Both describe the same trial. They answer different questions, and a source that quotes one without saying which is not necessarily being dishonest — but a source that quotes the larger number from one trial and the smaller from a competitor's is doing something specific.

Composite endpoints. "MACE" — major adverse cardiovascular events — typically bundles cardiovascular death, non-fatal myocardial infarction, and non-fatal stroke. Composites increase statistical power by counting more events, which is legitimate. The thing to check is which component drove the result. A composite that improves entirely because of its softest component, while mortality is unchanged, is a much weaker finding than the headline suggests. Look for the component breakdown; good papers report it.

Subgroup analyses. Treat with deep suspicion. In any trial with enough subgroups, some will show apparently striking effects by chance alone. Unless a subgroup was prespecified with a stated hypothesis and the analysis tested for interaction, a subgroup finding is a hypothesis for a future trial, not a result. This is the most common way a failed trial is presented as a success.

Missing data and dropouts. How many left, from which arm, and why. Differential dropout between arms is a warning sign. Ask how missing data were handled — the assumptions involved can move a result more than most readers imagine.


D.5 Reading the numbers

Absolute versus relative, which is the most consequential distinction in this appendix. Chapter 10's SELECT result, stated three ways:

Statement Figure
Relative risk reduction ~20% (hazard ratio ~0.80)
Absolute event rates ~8% → ~6.5%
Absolute risk reduction ~1.5 percentage points
Number needed to treat ~65–70 over about three years

All four describe the same finding in roughly 17,000 participants. The first is the one that gets quoted. The third and fourth are the ones a person deciding whether to take a drug actually needs. Neither is wrong and the difference in rhetorical force is enormous — which is why a source's consistent choice between them tells you what the source is for.

Number needed to treat is the most intuitive form: roughly how many people must be treated for about three years to prevent one event. It is the reciprocal of the absolute risk reduction. Always attach the time horizon — an NNT without a duration is meaningless, because treating for longer prevents more events.

Confidence intervals, which are more informative than p-values. A 95% confidence interval gives a range of effect sizes compatible with the data. Read it for two things: whether it crosses the line of no effect (1.0 for a ratio, 0 for a difference), and how wide it is. A hazard ratio of 0.80 with an interval of 0.73–0.87 is a precise estimate of a real effect. A hazard ratio of 0.80 with an interval of 0.45–1.42 is barely an estimate at all, and reporting it as "a 20% reduction" is reporting a point estimate the data do not support.

P-values, and what they do not mean. A p-value is the probability of observing data at least this extreme if there were no true effect. It is not the probability that the treatment works, it is not a measure of effect size, and 0.049 is not meaningfully different from 0.051. A statistically significant result can be clinically trivial, particularly in a very large trial, where tiny differences reach significance easily. Always look at the effect size next to the p-value.

Power. A trial too small to detect a real effect will produce a negative result regardless of whether the drug works. "No significant difference" in an underpowered trial means "we could not tell," not "there is no effect." This is a different statement from Chapter 28's nesiritide result, where a trial of roughly 7,100 patients was large enough that the negative finding is genuinely informative. The distinction between "tested and failed" and "not adequately tested" runs through this entire book, and trial size is where you check which one you are looking at.

Non-inferiority trials ask whether a new treatment is not worse than an existing one by more than a prespecified margin — usually because it has some other advantage. The margin is the whole design, and it is chosen by the investigators. A generous margin can make an inferior drug look acceptable. Check what the margin was and whether it was justified.


D.6 The things outside the paper

Registration. Look up the trial in a registry — ClinicalTrials.gov, the EU register, ISRCTN, or a national equivalent. Compare the registered primary endpoint with the published one. A change between registration and publication, without explanation, is among the most informative things you can find, and it takes about three minutes.

Funding and conflicts. Almost every large trial of an approved drug is industry-funded, because almost nobody else can pay for one. This is a reason to read carefully, not a reason to dismiss — Chapter 42 rated this claim form, and the conclusion was that a conflict of interest is grounds to verify rather than evidence of falsehood. Note who designed the trial, who held the data, who did the analysis, and who wrote the manuscript. Independent data monitoring and academic statistical oversight are meaningful safeguards.

Publication status. A press release is not a paper. A conference abstract is not a paper. A preprint is a paper that has not been reviewed. Topline results announced by a sponsor are a selected summary of an unreleased dataset, and Chapter 36's pipeline questions apply directly.

Reporting standards. Well-reported trials follow CONSORT and include a flow diagram accounting for every participant from screening to analysis. The flow diagram is worth thirty seconds of your time — it shows you dropouts, exclusions, and how many people the final numbers actually rest on.


D.7 Systematic reviews and meta-analyses

A meta-analysis pools results from multiple trials. Done well, it is the strongest form of evidence available. Done badly, it launders weak studies into a confident-looking summary number.

Four things to check:

Inclusion criteria. What was included and excluded, and was that decided in advance? A review that includes only positive trials will find a positive effect.

Heterogeneity. Were the pooled trials answering the same question in comparable populations? Pooling genuinely different studies produces a number that describes nothing in particular. Reviews report heterogeneity statistics; high heterogeneity is a reason to distrust the pooled estimate.

Publication bias. Positive results are more likely to be published than negative ones, so the published literature is a biased sample of the conducted literature. Good reviews test for this. This is the mechanism behind Chapter 6's selection argument, operating on the scientific record itself rather than on testimonials.

Quality of the included trials. A meta-analysis of twelve small unblinded studies is not equivalent to one large blinded trial, however impressive the combined participant count looks. Pooling does not fix design.


D.8 The fifteen-minute triage

When you need an answer rather than a full evaluation:

  1. Registry entry — does the registered primary endpoint match the published one? (3 min)
  2. Methods, four questions — who, what, against what, measured how. (4 min)
  3. Primary endpoint result with its confidence interval. Ignore the abstract's adjectives. (2 min)
  4. Absolute numbers. Convert any relative figure into absolute terms and an NNT if you can. (2 min)
  5. Flow diagram — how many enrolled, how many analyzed, dropout by arm. (1 min)
  6. Funding and conflicts. (1 min)
  7. Ask the one question the whole appendix builds to: does this trial support the claim I heard, or a narrower one? (2 min)

In practice, step 7 answers the question most of the time. The trial is usually real, competently run, and honestly reported. What has gone wrong happened afterward — in the press release, the article, the video, the thread — where a result about a specific population, a specific endpoint, and a specific comparator was restated without any of the three.


D.9 What this appendix cannot teach you

Some judgments require domain expertise: whether an endpoint is clinically meaningful in a particular disease, whether an entry criterion excludes the patients who matter most, whether a statistical method suits the data. Noticing that you have hit one of those is itself a skill, and the right response is to say the trial looks sound on the checkable dimensions and that you cannot assess the rest.

That sentence — this is as far as I can evaluate it — is not a failure. It is the difference between reading a paper and reading a headline about a paper, and the whole book has been an argument for being able to say it.


Related: Chapter 5 (the method) · Chapter 9 (Phase 2 optimism) · Chapter 10 (interim stopping, SELECT) · Chapter 16 (surrogate endpoints) · Chapter 28 (a negative outcome trial) · Chapter 36 (reading pipeline claims) · Appendix C (the dossier) · Appendix F (red flags) · Appendix H (worked evaluations)