Case Study 2 — The Dementia Signal in the Prescription Data
Why observational findings about new drugs are almost always too good
Type: Real, public, methodological · Tier 1 method facts, Tier 2 findings · Relevance: §10.7
Background: the observation
Large healthcare databases record what people were prescribed and what happened to them. Researchers have used such databases to ask whether people with type 2 diabetes who received GLP-1 receptor agonists went on to develop dementia at lower rates than those who received other diabetes drugs.
Several such analyses have reported lower dementia incidence in the GLP-1 group.
This is a real finding in real data, it was produced by competent investigators, and it is one of the reasons Alzheimer's trials were undertaken.
It is also almost exactly the shape of finding that has misled medicine repeatedly, and this case study is about why.
Confounding by indication, and its relatives
WHY PEOPLE WHO GET A NEW DRUG ARE DIFFERENT
① CONFOUNDING BY INDICATION
Clinicians choose treatments. A newer, more expensive drug goes
disproportionately to patients who are: younger, better insured, more
engaged with care, with fewer competing conditions, and — critically —
without the cognitive impairment that would make a complex regimen
inappropriate.
→ The group that received it was ALREADY less likely to be diagnosed
with dementia.
② HEALTHY ADHERER EFFECT
People who take medication consistently have better outcomes than people
who do not — INCLUDING when the medication is placebo. Adherence is a
marker for a great many health-relevant behaviors.
③ REVERSE CAUSATION / PROTOPATHIC BIAS
Early cognitive decline affects treatment decisions before it is
diagnosed. A patient beginning to struggle is less likely to be started
on a new injectable requiring self-administration.
→ The disease influenced the exposure, not only the reverse.
④ SURVEILLANCE AND DETECTION
Patients on newer therapies see clinicians more often. That could
INCREASE dementia diagnosis (more contact, more detection) or DECREASE
it (better management of vascular risk). The direction is unclear,
which is its own problem.
⑤ IMMORTAL TIME BIAS
Depending on how time is classified relative to prescription start,
periods during which a patient could not have had the outcome can be
misassigned to the treated group.
────────────────────────────────────────────────────────────────────────
STATISTICAL ADJUSTMENT HELPS WITH ① AND PARTLY WITH ④.
It cannot fix ② and ③, because you cannot adjust for what you did not
measure — and "the kind of person who gets started on a new injectable"
is not a measured variable.
The general problem, stated once: you can only adjust for confounders you thought to measure, and the most powerful confounder in this setting is a clinician's judgment about a patient, which is not in the dataset.
The historical base rate
This is the part that should govern how much weight the finding gets.
Observational analyses of drug benefits have a poor record. Repeatedly, a large observational literature has supported a benefit that a subsequent randomized trial did not confirm — and in several prominent cases the trial found harm where the observational data had found benefit. The pattern has recurred across hormone therapy, vitamin supplementation, and several drug classes.
The direction of the error is consistent: observational studies of drug benefits tend to overestimate them, because the people who receive treatments are systematically healthier in ways that are hard to measure.
Note the asymmetry. Observational data is much more reliable for detecting harms, particularly rare and distinctive ones, because nobody selects patients for a bad outcome. This is why pharmacovigilance works and why Chapter 5 §5.2 rated case reports as valuable for harm and near-useless for efficacy. The same dataset can be strong evidence about one thing and weak evidence about another.
⚠️ Hype Check — "studies show it lowers dementia risk"
What's true: analyses of large healthcare databases have reported lower dementia incidence among people prescribed these drugs. That is an accurate description of what was found.
What "studies show" conceals: the study design. The word "studies" flattens the distinction between a randomized trial and a database analysis, and that distinction is the entire question here.
What the finding does establish: that the association exists, and that a trial is worth running. That is not nothing — it is exactly how a research program should be generated.
What it does not establish: causation, direction, or magnitude. Five distinct biases push in the favorable direction, at least two of which cannot be adjusted away.
The honest sentence: "People prescribed these drugs have been observed to develop dementia at lower rates, but the people who get prescribed them differ from those who don't in ways that predict dementia independently — so trials are running, and until they report we don't know."
And note the reflex worth building: on meeting any claim that a drug prevents a disease, ask "randomized or observational?" before anything else. It is a single question and it sorts most of this literature.
Why this justifies 🔬 rather than ⚠️
Chapter 5's rating definitions do the work here.
⚠️ requires real human data that does not yet settle the question — Phase I/II trials, small or short RCTs, mixed randomized results, surrogate endpoints.
🔬 is for early-stage science proceeding properly, where mechanism may be established but clinical translation is unproven.
The Alzheimer's case has: a plausible mechanism, animal data, and observational epidemiology. It has no completed randomized efficacy trial. Observational association, however large the dataset, is not the kind of human data that ⚠️ describes — it is hypothesis-generating, which is what 🔬 is for.
This is a fine distinction and it is doing real work. Rating the Alzheimer's claim ⚠️ would place it alongside compounds that have actually been tested in randomized humans and produced ambiguous results. Those are different epistemic situations and the system distinguishes them deliberately.
The counterweight
Observational data is not worthless and this case study should not be read as saying so.
It generates hypotheses that would not otherwise be generated. It covers populations trials exclude — the old, the comorbid, the pregnant, the very ill. It detects harms trials are too small and too short to see. It answers questions no trial will ever be funded to answer. And for some questions, a randomized trial is impossible and observational data is all there will ever be.
The failure is not using observational data. It is using it for the one thing it is worst at, which is estimating the benefit of a treatment that clinicians chose to give to some patients and not others.
That specific inference is the one to distrust. Almost everything else observational data does, it does usefully.
Discussion questions
-
Work through the five biases and identify which could be addressed by statistical adjustment and which could not. What makes the difference?
-
Why is observational data more reliable for harms than for benefits? State the asymmetry precisely.
-
Reverse causation is listed as bias ③. Describe concretely how early cognitive decline could influence which diabetes drug someone is prescribed.
-
Suppose a randomized Alzheimer's trial reports a null result while the observational association remains. What should you conclude? What are the possible explanations, and how would you distinguish them?
-
This book rates the claim 🔬 rather than ⚠️ specifically because the human data is observational rather than randomized. Is that distinction too fine to be useful? Argue both sides.
-
Build the reflex. Write down the single question to ask first on meeting any claim that a drug prevents a disease. Then apply it to three health claims you encounter this week and record how many were randomized.