In the 1980s, nutrition science was as close to certain about beta-carotene as it has ever been about
In This Chapter
- The Hook: The trial they had to stop
- 2.1 The ladder
- 2.2 🚪 Healthy-user bias — the threshold concept
- 2.3 Confounding, and why "they adjusted for it" doesn't save you
- 2.4 Reverse causation
- 2.5 "Compared to what?" — the substitution question
- 2.6 Relative risk, absolute risk, and the 18% that terrified everyone
- 2.7 Surrogate endpoints: measuring the thing you can measure
- 2.8 The rest of the toolkit
- 2.9 Ninety seconds with an abstract
- 2.10 What good evidence actually looks like
- Project Checkpoint: Your Claim Filter
- Chapter Summary
- What's Next
Chapter 2 — How to Read Nutrition Research: Study Design, Confounders, Correlation vs. Causation, and Why Most Headlines Are Wrong
The Hook: The trial they had to stop
In the 1980s, nutrition science was as close to certain about beta-carotene as it has ever been about anything.
The evidence was not thin. It was overwhelming, and it came from several directions at once. People who ate more fruits and vegetables got less cancer — that finding was rock-solid, replicated everywhere anyone looked. People with more beta-carotene circulating in their blood got less lung cancer, and this held even after researchers adjusted for how much people smoked. Beta-carotene is a precursor to vitamin A, it's an antioxidant, and the mechanism made beautiful sense: oxidative damage contributes to carcinogenesis, antioxidants reduce oxidative damage, therefore antioxidants should reduce cancer.
Observational data, biomarker data, dose-response, mechanism. Everything lined up.
So two large randomized trials were launched to confirm it and quantify the benefit — because that's what a responsible field does before telling millions of people to take a pill. One in Finland, in male smokers, known as ATBC. One in the United States, in smokers and people exposed to asbestos, known as CARET.
Both groups of investigators expected to demonstrate a benefit. Reasonable people would have bet heavily on it.
Both trials found more lung cancer in the beta-carotene group.
CARET was stopped early. You do not stop a trial early for a null result; you stop it early when continuing would mean knowingly harming the people who volunteered. The supplement that thirty years of observational data said should prevent lung cancer appeared to be causing it in exactly the population most at risk.
I want you to sit with the size of that failure, because it is the single most important story in modern nutrition science and almost nobody outside the field knows it.
This was not a fringe hypothesis. It was not one bad study. It was a large, coherent, multiply-replicated body of evidence, with a clean mechanism, that pointed confidently in a direction that turned out to be backwards.
And here is the part that should genuinely alarm you: the observational data was not wrong. People with more beta-carotene in their blood really did get less lung cancer. That association was real, and it's still real today.
It just wasn't caused by the beta-carotene.
People with high blood beta-carotene are people who eat a lot of vegetables. People who eat a lot of vegetables — in the populations that were studied — also smoked less, exercised more, drank less, were wealthier, had more education, had better healthcare access, and did roughly forty other things that reduce lung cancer risk. Beta-carotene in the blood wasn't a cause. It was a marker of being the kind of person who does healthy things.
That phenomenon has a name, it is the most important concept in this chapter, and once you can see it you will not be able to stop seeing it.
This chapter is the one that makes every other chapter in this book legible. It's also the one that will still be useful to you in twenty years, when every specific claim in these pages has been revised. If you read one chapter of this book, read this one.
🏃 Fast Track: §2.1 (the ladder), §2.2 (healthy-user bias), and §2.5 (compared to what?) are the irreducible core — about twenty minutes. Do those three and the Project Checkpoint, and you'll have most of the practical value.
🔬 Deep Dive: §2.3 (why statistical adjustment does not save you), §2.6 (relative vs. absolute risk), and §2.8 (publication bias and the garden of forking paths) are where students and clinicians should spend real time. §2.9 is the hands-on drill — reading an abstract line by line — and §2.10 is the constructive half, what it looks like when evidence is actually good. Skipping §2.10 produces exactly the corrosive over-skepticism this chapter is trying to avoid.
2.1 The ladder
Not all evidence is equal, and the differences aren't subtle. Here is the hierarchy, from weakest to strongest as a basis for changing what you eat.
| Rung | Design | What it can tell you | What it cannot |
|---|---|---|---|
| 1 | Mechanism / in vitro | What could happen, biochemically. Generates hypotheses. | Almost nothing about living humans. Cells in a dish are not a person. |
| 2 | Animal study | Plausibility, dose ranges, mechanism in a whole organism | Whether it applies to humans. Doses are often wildly beyond human intake. |
| 3 | Case report / anecdote | That something happened once, to someone | Whether it generalizes to anyone else. Includes every testimonial ever filmed. |
| 4 | Cross-sectional study | A snapshot: these things co-occur right now | Which came first. Cannot distinguish cause from consequence. |
| 5 | Prospective cohort | Real outcomes over real timescales in real humans | Causation. Haunted by confounding (§2.2, §2.3). |
| 6 | Randomized controlled trial (RCT) | Causation — the randomization breaks confounding | Usually: long-term effects, hard endpoints, generalizability. Often small and short. |
| 7 | Systematic review / meta-analysis of RCTs | The weight of the trial evidence, pooled | Nothing the underlying trials couldn't. Garbage in, garbage out. |
| 8 | Consensus of independent expert bodies | The considered judgment of the whole evidence base | Nothing fast. These are slow, conservative, and usually right. |
How to place a study in about ten seconds. Read only the abstract and ask three things in order:
- Species? If it's not humans, you're at rung 1 or 2. Stop treating it as advice.
- Randomized? If yes, you're at 6 or above. If no, you're at 5 or below.
- What was measured — a disease, or a marker? A trial measuring cholesterol for twelve weeks is not a trial measuring heart attacks over twenty years.
That's it. Three questions, and you have placed the study more accurately than most of the coverage of it will.
📊 Diagram (described). Picture the ladder not as a straight vertical ladder but as a funnel lying on its side, narrowing to the right. At the wide left end sit millions of mechanistic and animal findings — cheap, fast, numerous, and mostly destined to go nowhere. Moving right, the funnel narrows sharply: a small fraction of those hypotheses ever get tested in humans at all. It narrows again at randomization, because trials are expensive and slow. It narrows again at hard endpoints, because measuring death takes decades. And it narrows almost to a point at replication.
Now overlay a second thing on that funnel: media attention, which is distributed almost exactly backwards. The wide left end — where the mouse studies are — generates the overwhelming majority of the headlines, because there are millions of findings there and each one is new. The narrow right end generates almost none, because a consensus statement that confirms what the last one said is not a story. The volume of coverage is inversely proportional to the strength of the evidence, structurally, at every point along the funnel.
🔄 Check your understanding. Place each of these on the ladder: (a) a testimonial from someone who lost 40 pounds on a diet; (b) a study in which 60 adults were randomly assigned to two diets for 12 weeks and had their blood pressure measured; (c) a study following 90,000 nurses for 25 years and recording who developed diabetes; (d) a study showing a plant compound kills cancer cells in a petri dish.
Answer
(a) rung 3 — case report/anecdote. (b) rung 6 — RCT, but note it's small, short, and measures a marker, so it's a weak 6. (c) rung 5 — prospective cohort. (d) rung 1 — in vitro.
The one most people misrank is (c) versus (b). The cohort has 90,000 people and 25 years and a real disease endpoint, which feels far more impressive than 60 people for 12 weeks. But no amount of size or duration fixes confounding — and the beta-carotene story is what happens when you forget that. Bigger observational studies give you a more precise estimate of a possibly biased quantity.
2.2 🚪 Healthy-user bias — the threshold concept
Here it is. The gate. If one idea from this book stays with you for the rest of your life, make it this one.
🚪 Threshold concept.
In free-living human populations, the people who do one healthy thing tend to do all the other healthy things too — and this is very difficult to separate from the effect of the thing you're studying.
This is healthy-user bias, and it is the single largest reason that nutrition headlines derived from observational studies are wrong.
Think about who, in a typical Western population, eats a lot of vegetables. Or takes a daily vitamin. Or eats breakfast. Or drinks green tea.
That person is, on average and to a statistically meaningful degree: less likely to smoke, more likely to exercise, more likely to have health insurance and use it, more likely to have a higher income, more likely to have more education, less likely to be food insecure, more likely to sleep adequately, more likely to attend screening appointments, more likely to take prescribed medication correctly, less likely to be socially isolated, and more likely to live somewhere with clean air and a park.
Every one of those things independently affects how long they live.
So when a study reports that vitamin-takers have lower mortality, you are looking at the combined effect of the vitamin plus all of that. And since the effect of the vitamin is probably small and the effect of all of that is large, what you're mostly measuring is the kind of person who takes vitamins.
The before-and-after
This is a threshold concept, so let me make the shift explicit.
| Before you understand healthy-user bias | After |
|---|---|
| "Studies show olive oil consumers live longer, so olive oil extends life." | "Studies show olive oil consumers live longer. Olive oil consumers also differ from non-consumers in dozens of ways that affect lifespan. How much of that gap is the oil?" |
| A big study is more convincing than a small one. | A big observational study is a more precise measurement of a possibly biased quantity. Size doesn't fix bias. |
| "They adjusted for confounders, so it's fine." | "They adjusted for the confounders they measured, with the precision they measured them at. What about the ones they didn't?" |
| Nutrition studies contradict each other randomly. | Nutrition studies contradict each other systematically — observational designs and trials disagree in a predictable direction, and that direction tells you something. |
The clearest demonstration, from outside nutrition
The most famous illustration of this phenomenon isn't a food at all. For years, observational studies consistently indicated that women taking postmenopausal hormone therapy had substantially lower cardiovascular risk. The association was large, replicated, and mechanistically plausible.
Then the Women's Health Initiative randomized it. The randomized results did not show the cardiovascular protection the observational data had promised, and revealed harms in some populations that led to a wholesale revision of clinical practice.
What had happened? In that era, the women who took hormone therapy were systematically different — they saw doctors more, were wealthier, were healthier at baseline. The therapy was a marker of being the kind of woman who received a lot of healthcare.
Same shape as beta-carotene. Same shape as vitamin E. Same shape as a great deal of what you have read about food.
🔬 Claim → Evidence → Verdict
The claim: "Breakfast is the most important meal of the day. Skipping breakfast makes you gain weight and slows your metabolism."
Where it comes from: A genuinely large body of observational evidence. Across many populations, people who eat breakfast have lower body weight, better diet quality, and better metabolic markers than people who skip it. The association is real, replicated, and substantial. (It also, not incidentally, has been enthusiastically promoted by the breakfast cereal industry for a century, though the observational finding stands on its own.)
What the evidence actually shows: Breakfast eaters differ from breakfast skippers in almost every way that matters. Breakfast skipping correlates with shift work, irregular schedules, higher alcohol intake, smoking, lower income, chaotic mornings, and — this one is important — with people who are already trying to lose weight by skipping meals, which is reverse causation. When randomized trials assign people to eat or skip breakfast and hold calories comparable, the dramatic weight effects largely fail to appear. Some people do better eating breakfast; some genuinely do better without it; the population-level effect on body weight is small.
📉 Evidence quality: Rung 5 observational evidence, strong and consistent — pointing at a conclusion that rung 6 evidence does not support. The classic signature of healthy-user bias plus reverse causation.
Verdict: 🟠 Probably false as a universal claim. Breakfast is a fine meal. There is no established metabolic penalty for skipping it, and "most important meal of the day" is a slogan, not a finding. (If "breakfast is the most important meal" is on your Chapter 1 Belief Inventory — and it's the single most common entry — leave it there. Don't correct the sheet. That's the point of sealing it.)
🍽️ On your plate. The practical rule: whenever you read that people who do X are healthier, immediately ask "what else is true about people who do X?" Spend fifteen seconds actually generating the list. If you can name four plausible confounders in fifteen seconds, so could the researchers — and the question becomes whether they measured them well enough, which §2.3 says is harder than it sounds.
2.3 Confounding, and why "they adjusted for it" doesn't save you
The obvious response to §2.2 is: fine, but researchers know about this. They statistically adjust for smoking, income, exercise, and the rest.
They do. And it helps. And it is nowhere near sufficient, for four reasons that every nutritional epidemiologist knows and almost no news article mentions.
1. You can only adjust for what you measured
If a study didn't collect data on sleep, it cannot adjust for sleep. And there is no study that measured everything, because "everything" includes things nobody has thought of yet.
2. You can only adjust as well as you measured it
This one is subtle and it is the killer.
Adjustment removes the confounding captured by your measurement. If you measure smoking as a yes/no question, you have not captured the difference between someone who smokes forty a day and someone who smokes two a week — and both answer "yes." The residual difference stays in your data, looking exactly like a diet effect.
Now recall from Chapter 1 that diet itself is measured by food frequency questionnaire, an instrument with large systematic error. So you are adjusting a badly measured exposure using badly measured covariates. What's left over after adjustment is called residual confounding, and in nutrition it is often larger than the effect being studied.
3. Adjustment can make things worse
If you adjust for something that sits on the causal pathway between your exposure and your outcome, you subtract part of the real effect. If you adjust for a "collider" — a variable caused by both the exposure and the outcome — you can manufacture an association that doesn't exist.
Deciding what to adjust for is a modelling judgment, not a mechanical procedure, and different defensible judgments produce different answers from the same data.
4. The multiverse problem
Which brings us to a genuinely unsettling finding about how this literature is produced.
Given one dataset and one research question, there are hundreds of defensible analytical choices: how to categorize intake, which covariates to include, how to handle missing data, whether to exclude early deaths, which subgroups to examine. When researchers have run every defensible combination on the same data and looked at the spread of results — an exercise sometimes called a multiverse or specification-curve analysis — the answers frequently range from clearly harmful through null to clearly protective.
Nobody is cheating. Every individual analysis is defensible. But a single published paper reports one point in that space, and you cannot tell from the paper which point.
🔍 Why this works. Here's the intuition that makes residual confounding click. Imagine trying to measure whether a small stone thrown into a river changes the water level. You can measure the river's level very precisely — but the river's level is also affected by rainfall, snowmelt, upstream dams, evaporation, and season. You can adjust for rainfall, if you measured rainfall. But you measured rainfall at one gauge, twelve miles away, once a day. The residual error in your rainfall measurement is enormous compared to the effect of one stone. The problem isn't that you failed to think of rainfall. It's that your correction is coarser than the thing you're trying to detect. Nutrition effect sizes are stones. Lifestyle confounding is weather.
🔄 Check your understanding. A study finds that people who take multivitamins have 8% lower all-cause mortality, adjusted for age, sex, smoking, BMI, exercise, income, and education. Why should you still be cautious?
Answer
Several reasons, any two of which are a good answer: (1) They adjusted for the confounders they measured — not sleep, stress, social connection, healthcare utilization, medication adherence, or the dozens of unmeasured behaviors that cluster with vitamin-taking. (2) The confounders they did adjust for were measured coarsely (self-reported exercise, a single income bracket), leaving residual confounding. (3) An 8% relative reduction in mortality is a small effect, well within the range that residual confounding routinely produces in this literature. (4) Vitamin-taking is close to a definitional marker of health-consciousness — it may be one of the purest healthy-user signals available.
Note what this does not say: it doesn't say the finding is false. It says the study cannot distinguish "vitamins help" from "the kind of person who takes vitamins lives longer." For that you need randomization — and the large randomized multivitamin trials have generally not found mortality benefits in well-nourished populations, which is Chapter 16's territory.
2.4 Reverse causation
Shorter, and easier, and constantly missed.
Sometimes the outcome causes the exposure, not the other way around.
- People who drink diet soda have higher rates of obesity and type 2 diabetes. Does diet soda cause obesity — or do people who are gaining weight or have been told to watch their sugar switch to diet soda?
- People with low cholesterol have higher rates of some cancers. Does low cholesterol cause cancer — or does undiagnosed cancer, in its early years, lower cholesterol?
- People who eat less meat have more illness in some populations. Does reducing meat cause illness — or do people who become ill reduce their meat intake because they've lost appetite, or been advised to?
- People who skip breakfast weigh more. Or: people who weigh more skip breakfast, because they're trying to lose weight.
The standard defence is to exclude the first few years of follow-up — a lag analysis — on the theory that if reverse causation is at work, the association should weaken when you ignore early events. It helps. It doesn't fully solve it, because some diseases develop over decades.
🍽️ On your plate. The test question: could the arrow point the other way? Ask it every single time, out loud if necessary. It takes three seconds and it catches a startling fraction of nutrition headlines — particularly any headline connecting a food to a condition that changes appetite or prompts dietary advice.
2.5 "Compared to what?" — the substitution question
We met this in Chapter 1. Now let's make it a tool.
Every dietary change is a substitution. You cannot remove a food and replace it with nothing; that would be a smaller diet, which is a different intervention. So the question "Is X bad for you?" is not merely hard — it is malformed. It has no answer, in the way that "is a number large?" has no answer.
The well-formed version is:
Compared to what, in whom, for how long, and measuring what?
Watch how it dissolves arguments.
| Malformed question | Well-formed version | Likely answer |
|---|---|---|
| Is butter bad for you? | Compared to olive oil, at 15% of calories, for LDL cholesterol, in adults, over 6 months? | Worse |
| Is butter bad for you? | Compared to refined-carbohydrate snacks, for cardiovascular events? | Not clearly |
| Is red meat bad for you? | Compared to legumes, for colorectal cancer risk? | Worse |
| Is red meat bad for you? | Compared to no protein source at all? | Nonsensical question |
| Are eggs bad for you? | Compared to a pastry breakfast, for satiety and glycemic response? | Better |
| Is fruit juice bad for you? | Compared to whole fruit, for fiber and satiety? | Worse |
| Is fruit juice bad for you? | Compared to soda, for micronutrients? | Better |
Notice that the same food gets opposite verdicts depending on the comparator, and both verdicts are correct. This is not scientists being evasive. It is the actual structure of the question, and almost all public nutrition argument consists of two people answering different comparisons at each other.
🧩 Productive struggle. Before reading on, take three minutes with this:
A large cohort study reports that people who eat the most whole grains have roughly 20% lower all-cause mortality than people who eat the least.
Write down what comparison this study is actually making. Not what the headline implies — what the data can support. Then write down at least three different real-world substitutions that could be hiding inside "eating fewer whole grains," and say whether you'd expect the same answer for each.
What I'd say
The study compares people who habitually eat more whole grains to people who habitually eat fewer, in the same population, over the same years. It does not compare "adding whole grains" to "not adding whole grains," because nobody was assigned anything.
Substitutions hiding inside "fewer whole grains": 1. Refined grains instead — white bread, white rice, pastries. Here you'd expect whole grains to look genuinely better, and trial evidence on intermediate markers supports that. 2. More meat instead — a low-grain, higher-animal-protein pattern. Different comparison, murkier answer. 3. More vegetables and legumes instead — a low-grain, high-plant pattern. Here you might see little difference or even favor the low-grain group, and the cohort would not distinguish this person from person 1. 4. Simply eating less food overall — food insecurity, illness, appetite loss. Now you have reverse causation on top of everything else.
All four people are in the "low whole grain" group, and the study reports their average. The average of four different substitutions is not a recommendation about any of them.
If you got two of these, you're already reading better than most health journalism.
2.6 Relative risk, absolute risk, and the 18% that terrified everyone
In 2015, IARC classified processed meat as a Group 1 carcinogen. Headlines worldwide put bacon in the same category as tobacco and asbestos, which is technically true and enormously misleading, and the reason is a distinction most people have never been taught.
Group 1 is a statement about how confident we are that something causes cancer. It says nothing about how much cancer. Tobacco and processed meat are both in Group 1. Tobacco causes an enormous amount of cancer; processed meat causes a small amount. The classification isn't wrong; it's answering a question people assumed it wasn't.
Then there's the number. The widely-quoted figure is roughly an 18% increased risk of colorectal cancer per 50 grams of processed meat consumed daily — about two rashers of bacon or one hot dog.
Eighteen percent sounds enormous. Let's do the arithmetic.
Step 1 — the baseline. Lifetime risk of colorectal cancer in a typical Western population is roughly 5% — about 5 people in 100. (Actual figures vary by country, sex, age, and family history; use your own population's data for anything that matters.)
Step 2 — apply the relative risk.
Increased risk = baseline × 1.18
= 5% × 1.18
= 5.9%
Step 3 — the absolute difference.
5.9% − 5.0% = 0.9 percentage points
In plain English: if 100 people ate an extra 50 g of processed meat every day for life, roughly one additional person among them would develop colorectal cancer. The other 99 outcomes would be unchanged.
| Framing | The same fact |
|---|---|
| Relative risk | "18% increased risk" — sounds alarming |
| Absolute risk | "5% becomes 5.9%" — sounds modest |
| Natural frequency | "About 1 extra case per 100 people, lifetime" — sounds like something you can actually weigh |
All three are accurate descriptions of the same finding. Headlines use the first almost exclusively, because it's the biggest-sounding true number available.
💡 Aha moment. Relative risk without a baseline is uninterpretable. A 50% increase in a very rare outcome is nothing; a 5% increase in a very common one can be enormous. Whenever you see a percentage increase in risk, the only correct first response is: increase from what? If the article doesn't say — and it usually doesn't — you have not been given enough information to have an opinion.
🔬 Claim → Evidence → Verdict
The claim: "Processed meat is a Group 1 carcinogen — the same category as smoking and asbestos. There is no safe amount."
Where it comes from: An accurate reading of the IARC classification, which does place processed meat in Group 1 on the basis of consistent evidence for colorectal cancer.
What the evidence actually shows: The classification is correct and the inference is wrong. Group 1 encodes confidence in causation, not magnitude of harm. The magnitude is small in absolute terms — on the order of one additional colorectal cancer case per hundred people over a lifetime, at 50 g/day. That's a real effect worth knowing about and reducing; it is not equivalent to smoking, which shifts lung cancer risk by more than an order of magnitude. "No safe amount" imports a precautionary framing that the classification doesn't carry.
📉 Evidence quality: Rung 8 for the classification itself. The magnitude estimate rests largely on rung 5 cohort data with the usual confounding caveats.
Verdict: 🟡 Unclear / it depends — the premise is accurate, the implication is not. Processed meat carries a small, real risk increase that scales with dose. Whether that's worth changing your diet over is a values question, and you're now equipped to answer it with real numbers instead of a category name.
2.7 Surrogate endpoints: measuring the thing you can measure
Almost every nutrition RCT measures a surrogate endpoint — a marker that stands in for the outcome you actually care about. LDL cholesterol instead of heart attacks. Blood glucose instead of diabetes. Inflammatory markers instead of disease.
This is unavoidable. Heart attacks take decades; LDL takes six weeks. Nobody can fund the alternative.
But surrogates can lie, and the history of medicine is littered with cases where a treatment moved a marker beautifully and did nothing — or harmed — on the outcome that mattered. Beta-carotene raised blood antioxidant levels exactly as predicted. High-dose vitamin E did what it was supposed to do biochemically, and large trials found no cardiovascular benefit and some safety signals at high doses.
A surrogate is trustworthy only to the extent that the causal chain from marker to outcome has been independently established. LDL cholesterol is a relatively good surrogate for cardiovascular risk, because that chain has been tested many ways, including by drugs and by genetics. "Inflammatory markers" is a much weaker surrogate, because the chain from a shifted cytokine to a clinical outcome is far less established — which is why "reduces inflammation" is one of the most abused phrases in wellness marketing. It sounds like a health outcome. It's a laboratory value.
| Surrogate | How trustworthy | Why |
|---|---|---|
| LDL cholesterol → CVD | Reasonably good | Chain tested by trials, genetics, multiple drug classes |
| Blood pressure → stroke/CVD | Good | Extensively validated |
| HbA1c → diabetes complications | Good for microvascular, weaker for macrovascular | Depends on the complication |
| "Inflammatory markers" → disease | Weak | Chain poorly established; markers move for many reasons |
| "Antioxidant capacity" → anything | Very weak | The beta-carotene lesson, unlearned repeatedly |
| Gut microbiome composition → health | Currently very weak | We can't yet say what a "good" composition is (Ch 27) |
🍽️ On your plate. When a study or a product claims a benefit, ask: did they measure a disease, or a number? If it's a number, ask whether that number has ever been shown to predict the disease when you change it deliberately. "Reduces inflammation," "boosts antioxidant status," "improves gut diversity," and "supports immune function" are all marker claims, and none of them is a health outcome.
2.8 The rest of the toolkit
Four more, briefly, because you'll meet all of them.
Publication bias
Studies that find something get published; studies that find nothing often don't. So the published literature is a biased sample of the research conducted, skewed toward positive results. This is why meta-analyses check for it (funnel plots, and formal tests), and why pre-registration — declaring your hypothesis and analysis plan before you look at the data — has become important.
The garden of forking paths
Test twenty things and, by chance alone, about one will hit the conventional significance threshold. A study measuring fifteen biomarkers across four subgroups has done sixty comparisons; finding two "significant" results is what you'd expect from noise.
Watch for: outcomes that appear in the abstract but weren't the study's stated primary outcome; findings reported only in a subgroup ("in women over 50…"); and results emphasized in the press release that appear nowhere in the paper's own conclusions.
Funding effects
Meta-research consistently finds that industry-funded studies report conclusions favorable to the funder more often than independently funded studies of the same question. This does not require fraud — it operates through study design, comparator choice, dose selection, outcome selection, and what gets published.
Apply this symmetrically. Supplement-industry funding is funding. So is a wellness brand's "research institute." So is a commodity board. Chapter 1's rule holds: follow the money in every direction or don't invoke it.
Effect size versus statistical significance
"Statistically significant" means probably not zero. It does not mean large, important, or worth doing anything about. In a study of 200,000 people, an utterly trivial difference will be highly statistically significant.
Always ask for the size, in units you can picture. Not "significantly improved," but "improved by 2 millimetres of mercury" — and then ask whether 2 mmHg matters to you.
🔄 Check your understanding. A press release says: "In a study of 40,000 adults, those consuming the most of Nutrient X had significantly better cognitive scores." Name three things you still need to know before this means anything.
Answer
Strong candidates: (1) Design — 40,000 adults suggests a cohort, so rung 5, so confounding and healthy-user bias apply immediately. (2) Effect size — "significantly better" in 40,000 people could be a difference too small to notice; ask for the actual magnitude on the actual scale. (3) Compared to what — what were low-consumers eating instead? (4) Was cognitive score the pre-registered primary outcome, or one of many measures? (5) Who funded it, and does anyone sell Nutrient X?
Any three of those, and you're reading better than the press release.
2.9 Ninety seconds with an abstract
Everything above is theory. Here's the practice, because the abstract is almost always all you can get — full papers sit behind paywalls, but abstracts are free on PubMed for essentially the entire biomedical literature.
Below is a constructed example — I've written it myself, in the standard style, to demonstrate the technique. It is not a real study and should not be cited as one. But every feature in it is one I've seen many times.
Background. Consumption of polyphenol-rich foods has been inversely associated with cardiometabolic risk in observational studies, though causal evidence is limited.
Objective. To assess the effect of daily supplementation with a standardized extract on markers of vascular function and glycemic control in healthy adults.
Design. Randomized, double-blind, placebo-controlled parallel trial. Forty-two healthy adults (aged 22–41) received either 900 mg/day of extract or placebo for 8 weeks. Outcomes included flow-mediated dilation, fasting glucose, fasting insulin, HOMA-IR, hs-CRP, IL-6, total cholesterol, LDL-C, HDL-C, and triglycerides.
Results. Flow-mediated dilation improved significantly in the intervention group compared with placebo (P = .04). No significant differences were observed for the remaining outcomes. In a post-hoc analysis restricted to participants with baseline BMI above 25, fasting insulin was significantly reduced (P = .03).
Conclusions. Daily supplementation improved endothelial function in healthy adults and may confer cardiometabolic benefit, particularly in overweight individuals. Larger trials are warranted.
Now the ninety seconds. Read it again with these annotations.
"Randomized, double-blind, placebo-controlled." Good. This is rung 6 — genuinely the strong half of the ladder, and better than most nutrition evidence. Credit where due.
"Forty-two healthy adults (aged 22–41)." Two problems immediately. Forty-two is small — small enough that a single unusual responder in either arm can move the result. And "healthy adults aged 22–41" is not you, unless you are a healthy adult aged 22 to 41. Findings in young healthy people frequently don't transfer to older people, sick people, or people with the condition the supplement is being sold to treat.
"900 mg/day of extract." Compare that to what you'd get from food. Standardized extracts routinely deliver doses that would require eating implausible quantities of the source food. If the study used a dose you cannot reach by eating, the study is not about the food.
Count the outcomes. Flow-mediated dilation, fasting glucose, fasting insulin, HOMA-IR, hs-CRP, IL-6, total cholesterol, LDL-C, HDL-C, triglycerides. That's ten. At the conventional threshold, you'd expect roughly one of ten to come up "significant" by chance alone even if the extract were inert. This is the garden of forking paths, visible in a single line of text.
"P = .04." One result, just under the threshold, out of ten outcomes. And notice what it is: flow-mediated dilation is a surrogate — a measure of how much an artery widens in response to increased blood flow. It's a plausible marker. It is not a heart attack.
"In a post-hoc analysis restricted to participants with baseline BMI above 25…" Stop here. This is the most important sentence in the abstract. Post-hoc means they went looking after seeing the data. Restricted to a subgroup means the sample is now smaller than 42 — perhaps twenty people, perhaps fewer. A subgroup finding, discovered after the fact, in a trial that already measured ten outcomes, is close to uninterpretable. It is a hypothesis for a future study, and nothing more.
"May confer cardiometabolic benefit, particularly in overweight individuals." And there it is — the conclusion has silently promoted the post-hoc subgroup finding into the summary sentence, which is the sentence the press release will quote, which is the sentence that becomes the headline.
What this study actually established: in 42 young healthy people, at a dose you cannot obtain from food, over 8 weeks, one of ten measured markers moved, marginally.
What the reel will say: "Clinically proven to improve blood vessel function and insulin resistance."
The red flags, collected
| In the abstract | What it means |
|---|---|
| Sample under ~50 | One unusual participant can drive the result |
| Many outcomes listed | Expect roughly 1 in 20 to hit significance by chance |
| P just under .05 on one outcome | Weak evidence; the number to want is the effect size |
| "Post-hoc," "exploratory," "subgroup" | Hypothesis-generating only. Not a finding. |
| Healthy young volunteers | May not transfer to the population being sold to |
| Extract, isolate, or standardized dose | Study is about the compound, not the food |
| Markers rather than disease | Surrogate endpoint (§2.7) |
| Conclusion broader than the results section | The gap between them is where the press release lives |
| No comparator stated for a dietary change | The substitution question is unanswered |
🔄 Check your understanding. Which single feature of the constructed abstract above would most reduce your confidence in the headline "Supplement Improves Insulin Resistance in Overweight Adults"?
Answer
That the insulin finding was post-hoc and in a subgroup. The trial was designed and powered for the whole group; splitting it afterward by BMI, having already measured ten outcomes, is close to guaranteed to produce something. It is not evidence that the supplement helps overweight adults — it's a suggestion that someone should design a trial to find out.
Partial credit for "the small sample" or "insulin was one of ten outcomes," both of which compound the problem. But the post-hoc subgroup is the one that makes the headline unsupportable rather than merely weak.
🍽️ On your plate. You can do this. It took you longer to read the annotations than it will take you to run the checklist next time, and you now have a real skill that most people writing about health don't have. Search PubMed for any nutrition claim you've heard this month, open the first abstract, and run the table. The first time you correctly call a study weaker than its headline, something changes permanently in how you read.
2.10 What good evidence actually looks like
If this chapter stopped here, it would have taught you to reject everything — and that would make you worse off, not better. Over-skepticism is just credulity with extra steps, and it has a real cost: people who conclude that nothing is known eat according to whoever spoke to them last.
So: what does convincing evidence look like?
It converges. Several independent kinds of evidence, with different weaknesses, point the same way. Cohorts and trials agree. Mechanism supports the finding rather than being invented afterward. Different populations, different countries, different research groups, different funding sources.
The specific power of convergence is that the biases don't overlap. Cohorts suffer from confounding; trials don't. Trials suffer from short duration and unrepresentative samples; cohorts don't. Mechanism suffers from being untested in humans; both of the others don't. When designs with non-overlapping weaknesses agree, the agreement is hard to explain away.
It has a dose-response relationship. More exposure, more effect. Not proof, but hard to produce by confounding alone.
It survives new methods. Mendelian randomization is the most useful recent addition to nutrition's toolkit: it uses genetic variants that affect a trait — say, how much alcohol someone tolerates, or how their body handles a nutrient — as a natural randomization. Because genes are allocated essentially at random at conception and don't change with lifestyle, they sidestep healthy-user bias. It's not perfect and it has its own assumptions, but when Mendelian randomization agrees with cohort data, confidence goes up substantially. When it disagrees — as it has for alcohol, which we'll get to in Chapter 12 — that's a serious signal.
It's boring, old, and unglamorous. As Chapter 1 argued: certainty and excitement are inversely related, because a claim only stays exciting while the evidence is weak enough to argue about.
And the field corrects itself in public. Here's a case worth knowing. PREDIMED, a large Spanish randomized trial of a Mediterranean dietary pattern, was one of the most influential nutrition trials ever run. Years after publication, irregularities in the randomization procedure at some study sites came to light. The paper was retracted and republished with a corrected analysis. The main conclusions largely held, but the episode is genuinely instructive in two directions at once: even flagship trials have flaws, and the mechanism for catching and correcting them worked, in the open, at a cost to the reputations of people who had every incentive to keep quiet.
That's what a functioning field looks like. Not one that never errs. One that publishes its errors.
🔬 Claim → Evidence → Verdict
The claim: "Observational nutrition studies are so confounded that they're worthless. Only RCTs count."
Where it comes from: Everything in §2.2 through §2.4, which is real. The beta-carotene reversal, the hormone therapy reversal, the vitamin E reversal, and the routine failure of observational nutrition findings to replicate in trials are a serious indictment, and the people making this argument are responding to genuine evidence.
What the evidence actually shows: Observational studies are the only design that can measure hard outcomes over decades in ordinary people eating ordinary food — the exact question we most want answered, and one no RCT will ever address. Discarding them means discarding nearly everything we know about long-term diet and disease. The better position is calibrated rather than binary: observational evidence is weak for small effects of single nutrients (where confounding is the same size as the signal) and considerably stronger for large effects and whole dietary patterns, for dose-response relationships, and when corroborated by trials, mechanism, and Mendelian randomization. Also worth noting: RCTs have their own failure modes — short duration, surrogate endpoints, unrepresentative volunteers, poor adherence — and "only RCTs count" would have you discard the evidence on smoking, which was never randomized and never will be.
📉 Evidence quality: The critique is well founded; the conclusion overshoots.
Verdict: 🟠 Probably false as stated. The right stance isn't "ignore cohorts," it's "know what cohorts can and can't carry, and require convergence before acting."
What we don't know
The honest limit of this chapter: we do not have a reliable way to quantify residual confounding in nutritional epidemiology. We know it's there. We know roughly which direction it usually points. We cannot say, for a given study, how much of a reported 20% risk reduction is real and how much is the kind of person.
This is the unresolved problem in the field, not a footnote to it. It's why serious people disagree about red meat, about eggs, about dairy, about moderate alcohol — not because they're reading different data, but because they're making different judgments about how much of the observed signal survives the confounding.
Anyone who tells you this is settled — in either direction — is telling you about their confidence, not about the evidence.
Project Checkpoint: Your Claim Filter
Component two of Your Nutrition Framework. This one is a physical object: a card you'll actually use.
Make a card — index card, phone note, whatever you'll have with you — with these six questions. Full version in Appendix D.
THE CLAIM FILTER
1. What kind of study? What rung? Species → randomized or not → disease or marker. Ten seconds.
2. Compared to what? Every dietary change is a substitution. If the comparison isn't stated, the claim isn't a claim yet.
3. Could the arrow point the other way? Reverse causation. Could the outcome be causing the exposure?
4. What else is true about people who do this? Healthy-user bias. Name four confounders in fifteen seconds. If you can, so could the researchers — did they measure them well?
5. How big, in absolute terms? Relative risk without a baseline is uninterpretable. Convert to natural frequency: how many people per hundred?
6. Who benefits if I believe this? In every direction — the company, the influencer, the researcher's career, the contrarian selling a book, and me.
How to use it. Take three claims — from your Belief Inventory, from your feed, from a family argument — and run each one all the way through. Write the answers down; don't do it in your head, because doing it in your head lets you skip the question you can't answer, which is always the important one.
What to expect. Most claims fail at question 2 or question 5, and they fail in a specific way: you can't answer. Not "the answer is bad" — you simply cannot determine what the comparison was or how big the effect was from the information you were given.
That is the single most useful output of this exercise. A claim you cannot evaluate is not a claim you should act on, and recognizing that state — I don't have enough information to have an opinion about this — is genuinely uncomfortable and genuinely the skill.
Keep the card. You'll use it in every remaining chapter of this book, and the Chapter 17 checkpoint is running your entire Belief Inventory through it.
Next checkpoint (Chapter 3): your digestion log — three days of what you ate and how you felt, with no judgment attached.
Chapter Summary
The eight-rung ladder — place any study in ten seconds by asking: species? randomized? disease or marker?
| Weak → | 1 mechanism · 2 animal · 3 anecdote · 4 cross-sectional · 5 cohort · 6 RCT · 7 meta-analysis of RCTs · 8 consensus | → Strong |
|---|---|---|
The five failure modes, and the question that catches each:
| Failure mode | The catching question |
|---|---|
| Healthy-user bias 🚪 | What else is true about people who do this? |
| Residual confounding | They adjusted for what they measured — how well did they measure it? |
| Reverse causation | Could the arrow point the other way? |
| Unspecified substitution | Compared to what? |
| Relative-risk inflation | How big in absolute terms? Per hundred people? |
Plus: surrogate endpoints (a number is not a disease), publication bias, the garden of forking paths, funding effects in every direction, and significance ≠ size.
What good evidence looks like: it converges across designs with non-overlapping weaknesses; it shows dose-response; it survives new methods like Mendelian randomization; it's boring and old; and the field that produced it corrects itself in public — as with the PREDIMED retraction and republication.
This chapter's verdicts:
| Claim | Verdict |
|---|---|
| Breakfast is the most important meal; skipping it causes weight gain | 🟠 Probably false |
| Processed meat is Group 1 — same as smoking, no safe amount | 🟡 Unclear / it depends |
| Observational nutrition studies are worthless; only RCTs count | 🟠 Probably false |
The one thing to remember: the people who do one healthy thing tend to do all the healthy things. Beta-carotene didn't prevent lung cancer. Eating vegetables was a marker of being the kind of person who doesn't get lung cancer. That confusion, repeated at scale for forty years, is most of what you've read about food.
What's Next
You now have the machinery. Everything from here forward is applying it.
Chapter 3 turns to the plumbing — what actually happens to a sandwich between your mouth and your bloodstream. It's the foundation for fiber, the microbiome, food intolerances, and clinical nutrition, and it quietly kills a remarkable number of myths on its own. It is very difficult to keep believing in detox protocols once you understand what a liver actually does.
It's also where the book stops being about epistemology and starts being about food. You've earned it.