Appendix D — Evidence Evaluation Toolkit
The book's reasoning tools, in one place, in the form you would actually use them.
⚠️ This is the appendix Chapter 38 §38.2 argues was the point of the whole book.
⚠️ Everything here appeared in a chapter first. What this appendix adds is that it fits on a few pages and can be used without rereading anything.
1. ⚠️ The ten-second placement test
From Chapter 2.
⚠️ Before anything else, place the claim on the ladder. Three questions, ten seconds:
⚠️ 1. Was it in humans? ⚠️ 2. Was it randomized? ⚠️ 3. Did it measure a disease, or a marker?
The eight rungs
| ⚠️ Rung | ⚠️ What it is | ⚠️ What it can show |
|---|---|---|
| 8 | ⚠️ Systematic review of randomized trials with hard outcomes | ⚠️ The strongest thing available in nutrition, and rare |
| 7 | ⚠️ Large randomized trial, hard outcome, long follow-up | ⚠️ Causation, in the population studied |
| 6 | ⚠️ Randomized trial, surrogate marker | ⚠️ Causation about the MARKER, not the disease |
| 5 | ⚠️ Large prospective cohort | ⚠️ Association. Confounding is the permanent problem |
| 4 | ⚠️ Case-control, cross-sectional | ⚠️ Association, more fragile |
| 3 | ⚠️ Metabolic ward / short controlled feeding | ⚠️ Mechanism in humans, briefly, in unusual conditions |
| 2 | ⚠️ Animal studies | ⚠️ Hypotheses. Not conclusions about people |
| 1 | ⚠️ Cell studies, mechanism, expert opinion, anecdote | ⚠️ Reasons to investigate. Nothing more |
⚠️ The most common error in nutrition reporting is presenting a rung 2 or 3 finding in rung 7 language.
⚠️ And the thing the ladder does not capture, which Chapter 2 spent a chapter on:
⚠️ HEALTHY-USER BIAS. People who do the healthy thing differ from people who don't in income, education, activity, smoking, healthcare access and a dozen other things. ⚠️ Every rung 4 and 5 finding has this problem, statistical adjustment reduces it and never removes it, and it is why beta-carotene looked protective in cohorts and caused harm in trials.
2. ⚠️ The six-question check
From Chapter 37 §37.7. The core tool. Stop as soon as one fails.
⚠️ Question 1 — What rung is it on?
⚠️ Apply the ten-second test above. ⚠️ In Theo's own log (Ch 37 CS1), 14 of 31 claims failed here alone.
⚠️ Question 2 — What is the comparison?
⚠️ Compared to WHAT — nothing, worse advice, or the best available advice?
⚠️ Most impressive results compare against doing nothing, which tests attention rather than the intervention (Ch 35 §35.11).
⚠️ Question 3 — What is the denominator?
⚠️ Per what? Per gram, per calorie, per serving chosen by whom?
⚠️ The unit is a hidden claim about what the thing is FOR (Ch 36 §36.3b). ⚠️ And relative risk without a baseline is not information (Ch 2, Ch 9).
⚠️ Question 4 — Does it REPLACE the boring answer, or ADJUST it?
⚠️ Claims that adjust are frequently right. Claims that replace are almost always wrong (Ch 37 §37.2b — this is a summary of what happened to fifty years of attempts, not a rule of nature).
⚠️ Question 5 — Is anything being sold, and does the claim generate a rule?
⚠️ Funding is not disqualifying and it is information (Ch 3). ⚠️ A claim that generates a personal rule set deserves the extra scrutiny of Chapter 34 §34.7.
⚠️ Question 6 — What would the claimant accept as disproof?
⚠️ If nothing would, it is not a claim about the world (Ch 35 CS2's "you haven't eliminated enough yet").
⚠️ For a claim arriving years from now, add three (Ch 38 §38.7):
⚠️ Has it been TESTED, or only proposed? · Is it new, or is it BACK — check Appendix E · And has anything CONVERGED on it?
3. The worksheet
⚠️ Copy this. One claim per sheet.
THE CLAIM, in one sentence, in my own words:
Where I encountered it: _ Who benefits if I believe it: _
| ⚠️ Question | ⚠️ Answer | ⚠️ Pass? | |
|---|---|---|---|
| 1 | What rung? | ||
| 2 | Compared to what? | ||
| 3 | What denominator? | ||
| 4 | Replaces or adjusts? | ||
| 5 | Anything sold? Generates a rule? | ||
| 6 | What would disprove it? | ||
| +1 | Tested or only proposed? | ||
| +2 | New, or back? (Appendix E) | ||
| +3 | Has anything converged on it? |
VERDICT I would give it: ⚠️ ✅ · 🟢 · 🟡 · 🟠 · ❌ · ⚗️
What I will do about it: __
⚠️ "Nothing yet, I'll see if it replicates" is a complete and usually correct answer (Ch 38 §38.7).
4. ⚠️ The recurring error patterns
⚠️ Named, so you can recognize them faster than you can analyse them.
| ⚠️ Pattern | ⚠️ What it looks like | ⚠️ Where |
|---|---|---|
| ⚠️ The beta-carotene template | ⚠️ Observed food → inferred compound → tested pill → nothing or harm | Ch 2, 13, 35 |
| ⚠️ Healthy-user bias | ⚠️ The people doing it differ in everything else too | Ch 2 |
| ⚠️ Relative risk with no baseline | "Raises risk 18%" — of what, from what? | Ch 2, 9 |
| ⚠️ A dose question argued as a presence question | "It contains X" as though quantity were irrelevant | Ch 18, 20 |
| ⚠️ Marker mistaken for outcome | ⚠️ Optimizing a number nobody has shown predicts disease | Ch 35 §35.2 |
| ⚠️ Mechanism mistaken for result | ⚠️ Plausible pathway, prediction never tested — or tested and failed | Ch 19 §19.5 |
| ⚠️ Denominator shopping | ⚠️ Per kg vs per calorie vs per serving, chosen to win | Ch 30 §30.5, Ch 36 §36.3b |
| ⚠️ The unfalsifiable framework | ⚠️ "You haven't eliminated enough yet" | Ch 35 CS2 |
| ⚠️ Over-extension | ⚠️ A real finding pushed past what it showed. The commonest myth shape | Ch 17 |
| ⚠️ A resource problem called a knowledge problem | ⚠️ "They just need educating" about people optimizing harder than the adviser | Ch 32 §32.11 |
| ⚠️ A structural failure called a personal one | ⚠️ "More discipline" applied to a plan, a budget or a cue | Ch 31, 32, 33 |
| ⚠️ Suppression claims | ⚠️ Compare against a DOCUMENTED episode (Ch 18 §18.12) before believing one | Ch 17 §17.9 |
5. ⚠️ Reading a study without reading a study
⚠️ Six things to look at, in order, when you have five minutes and a paper.
⚠️ 1. The abstract's last sentence. ⚠️ Compare it to the results section. Overreach lives here.
⚠️ 2. Who was studied. ⚠️ Number, age, sex, health status, country. ⚠️ Then ask whether that is you.
⚠️ 3. How long. ⚠️ Nutrition outcomes take decades; most trials run weeks.
⚠️ 4. What was measured. ⚠️ A disease, a marker, or a self-reported questionnaire? ⚠️ Dietary intake is usually self-reported and self-reported intake is unreliable (Ch 3).
⚠️ 5. The comparison arm. ⚠️ Question 2, and it is where the effect size usually comes from.
⚠️ 6. Funding and conflicts. ⚠️ Last, not first (Ch 3). ⚠️ It is information, not a verdict.
⚠️ And one thing that is not on the list: the journal's name. ⚠️ Prestigious journals publish findings that fail to replicate at rates that would surprise you (Ch 3).
6. ⚠️ The self-diagnostic
From Chapter 33 §33.3, and it belongs in an evidence toolkit for a reason.
⚠️ When the question is about your own eating rather than about a claim:
⚠️ NOT "am I doing this right?" — unanswerable in the moment, and it invites self-interrogation that helps nobody.
⚠️ INSTEAD: "WHAT STARTED THIS?"
⚠️ A CUE — a time, a place, a packet, a feeling, the end of something? ⚠️ An INTERVAL since eating? ⚠️ Or a CLAIM I read?
⚠️ Because the answer determines which tool works:
| ⚠️ A cue | ⚠️ Change the situation. Not the food, and not your resolve (Ch 33 §33.4) |
|---|---|
| ⚠️ An interval | ⚠️ Fix the earlier meal (Ch 33 §33.3) |
| ⚠️ A claim | ⚠️ Run the six questions before changing anything |
⚠️ Using the wrong tool is worse than doing nothing, and that is the whole of Chapter 33's threshold.
7. When to stop using this
⚠️ Two limits, both from Chapter 38 §38.9b.
⚠️ The filter is for claims that would change what YOU eat. ⚠️ Pointing it at other people converts a tool into a personality, and a person who has just learned to evaluate evidence is at their most insufferable.
⚠️ And the one exception, which is the only version that scales: ⚠️ if someone asks you how to tell whether a claim is worth acting on, teach them this appendix rather than giving them your answer.
⚠️ A filter can also become a wall. ⚠️ "Most of what gets sold to me about food is nonsense" is well supported by 136 ❌ and 🟠 verdicts in Appendix E — and is also the belief most likely to make you dismiss something true.
Where this came from
| Chapter 2 | ⚠️ The ladder, the placement test, healthy-user bias |
|---|---|
| Chapter 3 | ⚠️ Funding, publication bias, self-reported intake |
| Chapter 17 | ⚠️ Myth shapes and over-extension |
| Chapter 33 §33.3 | ⚠️ The cue/interval self-diagnostic |
| Chapter 35 §35.11 | ⚠️ Comparison against best available advice |
| Chapter 36 §36.3b | ⚠️ The denominator problem |
| Chapter 37 §37.7 | ⚠️ The six questions |
| Chapter 38 §38.7 | ⚠️ The three additions for a future claim |
| Appendix E | ⚠️ 468 claims already assessed — check here first |