Appendix D — Evidence Evaluation Toolkit

The book's reasoning tools, in one place, in the form you would actually use them.

⚠️ This is the appendix Chapter 38 §38.2 argues was the point of the whole book.

⚠️ Everything here appeared in a chapter first. What this appendix adds is that it fits on a few pages and can be used without rereading anything.


1. ⚠️ The ten-second placement test

From Chapter 2.

⚠️ Before anything else, place the claim on the ladder. Three questions, ten seconds:

⚠️ 1. Was it in humans? ⚠️ 2. Was it randomized? ⚠️ 3. Did it measure a disease, or a marker?

The eight rungs

⚠️ Rung ⚠️ What it is ⚠️ What it can show
8 ⚠️ Systematic review of randomized trials with hard outcomes ⚠️ The strongest thing available in nutrition, and rare
7 ⚠️ Large randomized trial, hard outcome, long follow-up ⚠️ Causation, in the population studied
6 ⚠️ Randomized trial, surrogate marker ⚠️ Causation about the MARKER, not the disease
5 ⚠️ Large prospective cohort ⚠️ Association. Confounding is the permanent problem
4 ⚠️ Case-control, cross-sectional ⚠️ Association, more fragile
3 ⚠️ Metabolic ward / short controlled feeding ⚠️ Mechanism in humans, briefly, in unusual conditions
2 ⚠️ Animal studies ⚠️ Hypotheses. Not conclusions about people
1 ⚠️ Cell studies, mechanism, expert opinion, anecdote ⚠️ Reasons to investigate. Nothing more

⚠️ The most common error in nutrition reporting is presenting a rung 2 or 3 finding in rung 7 language.

⚠️ And the thing the ladder does not capture, which Chapter 2 spent a chapter on:

⚠️ HEALTHY-USER BIAS. People who do the healthy thing differ from people who don't in income, education, activity, smoking, healthcare access and a dozen other things. ⚠️ Every rung 4 and 5 finding has this problem, statistical adjustment reduces it and never removes it, and it is why beta-carotene looked protective in cohorts and caused harm in trials.


2. ⚠️ The six-question check

From Chapter 37 §37.7. The core tool. Stop as soon as one fails.

⚠️ Question 1 — What rung is it on?

⚠️ Apply the ten-second test above. ⚠️ In Theo's own log (Ch 37 CS1), 14 of 31 claims failed here alone.

⚠️ Question 2 — What is the comparison?

⚠️ Compared to WHAT — nothing, worse advice, or the best available advice?

⚠️ Most impressive results compare against doing nothing, which tests attention rather than the intervention (Ch 35 §35.11).

⚠️ Question 3 — What is the denominator?

⚠️ Per what? Per gram, per calorie, per serving chosen by whom?

⚠️ The unit is a hidden claim about what the thing is FOR (Ch 36 §36.3b). ⚠️ And relative risk without a baseline is not information (Ch 2, Ch 9).

⚠️ Question 4 — Does it REPLACE the boring answer, or ADJUST it?

⚠️ Claims that adjust are frequently right. Claims that replace are almost always wrong (Ch 37 §37.2b — this is a summary of what happened to fifty years of attempts, not a rule of nature).

⚠️ Question 5 — Is anything being sold, and does the claim generate a rule?

⚠️ Funding is not disqualifying and it is information (Ch 3). ⚠️ A claim that generates a personal rule set deserves the extra scrutiny of Chapter 34 §34.7.

⚠️ Question 6 — What would the claimant accept as disproof?

⚠️ If nothing would, it is not a claim about the world (Ch 35 CS2's "you haven't eliminated enough yet").

⚠️ For a claim arriving years from now, add three (Ch 38 §38.7):

⚠️ Has it been TESTED, or only proposed? · Is it new, or is it BACK — check Appendix E · And has anything CONVERGED on it?


3. The worksheet

⚠️ Copy this. One claim per sheet.

THE CLAIM, in one sentence, in my own words:


Where I encountered it: _ Who benefits if I believe it: _

⚠️ Question ⚠️ Answer ⚠️ Pass?
1 What rung?
2 Compared to what?
3 What denominator?
4 Replaces or adjusts?
5 Anything sold? Generates a rule?
6 What would disprove it?
+1 Tested or only proposed?
+2 New, or back? (Appendix E)
+3 Has anything converged on it?

VERDICT I would give it: ⚠️ ✅ · 🟢 · 🟡 · 🟠 · ❌ · ⚗️

What I will do about it: __

⚠️ "Nothing yet, I'll see if it replicates" is a complete and usually correct answer (Ch 38 §38.7).


4. ⚠️ The recurring error patterns

⚠️ Named, so you can recognize them faster than you can analyse them.

⚠️ Pattern ⚠️ What it looks like ⚠️ Where
⚠️ The beta-carotene template ⚠️ Observed food → inferred compound → tested pill → nothing or harm Ch 2, 13, 35
⚠️ Healthy-user bias ⚠️ The people doing it differ in everything else too Ch 2
⚠️ Relative risk with no baseline "Raises risk 18%" — of what, from what? Ch 2, 9
⚠️ A dose question argued as a presence question "It contains X" as though quantity were irrelevant Ch 18, 20
⚠️ Marker mistaken for outcome ⚠️ Optimizing a number nobody has shown predicts disease Ch 35 §35.2
⚠️ Mechanism mistaken for result ⚠️ Plausible pathway, prediction never tested — or tested and failed Ch 19 §19.5
⚠️ Denominator shopping ⚠️ Per kg vs per calorie vs per serving, chosen to win Ch 30 §30.5, Ch 36 §36.3b
⚠️ The unfalsifiable framework ⚠️ "You haven't eliminated enough yet" Ch 35 CS2
⚠️ Over-extension ⚠️ A real finding pushed past what it showed. The commonest myth shape Ch 17
⚠️ A resource problem called a knowledge problem ⚠️ "They just need educating" about people optimizing harder than the adviser Ch 32 §32.11
⚠️ A structural failure called a personal one ⚠️ "More discipline" applied to a plan, a budget or a cue Ch 31, 32, 33
⚠️ Suppression claims ⚠️ Compare against a DOCUMENTED episode (Ch 18 §18.12) before believing one Ch 17 §17.9

5. ⚠️ Reading a study without reading a study

⚠️ Six things to look at, in order, when you have five minutes and a paper.

⚠️ 1. The abstract's last sentence. ⚠️ Compare it to the results section. Overreach lives here.

⚠️ 2. Who was studied. ⚠️ Number, age, sex, health status, country. ⚠️ Then ask whether that is you.

⚠️ 3. How long. ⚠️ Nutrition outcomes take decades; most trials run weeks.

⚠️ 4. What was measured. ⚠️ A disease, a marker, or a self-reported questionnaire? ⚠️ Dietary intake is usually self-reported and self-reported intake is unreliable (Ch 3).

⚠️ 5. The comparison arm. ⚠️ Question 2, and it is where the effect size usually comes from.

⚠️ 6. Funding and conflicts. ⚠️ Last, not first (Ch 3). ⚠️ It is information, not a verdict.

⚠️ And one thing that is not on the list: the journal's name. ⚠️ Prestigious journals publish findings that fail to replicate at rates that would surprise you (Ch 3).


6. ⚠️ The self-diagnostic

From Chapter 33 §33.3, and it belongs in an evidence toolkit for a reason.

⚠️ When the question is about your own eating rather than about a claim:

⚠️ NOT "am I doing this right?"unanswerable in the moment, and it invites self-interrogation that helps nobody.

⚠️ INSTEAD: "WHAT STARTED THIS?"

⚠️ A CUE — a time, a place, a packet, a feeling, the end of something? ⚠️ An INTERVAL since eating? ⚠️ Or a CLAIM I read?

⚠️ Because the answer determines which tool works:

⚠️ A cue ⚠️ Change the situation. Not the food, and not your resolve (Ch 33 §33.4)
⚠️ An interval ⚠️ Fix the earlier meal (Ch 33 §33.3)
⚠️ A claim ⚠️ Run the six questions before changing anything

⚠️ Using the wrong tool is worse than doing nothing, and that is the whole of Chapter 33's threshold.


7. When to stop using this

⚠️ Two limits, both from Chapter 38 §38.9b.

⚠️ The filter is for claims that would change what YOU eat. ⚠️ Pointing it at other people converts a tool into a personality, and a person who has just learned to evaluate evidence is at their most insufferable.

⚠️ And the one exception, which is the only version that scales: ⚠️ if someone asks you how to tell whether a claim is worth acting on, teach them this appendix rather than giving them your answer.

⚠️ A filter can also become a wall. ⚠️ "Most of what gets sold to me about food is nonsense" is well supported by 136 ❌ and 🟠 verdicts in Appendix E — and is also the belief most likely to make you dismiss something true.


Where this came from

Chapter 2 ⚠️ The ladder, the placement test, healthy-user bias
Chapter 3 ⚠️ Funding, publication bias, self-reported intake
Chapter 17 ⚠️ Myth shapes and over-extension
Chapter 33 §33.3 ⚠️ The cue/interval self-diagnostic
Chapter 35 §35.11 ⚠️ Comparison against best available advice
Chapter 36 §36.3b ⚠️ The denominator problem
Chapter 37 §37.7 ⚠️ The six questions
Chapter 38 §38.7 ⚠️ The three additions for a future claim
Appendix E ⚠️ 468 claims already assessed — check here first