Case Study 1 — Auditing a Finished Dossier
[constructed teaching example] The dossier below is invented for teaching. It is not a real person's document and it does not represent anyone's actual conclusions. What is not invented is the pattern of defects: these three are the ones readers produce most reliably, and they are planted here deliberately so you can practice finding them somewhere other than your own file, where finding them is harder.
The dossier belongs to a reader who did the work. All four entries are filled. Nothing is missing, nothing is lazy, and the ratings are not wild. It would pass a casual read, and it fails the audit in three places.
Read the dossier first. Try to find the defects before the walkthrough. Then check yourself against §40.3's four checks, run in order.
The dossier, as submitted
DOSSIER — four entries Last updated: 2026-07
================================================================
ENTRY 1 — SEMAGLUTIDE
1 IDENTITY GLP-1 receptor agonist; 31-residue backbone based on human
GLP-1 with three modifications. Trade names differ by indication.
Same name = same molecule? Yes from a pharmacy; not reliably for
compounded or gray-market material.
2 ORIGIN Analog of an endogenous human hormone. Class origin:
exenatide, from Gila monster venom. No evidentiary weight.
3 MECHANISM Activates the GLP-1 receptor. Direct agonist. Established.
4 PHARM Subcutaneous weekly; an oral formulation exists at very low
bioavailability. Route used matches route studied: YES.
5 EVIDENCE Large randomized trials in defined populations, replicated
across a trial program; one large cardiovascular outcome trial with a
hard endpoint. Conspicuously absent: decade-scale data.
6 RATING
Weight loss ✅ 2026-03
Glycemic control, adults with type 2 diabetes,
HbA1c ✅ 2026-03
MACE reduction, adults with established CVD and
overweight/obesity without diabetes, MACE ✅ 2026-03
Alzheimer's disease 🔬 2026-03
7 APPROVED Approved in many jurisdictions; indications and populations
differ by product and country.
8 CLAIMED Weight loss and diabetes (matches 7); informally, a long
list of other effects (does not). Overlap: PARTIAL.
9 RISKS GI effects common at studied use; labeled warnings apply.
Unstudied use: populations not represented in trials. Preparation
quality: substantial for compounded and gray-market material.
Unknown because nobody has looked: decade-scale outcomes.
10 STATUS Regulatory: approved, varies. Sport: check current WADA
List. Sold as: prescription; also compounded and gray-market.
Access/cost: contested. Quality: establishable for pharmacy product.
11 VERDICT Well evidenced for two specific claims, one a hard outcome.
Everything beyond those is a separate question. Confidence: HIGH.
As a question: "The cardiovascular result was in people with
established heart disease and no diabetes — does that include me?"
12 CHANGE MY MIND DOWN: long-term safety signal in post-marketing
surveillance; failed replication of the outcome trial. UP for an
unrated indication: a powered RCT with a hard outcome in that
population. Alzheimer's would be settled by an RCT in early disease,
cognitive and functional endpoints, placebo, multi-year, several
thousand participants. Underway? Yes.
ENTRY 2 — BPC-157
1 IDENTITY 15-residue synthetic peptide, GEPPPGKPADDAGLV. Laboratory
code, no generic name. Same name = same molecule? NO for
research-labeled material.
2 ORIGIN Described as derived from a sequence found in human gastric
juice. No evidentiary weight.
3 MECHANISM Various proposed; angiogenic and cytoprotective effects
described in animal models. Confidence in mechanism: PROPOSED.
4 PHARM Injectable in nearly all animal work. Sold orally, as a
nasal spray, and as a cream. Route used matches route studied: NO.
5 EVIDENCE Extensive rodent literature across many injury models. No
completed, peer-reviewed randomized human trial in the published
literature for tendon or soft-tissue healing. Replication: extensive
in animals, none in humans. Conspicuously absent: any human RCT.
6 RATING
Tendon and soft-tissue healing, adults, functional
recovery ❌ 2026-04
kind: evidence ABSENT
[margin note: the marketing around this is shameless — every clinic
site makes claims no one could support. Rated accordingly.]
7 APPROVED Not approved in any major jurisdiction. Which meaning?
Never submitted for the marketed indications.
8 CLAIMED Tendon healing, gut healing, joint repair, recovery from
almost any soft-tissue injury. Overlap with 7: NONE.
9 RISKS At studied use: no studied human use to speak of.
Unstudied use: everything. Preparation quality: substantial —
research-labeled material, no verified identity or concentration.
Unknown because nobody has looked: essentially all of it.
10 STATUS Regulatory: not approved; restricted from compounding in
some jurisdictions. Sport: prohibited under anti-doping rules —
verify against the current List. Sold as: research chemical.
Access: online. Quality: not establishable.
11 VERDICT No human evidence. Would not consider. Confidence: HIGH.
As a question: "Is there any human trial evidence for this?"
12 CHANGE MY MIND [blank]
ENTRY 3 — TB-500
1 IDENTITY Marketed as a fragment related to thymosin beta-4.
Laboratory code, no generic name. Same name = same molecule?
NO — sourcing unverifiable, and "TB-500" and "thymosin beta-4" are
used interchangeably by sellers although they are not the same thing.
2 ORIGIN Fragment of an endogenous human protein.
3 MECHANISM Actin-binding; proposed roles in cell migration and tissue
repair. Confidence: PROPOSED. [margin note: honestly a very elegant
mechanism — this one seems much more likely to be real]
4 PHARM Injectable in animal work. Route used matches route
studied: unclear.
5 EVIDENCE Animal work in several injury models. No completed
randomized human trial for soft-tissue or tendon healing.
Replication: animal only. Conspicuously absent: any human RCT.
6 RATING
Soft-tissue and tendon healing, adults, functional
recovery ⚠️ 2026-04
[margin note: training partner used it after a shoulder injury and
recovered well. Anecdote, obviously, but the mechanism fits.]
7 APPROVED Not approved. Which meaning? Never submitted.
8 CLAIMED Tissue repair, recovery, flexibility, healing.
Overlap with 7: NONE.
9 RISKS At studied use: no studied human use. Unstudied use:
everything. Preparation quality: substantial. Unknown because nobody
has looked: essentially all of it.
10 STATUS Regulatory: not approved. Sport: prohibited — verify
against the current List. Sold as: research chemical. Quality: not
establishable.
11 VERDICT Promising. Would want to see human data. Confidence: LOW.
As a question: "Has anyone run a human trial on this?"
12 CHANGE MY MIND UP: a human RCT showing benefit. DOWN: a human RCT
showing none.
ENTRY 4 — GHK-Cu (COPPER TRIPEPTIDE)
1 IDENTITY Copper-binding tripeptide, GHK, complexed with copper.
In cosmetics, an INCI-named ingredient. Same name = same molecule?
Formulation-dependent: concentration and vehicle vary widely and are
usually undisclosed.
2 ORIGIN Endogenous human peptide, identified in plasma.
No evidentiary weight.
3 MECHANISM Copper delivery and effects on extracellular matrix
signaling described in cell and tissue work. Confidence: PROPOSED
for the cosmetic claims.
4 PHARM Topical. Penetration of intact skin by a charged peptide
complex is the central question and is formulation-dependent.
Route used matches route studied: partially — studied topically, but
with formulations that are not the ones sold.
5 EVIDENCE Small, short, mostly industry-associated topical studies
with appearance endpoints; heterogeneous formulations; limited
independent replication. Conspicuously absent: adequately powered
independent trials with pre-specified objective endpoints.
6 RATING
Skin appearance, adults using a topical formulation,
investigator-rated appearance over weeks ⚠️ 2026-05
Wound healing in a clinical setting 🔬 2026-05
Hair growth in adults, measured hair density ❌ 2026-05
kind: evidence ABSENT for consumer formulations
7 APPROVED Cosmetic ingredient, not a drug. No efficacy approval
exists or is required for cosmetic appearance claims.
8 CLAIMED Firming, wrinkle reduction, repair, regeneration, hair
growth. Overlap with 7: NONE — cosmetic marketing is not an
approved indication.
9 RISKS Topical irritation; staining. Unstudied use: prolonged
daily use. Preparation quality: concentration undisclosed in most
consumer products. Unknown because nobody has looked: long-term.
10 STATUS Regulatory: cosmetic. Sport: not applicable. Sold as:
serums across a wide price range. Access: retail. Quality: label
rarely states concentration; not independently verifiable.
11 VERDICT Plausible modest cosmetic effect; the specific product I
own is not the formulation that was studied. Confidence: LOW.
As a question: "Is there a reason to prefer one of these over a
cheaper product with a longer track record?"
12 CHANGE MY MIND UP: an independent, adequately powered, randomized
trial of a disclosed consumer-strength formulation, with objective
endpoints over at least twelve weeks. DOWN: the same trial, null.
Underway? Not that I can find.
================================================================
That is a serious document. The reader read the book. Now audit it.
The audit
Check 3 first — the date check (passes)
Run it first because it is the cheapest, and because when it passes it tells you something about the other checks.
Every rating row in all four entries carries a date, and the dates are at month resolution rather than year. The document itself is dated. This check passes cleanly, and the pass is not trivial: it means the entries can be placed on a timeline, which is what §40.4 needs to look for direction and what §40.5 needs to build a version history.
Note also what a clean date check does not buy you. Three of these entries are defective, and every one of the defects sits inside a properly dated row. Dating is necessary and nowhere near sufficient, and readers who complete it sometimes conclude they have audited the file.
Check 2 — the population check (defect 1)
Read every Field 6 row and ask whether it names a population and an endpoint.
Entry 1 has four rows. Three of them are exemplary: glycemic control, adults with type 2 diabetes, HbA1c; MACE reduction, adults with established CVD and overweight/obesity without diabetes, MACE; and the Alzheimer's row, which is honestly 🔬 and refers to a defined disease.
The first row reads, in its entirety:
Weight loss ✅ 2026-03
Defect 1: an unpopulated rating.
This is not a quibble, and the surrounding rows prove it. The reader clearly knows how to write a populated row — they wrote three of them, and one of the three specifies a population with three qualifiers. They did not write the fourth because the weight-loss claim is the one they consider settled, and settled claims are the ones people stop specifying.
That is exactly backwards, and the entry's own Field 5 shows why. The book's evidence on this compound includes the same drug at the same dose producing materially different weight results in adults with overweight or obesity without diabetes versus adults with type 2 diabetes. A row reading "Weight loss ✅" points at both results and therefore at neither. Nothing in the entry recovers which one the reader meant.
Diagnosis: the defect tracks confidence, not carelessness. The most-certain row is the least specified one. Look for this pattern in your own file — your unpopulated rows will cluster on the claims you think are obvious.
Fix: name the population and endpoint the way the neighboring rows already do, and note that a second row may be needed, because there are two populations with two different results.
Check 1 — the same-evidence test (defect 2)
Find two entries with comparable evidence and different ratings.
Entries 2 and 3 are close to a controlled experiment. Line up their Field 5 entries:
| BPC-157 (Entry 2) | TB-500 (Entry 3) | |
|---|---|---|
| Human RCT for the claimed indication | none | none |
| Animal work | extensive, many models | present, several models |
| Replication in humans | none | none |
| Conspicuously absent | any human RCT | any human RCT |
| Route: used vs. studied | mismatched | unclear |
| Preparation quality | not establishable | not establishable |
| Field 7 overlap with Field 8 | none | none |
| Rating | ❌ | ⚠️ |
Defect 2: comparable evidence, different ratings, and the difference is not justified by design, population, or endpoint.
Both entries describe animal evidence and no completed human trial. Under the rating system, ⚠️ means promising but preliminary — real human data that does not settle the question. Entry 3 records no human data at all. By its own Field 5, Entry 3 cannot be ⚠️. It is ❌, kind: evidence absent — the same rating and the same kind as Entry 2.
The reader left the reasons in the margins, which is what makes this a teaching example rather than a mystery:
- Entry 2's margin: "the marketing around this is shameless... Rated accordingly." That is a rating moved by distaste. Rating rule 4.
- Entry 3's margins: "honestly a very elegant mechanism" and "training partner... recovered well." That is a rating moved by mechanism plausibility and by a single anecdote. Rating rule 3, plus Chapter 6's testimonial dynamic operating on someone the reader knows.
Note what this means: the ❌ on Entry 2 is a correct rating reached by a method that does not work. It happens to match what the evidence supports. Had the marketing been tasteful, the same method would have produced ⚠️ on identical evidence — which is precisely what it did one entry later. Being right for bad reasons is invisible from the inside, and the same-evidence test is what makes it visible.
Diagnosis: the two ratings differ by the reader's personal proximity to the compound, not by evidence. This is the more interesting version of §40.4's question, and it is worth naming precisely rather than filing under "generous" or "harsh": this reader is harsh toward what is sold to strangers and generous toward what is used by people they know.
Fix: move Entry 3 to ❌, evidence absent, dated. Move both margin notes out of Field 6 — the marketing observation belongs in Field 8, the training partner's experience belongs nowhere in a rating and can be recorded as a note in Field 5 explicitly labeled as an anecdote. Then reread Entry 2's ❌ and confirm you can defend it from Field 5 alone.
Check 4 — the Field 12 check (defect 3)
Is every entry falsifiable?
Entry 1: specific and strong. It names the readout that would move an unrated indication, specifies the Alzheimer's trial by population, endpoint, comparator, duration, and rough size, and records that such work is underway. Passes.
Entry 4: passes, and passes well — an independent, adequately powered randomized trial of a disclosed consumer-strength formulation, objective endpoints, minimum duration, and a note that nothing is underway. A reasonable person could not disagree about whether that had arrived.
Entry 3: weak but present. "UP: a human RCT showing benefit." This fails the §40.3 reasonable-person test — no population, no endpoint, no comparator, no duration, no size — and would be satisfied by a small unblinded study in an unrelated population. Flag it and sharpen it.
Entry 2: blank.
Defect 3: an entry whose Field 12 is empty — and it is the entry the reader is most confident about.
The Field 11 verdict reads "No human evidence. Would not consider. Confidence: HIGH." Confidence high, falsifiability zero. Whatever else it is, that is not a conclusion. Per §40.3, it is a belief.
And this is the defect readers resist hardest, because Entry 2's rating is correct. The evidence genuinely is absent. So the entry produces the right answer and the reader concludes the check is pedantry.
It is not, for two reasons the entry itself supplies.
First, the margin note already told us how this rating was reached. A rating reached by distaste and left unfalsifiable will not update when the evidence does. If a well-designed human trial reports next year, this reader has not written down what such a trial would have to look like to move them — and will therefore evaluate it after the fact, while holding a stated high-confidence position. That is the worst possible order.
Second, note the asymmetry across the file: the entries with real Field 12s are the ones the reader is uncertain about, and the blank one is the entry they are sure of. Field 12 gets written when it feels needed and skipped when it feels obvious, which inverts its purpose. It is most valuable exactly where it feels least necessary.
Fix: write it. Something like: UP — a randomized controlled trial in adults with a defined soft-tissue or tendon injury, functional recovery as the primary endpoint, placebo or standard-care comparator, at least three months of follow-up, adequately powered, registered before enrollment. DOWN — the same trial, reported null, which would move this from evidence-absent to evidence-present-and-negative, a stronger state of knowledge than the current one. Underway? Not that I can find.
What the audit found
Three defects, one clean check, and a pattern.
| Check | Result |
|---|---|
| Date check | Passes. Every row dated to the month. |
| Population check | Defect 1 — Entry 1's most confident row names no population or endpoint. |
| Same-evidence test | Defect 2 — Entries 2 and 3 have comparable evidence and different ratings, driven by distaste and by mechanism-plus-anecdote. |
| Field 12 check | Defect 3 — Entry 2, the highest-confidence entry in the file, is unfalsifiable. |
Now the pattern, which is worth more than the three fixes.
Every defect sits on the reader's most confident material. The unpopulated row is the claim they considered settled. The unjustifiable rating pair is the one where they had feelings. The blank Field 12 is the entry they were surest about. And the entry that passes every check cleanly — Entry 4 — is the one about a serum, where the reader had the least at stake.
That is the finding to carry into your own audit. The parts of your dossier that need auditing are not the parts you found difficult. They are the parts you found easy.
Questions
-
Entry 1's first Field 6 row is unpopulated while the three rows beneath it are exemplary. Using the entry's own Field 5, write the corrected row — or rows, if you think more than one is required. Justify your choice of how many.
-
The audit concludes that Entry 2's ❌ is the correct rating reached by a method that does not work. Explain how the same-evidence test made that visible when reading Entry 2 alone could not have. What would the reader have concluded if they had checked only their ratings against Chapter 37?
-
Entry 3 is rated ⚠️ on a Field 5 that records no human data. State the definition of ⚠️ from the rating system and show precisely where the entry contradicts it. Which rating rule does the fix invoke?
-
The reader's two margin notes break two different rating rules in opposite directions. Identify each rule, and explain why §40.4 insists that neither error is the more careful one.
-
A reader objects: "Entry 2's rating is right, so demanding a Field 12 is bureaucracy." Answer the objection using two specific consequences described in the walkthrough. Then say what the objection reveals about how the objector thinks ratings are produced.
-
Entry 4 passes all four checks and is about the compound the reader cares least about. Give the most plausible explanation for that relationship, and describe one concrete procedure you could add to your own audit to counteract it.