Case Study 23.1 — How to Evaluate a Literature You Cannot Fully Read
This is a methodological case study. It does not evaluate Semax or Selank; §23.2 and §23.3 already did that. It walks through the procedure a careful reader uses when the evidence for a compound sits behind a language, an indexing system, or a documentation practice they cannot get through — and it ends with that procedure written out in a form you can apply to any literature, in any field.
Read it as a transferable skill. The peptides are the occasion, not the subject.
The situation
You are trying to answer a question with a shape that recurs constantly outside this book: is there good evidence for X? — where the evidence for X, if it exists, is largely in a form you cannot directly assess.
Concretely, here is what a reader typically has at the start:
- A compound with a name, a proposed mechanism, and enthusiastic secondary coverage
- The knowledge that it is a registered medicine in one country and unapproved in others
- A handful of English-language items: a review or two, some translated abstracts, a great deal of forum discussion, and vendor pages citing sources they have plainly not read
- A strong intuition pulling in one direction, which varies by reader and which is the first thing that needs disarming
That last item deserves a moment. Before running any procedure, notice which way you are already leaning, because the procedure below is designed to be neutral and you are not. Readers who arrived excited about a compound will find the accessible evidence disappointing and will be tempted to fill the gap with the inaccessible evidence, imagined favorably. Readers who arrived skeptical will find the accessible evidence thin and will be tempted to treat the gap as empty. Both are filling the same blank with a prior.
What goes wrong, in two directions
Failure mode 1 — the parochial dismissal
The reader searches one database in one language, finds little, and concludes there is nothing.
The mechanics of this failure are worth stating precisely, because it does not feel like a failure while it is happening. It feels like diligence. The reader did search. The search returned few results. The conclusion — there is no good evidence — follows naturally and is stated with confidence.
What has actually happened is that a fact about a search has been converted into a fact about the world. The search covered the venues the reader's institution subscribes to, indexed by services that made commercial decisions about coverage, in a language the reader reads. Every one of those filters is real, none of them is a quality filter, and their combined effect is invisible in the result set. A database returns what it contains. It does not return a notice explaining what it does not contain.
Three tells that you have committed this error:
- Your conclusion would be unchanged if you learned that thirty relevant trials existed but were unindexed
- You cannot name a specific methodological problem with any specific study
- Your language slides from "I could not find" to "there is not"
Failure mode 2 — the credulous embrace
The reader learns that a substantial literature exists, that it supported a national registration, and that decades of clinical use stand behind it, and concludes that the compound is established.
This failure also feels like diligence, and like open-mindedness besides. The reader has resisted a parochial impulse. They have taken seriously work that others dismissed. This is genuinely admirable as a disposition and disastrous as a conclusion, because the reader has substituted the existence of a literature for the content of a literature.
Three tells:
- You can state how much evidence there is but not what any of it measured
- Your confidence rose when you learned the literature was large, before you learned anything about its design
- You find yourself explaining the absence of Western replication with a story about institutions rather than a hypothesis about funding, patents, or interest
Notice that these two failures are mirror images, and that each is available as a defense of the other. The parochial dismisser can point at the credulous embracer and say see, that is what taking it seriously leads to. The credulous embracer can point at the dismisser and say see, they will not even look. Neither observation is evidence about the compound.
What is actually knowable
Between "I can read everything" and "I can read nothing" there is a large middle region, and most real cases live in it. Here is what typically remains available.
Translated titles and abstracts. Frequently indexed even when full text is not. An abstract will not tell you whether an analysis was pre-specified, but it will often tell you the population, the design words used, and whether one primary endpoint is named or a list is offered.
Systematic reviews that searched broadly. The best ones report their search languages, their databases, what they identified, and what they could not obtain. A review that states "we identified n studies, of which k were available only in abstract" has measured the size of your gap for you. This is the single highest-yield item on the list, and it is routinely skipped by readers who go straight to primary sources they cannot read.
Registry records. International and national trial registries are public and searchable, and a registered protocol answers questions no abstract can. Absence from a registry is weakly informative for older studies — registration norms are recent everywhere — and much more informative for recent ones.
Regulatory documentation, where it is published. Some agencies publish substantial assessment reports; others publish little. This is a difference in document availability, and it should be recorded as such rather than converted into a judgment about the review that occurred.
Independent replication elsewhere. The most informative single item, and the most frequently over-read. Findings that generalize tend to travel. But the absence of independent replication is consistent with two very different worlds — one where the effect is not there, and one where nobody had a commercial reason to look — and the reader's job is to say so rather than to pick.
The characterizations in secondary sources. Even a review you consider unreliable will usually describe the underlying trials as small or large, single-center or multicenter, short or long. Those descriptors are hard to fabricate and easy to cross-check against each other. When three independent secondary sources all describe the primary literature as consisting of small single-center studies, you have learned something, even without reading one.
What is not knowable, and must be written down as such
The following cannot be recovered from abstracts, reviews, or registration status, and any confident statement about them is invention:
- Whether the reported analysis was the pre-specified one
- Whether blinding held, and whether anyone checked
- How many participants were lost, and what happened to them
- Whether adverse events were systematically collected
- Whether other trials by the same investigators went unreported
- The actual numbers, if only summaries are available
That is five or six of the most decision-relevant facts about any trial. A reader who cannot obtain them does not have a weak position on the compound. They have no position on those questions, and saying so is the correct output.
The distinction the whole procedure protects
One sentence, and everything above exists to defend it:
A limit on your ability to evaluate evidence is not a property of the evidence.
Run it in both directions, because it is only real if it costs you in both.
It means you cannot conclude the research is poor from I cannot assess the research. And it means you cannot conclude the research is sound from the research exists and I cannot assess it. The inaccessible literature is exactly as good as it is, and your inability to see it changes that number not at all — which is precisely why your rating has to describe your epistemic position honestly rather than pretending to describe the literature.
This is what the chapter means by an access-limited ⚠️. The symbol is the same one Cerebrolysin receives, and it means something different, and the four-line format exists so that the difference survives into your notes.
The procedure
Written out, so you can apply it to anything.
EVALUATING A LITERATURE YOU CANNOT FULLY READ — SEVEN STEPS
1 STATE THE CLAIM PROPERLY FIRST
Population + endpoint + timeframe. If the claim will not take that
shape (§23.1), stop — there is nothing to evaluate yet, and no
amount of literature will fix a defective claim.
2 NOTICE YOUR LEAN
Write one sentence: "Before looking, I expect ___, because ___."
Date it. You will check it in step 7.
3 SEARCH WIDER THAN YOUR DEFAULT
More than one database. Translated titles. National registries.
Systematic reviews that report their search languages. Record
WHAT YOU SEARCHED, not just what you found — the search is data.
4 APPLY THE SIX QUESTIONS, ALLOWING "?"
randomized · blinded and held · endpoint pre-specified · powered ·
population defined · reported regardless of outcome.
A "?" is a legitimate answer. Guessing is not.
5 SEPARATE THE TWO LEDGERS
Ledger A: what the accessible evidence shows.
Ledger B: what you could not assess, and the specific barrier
(language / indexing / full text / era / no registry).
Never let an entry migrate between ledgers silently.
6 RATE LEDGER A ONLY — AND LABEL THE RATING'S TYPE
Your rating describes what you can evaluate. Say so in the reason
line. "⚠️ — access-limited" and "⚠️ — conflicting evidence" are
different findings that share a symbol.
7 WRITE THE TWO TRIGGERS, AND CHECK YOUR LEAN
(a) What would clear the access flag? (a translation, a broad
systematic review, a registry entry — no new science needed)
(b) What would change the rating itself? (an independent
pre-registered trial — in either direction)
Then reread step 2. Did your conclusion land where your prior was?
If so, that is not proof you were wrong. It is a reason to check
your work once more.
Step 5 is where most readers fail, and the failure is almost always silent. An item from Ledger B — there are said to be many trials — drifts into Ledger A as there is substantial evidence, and no one, including the reader, notices the moment it crossed. Keeping them physically separate on the page is a crude fix that works.
Step 7(a) is the step that makes this a living entry rather than a verdict. An access flag can clear without any new science being done at all — someone translates a paper, a review widens its search languages, a registry posts a record. If you did not write the flag down, you will never notice when the thing that would resolve it arrives.
The output
Applied honestly, this procedure produces something that feels unsatisfying and is in fact a substantial result:
The accessible evidence for this claim is limited and does not meet the standard I apply elsewhere. A larger literature exists that I cannot assess; I have made no judgment about its quality, because I have not read it. My rating is ⚠️ and it is access-limited rather than conflict-limited. It would move on an independently conducted pre-registered trial in either direction, and the access flag would clear on a broad-search systematic review or on translations of the principal trials.
Read that paragraph again and notice how much it says. It states a position, distinguishes the position from a judgment it declines to make, identifies its own type, and names two separate falsifying conditions. Compare it to there's no real evidence and it's been used for decades in Russia, and it should be obvious which one a careful person wrote.
That paragraph is the deliverable. It is also, quietly, what expertise looks like from the inside: not a stronger opinion, but a better-specified one.
Discussion Questions
-
Failure mode 1 and failure mode 2 are described as mirror images. Are they equally common, in your experience? Are they equally costly? Make a case that one is worse than the other, then make the opposite case, then say which you actually believe and why.
-
Step 5 asks you to keep two ledgers and never let entries migrate silently. Describe a concrete situation — from this chapter or from anywhere else — in which you have seen a Ledger B item presented as a Ledger A item. What made the migration hard to notice?
-
The procedure treats "no independent replication elsewhere" as consistent with two very different worlds. Design a test, using only publicly available information, that would give you some traction on which world you are in. What would the test not be able to tell you?
-
Step 2 asks you to record your prior before looking, and step 7 asks you to check it. Some readers object that this invites you to discount conclusions you reached honestly, simply because they match what you expected. Is that objection right? What would a good response to it look like?
-
Consider a reader who follows this procedure carefully and arrives at "⚠️, access-limited" — and a reader who never looks at the literature at all and says "⚠️, probably not much there." They have written similar-looking conclusions. In what practical circumstances does the difference between them show up? Be specific.
-
The chapter insists that the standard in §23.9 is not culturally specific. Test that claim adversarially: is there anything in the seven-step procedure above, or in the six questions, that presupposes a particular research infrastructure, funding model, or publishing economy? If you find something, does it invalidate the standard or just complicate applying it?