Affiliate disclosure
Book titles on this page link to Amazon. As an Amazon Associate, DataField.Dev earns from qualifying purchases — at no additional cost to you.
Chapter 37 — Further Reading
This chapter is a table, so the useful further reading is not more tables. It is material that helps you maintain one: how evidence is graded, how ratings drift, how surrogate endpoints mislead, and where to look up a claim yourself when this snapshot has gone stale.
Everything below is a starting point, not a citation list for the ratings themselves. The reasoning behind each rating lives in its source chapter, and the full four-line callouts live in Appendix A.
Tier 1 — Start here
Evans, Thornton, Chalmers, and McPherson, Testing Treatments: Better Research for Better Healthcare. The single best plain-language introduction to why some evidence settles a question and some does not. Freely available online. If §37.3's distinction between evidence-absent and evidence-present-and-negative was new to you, this book is where to consolidate it, and it does so without requiring any statistics.
Ben Goldacre, Bad Science and Bad Pharma. The first is a general education in how claims go wrong in public; the second is specifically about how the evidence base for drugs is shaped by who funds trials and what gets published. Bad Pharma is the popular-level companion to §37.8's argument that the distribution of ratings is a fact about funding rather than about molecules. Read it critically — some of its specific reform proposals have since been implemented and others contested — but its central structural point stands.
Prasad and Cifu, Ending Medical Reversal. A book-length treatment of what §37.9 compresses into a paragraph: how treatments approved on surrogates and plausible mechanisms are later overturned by outcome trials. The nesiritide story in Case Study 37.1 is one instance of a pattern this book catalogues across medicine.
ClinicalTrials.gov and the WHO International Clinical Trials Registry Platform. Not reading, but the first place to go when you want to know whether a trial exists for a compound in this table. For an evidence-absent ❌, searching the registry takes about ninety seconds and tells you whether the row is about to move. This is the single most practical skill in this chapter.
Tier 2 — Going deeper
Guyatt, Rennie, Meade, and Cook, eds., Users' Guides to the Medical Literature. The standard manual for appraising a clinical study. Chapters 5 and 6 of this book are a compressed, peptide- specific version of what this volume does at length. Use it when you want to move from "this trial looks strong" to a defensible account of why.
The GRADE working group's methodology papers (Guyatt et al., Journal of Clinical Epidemiology, 2011 onward). A formal framework for rating certainty of evidence, with explicit criteria for downgrading and upgrading. This book's four tiers are a deliberately simpler instrument aimed at a general reader; GRADE is what a guideline committee actually uses. Comparing the two is instructive — in particular, GRADE's insistence that certainty and recommendation strength are separate things is the formal version of §37.8's safety callout.
Fleming and DeMets, "Surrogate End Points in Clinical Trials: Are We Being Misled?" Annals of Internal Medicine, 1996. Still the clearest short statement of the problem behind §37.9's claim that a ✅ resting on a surrogate is the least durable kind. Short, readable, and thirty years old without having aged.
Ioannidis, "Why Most Published Research Findings Are False," PLoS Medicine, 2005. The paper behind the base-rate reasoning in §37.9 and Chapter 9's phase 2 optimism. It is frequently over-quoted and under-read; the argument is narrower and more useful than its title suggests, and it explains why a ⚠️ resting on one early positive study is more likely to shrink than to grow.
Cochrane Library systematic reviews, and the Cochrane Handbook for Systematic Reviews of Interventions. When you want to know the current state of a question rather than the state of one trial. Several rows in this table would be rated by consulting a Cochrane review first and the primary trials second.
Regulatory assessment documents: FDA approval packages and EMA European Public Assessment Reports (EPARs). These are the most underused public documents in medicine. They contain the trial data a sponsor submitted, the reviewers' objections, and the exact population an approval covers — which is the population field in every ✅ row of the master table. Reading one EPAR end to end teaches more about what "approved" means than any secondary source.
Tier 3 — Primary sources and the raw material
The trial programs named in the source chapters. STEP and SUSTAIN and SELECT for semaglutide; SURPASS for tirzepatide; PROMID and CLARINET for somatostatin analogs; NETTER-1 for lutetium Lu 177 dotatate; PARADIGM-HF for sacubitril/valsartan. Every ✅ in this table has a program behind it, named in the chapter that issued the rating. Reading a single pivotal trial in full — methods first, results second, discussion last — is the highest-yield exercise available to a reader of this book.
PubMed, with a deliberate search strategy. The useful discipline is to search for the absence of evidence as carefully as its presence: filter to randomized controlled trials, filter to human studies, and notice how many Part III compounds return nothing. That empty result set is the ❌ in §37.3's left-hand column, seen directly.
The WADA Prohibited List and its S0 category. Worth reading once in the original, because Chapter 38's claim-form rating of "it isn't on the banned list" depends on a structural feature of the document that summaries omit: S0 prohibits any pharmacological substance not currently approved by a governmental health authority for human therapeutic use, which means unapproved compounds are prohibited by default rather than by enumeration.
Pharmacopeial monographs (USP, Ph. Eur.) for peptides that have them. The concrete answer to what "pharmaceutical grade" would have to mean, and therefore to why the phrase is ❌ when applied to a research-chemical product. A monograph specifies identity, assay, impurity limits, and test methods. Compare one against a vendor's certificate of analysis and the gap becomes self-explanatory.
Retraction Watch and PubPeer. For the uncomfortable but necessary habit of checking whether a paper you are relying on has been questioned, corrected, or withdrawn. This matters most for compounds whose preclinical literature comes from a small number of groups.
If you only do one thing
Take the three rows in this table you care about most, and look each one up yourself in ClinicalTrials.gov.
Not to check whether the rating is correct — to see, directly and in about five minutes, the difference between a compound with a completed randomized program, a compound with a single early-phase study, and a compound with nothing at all. The three states look completely different on the screen, and once you have seen them you will never again need this table to tell them apart.
That is the whole point of Chapter 37: not the list, but the ability to rebuild any row of it yourself when the list has gone out of date.