Appendix M — Answers to Selected Exercises
Worked solutions to the daggered (†) and odd-numbered exercises from every chapter. Where an exercise asks for a judgment rather than a fact, the answer given is a model response with its reasoning shown — not the only defensible answer. In a book about weighing evidence, an answer key that pretends every question has one right answer would be teaching the opposite of the lesson.
Chapter 1
Exercise 1.1
The general structure: a central alpha carbon bearing four groups — an amino group ($-NH_2$), a carboxyl group ($-COOH$), a hydrogen atom, and the R group (side chain). The first three are identical in all twenty amino acids. The R group is the only part that differs, and it determines whether the residue is hydrophobic, polar, acidic, or basic; large or small; rigid or flexible.
Exercise 1.2 †
The peptide bond forms between the carboxyl group of one amino acid and the amino group of the next. A molecule of water is released — this is a condensation reaction. The reverse, in which water is added back to break the bond, is hydrolysis, and it is what digestive proteases do.
Exercise 1.3
A residue is an amino acid after it has been incorporated into a chain. The word is used because in forming the bond, the amino acid gave up the atoms that became water — what is present in the chain is the remainder of the original molecule, not the whole of it. "A 15-residue peptide" and "a 15-amino-acid peptide" mean the same thing.
Exercise 1.5 †
The conventional boundary is approximately 50 residues — peptides below, proteins above.
It is a convention because no chemical property changes at residue 51. The bonds are the same, the amino acids are the same, and the folding principles are the same. Different sources place the line at 30, 40, 50, or 100. Insulin at 51 residues is described in the literature as both a peptide hormone and a small protein, and both descriptions are correct. The line is drawn for human convenience, and the fact that it is arbitrary is what allows it to be exploited rhetorically in both directions (§1.5).
Exercise 1.7
A disulfide bond is a covalent link between the sulfur atoms of two cysteine residues. It is the strongest of the interactions that hold peptide shapes together and can join separate chains. Insulin depends on them: two disulfide bonds join its A and B chains, and a third forms within the A chain. Without them insulin is not insulin — this is exactly why the recombinant production of insulin required solving a folding problem, not just a synthesis problem (Chapter 32).
Exercise 1.8 †
Proline's side chain is unusual: it loops back and bonds to the backbone nitrogen, forming a ring. That nitrogen therefore has no hydrogen available to donate to the hydrogen-bonding pattern that holds an alpha helix together, and the ring also restricts backbone rotation at that position. The result is that a proline in the middle of a helix breaks it. Proline functions as a structural comma — which is why a peptide like BPC-157, with four prolines including a run of three, has a rigid non-helical region (§1.4).
Exercise 1.9
- Aspirin: ~180 Da
- Insulin: ~5,800 Da
- Therapeutic monoclonal antibody (e.g. trastuzumab): ~148,000 Da
Roughly: a peptide is about thirty times an aspirin, and an antibody about thirty times an insulin.
Exercise 1.11 †
Three questions, each with what a satisfactory answer looks like:
1. What protects the peptide from digestion? A satisfactory answer names a specific mechanism — a permeation enhancer, an enteric coating, a stabilizing chemical modification, encapsulation — and explains why it works. An unsatisfactory answer is "it's absorbed sublingually so it bypasses the stomach," which is an assertion, not a mechanism.
2. What is the molecule's size, and is it small enough to cross the oral or sublingual mucosa? Sublingual absorption is genuinely possible for small molecules; for a peptide of a few thousand daltons it is a serious physical obstacle. A satisfactory answer gives the molecular weight and cites data. Most products cannot give the molecular weight at all.
3. What bioavailability has been measured, in whom, by what method? The gold standard answer is a pharmacokinetic study showing measured blood levels after oral versus injected administration. Note what the honest answer looks like for a genuine oral peptide: oral semaglutide's answer is "approximately 1%, established in a formal program, and here is the fasting protocol required to achieve even that." A product that cannot answer at all has not tested it.
Note that none of these questions requires knowing anything about whether the peptide works. They are chemistry questions, and they eliminate most oral peptide claims before efficacy arises.
Exercise 1.13
They are different molecules. Composition is not sequence. Gly-Ala-Ser and Ser-Ala-Gly
contain identical amino acids in identical proportions and are distinct compounds with distinct
primary structures — and therefore, potentially, entirely different secondary and tertiary structures
and entirely different biological activity.
This is precisely the point Sanger's work established (Case Study 1): a protein is not defined by what it is made of but by the order it is made in. It is also why "amino acid analysis" — which measures composition — is a weaker identity test than sequencing or mass spectrometry, a distinction that matters in Chapter 34.
Exercise 1.14 †
Why it might matter enormously: if the omitted residue is a receptor contact point, activity may be abolished. If it is structurally critical — a cysteine that forms a disulfide, a proline that enforces a turn — the molecule may misfold and be inactive or aggregate. And a 29-residue impurity in a 30-residue product is nearly impossible to remove by conventional purification, because the two molecules have almost identical physical properties. A product could be "98% pure" by one measure and contain several percent of a closely related inactive species.
Why it might not matter: if the residue is in a flexible terminal region not involved in binding or folding, the truncated molecule may retain full or partial activity. Some peptides tolerate substantial terminal truncation. In addition, if the deletion produces a completely inactive molecule that is also non-toxic, the practical consequence is a modest loss of potency — the dose is slightly weaker than labeled, which matters but does not harm.
What you would need to know: which residue was omitted and what it does; whether the truncated species is inactive, partially active, or antagonistic; whether it is immunogenic; and what fraction of the product it represents. The last is a quality-control question (Chapter 34), and it is the one most often unanswerable for gray-market material.
Exercise 1.17 †
A model answer, four sentences:
A peptide is a short chain of building blocks called amino acids — the same things proteins are made of, just far fewer of them. Your body makes hundreds of these naturally and uses them as messages: one tells your pancreas to release insulin, another tells your brain you've eaten enough. Drugs like Ozempic are laboratory-modified copies of one of those natural messages, engineered so they last about a week in your body instead of two minutes. But "peptide" only describes the kind of molecule — it says nothing about whether any particular one works, which is a completely separate question with a completely separate answer for each one.
Exercise 1.20 †
A model answer that does the hard part — being honest about the limit of what chemistry settles:
Here's what I can tell you and here's where I run out. Collagen is a protein, and hydrolyzed collagen has already been chopped up. When you swallow it, your digestion breaks it down further into amino acids, which go into a general pool your body draws on for whatever it's building. So the marketing picture — collagen goes in, collagen shows up in your skin — isn't what happens. That part I'm confident about.
Where I'd stop short of calling it useless: some very small fragments do get absorbed intact, and there's a real hypothesis that they might act as signals rather than as building material. There are also human trials reporting modest skin improvements. Most of them are small, short, and funded by people selling collagen, which is a reason for caution but not for dismissal.
So: the strong claim is wrong, the weak claim is unresolved, and if there's an effect it's probably small. Whether \$120 is worth a possible small effect is a values question, not a science question, and I genuinely don't have a strong view on it.
Note what this answer does: it separates the settled part from the unsettled part, it concedes what is real in the claim before criticizing it, and it declines to convert an evidence question into a verdict on the person asking.
Exercise 1.22 †
What Field 1 requires and this product does not supply: the actual name of the peptide; its length; its sequence; its molecular weight; whether a generic or research-code name exists; and whether the sequence is publicly documented anywhere.
What "proprietary peptide complex" does supply: that there may be more than one peptide; and that the seller has chosen not to disclose which.
What the omission suggests. Proprietary blends are legal in several product categories and are sometimes used to protect a genuine formulation. But note the asymmetry: a manufacturer with a well-characterized, high-quality peptide has every commercial incentive to name it, because naming it invites comparison they would win. Non-disclosure is most useful when disclosure would not help.
What it does not prove. That the product is ineffective, or that the contents are not what is implied. It establishes that the claim is unverifiable, which is a different and — for the purposes of this book — more important finding. An unverifiable claim cannot be rated at all; it fails before the evidence question is reached.
Exercise 1.25 †
-relin denotes an agonist at a releasing-hormone receptor — it stimulates. -relix denotes an
antagonist — it blocks.
The clinical consequence appears in hormone-sensitive cancer treatment. A GnRH agonist (such as leuprorelin) initially stimulates the receptor before continuous stimulation causes downregulation and suppression. That produces a transient surge in testosterone — the "flare" — which in prostate cancer can briefly worsen symptoms before the therapeutic suppression takes hold, and which sometimes requires additional medication to cover. A GnRH antagonist (such as degarelix) blocks the receptor immediately and produces suppression without a flare.
Same axis, same target, opposite initial effect, one letter of difference in the name. Chapter 27 covers this in full.
Exercise 1.27
What it suggests: generic names are assigned by international nomenclature authorities when a compound enters serious clinical development. A compound still identified only by a laboratory code decades after its discovery almost certainly never went through that process — meaning no completed regulatory development program, and by extension no completed pivotal trials.
What it does not prove: that the compound is inactive, unsafe, or was rejected by a regulator. Molecules go undeveloped for many reasons that have nothing to do with efficacy — unpatentability, a company failing, a therapeutic area falling out of favor, or simply nobody funding it. The naming pattern is strong evidence about regulatory history and no evidence about pharmacology, and conflating the two is an error in the skeptical direction that this book is as concerned about as the credulous one.
Exercise 1.28 †
The difference lives at primary structure, but its consequences propagate to tertiary structure and to receptor binding.
Two residues out of nine is over twenty percent of the molecule. In a peptide this short, essentially every residue is either making direct contact with the receptor or holding the shape that positions the others. Changing two can alter the surface the receptor sees enough to change which receptor is preferred — oxytocin and vasopressin act on related but distinct receptor families, and the selectivity between them is what separates uterine contraction from water retention.
What this implies about synthesis: enormously high stakes for fidelity. A synthesis process that occasionally substitutes or omits a residue is not producing a slightly weaker version of the target molecule; it may be producing a molecule with different receptor selectivity — potentially an active compound with a different physiological effect. This is the strongest single argument for rigorous identity testing (Chapter 34), and it applies with much greater force to short peptides than to long ones.
Exercise 1.30
Longest to shortest expected duration in the bloodstream:
- Therapeutic antibody — weeks. Too large for renal filtration, and antibodies are recycled by a dedicated salvage pathway.
- Semaglutide — about a week. Engineered specifically for this: DPP-4 resistance plus albumin binding via a fatty acid chain.
- Aspirin — hours for the parent molecule (its effect on platelets lasts far longer, but that is a different question — the drug is gone while the consequence persists).
- Native GLP-1 — one to two minutes. Cleaved by DPP-4 essentially on arrival.
Chapter 1 concepts sufficient to reason it out: size governs renal clearance; peptide bonds are enzymatically vulnerable; and chemical modification exists precisely to defeat both. The aspirin placement is the interesting one, because it separates drug duration from effect duration — a distinction Chapter 4 develops.
Exercise 1.32 †
The category does real work, just not evidentiary work. Three things it genuinely tells you:
Delivery constraints. Knowing something is a peptide tells you it is probably injected, probably has a short native half-life, probably cannot enter cells, and probably cannot cross into the brain unaided. That is a substantial amount of actionable information, and §1.1 showed it can eliminate a claim before any evidence is consulted.
Manufacturing and quality profile. Peptides are synthesized or expressed by specific methods with specific characteristic failure modes — deletion sequences, incomplete deprotection, residual solvents. Knowing the category tells you what to test for, which is Chapter 34's entire subject.
Metabolic fate. Peptides are broken down into amino acids rather than into novel chemical species. This is a genuine and underappreciated class-level safety property, and it distinguishes peptides from many small molecules.
The correct formulation is therefore not "the category is meaningless" but: the category predicts chemistry, delivery, and manufacturing — and predicts nothing about efficacy. Someone who abandons the word entirely loses three useful predictions in order to avoid one misuse.
Exercise 1.34 †
There is no single correct answer, but the exercise has a right finding, and most readers reach it: some of your peptides will be hard to pin down, and the difficulty is itself data.
Expect this pattern. Approved drugs (insulin, semaglutide, teriparatide) resolve in under a minute: generic name, brand names, sequence in UniProt, molecular weight, full label on DailyMed. Research-code compounds (BPC-157, TB-500) mostly resolve for identity — the sequences are published — but you will find inconsistencies in how vendors describe them, and TB-500 in particular is frequently described as "thymosin beta-4" when it is marketed as a fragment. Cosmetic ingredients (GHK-Cu, "acetyl hexapeptide-8") resolve to an INCI name and a length, and often no further.
What that suggests: the ease of completing Field 1 correlates strongly with how much formal development a compound has undergone. That correlation is not a rating and should not be treated as one — but it is a legitimate prior, and noticing it in your own dossier before Chapter 5 formalizes anything is the point of doing Field 1 first.
Chapter 2
Exercise 2.2 †
Two reasons lock-and-key fails:
Neither partner is rigid. Both the peptide and the receptor are flexible molecules. On approach they make loose contact and then both adjust conformation to improve the fit. The final bound shape may not have existed in either partner beforehand. This is why many short peptides are largely random coil in solution (Ch 1 §1.4) and only take a defined shape on binding — the shape is a property of the pair, not of the peptide.
Binding is reversible and dynamic. A key stays in a lock. A ligand binds, stays for some period, and dissociates. At any moment a fraction of receptors is occupied, and the occupying molecules are continuously exchanging. Raise concentration and the occupied fraction rises. Nothing is ever fully on.
(A third, worth crediting: fitting is not activating. An antagonist fits perfectly and does nothing.)
Exercise 2.5 †
The four mechanisms:
- Degrade the peptide — proteases and peptidases cleave the ligand. DPP-4 does this to GLP-1 within a minute or two.
- Clear it — the kidney filters small peptides out of blood.
- Desensitize the receptor — phosphorylation and beta-arrestin recruitment uncouple the receptor from its G protein within minutes. The ligand may still be bound.
- Downregulate — the cell internalizes receptors and reduces new receptor synthesis over hours to days.
Engineerable: 1 and 2. They are properties of the drug molecule and can be defeated by chemical modification — protease-resistant substitutions, albumin binding, cyclization, increased effective size. Semaglutide's ~7-day half-life is a total victory over both.
Not engineerable: 3 and 4. They are decisions made by the target cell in response to being stimulated. No modification of the drug persuades a cell to keep listening — and a longer-lasting agonist pushes harder on the machinery that shuts the signal down.
Exercise 2.10 †
A partial agonist activates a receptor to some submaximal ceiling — say 40% of maximum — regardless of dose.
Alone: it raises signaling from baseline toward 40%. It is functioning as an activator.
With a full agonist present: it competes for the same binding sites. Every receptor it occupies is a receptor the full agonist cannot occupy, and it activates that receptor only to 40%. Average activation across the receptor population therefore falls. It is functioning as an inhibitor.
What determines which: the level of activation in the system before the partial agonist arrives. Below its ceiling, it pushes up. Above its ceiling, it pulls down. The molecule did not change; the context did. This is a general and underappreciated point — several drug classes exploit it deliberately.
Exercise 2.12 †
Potency differs; efficacy is the same.
Both reach 100% of the maximum effect, so their ceilings are identical — same efficacy. Compound A achieves it at one-tenth the concentration, so its dose-response curve sits further left — greater potency.
On a dose-response plot: same height, different horizontal position. Note that potency alone says little about clinical value; a less potent drug given at a higher dose may be entirely equivalent in practice, and half-life or route may matter far more than either property.
Exercise 2.15 †
Three questions the affinity claim does not answer:
1. What is the efficacy? Affinity is how tightly it binds; efficacy is how much effect it can produce. A compound can bind thirty times more tightly and produce less effect — that describes a high-affinity partial agonist, or in the limit an antagonist. Binding tightly to a receptor and producing nothing is a real and common combination.
2. Does it reach the tissue, at what concentration, for how long? Affinity is measured in a controlled assay at a known concentration. In a person, the question is whether enough molecule arrives at the right place. A compound with spectacular affinity and negligible tissue penetration produces nothing. This is §2.9's step 2.
3. What happens to the receptor under sustained high-affinity binding? Tight, persistent binding is an efficient way to trigger desensitization and downregulation. Higher affinity can mean faster loss of response, not more of it. §2.7 and §2.8.
The general point: affinity is frequently the least clinically relevant of the three properties and the one most often quoted, precisely because a large multiple sounds impressive and requires no outcome data to generate.
Exercise 2.16 †
A model answer:
Think of the hormone as a phone call rather than a delivery. It doesn't bring anything into the cell — it just tells the cell something, and the cell does the work itself using its own energy. Because one call can set off a chain of instructions inside the cell, and each step in that chain triggers many copies of the next step, a single call can end up producing thousands of actions. So you need almost nothing of the hormone itself — the loudness of the response comes from the cell, not from the messenger.
Exercise 2.19 †
Three mechanistic explanations consistent with Chapter 2:
1. Concentration mismatch. A paracrine signal acts on immediate neighbors at high local concentration. Injected systemically and diluted into the whole blood volume, it may never reach the concentration at the target tissue that its local physiology depends on — while simultaneously being high enough elsewhere to cause off-target effects (§2.1, §2.2).
2. It activates receptors in tissues that were never meant to hear it. A local signal is local partly because the surrounding tissue does not express its receptor. Systemic delivery reaches every tissue that does express it anywhere in the body, producing effects that have no counterpart in normal physiology (§2.1).
3. Termination mismatch. A paracrine signal is released in a pulse and degraded almost immediately. Sustained systemic exposure produces continuous stimulation, which recruits desensitization and downregulation — so the target tissue stops responding while the off-target effects persist (§2.7, §2.8).
A fourth acceptable answer: rapid degradation and renal clearance may mean it never reaches the target tissue in meaningful quantity at all (§2.7).
Exercise 2.22 †
Insulin in type 1 diabetes replaces an absent signal. The beta cells producing it are destroyed; there is no intact homeostatic loop being overridden, so no counter-regulation is provoked and no persistent supraphysiological stimulation drives downregulation. The therapy restores something missing rather than pushing on something working.
GLP-1 agonist nausea fades while metabolic effects persist because tolerance is not a property of a drug — it is a property of each drug-tissue pair. The receptor populations mediating nausea (largely in the gut and brainstem) desensitize relatively quickly; those mediating insulin secretion and appetite regulation do so much less.
The general principle: expect tolerance when you override an intact regulated system; expect durability when you replace an absent one; and expect the two to occur at different rates in different tissues within the same patient on the same drug.
Exercise 2.25 †
Original claim: "modulates cellular signaling pathways to optimize recovery."
Rewritten to the specificity the evidence would need:
"In [which population], administration of [what dose, by what route, for how long] produced [what measurable change] in [which specific outcome], compared with [what comparator], measured [how], with [what effect size] — in [how many] participants, over [what duration]."
What each added specification requires somebody to have measured:
- Which population — a defined trial cohort, not "people"
- Dose, route, duration — a pharmacokinetic and dosing study
- Measurable change — an instrument with known reliability, not self-report alone
- Which outcome — a pre-specified endpoint, not one selected after the fact
- Comparator — a control group, ideally placebo, ideally blinded
- Effect size — actual numbers, and both relative and absolute where applicable
- N and duration — a powered study of adequate length
The original phrase requires none of these to have been done. "Modulates" is unfalsifiable — anything that binds anything modulates something. "Optimize" has no operational definition. "Recovery" names no measurable outcome. The sentence is constructed to be true regardless of the facts, which is exactly what makes it worthless. Chapter 5 formalizes this.
Exercise 2.27 †
A defensible position: it undermines this application of the concept while leaving the concept itself intact but unproven at scale.
For the concept surviving: biased agonism is a real, measurable phenomenon. GPCRs do couple to multiple downstream pathways, and compounds do differ in how they engage them. That is not in dispute.
Against this application: the specific hypothesis — that analgesia and respiratory depression map cleanly onto the G-protein and arrestin arms — has been challenged by subsequent work, including studies in animals with disrupted arrestin recruitment where respiratory depression persisted. An alternative account holds that apparent bias among these compounds reflects differences in intrinsic efficacy rather than pathway selectivity, which would mean the entire framework was mis-specified for this receptor.
What would settle it: a compound with rigorously demonstrated pathway bias, tested head-to-head against a conventional agonist at genuinely equianalgesic doses, with respiratory outcomes as a pre-specified primary endpoint. The equianalgesic condition is the crux — much of the disagreement turns on whether comparisons were made at doses producing equal analgesia or merely equal milligrams.
Note the shape of the answer: the concept and its application are separable, and evidence against one is not automatically evidence against the other. That distinction recurs throughout Part III.
Exercise 2.29 †
No single answer, but the exercise has characteristic findings.
Expect Field 3 to be easy for approved drugs — receptor identified, class known, downstream signaling characterized, endogenous role documented. Semaglutide, insulin, teriparatide, and octreotide all resolve quickly through the IUPHAR/BPS Guide or a drug label.
Expect it to be partial for research-code compounds. BPC-157 is the instructive case: as of this writing, no receptor has been definitively established for it, and the proposed mechanisms (nitric-oxide-related pathways, growth-factor and angiogenic signaling, and others) are inferred from downstream observations in animal models rather than from an identified receptor. Write "not established." That is not a criticism of the molecule; it is an accurate description of the literature, and a dossier that guesses here has already failed at its purpose.
Expect it to be nearly empty for cosmetic ingredients, where the "mechanism" cited is often a cell-culture observation with no established receptor and no demonstrated route to the tissue.
The finding to notice: how complete Field 3 is correlates with how much formal development a compound has undergone — the same correlation Chapter 1 found for Field 1. That is a legitimate prior and not a rating. Chapter 5 will tell you what to do with it.
Chapter 3
Exercise 3.2 †
The hypothalamic-pituitary portal circulation is a short, dedicated vascular connection running directly from the hypothalamus to the anterior pituitary, rather than through the general circulation.
Why it exists: it delivers releasing hormones to their target at high local concentration while keeping them essentially undetectable everywhere else. A hypothalamic peptide released into the whole blood volume would be diluted enormously and would also reach every other tissue expressing its receptor. The portal system solves both problems with anatomy rather than chemistry.
Two consequences follow. First, measuring a releasing hormone in peripheral blood is nearly meaningless — a genuine obstacle to studying these systems. Second, a drug mimicking a releasing hormone cannot use this route: injected, it arrives systemically diluted. That constraint shapes Chapter 15's entire subject.
Exercise 3.5 †
The GH axis. Stimulatory: GHRH (growth hormone-releasing hormone). Inhibitory: somatostatin (also called growth hormone-inhibiting hormone).
The presence of an explicit brake matters practically: any intervention that raises GH also tends to raise somatostatin through feedback, so the axis actively opposes attempts to elevate it. A drug acting on this axis is not pushing on a passive system — it is pushing on a system with a counter-lever.
Exercise 3.9 †
Step by step:
- The axis contains a sensor — typically at the hypothalamus and pituitary — that detects the circulating level of the final hormone.
- That sensor compares the level against a set point.
- When the level exceeds the set point, the sensor reduces output of the releasing hormone and the tropic hormone.
- Reduced tropic hormone means reduced stimulation of the target gland.
- The target gland reduces its production, and with sustained suppression may atrophy from disuse.
Why the system cannot distinguish sources: the sensor measures a concentration, not a provenance. A cortisol molecule from your adrenal gland and a cortisol molecule from a prescription are chemically identical and bind the same feedback receptors. There is no mechanism by which the sensor could tell them apart, and no evolutionary pressure that would have produced one — nothing in the environment supplied hormones from outside until pharmacology did.
Exercise 3.11 †
Why abrupt withdrawal is dangerous: long-term high-dose corticosteroid suppresses the HPA axis profoundly — reduced CRH, reduced ACTH, and adrenal atrophy from lack of stimulation. If the external steroid stops suddenly, the body has neither its own cortisol production nor the drug. Cortisol is required for blood pressure maintenance, glucose regulation, and the stress response; acute deficiency can be life-threatening, particularly under physiological stress such as illness or surgery.
Why tapering helps: gradually reducing the external dose lowers total circulating steroid slowly enough that the feedback sensor detects a deficit and begins increasing CRH and ACTH, restimulating the adrenal. The gland recovers responsiveness over weeks to months. Tapering does not accelerate recovery so much as it prevents the gap — it keeps total cortisol adequate while the axis restarts.
Exercise 3.14 †
Why a random GH measurement is uninterpretable: growth hormone is released in discrete pulses, predominantly nocturnal, and is near-undetectable between them. A sample drawn at a random time is overwhelmingly likely to fall in a trough. A low value therefore cannot distinguish deficiency from normal inter-pulse physiology, and an occasional high value cannot distinguish a normal pulse from excess.
What is measured instead: IGF-1, which is produced largely in the liver in response to GH and circulates at stable levels integrating GH exposure over time. Where IGF-1 is equivocal, dynamic stimulation or suppression testing is used — provoking GH release under standardized conditions and measuring the response, or suppressing it to test for excess.
This is a general principle worth extracting: when a signal is pulsatile, measure something downstream and stable rather than the signal itself. It applies well beyond endocrinology.
Exercise 3.17 †
Three reasons a sustained elevation might produce less benefit than expected:
- Desensitization and downregulation accumulate without the recovery windows that pulses provide. Net signaling can fall below what intermittent stimulation produces (Ch 2 §2.7).
- Feedback opposes it. A sustained elevation is more effectively detected by the feedback sensor than transient pulses are, so the axis suppresses harder.
- Downstream machinery may be pattern-sensitive. Some responses depend on the rate of change or on pulse frequency rather than on average level — GnRH is the demonstrated case, where frequency differentially favors LH versus FSH.
One reason it might produce a benefit the natural pattern would not: deliberate suppression. Continuous GnRH agonist delivery suppresses the reproductive axis, and that suppression is the therapeutic goal in hormone-sensitive cancer, endometriosis, and precocious puberty. The "wrong" pattern is the entire point. A pattern mismatch is not automatically a flaw — it is a question about what you are trying to achieve.
Exercise 3.18 †
A model answer:
Your body has a thermostat for hormones. It measures how much is floating around and adjusts its own production to keep the level about right. The problem is that the thermostat can't tell where the hormone came from — a hormone you were given counts exactly the same as one you made. So when you take some from outside, the thermostat sees plenty and turns your own production down. If you do that for long enough, the gland that used to make it gets out of practice, and starting it back up takes a while.
Exercise 3.22 †
Predictions, from §3.3 and §3.6:
Somatostatin inhibits growth hormone, insulin, glucagon, and multiple gut hormones. Therefore a somatostatin analog would plausibly cause:
- Effects on blood glucose — inhibiting both insulin and glucagon has opposing effects on glucose, so the net direction is not predictable from mechanism alone, but glucose dysregulation of some kind is likely.
- Gastrointestinal effects — inhibiting gut hormones that regulate motility, secretion, and gallbladder contraction would plausibly cause altered bowel habit, malabsorption, and gallbladder problems including gallstones.
- Suppressed growth hormone, with whatever consequences that carries in a given patient.
How to check: read the approved label. Somatostatin analogs are approved drugs (Chapter 27), and the adverse reactions section of an approved label is a public document listing what was actually observed in trials at what frequency.
This exercise is a rehearsal for a skill Chapter 27 formalizes: mechanism generates the hypotheses; the label reports the findings. Doing it in this order — predict, then check — is a genuinely good way to calibrate how well mechanism predicts.
Exercise 3.24 †
The natriuretic peptide system supports three different claims requiring three different studies:
As a biomarker — "does NT-proBNP concentration indicate cardiac stress and help diagnose heart failure?" Tested by diagnostic accuracy studies comparing the test against a reference standard. Answer: yes, strongly.
As an infused therapy — "does giving synthetic BNP improve outcomes in acute decompensated heart failure?" Tested by a randomized outcome trial. Nesiritide was approved and a subsequent large trial did not demonstrate the hoped-for benefits.
As a target for potentiation — "does inhibiting the enzyme that degrades natriuretic peptides, raising endogenous levels, improve outcomes in chronic heart failure?" Tested by a different randomized outcome trial. Sacubitril/valsartan succeeded.
What this implies: a rating must attach to a claim, not to a molecule or a system. Any single verdict on "natriuretic peptides" would necessarily be wrong about at least two of these three. This is the strongest demonstration in Part I of the principle stated in §1.9 and formalized in Chapter 5, and it is not a contrived example — it is standard cardiology.
Exercise 3.26 †
Original: "optimize the HPA axis for better stress resilience."
Rewritten to the required specificity:
"In [defined population], [compound] at [dose, route, duration] produced [specific change in a named measurable — e.g. morning cortisol, cortisol awakening response, ACTH stimulation response] and improved [a named clinical outcome — e.g. a validated stress or fatigue scale score] compared with [placebo/comparator], in [n] participants over [duration], with [effect size]."
Mapping to §2.9's steps:
- Naming the compound and its receptor action → step 1
- Dose, route, duration → steps 2 and 3
- A measurable change in a hormone level → step 4 (evidence the machinery responded)
- Persistence over the stated duration → step 5 (the system did not fully compensate)
- Improvement in a named clinical outcome versus a comparator → step 6
- Adverse events reported alongside → step 7
The original phrase requires none of these. "Optimize" has no operational definition — it is compatible with raising cortisol, lowering it, or changing nothing. "Stress resilience" names no measurable outcome. As written, the claim cannot be false, which is precisely what makes it worthless.
Exercise 3.28 †
It strengthens §2.9's argument, and it does so in an unusually pointed way.
Note carefully what was and was not understood in the GnRH case. The receptor was identified. The ligand was characterized. The agonist bound and activated it — all of §2.9's step 1, fully in hand. The mechanism was not misunderstood in any particular.
And the clinical outcome ran in the opposite direction from the naive prediction, because a variable that mechanism-at-the-receptor-level does not capture — the temporal pattern of delivery — turned out to determine the sign of the effect.
That is a stronger result than "mechanism is incomplete." It shows that mechanism can be correct and complete at the level it describes and still fail to predict direction, because the relevant causal variable lives at a different level of description entirely.
The honest counterweight: once the pattern-dependence was understood, it became part of the mechanism, and it is now used deliberately and predictably. Mechanism absorbed the surprise. But it absorbed it after the fact, which is exactly the pattern §2.9 describes — mechanism explains beautifully in retrospect and predicts unreliably in prospect.
Exercise 3.29 †
No single answer. Characteristic findings:
The feedback line will be the most informative one you write. For pituitary hormones and sex steroids, suppression is emphatic and well documented. For gut hormones, largely absent. For most research-code compounds in the performance space, the honest entry is "not established" or "not studied" — nobody has run the experiment that would answer it, which is itself a significant finding for a compound people take continuously for months.
The pattern line will surface the questions Part III has to answer. Growth hormone: strongly pulsatile, delivered as an injection producing a sustained elevation — mismatch. CJC-1295: modeled on a pulsatile hypothalamic hormone, engineered explicitly for extended duration — mismatch, and this one is central to Chapter 15's analysis. Semaglutide: GLP-1 is meal-triggered rather than strictly pulsatile, and the drug is continuous — a mismatch that appears, empirically, to be tolerated well, which is worth noticing because it shows a mismatch is a question rather than a verdict.
Where the question does not apply, say so. For a wholly synthetic compound with no endogenous counterpart, there is no axis to suppress. Recording "not applicable, no endogenous counterpart" is a complete and correct answer, and it is different from "unknown."
Chapter 4
Exercise 4.2 †
Bioavailability — the fraction of an administered dose that reaches the general circulation intact. Intravenous is 100% by definition; every other route is less.
Half-life — the time for drug concentration in the body to fall by half.
Clearance — the volume of blood cleared of drug per unit time.
Clearance is the underlying process; half-life is the observed quantity. Clearance describes what the kidneys, liver, and enzymes are doing. Half-life is what you measure in a plasma sample. Two drugs can share a half-life with very different clearances if their distribution volumes differ — but for the purposes of this book, half-life is the number that predicts behavior and clearance is the explanation for it.
Exercise 4.5 †
| Position | Change | Problem solved |
|---|---|---|
| 8 | Ala → Aib (2-aminoisobutyric acid, non-natural) | Abolishes the DPP-4 cut site. Defeats proteolysis. |
| 34 | Lys → Arg | Leaves a single lysine as a fatty-acid attachment site, so the conjugation reaction produces one defined product rather than a mixture. Manufacturing control — no pharmacological benefit to the patient. |
| 26 | C18 diacid fatty acid attached via a spacer to the remaining lysine | Binds albumin (~66,000 Da); the complex is far too large to be filtered by the kidney and is also shielded from proteases. Defeats renal clearance. |
Result: half-life from ~1–2 minutes to ~1 week. Receptor activity unchanged.
Exercise 4.9 †
A 6-hour half-life means concentration halves every 6 hours.
- 6 h → 50%
- 12 h → 25%
- 18 h → 12.5%
- 24 h → about 6%
So roughly 6% remains after a day.
Expected dosing interval: something on the order of the half-life — likely twice daily or three times daily. A once-daily schedule would produce large peak-to-trough swings, with the drug at ~6% of peak before the next dose. Whether that matters depends on the therapeutic window: for a drug where the effect persists after the drug is gone, wide swings may be acceptable; for one where effect tracks concentration, they are not.
Exercise 4.12 †
The variable route is the more serious problem, and it is not close.
A route delivering 1% consistently is a scaling problem. You put 100 times more drug in the dose, accept the cost, and the patient reliably receives what was intended. Oral semaglutide is exactly this and it is an approved product.
A route delivering somewhere between 0.2% and 4% is a safety problem, because the dose the patient actually receives is unpredictable and may vary twentyfold. If you dose for the low end, patients at the high end are overdosed. If you dose for the high end, patients at the low end are untreated. There is no dose that is correct for everyone, and — critically — no way for the patient or clinician to know which end they are on today.
This is why §4.4 argues that consistency matters more than magnitude, and it is the central pharmacological failure mode of both intranasal delivery (Chapter 21) and inhaled insulin (Case Study 2), where absorption depended on the condition of the mucosa or the lung.
Exercise 4.14 †
See 4.5 for the three modifications.
The one with no pharmacological benefit is position 34 (Lys→Arg).
It exists because native GLP-1 has two lysine residues, and the fatty acid in modification 3 attaches to a lysine's free amino group. With two lysines available, the conjugation reaction would produce a mixture: some molecules modified at one site, some at the other, some possibly at both. That mixture would be a quality-control problem — two isomers with potentially different pharmacokinetics, difficult to separate because they are nearly identical in physical properties.
Removing one lysine leaves a single attachment point, so the reaction produces one defined product.
Why this matters beyond trivia: it demonstrates that a drug molecule's structure encodes manufacturing decisions as well as pharmacological ones. When Chapter 34 asks what "purity" means and how you would detect a closely related impurity, this is the kind of problem being prevented. The molecule was designed so the problem could not arise.
Exercise 4.17 †
Proteolysis defeated, renal filtration not: half-life improves substantially but remains short — hours rather than minutes. The molecule is no longer being enzymatically destroyed, but it is still small enough to be filtered out of blood on passage through the kidney, and glomerular filtration is continuous and fast. Exenatide is precisely this case: DPP-4-resistant, ~4,200 Da, half-life of about two to three hours. A roughly hundred-fold improvement that stops well short of once-weekly.
Renal filtration defeated, proteolysis not: the molecule is retained in circulation but is being cleaved while it is there. The outcome depends on how fast the relevant protease works. For a DPP-4 substrate like GLP-1, cleavage is rapid enough that albumin binding alone would help considerably less than it appears — the drug would be retained as an inactive fragment. Albumin binding does provide some protease shielding, which is why the two strategies are not fully independent, but the cut site remaining is a liability.
The general point: the exits are largely independent, and the observed half-life is set by whichever remains open. This is why successful modifications come in combinations.
Exercise 4.19 †
Three questions, ordered by how much they narrow the possibilities:
1. What is the measured absolute bioavailability in humans by this route, compared with injection? This single question resolves most cases. An answer requires a pharmacokinetic study measuring plasma concentrations after oral and after injected administration in the same people. If the answer is "we haven't measured it," everything else is speculation — and note that this study is neither difficult nor expensive by pharmaceutical standards, so its absence is a choice.
2. What specifically does the delivery system do, and against which of the five barriers? A real answer names a mechanism: an enteric coating (barrier 1), a protease inhibitor (barrier 2), a permeation enhancer (barrier 4). "Liposomal" is not an answer — it names a formulation type, not a solved barrier, and liposomes face their own stability problem in the gut.
3. What is the molecular weight? If it is a few thousand daltons, barrier 4 is a serious physical obstacle regardless of what protects it from enzymes. This question is free and the answer is usually publicly available.
Exercise 4.21 †
A model answer:
Your digestive system is built to take proteins apart — that's how you get amino acids out of food. Insulin is made of exactly the same stuff as the protein in a steak, so if you swallow it, your stomach and intestines treat it as lunch. Even the small amount that somehow survived would be too big to squeeze through the wall of your gut into your bloodstream. And anything that did squeeze through goes straight to your liver, which removes most of it before it reaches the rest of you.
Exercise 4.25 †
Original: "protected by a proprietary delivery matrix for maximum absorption."
The specific claim that would need to be true:
"In [n] healthy human volunteers, oral administration of [dose] of [named peptide] in this formulation produced plasma concentrations corresponding to an absolute bioavailability of [x]% relative to subcutaneous administration of the same compound, with a coefficient of variation of [y]%."
The study that would establish it: a crossover pharmacokinetic study — the same participants receive the oral formulation and an injected reference on separate occasions, with serial blood sampling and measurement of the peptide by a validated assay. Absolute bioavailability is calculated by comparing exposure between the two routes.
What the original phrase requires: nothing. "Proprietary" means undisclosed. "Delivery matrix" names no mechanism and no barrier. "Maximum absorption" has no referent — maximum relative to what? The sentence is unfalsifiable, which is the same failure Chapter 2's "modulates cellular signaling pathways" example had, in a different vocabulary.
Exercise 4.28 †
The strongest objection: the delivery filter, applied confidently and early, would have rejected oral semaglutide.
Before it existed, every element of the filter argued against it. The route was one the molecule manifestly cannot survive. There was no measured bioavailability because it had not been achieved. The half-life question was fine but irrelevant if nothing was absorbed. And the reasonable conclusion from Chapter 4's own §4.1 is that oral peptide delivery is effectively impossible.
It was achieved anyway — by accepting a ~1% yield, adding a permeation enhancer, absorbing from an unexpected site (the stomach rather than the intestine), and building the entire product around the loss.
What this shows: the filter is a prefilter on claims, not a verdict on possibilities. It correctly identifies that a claim is implausible as stated and by that route — and that remains useful, because most products making such claims genuinely have not solved anything. But "nobody has solved this" and "this cannot be solved" are different statements, and the filter only supports the first.
The correct use, therefore: treat a delivery failure as a burden-shifting device, not a conclusion. It moves the burden onto the seller to explain what they solved. Oral semaglutide's developers could answer in detail. Almost nobody selling an oral peptide can. The filter's value is that it makes the burden explicit and cheap to apply, not that it is infallible.
Note also that this is Chapter 2's asymmetry appearing again, one level up: mechanism is strong evidence against, and the strongest version of "strong" is still not "conclusive."
Exercise 4.29 †
No single answer. Characteristic findings:
The modification line sorts your dossier almost perfectly. Approved peptide drugs have documented, usually elegant modifications with a stated purpose. Semaglutide has three. Liraglutide has one fatty acid. Insulin analogs have specific substitutions with named purposes (glargine's isoelectric shift, lispro's reversed residues to prevent hexamer formation). Octreotide is cyclized and contains D-amino acids.
Research-code compounds typically read "none." BPC-157 is described as the native gastric-juice sequence. TB-500 is a fragment of a natural peptide. Neither has a half-life-extending modification — which raises the §4.30 question immediately: what is the native half-life, and does the dosing pattern people use make sense against it?
And the bioavailability line will mostly read "not measured." This is worth writing exactly that way rather than "unknown." A compound whose oral bioavailability has been measured at 1% and one where nobody has ever run the study are in completely different epistemic positions, and collapsing them loses the distinction that Chapter 5 is about to make central.
Expect the filter score (exercise 4.31) to correlate strongly with your eventual ratings — and expect it to correlate imperfectly, which is the more interesting result. A compound can pass the delivery filter completely and still have no evidence that it works. Delivery is necessary and not sufficient, exactly as mechanism is.
Chapter 5
Exercise 5.2 †
Strongest to weakest: systematic review / meta-analysis · large randomized controlled trial · small or short RCT · cohort study · case-control study · case series / case report · animal study · in vitro / cell study · mechanism / plausibility · expert opinion · anecdote / testimonial.
The essential caveat: the ladder ranks what a result licenses you to believe about the next person, not the worth or difficulty of the research. Every drug starts at the bottom and climbs. A brilliant mechanistic paper is better science than a sloppy RCT and worse evidence for a clinical claim.
Exercise 5.5 †
Population (who was studied) · comparator (compared with what) · randomization (how they were assigned) · blinding (who knew) · endpoint (what was measured, pre-specified or not) · duration and size (for how long, in how many).
Read them in that order. Population first, because it determines whether the trial is about anyone resembling the person you care about, and no amount of methodological quality fixes a population mismatch.
Exercise 5.9 †
An estimand is the precise quantity a trial's analysis is estimating — in particular, how participants who discontinued treatment are handled.
Treatment-regimen estimand: "What happens if you are assigned this drug?" Counts everyone as randomized, including those who stopped. Reflects real-world prescribing, where some people will not tolerate a drug. Generally the primary analysis for regulatory purposes.
Efficacy estimand: "What happens if you actually take it, as prescribed, for the full period?" Restricted to those who remained on treatment. Reflects the biological effect more directly and is systematically larger.
Both are pre-specified, both are reported, and neither is a trick. The failure is in reporting a number without naming which — and because the efficacy figure is always larger, an unnamed figure is systematically optimistic rather than randomly imprecise.
Exercise 5.12 †
As evidence of efficacy, a case report is nearly worthless because it has no comparison group. One person improving after a treatment is compatible with the treatment working, with natural recovery, with regression to the mean, with expectation, and with any of several other things that changed at the same time. There is no way to distinguish among them from a single case.
As evidence of harm, a case report can be genuinely valuable — occasionally decisive.
The asymmetry that does the work: rare, distinctive events. If a drug causes a characteristic adverse event that essentially never occurs otherwise, a handful of reports is strong evidence, because the base rate of the event is near zero. Efficacy claims have no such asymmetry — improvement is common regardless of treatment, so an improvement carries almost no information.
The formal version: evidence strength depends on the likelihood ratio. For a distinctive rare harm, the ratio is enormous. For "the person got better," it is close to 1.
Exercise 5.15 †
Reasons a rat transected-tendon result may not translate:
Injury mechanism. A surgical transection is a clean, acute, complete division with a known time zero. Human tendinopathy is typically chronic, degenerative, partial, and developed over months or years of overuse. These are arguably different pathologies, not the same one at different scales.
Healing biology and timescale. Rodent healing is faster and differs in inflammatory and remodeling dynamics. A 14-day endpoint in a rat does not correspond to any particular human timepoint.
Endpoint. The rat study measures histology and load-to-failure. Humans want to know about pain, function, return to activity, and recurrence — none of which a rat can report.
Dose. The doses used in the animal study, scaled to a human, may be unachievable, unsafe, or simply untested.
Age and comorbidity. Laboratory rodents are typically young, healthy, and genetically uniform. Human tendon injury clusters in older, less uniform, often comorbid populations.
Delivery. Route, formulation, and local concentration at the injury site may differ entirely.
Publication filter. The published animal result is drawn from an unknown number of experiments run, because animal work is largely unregistered.
Exercise 5.18 †
Six questions, in order of importance:
1. Compared with what? Was there a control arm, and what was in it? "Improvement in the primary endpoint" with no comparator tells you almost nothing.
2. Was the primary endpoint pre-specified? Check a registry. If 20 outcomes were measured and this one reached p = 0.04, the result is likely noise.
3. What was the endpoint, and is it a surrogate? If so, is the arrow to a clinical outcome established?
4. Was it blinded, and does the endpoint involve judgment? Unblinded plus subjective is a much larger problem than either alone.
5. What was the effect size, and the confidence interval? p = 0.04 with 38 participants means the interval is wide and probably includes values too small to matter.
6. Who was studied, and who funded it?
Note that "p = 0.04" appears nowhere in the top three. Significance is the least informative number quoted, and it is the one most often quoted.
Exercise 5.21 †
Both arms in nearly all obesity pharmacotherapy trials receive structured lifestyle support: dietary counseling, activity guidance, and regular study contact.
Effect one — makes the drug look smaller. The placebo arm is receiving a real intervention and loses meaningfully more weight than an untreated population would. The measured drug effect is therefore the increment above an active comparator, not above nothing.
Effect two — makes the drug look larger than practice will deliver. A patient who receives a prescription and no structured support is not receiving what the trial delivered. The trial's total result includes a component the patient will not get.
Which dominates is not knowable from the trial. What is knowable: the result describes drug-plus-support versus support alone. That is not the comparison most readers have in mind, and it is why "15% weight loss" is a statement about a protocol rather than about a molecule.
Exercise 5.23 †
| Claim | Surrogate | Hard endpoint it stands in for | Arrow established? |
|---|---|---|---|
| (a) raises growth hormone | serum GH | body composition, strength, function | No — Ch 15 |
| (b) lowers LDL cholesterol | LDL-C | cardiovascular events | Largely yes — a substantial randomized literature supports it |
| (c) increases lean body mass | DXA lean mass | strength, function, independence | No — Ch 16 shows these dissociate |
| (d) reduces an inflammatory marker | e.g. CRP | almost anything | Rarely — one of the least reliable |
| (e) shrinks a tumor | tumor response | overall survival | Sometimes — oncology has learned these can diverge |
| (f) improves skin elasticity on an instrument | cutometer reading | looking better | No — Ch 30 |
(b) is the important one for calibration: surrogates are not inherently invalid, and LDL is the best-supported example in medicine. The question is always whether the specific arrow has been tested.
Exercise 5.26 †
Control event rate 4% over two years; 30% relative risk reduction.
- Treated event rate: 4% × (1 − 0.30) = 2.8%
- Absolute risk reduction: 4% − 2.8% = 1.2 percentage points
- NNT = 1 ÷ 0.012 ≈ 83
So roughly 83 people must be treated for two years for one additional person to avoid the event.
Honest one-sentence version: "Over two years, this treatment reduces the chance of the event from about 4 in 100 to about 3 in 100 — a 30% relative reduction, which means roughly one person benefits for every eighty treated."
Exercise 5.28 †
Maximally impressive (accurate): "Semaglutide cuts cardiovascular events by 20% in high-risk patients."
Maximally underwhelming (accurate): "Semaglutide reduced cardiovascular events by 1.5 percentage points over three years — meaning about 98.5 in 100 patients saw no difference in this outcome."
Honest: "In adults with established cardiovascular disease and overweight or obesity but without diabetes, semaglutide reduced major cardiovascular events from about 8% to about 6.5% over roughly three years — a 20% relative reduction, or about 1.5 percentage points, meaning roughly one event avoided for every 65 to 70 people treated over that period."
Note what the honest version costs: it is three times as long and cannot be a headline. That is the actual structural reason relative risk dominates public communication — not malice, but that the honest version does not fit the format.
Exercise 5.31 †
See 5.9. Applied to SURMOUNT-1 (tirzepatide 15 mg, 72 weeks, adults with obesity without diabetes):
- Treatment-regimen estimand: about −21%. Counts all randomized participants including discontinuers.
- Efficacy estimand: about −22.5%. Restricted to participants who remained on treatment.
The efficacy estimand is systematically larger because the participants excluded from it are disproportionately those who stopped — and people stop for reasons correlated with doing poorly: side effects they could not tolerate, or lack of response. Removing them removes the lower tail.
The gap is therefore not just a statistical artifact; its size is informative, telling you something about tolerability and adherence. A large gap means many people did not stay on the drug, which matters clinically.
Exercise 5.34 †
"In a clinical study, 87% of users reported improvement within 30 days."
Red flags:
- "Users," not "participants." Suggests no randomization, possibly no study protocol at all.
- No control group is implied or stated. Compared with what? Most conditions improve over 30 days.
- "Reported" — self-reported, subjective, unblinded. Expectation drives this directly.
- "Improvement" is undefined. Improvement in what, measured how, by how much?
- "Clinical study" is not a protected term and does not imply randomization, blinding, or registration.
- 87% is a proportion of responders, not an effect size, and tells you nothing about magnitude.
- No n, no duration beyond 30 days, no comparator, no funding disclosure, no registration.
What design would make the sentence mean something: a randomized, double-blind, placebo-controlled trial, pre-registered, with a pre-specified validated outcome measure, adequately powered, reporting the difference between arms with a confidence interval — not the proportion in one arm who felt better.
Note that even a perfectly conducted version of this study would report something like "the treatment group improved by X points more than placebo (95% CI ...)", not "87% of users reported improvement." The form of the sentence tells you the design before you read the content.
Exercise 5.37 †
A model answer:
Studies aren't all the same kind of thing. Most of those hundred are probably experiments on rats or on cells in a dish — which are real science, and are the reason anyone thought the compound was worth investigating, but they can't tell you what happens in a person. Roughly nine out of ten things that look great in animals don't work out in humans. So the question isn't how many studies there are, it's how many are the kind that could actually answer the question — and for a lot of popular compounds, that number is zero. A hundred studies that don't test the claim is still zero tests of the claim.
Exercise 5.40 †
(a) Semaglutide for weight loss in adults with obesity — see §5.12. ✅.
(b) BPC-157 for tendon healing in humans — see §5.12. ❌.
(c) A topical peptide serum for wrinkle reduction:
Claim: A topical peptide serum visibly reduces wrinkles in adults. Rating: ⚠️ to ❌ depending on the specific claim and the specific peptide. Why: Some cosmetic peptides have real cell-culture activity and some small human studies report instrument-measured changes. But the studies are typically small, short, industry-funded, and use surrogate endpoints (instrument readings) rather than blinded assessment of appearance — and the underlying delivery problem (Ch 4, Ch 30) is unresolved for most of them. What would change it: an adequately powered, randomized, blinded trial with blinded photographic assessment against vehicle control, reporting a difference a person would notice.
(d) A compound in Phase 2 trials for a serious disease:
Rating: ⚠️ at most, and 🔬 if the Phase 2 has not reported. Why: Phase 2 establishes a signal, not an effect size. Phase 2 results are systematically optimistic, and roughly nine in ten compounds entering human trials never reach approval. What would change it: a completed, adequately powered Phase 3 in the target population with a pre-specified clinical endpoint.
Exercise 5.43 †
No answer key, by design — but a note on what most readers report.
The common experience is that writing the honest rating is easy and writing the falsifier is hard. People find themselves unable to specify what result would change their mind, and discover that the reason is that they had not been holding a position about evidence at all. They had been holding a hope, and hopes do not have falsifiers.
That discovery is the single most valuable moment in this chapter, and it is worth sitting with rather than resolving. Nothing about it means the hope is wrong. It means you now know which of your beliefs are load-bearing and which are decorative.
Exercise 5.45 †
No single answer. Characteristic findings:
Most dossier entries will need more than one rating. Semaglutide will need at least three (weight loss, cardiovascular, and whatever else the reader cares about). Growth hormone needs two (approved deficiency indications versus anti-aging use). This is the point of rule 6, and readers who end up with exactly one rating per molecule have not applied it.
The falsifier line will be the hardest, and it should be. "More research" is not a falsifier. A usable one names a design, a population, an endpoint, and a duration — something you could hand to a trialist. If you cannot write it, you do not yet know what you believe well enough to defend it.
And expect the distribution to be informative. Count your ✅, ⚠️, ❌, and 🔬. If the compounds you personally use or want to use are distributed more favorably than the ones you don't, that is not evidence you chose well. It is the drift rule 4 warns about, and noticing it now is much cheaper than noticing it in Chapter 37.
Chapter 6
Exercise 6.2 †
Regression to the mean — people begin treatments when they feel worst, and extreme states are followed on average by less extreme ones regardless of intervention. Arithmetic, not psychology.
Natural history — most injuries heal and most acute conditions resolve. The question is never "did it improve" but "did it improve faster than it would have," and an individual has no access to the counterfactual.
Expectation — someone who researched, paid for, injected, and told friends about a compound is not a neutral observer of their own outcome. Expectation effects on subjective endpoints are large and operate below awareness.
Co-intervention — people who start a peptide frequently also change sleep, diet, training, and attention simultaneously. Any of these could produce the improvement.
Survivorship — the structural one. Those for whom something happened post; those for whom nothing happened mostly do not write about the absence of an event. The visible sample is filtered before any reader arrives.
Exercise 6.5 †
Nobody.
Researchers need findings to be important (career and funding incentives favor positive results). Testimonial posters gain status from a story. Vendors need the compound to work and purity to be the main risk. Compounders need equivalence. Media needs whichever register performs. Influencers need affiliate revenue and audience trust. Aggregators need the page to seem to have answered the question. And the manufacturer of an approved competitor needs their compound to work and unapproved ones to be unsafe.
The asymmetry is structural rather than moral. "We don't know yet" sells nothing, ranks poorly, ends conversations, and satisfies nobody who arrived with a question. It is not that the participants are unusually venal; it is that accurate uncertainty has no constituency.
And this book is in the diagram too. It sells nothing, which removes one class of bias and leaves others — principally a preference for rigor that can shade into excessive conservatism, and an author who finds ❌ more satisfying to write than ⚠️.
Exercise 6.8 †
No single answer, but the exercise reliably produces the same finding.
What you will typically observe: the preclinical paper's limitations paragraph states, in the authors' own words, that findings are limited to the model used, that doses may not translate, that functional outcomes were not assessed, and that clinical trials are required. It is usually the last paragraph of the discussion, it is written by the people best placed to know, and it is stated plainly.
What the downstream page typically contains: the finding, without the species, the model, the dose, the endpoint, or the duration — and never the limitations.
The instructive part is that nobody removed it maliciously. Each retelling compressed slightly. The authors' qualifiers survive one step, sometimes two, and are gone by step three. This is the mechanism of the entire chapter, observable in one afternoon on any compound you choose.
Exercise 6.10 †
Working through §6.3 for an eight-month shoulder injury resolving in three weeks:
Regression to the mean: the person almost certainly started the compound during a bad stretch — that is when people act. Symptom severity fluctuates even in chronic conditions, and a bad stretch is statistically likely to be followed by a better one.
Natural history: eight months is a long time, and some tendon problems do resolve, sometimes abruptly, after long plateaus. The improvement arriving three weeks after starting something does not establish that the something caused it.
Expectation: shoulder pain is a subjective endpoint, and expectation effects on subjective endpoints are among the largest in medicine. The person has invested money, effort, and identity in this working.
Co-intervention: somebody who takes a compound for an injury usually also modifies training, adds rehabilitation exercises, sleeps more, and pays closer attention to the joint. Any of these might have been the active ingredient.
Survivorship: you are reading this account because it had a satisfying ending. The people who tried the same compound for the same injury and saw nothing are not writing posts.
The essential framing: none of this requires the person to be lying, exaggerating, or foolish. Their shoulder really did get better. The claim being evaluated is not "did this happen" but "was this caused by the compound, and would it happen to the next person" — and a single case cannot answer either.
Exercise 6.13 †
| Element | What it contributes |
|---|---|
| Links to animal studies | Scientific legitimacy; implies an evidence base without characterizing its rung |
| Purity certificate | Implies the important question is purity, and that it is settled — reframing the decision away from efficacy |
| Dosage calculator | Presupposes human administration; provides operational detail no reagent buyer needs |
| Bacteriostatic water, syringes | Completes the human-use kit; makes the intended use unambiguous |
| Customer reviews describing outcomes | Social proof; supplies the efficacy claim the vendor never makes |
| "We make no medical claims" | Legal insulation |
What the assembly communicates: that this compound works for the conditions in the linked studies, that quality is the main thing to worry about, that quality is confirmed, and that here is how to use it. Every one of those propositions has been conveyed and none has been stated.
The diagnostic question: what is this page for? A supplier selling reagents to laboratories needs none of that infrastructure — no dosage calculator, no injection supplies, no outcome reviews. The infrastructure reveals the intended use more reliably than the disclaimer denies it.
Exercise 6.16 †
Miracle register: "The peptide that could end chronic injury — what the research shows."
Menace register: "The unapproved drug being injected by amateurs, with no safety data."
Both are technically accurate — there is research, it is unapproved, and there is no human safety data.
Honest version: "A peptide with substantial and reasonably consistent evidence of accelerated healing in rodent injury models has attracted a large following among athletes, despite having no completed randomized human trials for any indication and no published human safety data. Its supporters point to the animal literature; its critics point to the absence of human evidence. Both are correct, and roughly nine in ten compounds that look this good in animals fail somewhere in human development."
Word count: roughly 70, versus 11 and 13. And it ends in "we don't know," which is why it cannot be a headline. The structural point is that no individual journalist has to do anything wrong for the short versions to win.
Exercise 6.19 †
BPC-157. Genuinely true: a substantial, often well-conducted rodent literature across tendon, gut, vascular, and neuro models, from multiple groups. Distortion enters: at the extrapolation step — volume of animal work presented as though it accumulated into human evidence, when a hundred studies on the bottom rung remain on the bottom rung.
NAD+ precursors. Genuinely true: real and central biochemistry; the precursors demonstrably raise NAD+ in humans. Distortion enters: at the surrogate step — a moving biomarker presented as a demonstrated benefit, plus model-organism healthspan results presented as human ones.
GHK-Cu. Genuinely true: decades of genuine wound-healing and matrix-biology science. Distortion enters: at the delivery step — laboratory and wound findings cited as cosmetic efficacy on intact skin, with the penetration question, which is decisive, unaddressed.
The pattern: the underlying science is real in all three; the distortion enters at one identifiable step; and the step differs. Hype is a real finding with a step removed.
Exercise 6.22 †
No single answer. A good response identifies a compound, names the genuine underlying finding specifically (with rung), names the claim being made, and locates the missing step precisely — species extrapolation, surrogate-for-outcome, delivery, dose, or population.
The most common error in student answers is stopping at "there's no evidence." That is a conclusion, not an analysis. The exercise asks where the distortion enters, which requires taking the real finding seriously first. An answer that cannot state what is genuinely true has not done the work and will not persuade anyone who believes the claim.
Exercise 6.24 †
Why mechanistic understanding increases susceptibility:
Chapter 2's seven-step gap separates receptor activation from patient benefit. A reader who understands receptor pharmacology can follow step 1 in vivid detail — they can picture the binding, the cascade, the downstream effect. Steps 2 through 7 are abstract by comparison; they involve absorption, concentration, compensation, and outcomes, none of which produces a mental image.
Vividness is not evidence, but it feels like it. The available-to-imagination part of the argument is step 1, so the reasoning fills in the rest. Someone with no mechanistic knowledge cannot make this error as forcefully, because they cannot picture step 1 either.
And knowing about the bias does not switch it off. This is the crucial and uncomfortable part. The mechanism remains vivid after you have been warned.
The procedural defense: write out the seven steps explicitly and mark each E (evidence), A (assumed), or U (unknown), as in exercise 2.30. This converts a vivid impression into a countable list, and the count is not susceptible to vividness. Procedure works where attitude does not.
Exercise 6.27 †
A model answer:
That's a really specific account and I don't think he's making it up — his shoulder probably did get better. The tricky part is that we can't tell from one person whether the compound did it. People usually start something when they're at their worst, injuries often improve on their own eventually, and anyone who's spent money and effort on something tends to notice improvement. Plus you're only seeing the videos from people it seemed to work for — the ones who felt nothing don't post.
That's not a knock on him. It's just that one story, even a completely honest one, can't answer the question. What would answer it is a trial where people with the same injury got either the compound or a placebo and nobody knew which — and as far as I can find, nobody's run one for this.
Note what the answer does: concedes the person's honesty first, explains the mechanisms without naming them as biases, and ends with what would settle it rather than with a verdict on the friend.
Exercise 6.31 †
A defensible ranking, with reasoning:
1. Stage ⑤ (search results) — highest leverage, hardest to change. This is where nearly every reader actually encounters the compound. Improving what ranks would change more minds than anything else. Who could intervene: search and video platforms, through ranking policy. Whether they will is a different question, and the incentives point the other way.
2. Stage ② (testimonials) — high impact, some tractability. Platform-level context labels, or community norms that ask for a comparison group, would help. Some communities do this already.
3. Stage ④ (media) — moderate impact, moderate tractability. Style guides requiring absolute alongside relative figures, and requiring the evidence rung to be stated, would improve coverage measurably. Some outlets have adopted versions of this.
4. Stage ③ (gray market) — high impact on harm, low tractability. Enforcement is jurisdictionally fragmented and the market is adaptive. Chapter 38.
5. Stage ① (preclinical) — low leverage. The papers are already honest. Preclinical registration would help the literature's reliability considerably (Chapter 5, Case Study 2) but would not affect what a reader encounters at stage ⑤.
The instructive result: leverage and tractability are inversely related here. The place where intervention would help most is the place where the intervening parties have the least incentive.
Exercise 6.33 †
No answer key — but a note on what to check.
Your falsifier must be handable to a trialist. "More research" is not a falsifier. "A randomized, placebo-controlled trial in people with chronic tendinopathy, n ≥ 100, at least 12 weeks, with a pre-specified functional endpoint, showing no difference from placebo" is.
The "what would NOT change my mind" line is doing more work than it appears to. It inoculates you against evidence that feels decisive and isn't — another animal study, a testimonial, a purity certificate, an expert's opinion. Most people who change their mind about a health claim do so on the strength of something in that category, and having named them in advance is the only reliable defense.
And check achievability. If the only thing that would change your mind is a study nobody will ever run — because there is no patent, no regulatory obligation, and no funder — then your position is unfalsifiable in practice, whatever it looks like on paper. That is worth knowing about yourself before Chapter 17, and it is not a reason to abandon the position. It is a reason to hold it with the appropriate grip.
Chapter 7
Exercise 7.2 †
GLP-1 is made by L cells, concentrated distally — in the ileum and colon, though present throughout the intestine.
GIP is made by K cells, concentrated in the upper small intestine — duodenum and proximal jejunum.
Why the distribution matters: most nutrient absorption occurs in the upper small intestine, so in normal anatomy relatively little undigested nutrient reaches the distal L-cell territory. Nutrients arriving there signal that a large meal has come in — the ileal brake. This is also the anatomical basis for Case Study 2: a gastric bypass delivers nutrients to L-cell territory early and largely undigested, which is why the operation dramatically increases GLP-1 responses.
Exercise 7.4 †
| Action | Tissue | Effect |
|---|---|---|
| ① Glucose-dependent insulin secretion | pancreatic beta cell | amplifies insulin release that glucose is already driving |
| ② Glucagon suppression | pancreatic alpha cell | reduces hepatic glucose output |
| ③ Slowed gastric emptying | stomach | flattens post-meal glucose rise; prolongs fullness; source of the nausea |
| ④ Reduced appetite | brain (hindbrain and hypothalamus) | reduces food intake |
Note that ①–③ all lower glucose by different routes, and ③–④ both reduce intake. It is a coordinated program rather than a single effect, which is why one receptor produces so many therapeutic effects.
Exercise 7.8 †
Glucose-dependence means GLP-1 amplifies insulin secretion only when glucose is already driving it. Glucose entering the beta cell is the primary trigger; GLP-1's cAMP signal increases the cell's responsiveness to that trigger. It turns up the gain rather than switching the system on.
The safety consequence: at normal or low blood glucose there is little primary signal to amplify, so little insulin is released. GLP-1 receptor agonists used alone therefore rarely cause hypoglycemia — unlike insulin and unlike sulfonylureas, which drive insulin release regardless of glucose. This is the main reason the class succeeded where earlier insulin secretagogues were limited, because hypoglycemia is what constrains how aggressively glucose can be lowered.
Exercise 7.10 †
Alone, they rarely cause hypoglycemia because of glucose-dependence: at normal glucose, there is little insulin secretion to amplify.
In combination they can contribute because insulin and sulfonylureas drive insulin release independently of glucose. A GLP-1 agonist added on top does not remove that glucose-independent drive; it adds further glucose-lowering through glucagon suppression and slowed gastric emptying, which can push glucose low enough that the glucose-independent agent's insulin becomes excessive.
This is why adding a GLP-1 receptor agonist to insulin or a sulfonylurea frequently requires reducing the dose of the other agent — a clinical decision belonging to the prescriber.
Exercise 7.13 †
The puzzle: DPP-4 is abundant on capillary endothelium in the gut wall itself, so much GLP-1 is inactivated before leaving the intestinal circulation, and hepatic first-pass extraction removes much of the remainder. Very little intact GLP-1 reaches the systemic circulation — yet GLP-1 has systemic effects.
The proposed resolution: much of endogenous GLP-1's action is local. Vagal afferent nerve endings in the gut wall carry GLP-1 receptors and sit within a very short distance of the L cell. Locally released GLP-1 activates them before DPP-4 can act, and the vagus relays the signal neurally to the hindbrain. The hormone functions more as a paracrine signal with a neural relay than as a classical endocrine one.
What it implies about injected agonists: a long-acting DPP-4-resistant agonist produces sustained systemic exposure and reaches receptors directly — including in the area postrema and NTS, where the blood-brain barrier is incomplete. This is not simply more of the natural signal; it may be a different signal acting partly through different routes. That is a candidate explanation for why the drug effect (weight loss of ~15%) so far exceeds anything endogenous GLP-1 achieves, and it is a reason to be cautious about describing these drugs as "restoring" natural signaling.
Exercise 7.15 †
The four proposed explanations for why adding GIP agonism helps:
- Central appetite effect. GIP receptor agonism in the brain may contribute to appetite reduction independently of peripheral metabolic effects.
- Restoration with improved glycemia. GIP's blunted insulinotropic response in diabetes may recover once glucose control improves, so GIP agonism becomes useful after GLP-1 agonism has acted.
- Functional antagonism through desensitization. Sustained GIP receptor agonism may desensitize the receptor, so a chronic agonist behaves functionally like an antagonist — which would reconcile the animal knockout data (where GIP receptor deletion protected against obesity) with the clinical results.
- Outweighed. The peripheral fat-storage effect may simply be smaller than the combined benefit.
A defensible assessment: explanation 1 is the most plausible on current evidence, because the appetite effect is where tirzepatide's advantage is largest and because central GIP receptors are documented. Explanation 3 is the most interesting and the least established — it would be a striking result if true, and Chapters 2 §2.7 and 3 §3.5 make it entirely coherent rather than absurd. Explanation 2 is testable and under-tested. Explanation 4 is the least informative because it does not predict anything.
The honest conclusion: the mechanism is not understood, and the fact that both GIP agonists and GIP antagonists are in development is the strongest available evidence of that.
Exercise 7.18 †
GLP-1 agonists do not suppress endogenous GLP-1 because L cell release is triggered by nutrients in the gut lumen, not by a sensor comparing circulating GLP-1 against a set point. There is no upstream tier measuring the total and turning itself down.
Exogenous growth hormone does suppress endogenous GH because the GH axis is a hierarchical feedback system: elevated GH and the IGF-1 it produces are detected and inhibit GHRH release while raising somatostatin. The system measures total and corrects.
The general rule: expect feedback suppression where there is an upstream sensor comparing total circulating hormone against a set point — and do not assume one exists. Getting this right requires knowing the specific architecture; not every hormone system is an axis.
Exercise 7.21 †
Where it sits on the ladder: low. Self-reported subjective experience in unblinded, highly motivated patients, with no validated measurement instrument, and a term that originated in patient communities. On Chapter 5's ladder this is between anecdote and case series — and it is a large volume of consistent anecdote, which is not the same as evidence.
What would be required to study it properly: first, a validated instrument — a scale for intrusive food-related cognition with demonstrated reliability and validity. Then a randomized, blinded design measuring it as a pre-specified endpoint. Blinding is genuinely difficult here, because the drug produces noticeable gastrointestinal effects that let participants guess their assignment (Chapter 5 §5.5), so an active comparator producing similar side effects would strengthen it considerably.
Why dismissing it would be both rude and scientifically wrong: the observation came from patients, consistently, unprompted by any pharmaceutical hypothesis, and it describes something that the mechanism makes entirely plausible — GLP-1 receptors are present in brain reward and appetite circuitry. Chapter 5's rule is that a claim's origin does not determine its truth value. A consistent, unprompted, mechanistically plausible patient report is an excellent reason to build the instrument and run the study. It is not a substitute for having done so.
Exercise 7.22 †
A model answer:
If you drink a sugary drink, your pancreas releases a certain amount of insulin. If you instead get exactly the same amount of sugar injected into a vein, so your blood sugar rises identically, your pancreas releases considerably less insulin. The difference is that eating triggers your intestine to send hormones ahead of the food, warning the pancreas that sugar is on the way — so your body prepares rather than just reacting. Those hormones are called incretins, and one of them, GLP-1, is what drugs like Ozempic are copies of.
Exercise 7.26 †
Two claims this should make you more cautious about:
"These drugs restore natural satiety signaling." If endogenous GLP-1 acts largely locally through vagal afferents in brief meal-triggered bursts, and an injected weekly agonist produces sustained systemic exposure reaching brain receptors directly, then the drug is not restoring the natural signal. It is producing a different one. "Restore" also implies a prior deficiency that has not been established.
"The side effects are just your body adjusting to normal signaling." Nausea arises substantially from slowed gastric emptying and from hindbrain receptor activation at concentrations physiology never produces. It is not the natural signal being re-established; it is one of the four actions overshooting at supraphysiological exposure.
A third, if you want it: "Because it works through a natural pathway, long-term safety is predictable." Sustained supraphysiological agonism through partly different routes is not a condition physiology has ever encountered, so physiological experience is a weak guide to what decades of it will do. That is not a prediction of harm — it is a statement that the reassurance is weaker than it sounds.
Exercise 7.28 †
No single answer. Characteristic findings:
The "load-bearing action" line is where most dossier entries improve most. For a multi-action peptide, identifying which action produces the claimed benefit forces a precision that most sources avoid. Frequently the honest entry is "not established which," and writing that is more informative than listing four actions.
The "routes" line will separate approved drugs from research compounds sharply. For GLP-1 agonists you can say something substantive about neural versus humoral routes and about how drug and endogenous signaling may differ. For most research-code compounds, the route is unstudied — and for several, the receptor itself is unidentified, which makes the route question unanswerable in principle.
And the "what is NOT known" line will be longest for the best-evidenced compound in your dossier. This is the calibration that matters. GLP-1's mechanism has genuine unresolved questions precisely because enough work has been done to identify them. A compound with a short "not known" line usually has one because nobody has looked, not because everything is settled.
Chapter 8
Exercise 8.2 †
SUSTAIN — semaglutide in type 2 diabetes. Established substantial A1C reduction against placebo and active comparators, weight loss as a secondary finding, and — in SUSTAIN 6, a cardiovascular outcomes trial — a statistically significant reduction in major adverse cardiovascular events in a study designed only to rule out harm.
STEP — semaglutide 2.4 mg in weight management. About −15% from baseline at 68 weeks in adults with overweight or obesity without diabetes versus roughly −2.4% on placebo; about −10% in adults with type 2 diabetes. Both arms received lifestyle intervention.
PIONEER — oral semaglutide. Established that an orally administered peptide could produce clinically meaningful glycemic control, at roughly 1% bioavailability with a strict fasting protocol.
SELECT — cardiovascular outcomes in adults with established cardiovascular disease and overweight or obesity without diabetes. Roughly 17,000 participants over about three years: a 20% relative reduction in MACE, roughly 8% to 6.5% absolute, about 1.5 percentage points.
Exercise 8.5 †
Relative: a 20% reduction in major adverse cardiovascular events; hazard ratio about 0.80.
Absolute: roughly 8% to roughly 6.5% over about three years — about 1.5 percentage points.
Population and duration: adults with established cardiovascular disease and overweight or obesity, without diabetes; roughly 17,000 participants; about three years of follow-up.
Approximate NNT: 1 ÷ 0.015 ≈ 67 — roughly 65 to 70 people treated for about three years for one additional person to avoid a major cardiovascular event.
All four elements are required. The relative figure alone overstates personal benefit by roughly a factor of ten; the absolute figure alone understates population importance; and either without the population is a claim about people who were not studied.
Exercise 8.8 †
The difference: STEP 1 (no diabetes) produced about −15%; STEP 2 (type 2 diabetes) about −10%. Same molecule, same 2.4 mg dose, same 68 weeks.
Contributing explanations (none individually settled, and probably operating together):
- Altered incretin responsiveness in type 2 diabetes. The incretin effect is substantially reduced in the disease (Ch 7 §7.1), and while GLP-1's insulinotropic action is largely preserved, the broader metabolic response may differ.
- Concurrent medications. Many participants with type 2 diabetes take insulin, sulfonylureas, or other agents that promote weight gain, working against the drug.
- Differences in metabolic state — including insulin resistance, which affects how the body responds to reduced intake.
The error produced by quoting one for the other: a population swap (Ch 5 §5.9's Trap 3). A person with type 2 diabetes told to expect 15% has been given a figure from a population they are not in, and will conclude the drug failed when it performs exactly as the relevant trial predicted. This is the most common error in popular coverage of this drug and it is entirely avoidable.
Exercise 8.11 †
Weight is a surrogate endpoint: it is measured because it is believed to predict cardiovascular events, diabetes, and mortality. The association between adiposity and cardiovascular risk is one of epidemiology's most robust, so it is a better-supported surrogate than most.
And it was insufficient, for a reason specific to this field: obesity pharmacology has repeatedly produced drugs that improved weight and harmed patients. Fenfluramine reduced weight and damaged heart valves. Sibutramine reduced weight and increased cardiovascular events. Rimonabant reduced weight and produced psychiatric harm. In each case the surrogate moved correctly while the outcome did not.
SELECT measured the outcome directly — cardiovascular death, myocardial infarction, and stroke, in a population without diabetes, over about three years. That converts the claim from "improves a risk factor" to "reduces events," which is categorically stronger and which cannot be undermined by the historical failure mode.
Chapter 5 §5.6's rule applies: a surrogate is only as good as the evidence that changing it changes the outcome, and that evidence is a separate research program. SELECT is that program.
Exercise 8.13 †
A model answer, under 100 words:
In studies on rats, this class of drug caused a rare kind of thyroid tumor. Rat thyroid cells carry far more of the relevant receptor than human thyroid cells do, which is a good reason to think the finding may not apply to people — and no increase has been seen in humans. But the tumor is rare enough that a small increase would be hard to detect, so the warning stays and monitoring continues. The one situation where it genuinely matters: if you or a close relative has had medullary thyroid cancer, or a condition called MEN2, this drug isn't for you.
Exercise 8.16 †
Why "the mechanism overshooting": slowed gastric emptying is one of GLP-1's four intended actions (Ch 7 §7.3), and it is the action that produces both prolonged fullness — a therapeutic effect — and nausea, early satiety, and vomiting when it goes too far. Hindbrain receptor activation in the area postrema contributes, and that region is also the vomiting trigger zone. The benefit and the side effect share a mechanism.
What follows for a patient:
Titration makes sense. Starting low and increasing stepwise allows accommodation, and this is a pharmacological rationale rather than an administrative one.
Attenuation is expected. Tolerance is a property of drug-tissue pairs (Ch 2 §2.8); the gastric receptors desensitize faster than the ones producing metabolic benefit, which is why the nausea typically fades while the effect persists.
And "this side effect means it's working" is nearly true and should not be said — it is a plausible-sounding framing that would discourage reporting genuinely problematic symptoms. The accurate version: the nausea comes from the same action that produces fullness, it usually settles, and severe or persistent symptoms are a reason to contact a clinician rather than to endure.
Exercise 8.18 †
Lean mass is a surrogate for strength and physical function — for what a person can actually do.
Why trials finding improved physical function complicate the concern: if lean mass loss were producing meaningful functional decline, you would expect function to worsen. Several trials measuring physical function on validated instruments found it improved — which is what you would predict if the functional benefit of losing substantial fat mass outweighs the cost of losing some lean mass.
What this does not settle:
- Population. Trial populations skew younger and healthier than the people most vulnerable to lean mass loss. In an older adult with low baseline muscle, the arithmetic could easily run the other way.
- Duration. Function measured over 68 weeks says little about function after ten years of continuous therapy.
- The instruments. General physical function measures may not detect the specific losses that matter for frailty.
The honest position: lean mass loss is real and is a property of weight loss generally rather than of this drug class; the functional consequence is not established and appears net favorable in studied populations; and the population where it would matter most is the least studied. What an individual should do about it — protein, resistance training, monitoring — belongs with their clinician.
Exercise 8.20 †
Semaglutide overrides an intact system. Appetite regulation is functioning; it is defending a set point; the drug outweighs it while present. The set point was never removed.
Compare insulin in type 1 diabetes, which replaces an absent signal — the beta cells are destroyed, there is no intact system being overridden, no counter-regulation is provoked, and the effect persists indefinitely.
When semaglutide stops, the override is removed and the intact system resumes. Weight substantially returns, and cardiometabolic improvements move back toward baseline alongside it.
Chapter 2 §2.8 stated the rule: expect tolerance and reversal when you override an intact regulated system; expect durability when you replace an absent one. Chapter 13 shows a long history of appetite drugs defeated by exactly this mechanism.
Exercise 8.23 †
A defensible answer: it is not a treatment failure. It is a system failure, and the distinction matters for what should be done about it.
Why not a treatment failure: the drug did what the evidence says it does, for as long as it was taken. Regain on discontinuation is expected physiology, is documented in the trials, and is the normal behavior of chronic therapy. Nothing about the pharmacology underperformed.
Whose failure, then: the answer depends on your view of what a health system owes. A defensible position is that if a therapy is effective, its benefit ends on discontinuation, and discontinuation is driven by cost rather than by tolerance or preference, then the outcome was determined by an access decision rather than a clinical one. The person did not fail the treatment and the treatment did not fail the person; the financing did.
The complication worth acknowledging: health systems have finite resources and must make coverage decisions, and "this drug works and is expensive and must be taken indefinitely" is a genuinely hard problem rather than a simple injustice. A drug that must be continued for life in a large population at current prices poses budget questions that do not have obvious answers. Chapter 12 takes this up properly.
What the answer should avoid: framing regain as a personal failure. The physiology is doing exactly what Chapters 2 and 3 predict, and attributing it to willpower is both inaccurate and — given the moral framing that surrounds weight — actively harmful.
Exercise 8.24 †
What I would ask: "In whom, and over what period?" And, if they know: "was that the analysis counting everyone, or only the people who stayed on it?"
What I would say:
That figure comes from a trial in adults who were overweight or had obesity and didn't have diabetes, over about 68 weeks, with both groups also getting structured diet and activity support. In people with type 2 diabetes the same drug and dose produced about 10%, which is still a lot but is meaningfully different. So the number is real — it's just a number about a specific group of people over a specific time, and which group matters.
The other thing worth knowing: the weight comes back when you stop. Not instantly and not all of it, but substantially. That doesn't mean the drug doesn't work — it works the way a blood pressure medication works, for as long as you take it — but it does change what starting it means.
Exercise 8.28 †
The argument: if semaglutide reduces cardiovascular events in high-risk people, it should also reduce them in lower-risk people, since the mechanism does not care about baseline risk.
Why it is weaker than it sounds:
Absolute benefit scales with baseline risk (Ch 5 §5.8). Even granting an identical 20% relative reduction, a population with a 2% three-year event rate would see about 0.4 percentage points of absolute benefit — an NNT of roughly 250 rather than roughly 67. The adverse effects do not scale down; they occur at the same rate regardless of baseline risk. So the benefit-to-harm ratio is substantially less favorable in a lower-risk population, even if the relative benefit is identical.
And the relative reduction may not be identical. Populations differ in the mechanisms driving their risk. If the benefit is mediated by something specific to established atherosclerotic disease, it may not transfer at all. Case Study 1's six candidate mechanisms are unresolved, which means we cannot predict transferability from mechanism.
What would settle it: a randomized cardiovascular outcomes trial in a primary-prevention population — no established cardiovascular disease — powered for a hard endpoint. Such a trial would be substantially larger and longer than SELECT, because the event rate is lower.
Note that this is the reasoning behind the chapter's refusal to rate primary prevention. Declining to extend a rating is not caution for its own sake; it is the recognition that the arithmetic of absolute benefit changes the answer.
Exercise 8.30 †
No single answer. Characteristic findings when readers compare their entry to the model:
Fields 1–4 will be comparable in length for most compounds, because identity, origin, mechanism, and pharmacology can usually be assembled for anything with a published sequence.
Field 5 is where entries diverge dramatically, and the divergence is the finding. For semaglutide, Field 5 lists four trial programs and then a "what does not exist" line of four items. For a research-code compound, Field 5 often has no human trials to list and a "what does not exist" line longer than everything above it.
Field 6 will need multiple ratings for anything approved and frequently exactly one — or none — for anything not.
Field 9 (Risks) is the most revealing. For an approved drug it can be assembled from the label, with rates. For an unapproved compound the honest entry frequently reads "not systematically studied in humans," which is a very different statement from "no known risks" and should be written as the former.
And Field 12 will be hardest for the compound you care most about. That was true in Chapter 6 and it remains true here.
Chapter 9
Exercise 9.2 †
SURPASS — type 2 diabetes. Supported approval for glycemic control in 2022 (Mounjaro). SURMOUNT — obesity. Supported approval for weight management in 2023 (Zepbound).
Same molecule, two brand names, two indications — the identical pattern to Ozempic and Wegovy, and it confuses patients for the identical reason.
Exercise 9.5 †
Retatrutide targets the GLP-1, GIP, and glucagon receptors — a triple agonist.
Its reported result of approximately −24% body weight at 48 weeks at the highest dose is from a Phase 2 trial. This book states the phase every single time the figure appears, because the number is larger than an approved drug's Phase 3 number and the comparison is a category error.
Exercise 9.8 †
Treatment-regimen estimand — the question: "What happens if a doctor prescribes this to someone?" It counts everyone who was randomized, including participants who stopped taking the drug. This is the question a prescriber and a patient are actually asking, because "I might be one of the people who can't tolerate it" is part of what they are deciding about.
Efficacy estimand — the question: "What happens to someone who takes this, as prescribed, for the whole period?" It restricts to participants who remained on treatment. This is closer to the question a pharmacologist asks about the molecule's biological effect.
The framing that matters: these are not two analyses of the same question, one better than the other. They are answers to two different questions, both of which people genuinely ask. Neither is "the" result.
Exercise 9.11 †
(a) A patient deciding whether to start — treatment-regimen. They are deciding whether to be prescribed the drug, and the probability that they will be among those who stop is part of what they are deciding about. Giving them the efficacy figure conditions on an outcome they cannot guarantee.
(b) A pharmacologist studying the effect on fat mass — efficacy comes closer, since the question is about the drug's action rather than about a prescribing decision. Even this is imperfect, because the people who remained are a selected group.
(c) A health system forecasting population outcomes — treatment-regimen, emphatically. The system will pay for everyone prescribed, including those who stop, and needs to know what the whole cohort achieves.
(d) A marketing department — the efficacy figure, because it is larger. This is not a criticism; it is a prediction, and it is why an unnamed figure should be assumed to be this one.
Exercise 9.14 †
"Which is better?" decomposes into at least six questions with different answers:
- Better for what? Weight, glycemic control, cardiovascular outcomes, and quality of life are different endpoints with different evidence.
- Better in whom? Diabetes versus no diabetes changes the answer by roughly a third for weight, consistently, for both drugs.
- Better tolerated? Both are gastrointestinal-dominant; comparative tolerability at equipotent doses is a genuine and under-answered question.
- Better evidenced? Semaglutide has SELECT — a hard-endpoint cardiovascular outcomes trial in a population without diabetes. Tirzepatide's outcomes program is ongoing. On this specific question the evidence is not equivalent, and greater weight loss does not substitute.
- Better available? Both have experienced supply constraints. Not a pharmacological property, and it determines what actually gets taken.
- Better affordable? Coverage, list versus net price, and jurisdiction. Chapter 12.
Which matters most depends on the person, which makes this a conversation with a clinician rather than a fact about molecules.
Exercise 9.17 †
Five mechanisms; any three suffice:
Small samples produce high variance. A trial of 80 people yields a much wider spread of possible estimates than one of 2,000, and the unusually favorable estimates are the ones that generate excitement and funding.
Regression to the mean, via selection. Compounds advance to Phase 3 because they produced a good Phase 2 result. That selection enriches the Phase 3 population of compounds for overestimates, so the Phase 3 figure is expected to be smaller on average even if the compound works exactly as well as it truly does. This is the least intuitive mechanism and the most important.
Best-dose selection. Phase 2 studies several doses and the best-performing one is quoted and carried forward. Selecting a maximum from several noisy estimates produces a biased maximum.
Favorable populations. Phase 2 cohorts are frequently narrower, healthier, more adherent, and more motivated than Phase 3 cohorts, and far more than real-world populations.
Shorter duration. Effects that attenuate look better measured earlier; adverse effects that accumulate look better too.
None requires misconduct. All five operate on competent investigators running well-designed trials in good faith, which is exactly what makes the pattern reliable.
Exercise 9.20 †
No single answer. What a good response records:
- The phase — and whether it was stated at all. Frequently it is not, or appears only in a parenthesis.
- The estimand — almost never named outside the trial publication itself.
- The dose — often present, often as "up to."
- The population — usually present in general terms ("adults with obesity") and rarely with the diabetes status, which is the qualifier that matters most.
- The endpoint — usually present; whether it is a surrogate is almost never flagged.
- The source type — press release, conference abstract, or peer-reviewed publication. Frequently ambiguous by design.
The characteristic finding is that announcements answer questions 3, 4, and 5 reasonably well and questions 1, 2, and 6 poorly — which is exactly the pattern you would predict, since those three are the ones whose honest answers would reduce the impressiveness of the number.
Exercise 9.22 †
The two reasons GIP was written off:
- Its insulinotropic effect is substantially blunted in type 2 diabetes, while GLP-1's is largely preserved — which removed the obvious therapeutic rationale in the disease it would be used to treat.
- GIP appears to promote fat storage. GIP receptors are present on adipocytes, and GIP receptor deletion in animal models has been reported to protect against diet-induced obesity.
The four candidate explanations, evaluated:
Central appetite effect — most plausible on current evidence. Central GIP receptors are documented, and tirzepatide's advantage over GLP-1 agonism alone is largest on weight, which is consistent with an appetite mechanism rather than a peripheral metabolic one.
Restoration with improved glycemia — testable and under-tested. It predicts that GIP agonism should contribute little early and more later, which is a measurable prediction.
Functional antagonism through desensitization — the most interesting and least established. It would reconcile the animal knockout data with the clinical results, and Chapters 2 §2.7 and 3 §3.5 make it entirely coherent rather than absurd. It also predicts that a GIP antagonist should produce similar effects — which is being tested, since both are in development.
Outweighed — the least informative, because it predicts nothing and cannot be distinguished from the others by any experiment.
The honest conclusion: unresolved, and the simultaneous pursuit of GIP agonists and antagonists is the strongest available evidence of that.
Exercise 9.25 †
A model answer:
In a long trial, some people stop taking the drug — side effects, life gets in the way, whatever. So there are two fair ways to report the result. You can count everyone who was assigned the drug, including the ones who quit, which tells you what happens if a doctor prescribes it. Or you can count only the people who actually stuck with it, which tells you what the drug does if you take it. Both are honest, and the second number is always bigger — because the people who quit are usually the ones it was going worst for. So when you see a headline number, it's almost always the bigger one, and nobody mentions which.
Exercise 9.29 †
The argument to evaluate: semaglutide carries ✅ for both weight and cardiovascular outcomes; tirzepatide carries ✅ for weight only; therefore semaglutide is better.
What is right about it: on the specific question of cardiovascular outcome evidence, semaglutide genuinely has something tirzepatide does not. SELECT is a hard-endpoint trial in roughly 17,000 people over about three years, and tirzepatide's outcomes program is ongoing. That is a real and substantial difference in the evidence base, and it should influence a clinical decision for a patient whose primary concern is cardiovascular risk.
What is wrong about it:
"Not rated" is not "rated negatively." Tirzepatide's cardiovascular status is unknown, not unfavorable. The absence of a completed outcomes trial is a fact about the literature, exactly as Chapter 5 insists.
The ratings are not commensurable as a score. Counting ✅s treats the rating system as a scoreboard, which is precisely the compression rule 6 forbids. A drug with three ✅ claims and one serious harm is not "better" than one with two ✅ claims and none.
And it ignores what tirzepatide's ✅ covers. On weight specifically, tirzepatide produced the largest effect of any approved pharmacotherapy. For a patient whose primary goal is weight reduction, that is the relevant comparison.
The defensible conclusion: these are different advantages. Semaglutide has outcome evidence; tirzepatide has a larger effect on the surrogate. Which matters depends on the patient's situation, and "better" without a specified purpose is not a question with an answer.
Exercise 9.31 †
No single answer. Characteristic findings:
Most readers cannot find two peptides in their dossier used for the same purpose, which is itself worth noticing — the compounds people care about tend to be spread across indications rather than competing.
Where a comparison is possible, the last row does the work. Two compounds almost always differ in trial duration, population, comparator, endpoint, or funder, and attributing the whole apparent difference to the molecules is the standard error. In the semaglutide/tirzepatide comparison, the differences that are not about the molecules include four weeks of trial duration, slightly different inclusion criteria, and — most importantly — a head-to-head conducted in diabetes at a comparator dose below semaglutide's maximum.
And the "cardiovascular outcomes" row frequently ends the argument. For pairs where one compound has hard-endpoint data and the other does not, that row is usually more decisive than any difference in effect size — and it is the row most often missing from comparisons made outside this book.
Chapter 10
Exercise 10.2 †
FLOW studied semaglutide against placebo in adults with type 2 diabetes and chronic kidney disease, with a composite kidney endpoint (progression of kidney disease, kidney failure, and death from kidney or cardiovascular causes).
It was stopped early for efficacy, on the recommendation of its independent data monitoring committee, because an interim analysis crossed the pre-specified benefit boundary.
Exercise 10.5 †
The apnea-hypopnea index is the number of apnea and hypopnea events per hour of sleep — the standard objective measure of obstructive sleep apnea severity.
Why sleep apnea is the clearest weight-mediated case: the airway collapses because soft tissue surrounding it makes it collapsible. Reducing that soft tissue reduces collapse. No direct drug effect on the airway is required for the mechanism to work, and the association between adiposity and sleep apnea is one of the strongest in the condition's epidemiology.
Compare the cardiovascular result, where benefit appeared earlier and larger than weight loss alone comfortably explains. Two indications, two different answers to "is it just the weight?"
Exercise 10.7 †
Good news: a treatment is working clearly enough that an independent committee judged continuing to assign participants to placebo no longer defensible. Participants benefit sooner, and the result reaches everyone else sooner.
Statistical complication: the trial stopped because an interim analysis crossed a boundary — and interim analyses that cross boundaries are, on average, the ones at which random variation happened to favor the treatment most. The stopping rule selects for chance-favorable looks. Consequences: the point estimate is biased upward, the confidence interval is wider because fewer events accrued, and secondary endpoints are badly underpowered.
The right reading: direction trustworthy, magnitude held loosely.
Exercise 10.11 †
Four problems with a liver biopsy endpoint:
- Invasive. A needle into the liver carries real if small risk, so it cannot be repeated often, which limits both trial design and follow-up.
- Sampling error. A biopsy takes a tiny fraction of the organ, and the disease is patchy. Two biopsies from the same liver can disagree.
- Reader variability. Histological scoring is a pathologist's judgment, and pathologists disagree with each other and sometimes with themselves.
- It is still a surrogate. Histology stands in for cirrhosis, liver failure, and death — the outcomes that actually matter — which accrue over many years.
The fourth is what makes it a surrogate. Problems 1–3 make it a difficult measurement; problem 4 means that even a perfect measurement would not directly answer the clinical question.
Exercise 10.15 †
| Indication | More plausibly | Reason |
|---|---|---|
| Sleep apnea | weight-mediated | Less soft tissue around a collapsible airway is a complete mechanical explanation |
| HFpEF symptoms/function | probably substantially weight-mediated | People who weigh less walk further and report less breathlessness; the trials cannot separate this |
| Cardiovascular events | probably not entirely weight-mediated | Benefit appeared earlier and larger than weight change alone comfortably explains |
| Chronic kidney disease | unclear | Both weight and direct renal receptor effects are plausible; the trial does not distinguish |
| MASH | probably substantially weight-mediated | Liver fat responds strongly to weight loss by any means |
| Alzheimer's | unknown, and no established benefit to explain | |
| Addiction | probably not weight-mediated | The reported effect is on appetitive behavior rather than on body composition |
The exercise's value is that it forces separate judgments where a single narrative would flatten them.
Exercise 10.18 †
A design that would help: a randomized trial comparing the drug against an intervention producing matched weight loss by another means — intensive lifestyle intervention, a different weight-loss drug with no GLP-1 activity, or a calorically matched protocol — with cardiovascular events as the endpoint.
If the drug arm shows benefit beyond the matched-weight-loss arm, that supports a direct effect. If the two are equivalent, it supports weight mediation.
Limitations, and they are severe:
- Achieving matched weight loss is very hard. No other intervention produces comparable weight loss reliably, so the comparator arm will likely lose less, reintroducing the confound.
- Enormous size and duration. Cardiovascular events in a population healthy enough to randomize this way accrue slowly.
- Ethical difficulty. Once a drug has demonstrated cardiovascular benefit, randomizing high-risk patients away from it becomes hard to justify.
- Mediation analysis is the fallback and it is observational — it asks whether the benefit statistically tracks weight change, which is subject to all the usual confounding.
The honest conclusion: the question may not be cleanly answerable, which is itself worth knowing. Some mechanistic questions are practically undecidable, and treating "unresolved" as a temporary state that a better study will fix is sometimes wrong.
Exercise 10.20 †
The Alzheimer's claim has: a plausible mechanism (metabolic dysfunction associated with Alzheimer's risk; brain GLP-1 receptors), animal data reporting neuroprotection, and observational epidemiology reporting lower dementia incidence.
None of these is randomized human efficacy data.
Mechanism is Chapter 2 §2.9 — necessary and never sufficient. Animal data is Chapter 5 §5.3 — a reason to run a trial, not a preview of its result. Observational epidemiology is subject to confounding by indication, healthy adherer effects, protopathic bias, and surveillance effects, at least two of which cannot be adjusted away.
⚠️ requires real human data that does not settle the question — meaning randomized trials that were small, short, mixed, or surrogate-based. 🔬 is for early-stage science proceeding properly, where translation is unproven.
A program supported by mechanism, animals, and observation, with randomized efficacy results not yet established, is 🔬. Rating it ⚠️ would place it alongside compounds that have actually been tested in randomized humans and produced ambiguous results — a materially different epistemic situation.
And the field's history matters: Alzheimer's has an unusually poor record of mechanistically attractive hypotheses failing in trials, which is a reason for the more conservative rating rather than the less.
Exercise 10.23 †
Upward compression: treating established indications as licensing unestablished ones. "It's proven to reduce heart attacks, so the Alzheimer's research is probably right."
Downward compression: treating unestablished indications as discrediting established ones. "They're claiming it cures everything, so I don't believe the heart result either."
What they have in common: both attach a rating to a molecule rather than to a claim. Each assumes that evidence for one claim transfers to another, differing only in the direction of transfer. The claims rest on entirely different evidence bases — different trials, different populations, different endpoints, different stages — so neither inference is available.
Rule 6 prevents both: one molecule, many ratings, each attached to a claim with a population and an endpoint.
Exercise 10.26 †
A model answer:
Each of those is a separate question that needs its own study. The heart result comes from a trial of about seventeen thousand people followed for three years, where they counted actual heart attacks and strokes. The Alzheimer's work is at a much earlier stage — there's a good reason to think it might help, and trials are running, but nobody has the answer yet.
That's not a knock on the Alzheimer's research. That's how it's supposed to go: you notice something promising, then you test it properly. The mistake would be assuming that because one thing was proven, the other must be true too — they're different claims resting on completely different evidence.
Exercise 10.30 †
When a mechanical explanation is NOT a criticism: when the claim being rated is about an outcome, and the outcome occurred. "Tirzepatide reduces apnea-hypopnea index in adults with obesity and sleep apnea" is established by randomized trials. That the mechanism runs through weight loss does not make the reduction less real, and the person whose sleep apnea improved does not care by what route.
When a mechanical explanation IS a criticism: when it reveals that the claim has been described in a way the mechanism does not support. If a drug is marketed as "treating sleep apnea" in a way that implies a specific airway action, and the mechanism is generic weight loss, then the marketing has overclaimed — because it implies the drug would help sleep apnea that is not weight-related, which it would not.
The general principle: mechanism does not bear on whether an established outcome occurred. It bears on how far the result can be extrapolated, and on whether the claim as stated is supported. Naming the mechanism usually sharpens a claim rather than diminishing it — and occasionally reveals that the claim as marketed was broader than the evidence.
Exercise 10.32 †
No single answer. Characteristic findings:
Most readers discover more claims than they expected. A compound they think of as having one use turns out, on listing, to carry three or four claims picked up from different sources.
The ratings almost always diverge. If every row gets the same rating, either the compound genuinely has one claim — which is worth noting explicitly — or the reader has been rating the molecule rather than the claims.
And upward compression shows up most often in the "risks" adjacent claims. Readers who rated a compound ✅ for its established use frequently wrote about its safety in general terms, when the safety data covers only the studied population, dose, and duration. That is the same compression error operating on Field 9 instead of Field 6, and it is worth catching.
Chapter 11
Exercise 11.2 †
Insulin is 51 amino acids in two chains: an A chain of 21 residues and a B chain of 30. They are joined by two interchain disulfide bonds, and there is a third disulfide within the A chain. Molecular weight approximately 5,800 Da.
Why the bonds matter: remove them and you have two inactive peptide chains. The disulfides are the structure rather than a decoration, which is why recombinant production required solving a folding problem (Ch 32) and why storage conditions affect activity.
Exercise 11.5 †
Endogenous insulin regulation self-corrects because it has two levers. When glucose falls, beta cells stop releasing insulin and alpha cells release glucagon. Both act to restore glucose.
Injected insulin removes the first lever. The dose is already administered and cannot be withdrawn. Counter-regulation still operates — glucagon is still released — but it is now pushing against a fixed input it cannot influence, so glucose may continue falling.
Compounding this: insulin has no glucose-dependence. Unlike GLP-1, which amplifies a response glucose is already driving (Ch 7 §7.4), insulin operates switches directly. It moves glucose out of the bloodstream whether or not there is enough of it there.
And in long-standing type 1 diabetes it gets worse, because the counter-regulatory response itself becomes impaired — glucagon release in response to hypoglycemia is blunted, and the adrenaline response producing warning symptoms can fade, giving hypoglycemia unawareness.
Exercise 11.9 †
The general principle: the strength of evidence required scales with the plausibility of alternative explanations.
Randomization exists to exclude explanations that could be confused with a treatment effect — natural history, expectation, regression to the mean, and chance.
For insulin in 1922, none of those was available. Untreated type 1 diabetes did not remit. Patients did not spontaneously recover from ketoacidotic wasting. There was no plausible account on which a dying fourteen-year-old recovers, gains weight, and lives thirteen more years because of expectation.
Which is rare. For essentially every other claim in this book, at least one alternative explanation is live — which is why Chapter 5 exists. A field reasoning from insulin to "we don't need trials" would be catastrophically wrong, and Chapter 5's CAST case study is what that error looks like when the assumption happens to be false.
Exercise 11.12 †
See the diagram in §11.4 and answer 11.5. The essential structure:
Two levers, natural: glucose falls → insulin secretion stops and glucagon is released → glucose restored.
One lever, injected: glucose falls → the injected dose remains → glucagon is released → but it opposes a constant rather than a variable.
The key insight is that hypoglycemia on insulin therapy is not a dosing failure that better technique eliminates. It is structural: any therapy that supplies a hormone whose correct dose varies hour to hour, and which cannot be withdrawn once given, will produce excursions. That is why the entire subsequent history of insulin — analogs, pumps, sensors, closed loops — is about narrowing the gap between delivered dose and actual need rather than about improving the molecule.
Exercise 11.16 †
| Type 1 | Type 2 | |
|---|---|---|
| Core defect | autoimmune destruction of beta cells | insulin resistance with progressive beta cell dysfunction |
| Insulin production | essentially absent | present, often high initially, declining over time |
| Nature of insulin therapy | replacement of an absent hormone | overcoming resistance to a hormone already present |
| Required for survival | yes | no — one option among several |
| Typical doses | physiological replacement | frequently much higher |
| Underlying condition corrected | no, but the deficiency is fully replaced | no — resistance persists |
Exercise 11.19 †
One rule (Ch 2 §2.8): expect tolerance and reversal when a therapy overrides an intact regulated system; expect durability when it replaces an absent one.
Insulin in type 1: the beta cells are destroyed. There is no intact system to override, no feedback sensor measuring exogenous insulin, no set point being defended. Nothing adapts. The therapy works identically at year fifty.
GLP-1 agonists for weight: appetite regulation is intact and defending a set point. The drug outweighs it while present. Stop, and the intact system resumes.
What the rule predicts for an unfamiliar compound: ask whether the thing it supplies is absent or present-but-being-opposed. If absent — a replacement — expect durability and no rebound. If present — an override — expect adaptation during use and reversal after stopping.
This is the single most portable prediction in Part II, and Part III is where it earns its keep: several compounds there are marketed as replacements ("restoring youthful levels") and are pharmacologically overrides.
Exercise 11.21 †
Insulin detemir carries a fatty acid enabling albumin binding. So does semaglutide.
Detemir's problem: basal insulin needs to be absorbed slowly and steadily from the injection site over roughly a day, without a pronounced peak — because a peak in background insulin causes hypoglycemia, characteristically at night. Albumin binding provides a circulating reservoir that releases gradually.
Semaglutide's problem: a peptide small enough to be filtered by the kidney disappears in hours. Albumin binding makes the complex too large to filter, extending half-life to about a week (Ch 4 §4.5).
Why the same strategy serves both: albumin binding does two things — it protects from proteases and it defeats renal filtration — and it slows the rate at which free drug becomes available. A drug needing a flat profile over a day and a drug needing a week of duration are both solving "release this slowly and keep it around," and one mechanism does both.
This is Chapter 33's point in miniature: the modification toolkit is general, and the same tool solves different problems in different molecules.
Exercise 11.25 †
A defensible position: yes, it belongs, and excluding it would distort the picture.
The case for including it. The rating system attaches to claims about outcomes in populations, not to molecule types. "Automated insulin delivery improves glycemic control and reduces hypoglycemia in type 1 diabetes" is exactly that kind of claim, tested by randomized trials, and it can be evaluated by the same six criteria as any drug claim.
And the substantive reason: the most significant progress in insulin therapy over the past two decades has been in delivery and control, not in molecules. A book that rated only peptides would imply insulin has barely advanced since the analogs, which is false in precisely the way that matters to someone living with the condition. Chapter 4 argued that the delivery system is part of the drug; this rating applies that consistently.
The case against, which deserves acknowledgment: it stretches the book's scope, and a reader could reasonably want a peptide book to stay on peptides. The counter is that the book's actual subject is evidence evaluation, with peptides as the curriculum — and a method that could not evaluate a device claim would be a narrower method than advertised.
Exercise 11.27 †
The accurate version, with three complications:
What is true. Insulin's original patent was sold for a dollar in 1923 with the stated intention that it not be a source of profit. U.S. list prices rose substantially over recent decades, far beyond what manufacturing cost explains. People have rationed insulin for cost reasons, and deaths from doing so are documented. U.S. prices have far exceeded those in most other high-income countries for identical products.
Complication 1 — different products. The 1923 patent covered the 1923 preparation. Every insulin in current use is a later, separately developed and separately patented product: recombinant human insulin, then each analog. "They gave away the patent and then charged for it" describes different molecules. This does not excuse the pricing; it means the causal story is not a simple betrayal.
Complication 2 — list is not net. In the U.S. system, rebates flow between manufacturers, pharmacy benefit managers, and insurers, so the list price is not what most insured payers pay. It is very close to what an uninsured person pays. The structure that lowers net prices for insurers therefore raises the effective price for the least protected group.
Complication 3 — biosimilar entry was slow for regulatory reasons. Demonstrating that a copy of a biologic is comparable is much harder and more expensive than demonstrating chemical identity for a small-molecule generic. That delayed competition, and the delay was not purely a pricing decision.
And it has moved. Out-of-pocket caps, direct-purchase programs, and biosimilar entry have reduced costs for many people. Better than it was; not resolved; and any specific figure would date quickly.
Exercise 11.31 †
A model answer:
Insulin is a protein — the same kind of stuff as the protein in food. Your digestive system's entire job is to take proteins apart into their building blocks, so a swallowed insulin pill would mostly be digested as lunch. Even the small amount that survived would be too large to get through the wall of your gut into your bloodstream, and whatever did get through would go straight to your liver, which would clear most of it.
People have been trying to solve this for a hundred years. It's not that nobody thought of it.
Exercise 11.34 †
No single answer. What a good response contains:
The "what its history predicts" line is the one that matters, and for insulin it yields three transferable claims, none of them pharmacological:
- Replacement therapies are durable; overrides are not. Applies directly to every compound in Part III.
- The hard problem is usually delivery and control, not the molecule. Insulin's molecule was solved in 1922 and the dosing problem is still not solved in 2026. Chapter 36 is about this.
- A life-sustaining peptide becomes a pricing question eventually, and the mechanism is market structure rather than manufacturing cost.
The characteristic finding when readers do this for their own compound: for approved drugs, Field 5 at depth is rich and the "still open" line is longer than expected. For research compounds, the "first evidence" line and the "accumulated" line are frequently the same line, which is itself the answer to the question the exercise is asking.
Chapter 12
Instructor use. These are model responses, not the only acceptable ones. For the Judgment and ethics items the "solution" is a description of what a strong response contains, since the content of the judgment is the student's.
Exercise 12.3
Define fill-finish and explain why it, not synthesis, was the binding constraint.
Fill-finish is the manufacturing stage in which sterile drug product is transferred into its final container — here, a cartridge inside an injector pen — under conditions that maintain sterility throughout.
It was the constraint for three reasons. First, an aseptic filling line is a building rather than a machine: classified clean-room space, validated air handling, environmental monitoring, personnel-flow design, and bespoke equipment. Second, the line must be validated (a media-fill campaign demonstrating it produces sterile product repeatedly) and then approved by a regulator for that specific line, at that specific site, for that specific product — regulators approve processes, not capacity in the abstract. Third, the injector pen is a precision mechanical assembly with a dose-setting mechanism that must not fail, manufactured to tolerances closer to consumer electronics than to a vial, at a scale of hundreds of millions of units.
Peptide synthesis, by contrast, scales with reactors and time. It was never the bottleneck.
Look for: the student distinguishing "can be built with money" from "can be validated and approved with money."
Exercise 12.6
Semaglutide sodium, acetate, and base.
The approved products contain semaglutide base. Regulators stated that semaglutide sodium and semaglutide acetate are not the same substance as semaglutide base and that their safety and effectiveness had not been established. A salt form is a different chemical entity, potentially with different solubility, stability, and behavior — so the regulator's position was that a compounder using a salt form was not compounding semaglutide at all.
The reason salt forms appeared is economic rather than pharmacological: ingredient-sourcing rules constrain which substances a compounder may lawfully use, and a salt form was, for some suppliers, the version obtainable.
Common error to correct: students treating this as a naming technicality. It is a different substance, and that is the whole point.
Exercise 12.9
Chapter 11's five features, and which GLP-1 agonists share.
① Cannot be stopped (interruption rapidly life-threatening) — not shared. ② No therapeutic alternative — partly shared; other agents, bariatric surgery, and behavioral programs exist and are not equivalent. ③ Perfectly inelastic demand — shared. ④ Patient is the residual payer, because rebates lower net prices for insurers rather than list prices for the uninsured — shared. ⑤ Coverage discontinuous at predictable life transitions — shared, and arguably worse, because an employer can drop an entire weight-management category at renewal, affecting a whole workforce at once, which essentially never happens with insulin.
Extension worth drawing out: Chapter 11's observation that ①–③ are properties of the disease and ④–⑤ of a health system. That is why type 1 diabetes is identical everywhere but the rationing deaths are not.
Exercise 12.12
Rewrite the SELECT headline; say what was withheld.
Model rewrite (24 words): "In adults with existing heart disease and obesity, semaglutide cut major cardiac events from about 8% to 6.5% over three years."
What the original withheld: the baseline risk, without which "20 percent" is uninterpretable; the absolute difference of about 1.5 percentage points; the number needed to treat of roughly 65–70; and the population — established cardiovascular disease, overweight or obesity, no diabetes.
Who is harmed: two groups. People in lower-risk populations, who will infer a benefit larger than any they can expect, since the same relative reduction produces a smaller absolute benefit and a larger NNT as baseline risk falls. And anyone weighing three years of a chronic therapy, for whom the absolute number is the decision-relevant one.
Look for: the student stating both figures without being prompted, and naming the population.
Exercise 12.15
The capacity announcement.
The commentator has assumed that capital converts into supply on the timescale of capital. It does not. The announcement funds construction; supply requires construction plus validation plus line-by-line, product-by-product regulatory approval, and the last two are measured in years and cannot be compressed by spending more.
The question they should have asked: when is the line expected to be validated and approved for this product at this site? Secondarily: does the announcement cover pen assembly capacity as well as fill capacity, since those are different industries with different supply chains?
Exercise 12.18
What the intake form does not ask.
Three among several defensible answers:
- Eating disorder history. A person with a restrictive eating disorder presents as an ideal candidate on paper — motivated, compliant, weight-focused. Appetite suppression removes the friction restriction must overcome, and the prescription supplies a framework that makes the behavior legible as treatment. §12.10.
- Pregnancy status and pregnancy plans. Clinically relevant and routinely part of an in-person assessment.
- Whether the self-reported height and weight are accurate, and whether the person can tell which answers qualify them. A form whose only failure mode is "not qualifying" is a sorting mechanism, not an assessment; when the qualifying answers are obvious and the inputs are self-reported, eligibility criteria become advisory.
Also acceptable: pancreatitis history, gallbladder disease, full medication list with interaction review, whether a regular clinician will be informed, and whether any pathway exists to decline.
Look for: the student noticing that the fourth item — "is there a path to no?" — is a property of the system rather than of the questionnaire.
Exercise 12.21
Why the same relative reduction gives a larger NNT in a lower-risk population.
Absolute risk reduction is the relative reduction applied to the baseline risk. If the baseline risk is lower, the same proportional reduction removes a smaller number of percentage points. The number needed to treat is the reciprocal of the absolute risk reduction, so a smaller absolute reduction produces a larger NNT.
Implication for extrapolation: SELECT enrolled a deliberately high-risk population — established cardiovascular disease. Applying its relative reduction to a population without established cardiovascular disease, even granting that the relative effect holds (which is itself an assumption), yields a smaller absolute benefit and a larger number needed to treat. This is why the rating carries the population inside it, and why "20 percent reduction" travels much further than it should.
No arithmetic required, and students should not invent numbers.
Exercise 12.24
"The weight all comes back when you stop."
Model response:
"You're right — it usually does. That's been measured, and it's not in dispute.
Here's why it isn't a mark against the drug. This medication doesn't fix anything permanently; it holds something down while you're taking it. Your body has a weight it defends — when you lose weight, your hunger goes up and your body burns slightly less. The drug pushes back against that system while it's in you. Take it away, and the system that was there the whole time carries on doing what it was doing.
Compare it to blood pressure medication. Nobody says a blood pressure pill failed because your blood pressure goes back up when you stop taking it. That's just what it means for something to be ongoing treatment. We only apply the 'it should stick after you stop' standard to weight — and that expectation is about how we feel about weight, not about how the drug works."
Look for: the comparison drawn from outside weight medicine, and the student conceding the factual claim before disputing the inference.
Exercise 12.27
The three users.
There is no correct verdict. A strong response contains all of the following:
- Three separate paragraphs, each naming the counterfactual explicitly — what the student thinks that person should have done instead. This is the hinge of the exercise. A judgment without a stated alternative is not a judgment; it is a mood.
- For user one: recognition that the exception was designed for precisely this case, and that the alternative was no therapy.
- For user two: recognition that the exception was not designed for a price barrier, and that the price barrier was real. The strongest responses say plainly what the alternative was, including if it is "gone without."
- For user three: recognition that the desire is not illegitimate but the transaction differs — no medical endpoint on the benefit side, and during a shortage the supply consumed came from a pool patients with diabetes were denied.
- The final question answered honestly. If the three verdicts came out identical, the student judged the category, not the circumstances — Chapter 5's rule that a rating attaches to a claim with a population, not to a category, applied to people.
Grade the structure, not the conclusion.
Exercise 12.31
All four positions, then your own view.
The assessable skill is separation, not conclusion. A strong response:
- States each position in language its advocates would accept, with no tell in the phrasing. A common failure is a "behavior" paragraph that reads as a straw man and a "disease" paragraph that reads as the author's own view.
- Includes each position's best evidence: for disease, metabolic adaptation, monogenic obesity, treatment response; for behavior, energy balance, documented durable changers, self-efficacy; for food environment, decades-scale population change in genetically stable populations; for weight-neutral, fitness gains partly independent of weight, stigma effects, the withdrawn-drug record.
- Keeps the author's own view in a separately labelled paragraph, and does not let it leak upward.
- Bonus credit for noticing the positions are not mutually exclusive, and for correctly handling heritability (within-population variance says nothing about what moved a population average).
- Bonus credit for invoking the chapter's logical point: a drug that works on a system is not evidence that the system caused the condition — so the treatment-response argument does not settle the question in favor of Position 1.
The most common failure is a student who cannot write 150 fair words about the position they find distasteful. That is the exercise.
Chapter 13
Exercise 13.2 †
Leptin is a hormone produced by adipose tissue in proportion to fat mass. Its level tracks how much energy is stored, which makes it a report on the body's energy reserves rather than a meal-triggered signal.
That distinction matters: gut hormones like GLP-1 answer "is food arriving now?"; leptin answers "how much do we have in the bank?" Both feed the same hypothalamic circuitry, on different timescales.
Exercise 13.5 †
Two neuron populations in the arcuate nucleus, with opposite effects, both responsive to leptin:
- POMC neurons — stimulated by leptin. POMC is cleaved to α-MSH, which activates MC4R and reduces hunger.
- AgRP/NPY neurons — inhibited by leptin. They release AgRP, which blocks MC4R and increases hunger.
They converge on one receptor: MC4R. An accelerator and a brake meeting at a single point.
Why the architecture matters: breaking any component — POMC, the leptin receptor upstream, or MC4R itself — produces severe early-onset obesity. These are the monogenic obesities, and MC4R is the most common known monogenic contributor.
Exercise 13.9 †
What the model predicted: obesity is leptin deficiency; supply leptin; appetite and weight normalize. Structurally identical to insulin for type 1 diabetes, which is the cleanest kind of medicine there is.
Why the prediction failed: people with common obesity do not have low leptin. They have elevated leptin — appropriately, since it tracks fat mass. The signal was being sent loudly and not acted upon.
What was actually wrong — and this is the precise part: not the biology, and not the hormone. Leptin does exactly what the model said it does. The model was wrong about which disease it explained. Leptin deficiency is a real condition and leptin treats it dramatically. It is simply not the condition that most people with obesity have.
The general form of the error: an animal model of a deficiency is a model of that deficiency, not of the common human disease that superficially resembles it. And the resistance phenomenon was visible in the db/db mouse from the 1960s — the field had the counter-evidence the whole time and attended to the deficiency model instead.
Exercise 13.13 †
Tirzepatide: one molecule, two receptors. A single pharmacokinetic profile; both receptors engaged in the same temporal pattern always; a ratio of activity fixed by the molecule's structure.
CagriSema: two molecules, one injection. Two pharmacokinetic profiles; a ratio that can be adjusted by changing the amount of each component.
Advantage of the single molecule: guaranteed co-engagement. The two receptors are always activated together in the same proportion, so there is no drifting ratio across the dosing interval and no possibility of one component outlasting the other.
Advantage of the two-molecule design: tunability. A fixed ratio is a bet that cannot be revised without designing a new molecule; a combination can be re-proportioned in development, or in principle titrated separately.
Neither is obviously better, and which wins is an empirical question the respective programs are answering.
Exercise 13.16 †
Setmelanotide works because the patients have an identified broken step upstream — POMC is not made, or the leptin receptor cannot signal — and the drug acts downstream of the break, activating MC4R directly.
In common obesity the pathway is intact. There is no broken step to bypass. Agonizing a receptor that is already being appropriately signaled by an intact upstream system is a different intervention entirely, and nothing about setmelanotide's efficacy in deficiency predicts what it would do there.
The general principle: a proven mechanism licenses a claim only within the population where the mechanism's premise holds.
Why this is the strongest form of Chapter 2's warning: the usual version says a plausible mechanism is not evidence. Here the mechanism is not plausible — it is demonstrated, in humans, by a successful trial. And it still does not extend, because the demonstration was conditional on a defect that most patients do not have.
Exercise 13.19 †
The structural argument: appetite is regulated by a redundant network of at least a dozen partly overlapping signals converging on hypothalamic circuitry, defending a set point, with counter-regulation opposing displacement. Such a system is robust against the loss or addition of any single node.
Applied:
Leptin — targeted a signal that was already maximal. Adding more of a message being ignored changes nothing. The failure was not redundancy exactly; it was resistance, which is redundancy's cousin: the system had already routed around the signal.
PYY — added one satiety signal to a system that already had several. Redundancy absorbed it. The effect in infusion studies was modest and short-lived, and nausea was dose-limiting.
Ghrelin blockade — removed one hunger signal from a system with several others. Redundancy absorbed it. Ghrelin also has roles beyond appetite, so blockade was not clean.
The common structure in all three: a single-node intervention against a multi-node system.
Exercise 13.22 †
① Half-life. Endogenous GLP-1 is destroyed by DPP-4 in roughly one to two minutes. Whatever a supplement raises is gone almost immediately. Semaglutide's half-life is about a week.
② Magnitude. A dietary intervention produces a modest increase in a postprandial peak. A therapeutic agonist produces sustained supraphysiological receptor occupancy. These differ by orders of magnitude.
③ Route. Endogenous GLP-1 acts largely locally, via vagal afferents, because so little survives to circulate. Injected agonists act systemically, including at brain regions with a leaky blood-brain barrier. Raising the endogenous signal amplifies the local route, not the systemic one.
④ The decisive one, requiring no measurement: every human being has had normal GLP-1 physiology their entire life, and it has never produced fifteen percent weight loss in anyone. The drug effect exists precisely because it is not physiological. No amount of enhancing a physiological signal produces a supraphysiological result, because the physiological signal is what everyone already has.
Exercise 13.25 †
A model answer:
They did find the hormone — leptin, in 1994 — and it was a genuinely huge discovery. The catch is what it turned out to mean. Leptin is made by fat tissue and tells your brain how much energy you've stored, so people with more body fat make more of it, not less. So obesity isn't usually a leptin shortage; it's more like the brain has stopped listening to a signal that's being shouted.
Giving people extra leptin mostly didn't help, for that reason. But there are a few dozen people in the world born unable to make leptin at all, and for them it's transformative — genuinely one of the most dramatic treatments in medicine.
So "obesity is a hormone problem" isn't wrong exactly, it's just that finding the hormone turned out to be the start of the problem rather than the end of it.
Exercise 13.29 †
A defensible ranking by what the field learned, with the note that this deliberately differs from ranking by drugs produced:
1. Leptin. Produced no general drug and generated more knowledge than anything else in the chapter: adipose tissue as an endocrine organ, the hypothalamic appetite circuitry, the monogenic obesities, the defended set point as a concrete concept, and — indirectly — setmelanotide.
2. MC4R agonism. Produced a drug for a tiny population and confirmed causality in the appetite circuit, which is a real and hard-won result. It ranks below leptin only because it was made possible by leptin.
3. Ghrelin. Produced no weight drug but substantially clarified what a "hunger hormone" is and is not, and its agonists found uses in cachexia and in GH release that fit the physiology better than antagonism did. A useful negative result.
4. PYY. Least productive: the failures were largely practical — delivery, dose-limiting nausea, modest effect — rather than conceptually illuminating.
The point of the exercise is noticing that this ordering is nearly the inverse of a ranking by commercial success, and that a field evaluated only on the second measure would look far less productive than it was.
Exercise 13.31 †
No single answer. Characteristic findings:
The "gap" line sorts your dossier immediately. For approved drugs the gap is usually specific and small — an off-label population, or a broader claim than the label supports. For unapproved compounds the gap is frequently the entire entry: there is no approved indication, so every claimed use is outside it.
The "who makes each claim" line is where epistemic laundering becomes visible (Ch 6 §6.4). Readers consistently find claims that no manufacturer and no published paper makes, circulating as community consensus. Recording the claim's source alongside the claim is what exposes this, and almost no product page does it for you.
And most readers cannot find a compound whose marketing is narrower than its science (exercise 13.33). Setmelanotide is one. If your dossier contains none, that is not a failure of the exercise — it is information about how you selected your peptides, which is worth noticing before Chapter 37 asks you to audit your own drift.
Chapter 14
Solutions to the daggered (†) exercises. Item numbering follows exercises.md.
Exercise 14.A1
Question. Name the hypothalamic hormone that stimulates pituitary growth hormone release and the one that inhibits it. Say which is the accelerator and which is the brake, and state one reason it matters that the axis has both.
Solution.
GHRH (growth hormone releasing hormone) stimulates — the accelerator. Somatostatin inhibits — the brake.
Why it matters: most endocrine axes are built from an accelerator plus negative feedback from the end product. This axis adds a dedicated inhibitory hormone released from the hypothalamus on its own schedule. The practical consequence is that there are two distinct ways to raise growth hormone: press the accelerator, or lift the brake. Those are different pharmacological interventions with different consequences — different receptors, different pulse shapes, different interactions with the rest of the system — even though both raise a measured growth hormone level. Chapter 15's secretagogues divide along exactly this line, and a student who cannot name the brake will not be able to tell those compounds apart.
Full credit also for noting that the presence of an explicit brake is itself weak evidence that unregulated elevation of this axis is something the body actively guards against.
Exercise 14.A9
Question. Write out this chapter's two ratings in full, each with its population and its endpoint.
Solution.
Rating 1 — ✅ Strong clinical evidence (as of this writing, 2026). Claim: Growth hormone replacement improves growth outcomes in children with documented growth hormone deficiency, and improves body composition, bone mineral density, lipid profile, and quality of life in adults with documented growth hormone deficiency. Population: children and adults with documented growth hormone deficiency. Endpoints: growth (children); body composition, bone mineral density, lipids, quality of life (adults).
Rating 2 — ❌ Hype outpaces evidence (as of this writing, 2026). Claim: Growth hormone administered to healthy adults without documented growth hormone deficiency slows or reverses aging, or produces meaningful improvements in strength, function, healthspan, or well-being. Population: healthy adults without deficiency. Endpoints: aging, strength, function, healthspan, well-being.
Credit requires both the population and the endpoint in each. An answer of the form "growth hormone: ✅ for deficiency, ❌ for anti-aging" gets partial credit only — it drops the endpoints, which is where the whole argument lives. Note also that an answer must not compress the two into a single molecule-level verdict; that is the error Rule 1 exists to prevent.
Exercise 14.B2
Question. A user of exogenous growth hormone finds his own pituitary output has fallen and concludes the drug damaged his pituitary. Explain what is actually happening and whether it is evidence of harm.
Solution.
Nothing has been damaged. This is the feedback loop from §14.1 operating exactly as designed.
The sequence: exogenous growth hormone raises circulating IGF-1. IGF-1 inhibits pituitary growth hormone release and increases hypothalamic somatostatin output. The pituitary, receiving both a stronger inhibitory signal and a suppressed releasing signal, reduces its own secretion. Endogenous production falls.
This is expected physiology, not a malfunction. A control loop that did not respond this way would be the abnormal finding.
The Chapter 3 frame is the right one to name: this is an override of an intact system, not replacement of an absent one. Chapter 5's rule predicts adaptation whenever an intact regulatory system is overridden, and suppression of endogenous output is precisely that adaptation.
Two honest qualifications belong in a complete answer. First, "expected" is not the same as "consequence-free" — the practical question of what happens on discontinuation, and over what timescale recovery occurs, is a real one and belongs in a conversation with an endocrinologist rather than in a book. Second, none of this reasoning applies to a person receiving replacement for documented deficiency: their pituitary was not producing adequately in the first place, so there is no intact loop to suppress. Same observation, different meaning, depending entirely on which population you are in.
Exercise 14.B4
Question. Apply replacement-versus-override to three cases.
Solution.
(a) 9-year-old with congenital growth hormone deficiency — REPLACEMENT. The signal is absent and is being supplied from outside. The prediction is durable benefit, because the system is built to receive this signal and has not habituated to an artificial one. This is what the literature shows, and the endpoint (growth) is not even a surrogate — it is the thing itself.
(b) 45-year-old post-resection of a pituitary adenoma — REPLACEMENT. Same logic. The pituitary's capacity has been damaged or removed; supplying the hormone restores an absent signal. Note the diagnostic point from §14.3: this person has a structural reason to have pituitary failure, which is what makes the deficiency diagnosis credible in the first place. Pre-test probability is doing real work.
(c) Healthy 58-year-old with an age-appropriate IGF-1 — OVERRIDE. The pituitary works. Nothing is being restored. Exogenous hormone pushes a functioning control loop away from where it is currently sitting, and the loop responds by suppressing endogenous output (see B2). The prediction is adaptation, and the demonstrated effects are on surrogates rather than outcomes.
The most common error here is treating (c) as a mild version of (a) and (b) — a continuum with a threshold somewhere on it. That is the exact move §14.6 identifies as step five in how the industry was built: reframing an override as a replacement, which silently imports the ✅ from a different population. A decline is not an absence.
Exercise 14.B7
Question. Write one sentence stating what the systematic review supports and one stating what it does not, without converting "no evidence of effect" into "evidence of no effect."
Solution.
Supports: In healthy older adults, growth hormone administration produces measurable changes in body composition — increased lean mass, decreased fat mass — accompanied by higher rates of edema, arthralgia, carpal tunnel syndrome, gynecomastia, and impaired glucose metabolism than in controls, without a convincing improvement in the strength and functional measures that were assessed.
Does not support: It does not establish that growth hormone has no functional effect under any exposure, duration, or subgroup — the pooled trials were mostly short, modest in size, and not designed to detect fracture, disability, or mortality differences — so this is an absence of demonstrated benefit on the endpoints studied, not a demonstration of absent benefit.
The discipline being tested is the second sentence. A student who writes "growth hormone does not improve strength" has stated something the review cannot support, and has handed a valid objection to anyone arguing the other side. The correct posture is that the burden has shifted decisively after three decades of looking — which is a strong claim, and a different claim.
Exercise 14.C5
Question. A friend asks: "Is growth hormone safe?" Answer in four sentences, fair to three different readers.
Solution. Model answer:
"Safe" isn't a property of a drug — it's a judgment about a specific person, exposure, and purpose, so the honest answer depends on which situation you're in. For someone with documented deficiency, regulators have repeatedly judged that the benefits outweigh the known risks, which are real — swelling, joint pain, carpal tunnel, effects on blood sugar — and are managed by monitoring. For a healthy adult taking it to feel younger, the picture is different: the same side effects show up more often than in controls, long-term safety in healthy people has never been established, and the best long-term human evidence about chronically raised growth hormone — the disease acromegaly — shows serious harm at much higher exposures, which tells you the direction of concern without telling you the size of the risk. And if you're already taking it, the useful things are concrete: glucose and HbA1c, blood pressure, IGF-1 against an age-referenced range, any new hand numbness or joint pain, and staying current on age-appropriate cancer screening — with an endocrinologist who isn't selling it.
Note that the answer never says "yes" or "no," never moralizes, and never gives a protocol. Full credit requires all three.
Exercise 14.D1
Question. Identify the problems and the truths in: "hGH is the master hormone... it drops 14% per decade starting in your thirties. Restoring your levels to where they were at 25 is not experimental — it's replacement therapy, the same as thyroid or testosterone."
Solution.
What is true. The decline is real and the figure is in the right neighborhood — roughly 14% per decade after about 30, stated approximately. Growth hormone does act on a wide range of tissues, both directly and via IGF-1. Neither of those is fabricated, and that is what makes the passage effective.
The problems, named:
- "Master hormone" — not a technical term. It creates an impression of centrality that licenses the rest of the argument without asserting anything checkable.
- "Restoring your levels to where they were at 25" — the age-adjusted-versus-youthful reference range move (§14.4). Comparing a 55-year-old to a 25-year-old's range and calling the result "low" is a choice of comparison group presented as a clinical finding.
- "It's replacement therapy" — the central error. This is an override of an intact system, not replacement of an absent one (§14.3). The phrase imports the ✅ rating from a population the reader is not in.
- The thyroid and testosterone analogy — smuggles in the same reframe by association. Thyroid replacement treats documented hypothyroidism; the comparison would only hold for someone with documented deficiency, which is precisely the population the passage is not addressing.
- "Not experimental" — technically true and rhetorically false. The molecule is not experimental; the claim being made about it in this population is unsupported. Chapter 5's Rule 1 again: the rating attaches to the claim.
- Missing entirely: any endpoint. Nothing in the passage says what will improve or how it would be measured.
Exercise 14.D4
Question. Identify the problems and the truths in: "There's zero evidence that growth hormone causes cancer. That's a myth pushed by people who want to keep this away from you."
Solution.
What is true. No causal claim is established. There is no adequately powered long-term randomized trial in humans demonstrating that administered growth hormone causes cancer, and the observational IGF-1 associations are heavily confounded — IGF-1 tracks nutrition, body size, insulin, activity, and socioeconomic position, all of which independently relate to cancer risk. Someone who says "growth hormone causes cancer" is overclaiming, and this passage is right to push back on that.
The problems:
- "Zero evidence" is false. There is evidence — mechanistic (proliferation, apoptosis inhibition), observational (IGF-1 associations), and from the natural experiment (increased colonic polyps and multi-system harm in acromegaly). None of it establishes causation. "Evidence that does not establish causation" and "zero evidence" are different states, and the passage collapses them.
- It converts an open question into a settled one. §14.7's fourth sentence: not knowing is not the same as knowing it is fine. The absence of a long-term randomized trial means the question is open, not closed in the reassuring direction.
- It attacks motive rather than evidence. "People who want to keep this away from you" is an argument about who is speaking, not about what is true. It also inoculates the reader against future correction, which is its actual function.
- It ignores the natural experiment. Acromegaly is the most informative human evidence available about chronic elevation of this axis, and a passage claiming "zero evidence" has either not encountered it or has chosen not to mention it.
The correct statement is the four-sentence version in §14.7: the mechanism points the wrong way; an observational association exists; confounding is substantial and causation is unestablished; and not knowing is not the same as knowing it is fine. Note that this exercise and D-item overclaiming in the other direction are the same error with opposite signs.
Exercise 14.D6
Question. Identify the problems in: "Growth hormone is a peptide, and peptides break down into amino acids, so worst case nothing happens."
Solution.
What is true. Growth hormone is indeed proteolyzed into amino acids and fragments that enter normal metabolism. It does not accumulate in tissue the way some small molecules do, and it does not generate the reactive metabolites behind certain small-molecule toxicities. As a class property, this is a genuine advantage, and Chapter 1 said so.
The problems:
- Effects happen before breakdown. The molecule acts at a receptor, and the receptor does not care that the ligand will eventually be recycled. Growth hormone's effects on glucose, fluid balance, and IGF-1 production all occur while the molecule is intact.
- The relevant exposure is chronic, not acute. Every risk in §14.7 and every finding in §14.8 concerns sustained signaling, not the fate of a single molecule. Metabolic clearance says nothing about what repeated signaling does over years.
- "Peptide" is doing no work. This is Chapter 1's central thesis. Insulin is a peptide; botulinum toxin is a peptide product; the category carries no evidentiary weight about safety.
- It says nothing about what else is in the vial. Purity, identity, sterility, endotoxin, and concentration are separate questions, and in the unregulated market they are where documented harm has actually occurred (Chapters 19 and 34). Growth hormone is a large recombinant protein that is difficult to manufacture correctly, which makes it a particularly poor candidate for informal supply.
- "Worst case nothing happens" is not a risk assessment. It is a reassurance with no measurement attached, and it is unfalsifiable as stated.
Exercise 14.E2
Question. Growth hormone is approved for idiopathic short stature. Argue both sides. What does this indication share with the anti-aging claim, and what makes it different?
Solution. This is a judgment item; the model answer below is a defensible position, not the only one.
What it shares with the anti-aging claim. Both treat a position on a distribution rather than an identified pathology. In neither case is a hormone absent; in both, the person is at one end of a normal range. Both raise the question of whether the endpoint being improved is one that matters — for idiopathic short stature, adult height is a real, objective, permanent outcome, but whether increased adult height translates into improved well-being is a genuinely contested empirical question that has been studied with mixed results. And in both, the framing does moral work: calling shortness a condition, or calling an age-appropriate IGF-1 "suboptimal," makes treatment sound like correction.
What makes it different — and this is the stronger half of the argument.
- The endpoint is not a surrogate. Adult height is the thing itself, measured directly, and it is durable. Lean mass is a stand-in for strength and function, and the stand-in did not deliver.
- The effect is demonstrated on that endpoint. Trials show increased adult height. There is no equivalent demonstration on any anti-aging outcome.
- The population is monitored, and the exposure is bounded. Treatment occurs under pediatric endocrine follow-up and ends at growth plate closure. Anti-aging use is typically unmonitored and open-ended, and duration is where §14.8's concerns live.
- Regulators saw the full dataset and accepted it. That is not dispositive, but it is not nothing, and it is a bar the anti-aging claim has never attempted.
A defensible conclusion. The indication is contested for good reasons, and someone who is uncomfortable with it is not being unreasonable. But the discomfort is largely ethical — about medicalizing normal variation — while the anti-aging claim's problem is largely evidentiary. Those are different objections, and conflating them weakens both. Students who notice that distinction have understood the item.
Exercise 14.E4
Question. Write what you would say to a person with no deficiency who has used growth hormone for two years, feels better, sees less body fat, and is asking what to pay attention to.
Solution. This item is graded on posture as much as content. A strong answer does five things:
- Takes the reported experience seriously. The body composition change is real and replicated; this person is not imagining it. An answer that opens by doubting them has lost the reader and, as §14.6 explains, is also factually wrong.
- Names which claim they are relying on, plainly and without judgment. They are in the ❌ row. The honest description is: a real drug with real effects, taken for a benefit that has not been demonstrated on any measured outcome. That is a decision a person is entitled to make.
- Distinguishes what they can see from what matters. The visible effects and the consequential effects are not the same effects — the central lesson of Case Study 14.2. Nothing they can observe in a mirror bears on cardiac structure or glucose trend.
- Gives concrete, monitorable specifics — not a protocol. Glucose and HbA1c (insulin antagonism). IGF-1 against an age-referenced range. Blood pressure. New hand numbness or tingling, new joint pain, swelling. Age-appropriate cancer screening kept current — because the question is open, not because a causal link is established.
- Names what a clinician adds, and which clinician. Testing that distinguishes the two rows; knowledge of contraindications like active malignancy, diabetes, untreated sleep apnea, and certain cardiac conditions; trend rather than snapshot; and a conversation about what happens on stopping, including feedback suppression of endogenous production. An endocrinologist, preferably one whose practice is not built around selling the intervention.
Automatic loss of credit for: any dose, schedule, cycle, source, or "if you're going to do it anyway, at least..." construction; any moralizing; and any implication that the ❌ means the drug does nothing.
Exercise 14.F2
Question. For one dossier compound, find an item you had listed as a documented effect that is actually only a mechanistic concern, and move it to 9c.
Solution. There is no single correct answer, but the method is gradeable, and it is the point of the exercise.
The procedure. For each item in 9a and 9b, ask: what is the source? There are only four acceptable answers — an approved label, a controlled trial, a systematic review or pooled analysis, or post-marketing surveillance. If the honest answer is "it follows from how the molecule works," the item is a mechanistic concern and belongs in 9c, phrased as an open question.
Worked example on growth hormone. A student's first-draft entry might list "increased cancer risk" under 9b, serious/monitored. Audit it: is there a label warning stating that growth hormone causes cancer? No. A randomized trial demonstrating it? No. What exists is a mechanism (proliferation, inhibition of apoptosis), an observational association with IGF-1 that is heavily confounded, and a natural experiment at far higher exposure showing multi-system harm including increased colonic polyps. That is a genuine, substantive concern — and it is 9c, not 9b. Rewritten:
9c — Long-term malignancy risk is unresolved. Mechanism is adverse; observational IGF-1 associations exist but confounding is substantial and causation is not established; acromegaly shows real harm at much higher, decades-long exposure. Open question, not a documented effect.
Note that moving the item does not weaken it. It sharpens it, and it makes the entry defensible to someone arguing from either direction.
The mirror-image error, also worth catching. Items in 9c that actually belong in 9a — real, documented, label-listed effects that a student softened into "some people report." Edema and arthralgia are documented effects of growth hormone, not rumors. Chapter 5's rule cuts both ways: mechanism is necessary and never sufficient, and that applies to harms exactly as it applies to benefits.
Chapter 15
Solutions to the exercises in exercises.md. Items marked † are open or argumentative; for those,
what follows is a marking guide rather than an answer.
Exercise 15.1
Expected diagram: hypothalamus at top, with two arrows to the anterior pituitary — GHRH (+) as the accelerator and somatostatin (−) as the brake. Pituitary releases growth hormone, drawn as bursts rather than a line. GH acts on the liver, which produces IGF-1. IGF-1 acts on peripheral tissues and also runs a feedback arrow back up, inhibiting GH release and raising somatostatin.
Stars belong at the GHRH receptor (sermorelin, CJC-1295, tesamorelin) and the ghrelin receptor / GHS-R (GHRP-2, GHRP-6, ipamorelin, MK-677) — both on the anterior pituitary. Full credit requires noting that no compound in this chapter acts on the target tissue.
Exercise 15.2
(i) It might mean a genuinely suppressed or low-output axis. (ii) It might equally mean the sample was drawn between pulses, since GH secretion is strongly pulsatile and falls to near-undetectable in the troughs — a healthy person will produce this result routinely. (iii) Measure IGF-1, because it integrates GH exposure over a longer window and does not pulse.
Exercise 15.3
Model answer: growth hormone is secreted in a pulsatile pattern, so its concentration at any given instant depends mostly on where in a burst cycle the sample was taken; IGF-1 is produced by the liver in response to accumulated GH exposure and remains relatively steady, so a single IGF-1 value is informative where a single GH value is not.
Exercise 15.4
GHRH receptor: sermorelin, tesamorelin, CJC-1295. Ghrelin receptor (GHS-R): GHRP-6, MK-677, ipamorelin.
Exercise 15.5 †
Marking guide. A strong answer notices that IGF-1 has two roles that pull in opposite directions: it is the evidence the intervention is working and it is the signal that suppresses further GH release. Over four weeks, a rise plausibly reflects the intervention's direct effect. Over many months, the observed value reflects an equilibrium between the stimulus and the feedback resisting it, so a stable IGF-1 could mean sustained effect, or an attenuating effect being propped up by a compensating factor, or a system that has reset. Credit for noting this cannot be distinguished from IGF-1 alone. Best answers connect to downregulation (§15.6) and to the absence of long-duration data for most compounds.
Exercise 15.6
For the user: pressing the accelerator does not disengage the brake, and successful stimulation raises IGF-1, which raises somatostatin — so the system actively resists the intended change and returns are expected to diminish. For the trial reader: a short trial captures the initial response before equilibration, so early results systematically overstate what a long exposure would show. Any trial duration shorter than the intended use is measuring a different thing.
Exercise 15.7
Peptides: sermorelin, ipamorelin, tesamorelin, GHRP-6. Not a peptide: MK-677 (ibutamoren) — a small, orally active, non-peptide ghrelin receptor agonist. The distinction changes expectations because it is not made of peptide bonds, so digestive proteases have nothing to cut; it is orally active and long-acting, which makes it easier to obtain and easier to take continuously.
Exercise 15.8
Model answer: DAC is a chemical group attached to the peptide that forms a covalent bond with albumin, the most abundant protein in blood plasma. Because albumin remains in circulation for weeks, a peptide tethered to it is largely shielded from kidney filtration and from many proteases. The practical result is that a single administration produces a receptor-relevant signal lasting days rather than minutes.
Exercise 15.9 †
Marking guide. Must identify (i) CJC-1295 with DAC — albumin-tethered, duration in days; (ii) the modified GRF(1-29) fragment without DAC — peptidase-resistant substitutions only, duration far shorter. The key reasoning: aggregation of informal reports requires that the reports be about the same thing. Here they are not, and the reports do not reliably say which molecule is involved, so pooling them mixes two pharmacologies. This is why more anecdotes cannot fix the problem — the defect is in the labeling of the exposure, not the quantity of observations. Strong answers connect this to why formal trials specify the investigational product precisely.
Exercise 15.10
Selective for growth hormone release via the ghrelin receptor; against appetite stimulation, cortisol release, and prolactin release — the three activities associated with GHRP-6 that the design aimed to reduce. Credit for adding that this is a statement about receptor and pathway engagement, not a safety claim.
Exercise 15.11
It is evidence about regulatory history: a compound still known only by a laboratory code decades after discovery almost certainly never completed formal drug development, because naming authorities assign stem names to compounds in serious clinical development. It is not evidence about pharmacology — not about whether the molecule binds its receptor, raises GH, or would work if tested.
Exercise 15.12 †
Marking guide. Strongest version of the argument: a generic name and a formal diagnostic use show that regulators have reviewed the compound, that its pharmacology is characterized well enough to be relied on clinically, and that it demonstrably does what it is claimed to do at the pituitary — so this is not an unstudied gray-market molecule.
Dismantling: every element of that is true and none of it touches the claim at issue. A diagnostic use establishes that the compound reliably probes the axis — that is a step-2 fact. The popular claims are step-6 claims about body composition, recovery, and aging. An approval to test an axis is not evidence of benefit from stimulating it; a stress test is not a therapy. Credit for identifying this as borrowing credibility across an indication boundary, the same error as citing tesamorelin's approval for general fat loss.
Exercise 15.13
A secretagogue is an instruction, not an ingredient. With no pituitary, there is nothing to instruct, so no amount of signal produces hormone. Generalization: every "stimulate rather than replace" strategy depends on the productive capacity of the tissue being stimulated remaining intact, which means these strategies are systematically least useful in exactly the patients whose deficiency is most severe.
Exercise 15.14
Four bullets: (i) pulsatility is preserved, because the pituitary releases in bursts regardless — the stimulus raises amplitude rather than replacing the rhythm; (ii) the gland stays in use, with no suppression or disuse to recover from; (iii) regulation is retained, since somatostatin, IGF-1 feedback, and hypothalamic sensing remain in the circuit — a secretagogue is a request rather than a fact; (iv) there is a ceiling, since the pituitary can only release what it has. Deduct for smuggling in counter-arguments; the exercise tests whether the student can steelman.
Exercise 15.15
(i) Feedback still applies and opposes you: elevated IGF-1 inhibits GH release and raises somatostatin. (ii) An extended-duration analog delivers a sustained signal to a pulsatile axis, which is a different signal rather than a more natural one. (iii) Even granting the entire argument, it establishes only the shape of a hormonal signal, not an outcome for the person.
Exercise 15.16 †
Marking guide. For: the loop constrains magnitude, so extreme sustained elevations achievable by direct injection are not achievable this way; the system retains the ability to decline. Against: the loop's function is to return the system to its set point, so it resists the intended change by the same mechanism that limits the extreme; "preserved regulation" describes an opponent, not a guardrail; combined with downregulation it predicts attenuation. Award credit for the decision step naming a specific deciding fact — most commonly that the loop constrains the intended effect and the excess by one mechanism, so it cannot be claimed as a benefit selectively.
Exercise 15.17
Model answer, non-technical: your pituitary is built to hear short shouts with silence in between, and the silence is part of what it hears. A long-acting compound replaces the shouts with a continuous hum. That is not the same message delivered for longer — it is a different message, and how the gland responds to it over months is a question that has to be answered by measuring, not by assuming.
Exercise 15.18 †
Marking guide. Does establish: continuous stimulation of a pulsatile axis can produce the opposite of the intended effect, and this is not speculation because it is the mechanism of an approved drug class; therefore "sustained stimulation is at least as good as pulses" is a claim requiring evidence rather than a safe default.
Does not establish: that extended-duration GHRH analogs suppress the GH axis. Different receptor, different axis, different signaling context; no one has demonstrated the reversal here. Mark down heavily for any answer that asserts suppression — that is the exact mechanism-to-conclusion overreach the book objects to when it runs in the pro-compound direction, and the point of the exercise is to catch students being inconsistent about it.
Exercise 15.19
(i) Negative feedback — elevated IGF-1 inhibits GH release and raises somatostatin. (ii) Receptor downregulation — persistent stimulation often reduces receptor number or responsiveness. For the compounds in this chapter, neither has been characterized over realistic durations of use; tesamorelin and MK-677 have the longest human exposure data, and even those do not fully answer it for the others.
Exercise 15.20
A surrogate endpoint is a measurable stand-in for the outcome actually of interest, used on the hypothesis that changing it will change the outcome. Examples from outside the chapter: LDL cholesterol as a surrogate for cardiovascular events; bone mineral density as a surrogate for fractures; HbA1c as a surrogate for diabetic complications; viral load as a surrogate for clinical progression.
Exercise 15.21
Steps 1–3 (receptor binding → GH release → IGF-1 rise) above the line, all ESTABLISHED. Steps 4–7 (intended tissue change → measurable composition change → functional outcome → net benefit over years) below it. The line sits between step 3 and step 4.
Exercise 15.22
(a) mechanism · (b) outcome · (c) mechanism · (d) outcome — though note it is an intermediate outcome and itself a surrogate for hard endpoints · (e) mechanism, dressed as an outcome, which is why it is effective marketing · (f) outcome.
Exercise 15.23 †
Marking guide. The strongest essays make the point that a clean success would have been consistent with the surrogate-to-outcome arrow holding automatically, and a clean failure (no surrogate movement) would have said nothing about the arrow at all. This trial is informative precisely because it decoupled them: the mechanism performed exactly as designed for two years, and the functional outcome still did not follow. That is direct evidence about the inference, not just about the compound — which is what makes it transferable to compounds that have not been invented yet. Credit also for noticing that the result is unfavorable to the sponsor and was reported anyway, and for the fluid-retention caveat on the fat-free mass gain.
Exercise 15.24
(i) An increase in contractile muscle tissue. (ii) Sodium and water retention, which GH axis stimulation causes and which body composition scans register as fat-free mass. Only the first would satisfy someone hoping to be stronger — which is why functional endpoints, not composition endpoints, settle the question.
Exercise 15.25 †
Marking guide. Acceptable rate-limiting factors include total sleep and sleep quality, total energy and protein intake, training load management and adequate recovery time between sessions, an unresolved structural injury, and untreated illness or medication effects. For each, the expected observation is that raising growth hormone does not address the limiting factor at all. The general point to reward: adding more of a non-limiting input produces nothing, which is the most common reason a mechanistically sound intervention fails to register in practice.
Exercise 15.26
Verbatim: Tesamorelin for reduction of excess visceral abdominal fat in HIV-associated lipodystrophy: ✅. MK-677 for body composition in healthy adults: ⚠️. CJC-1295 + ipamorelin for improved body composition, recovery, or function in healthy adults: ⚠️→❌. Sermorelin / GHRP-2 / GHRP-6 for the popular claims: ❌.
Exercise 15.27
Tesamorelin: post-approval evidence that the visceral fat reduction fails to translate into meaningful health benefit, or that long-term harms outweigh it in this population. MK-677: adequately powered trials showing the composition change produces a functional or clinical outcome with metabolic effects characterized over multi-year exposure (toward ✅), or showing the added mass is predominantly fluid or the metabolic cost dominant (toward ❌). CJC-1295 + ipamorelin: a randomized, placebo-controlled trial in healthy adults with a pre-registered functional or composition endpoint — noting that further surrogate data would change nothing. Sermorelin / GHRP-2 / GHRP-6: randomized, placebo-controlled human trials with the relevant functional endpoints.
Exercise 15.28
Both engage the GHRH receptor and both raise growth hormone. The difference is what was tested: tesamorelin was carried through randomized phase 3 trials with a pre-specified objective endpoint in a defined population and reviewed by a regulator; sermorelin has never been tested against the outcomes it is currently sold for. The ratings describe the state of evidence for two different claims, not the relative quality of two molecules.
Exercise 15.29 †
Marking guide. For the split: rule 1 attaches ratings to claims, and "raises IGF-1" and "improves recovery" are two claims with two evidence bases; collapsing them destroys the information the reader needs and would also require the book to deny something true. Against: a reader wants a verdict, and a two-symbol rating can be read selectively by anyone who wants to see the ⚠️; the notation also risks implying a trajectory rather than a distinction. Deciding: accept either conclusion if defended. Reward students who notice the arrow could be misread as "currently ⚠️, heading toward ❌" rather than "surrogate ⚠️, outcome ❌," and who propose clearer notation.
Exercise 15.30
Rule 4 — never downgrade with distaste. If ignored, MK-677 would be pushed to ❌ on the basis of the market it is sold in and the claims made for it, which would mean the rating no longer described the evidence. It would also make the system unfalsifiable: a compound with a real two-year randomized trial would be rated below one without, purely on cultural association.
Exercise 15.31 †
Marking guide. The advocate's paragraph should be genuinely strong: mechanism narrows the space of plausible claims, guides which trials are worth funding, predicts side effects correctly (the glucose signal was predicted by GH physiology before it was measured), and distinguishes a compound with a coherent target from one with none. All of that is true and the book agrees with it.
Where it goes wrong: every one of those functions is prospective — mechanism tells you what to test and what to watch for. None of them is retrospective evidence that the effect occurred. The rating system measures what has been demonstrated in humans, and mechanism by construction is the thing that precedes demonstration. Credit for citing Chapter 2 §2.9 — roughly nine in ten compounds entering human trials never reach approval, and nearly all of them had good mechanisms.
Exercise 15.32
Claims present: (i) "restores growth hormone to youthful levels" — partly supportable as a statement about a laboratory value; unfalsifiable as written, since "youthful levels" is undefined and GH is pulsatile. (ii) "verified by laboratory testing" — supportable, and irrelevant to the outcome question; verifying a surrogate is not verifying a benefit. (iii) "peptide protocol" — note MK-677 would not qualify. (iv) "without the risks of synthetic HGH" — unsupported; no comparative safety trial exists, and the on-target risks of raising GH and IGF-1 are shared by construction. (v) Implied but unstated: that any of this produces a benefit. That claim is the load-bearing one and it is never made explicitly, which is itself worth noting.
Exercise 15.33 †
Marking guide. This item is assessed on tone as much as content. A passing answer concedes plainly that these compounds do raise GH and IGF-1, that a lab report showing it is real, and that feeling better is data about something even if not about the compound. It then introduces the surrogate/outcome distinction as a question rather than a correction, and names something concrete and useful — the glucose and insulin sensitivity question, the identity-of-molecule question, a baseline worth having. Mark down any answer that opens by telling the person they have been fooled; that sentence ends the conversation and the exercise is explicitly about noticing it in your own draft.
Exercise 15.34 †
Marking guide. Expect: population — healthy trained adults, since that is who uses it; duration — long enough for adaptation to appear, so months rather than weeks; primary endpoint — a functional measure (strength, work capacity, time to return to training) rather than a composition measure; comparator — placebo, with the practical difficulty that the compounds have noticeable effects that may unblind; pre-specified analysis distinguishing muscle from fluid — e.g. pairing composition imaging with a direct functional measure, or a hydration-sensitive assessment. On funding: no patent position, no regulatory pathway being pursued, an existing market that does not require the evidence, and a plausible outcome unfavorable to sellers — award credit for recognizing that the incentive structure, not the science, is the binding constraint.
Exercise 15.35
Assessed against the worked demonstration in the chapter. Required fields: direct target (receptor and tissue), immediate effect, intermediate steps, claimed final effect, step count, opposing forces, and evidence stops at. The most commonly omitted line is opposing forces; a complete answer names IGF-1 feedback, possible downregulation, the sustained-versus-pulsatile mismatch where applicable, and fluid retention confounding composition measures.
Exercise 15.36
Assessed on completeness and on the final comparison. The pedagogical payoff is the ranking step: most students find their confidence ordering does not match their evidence ordering, and the useful observation is which direction the mismatch runs — typically more generous toward compounds already in use. Reward honest reporting of that mismatch over tidy answers.
Chapter 16
Covers exercises.md (34 items, Sections A–G). Quiz answers live in quiz.md behind the
<details> key and are not duplicated here.
† items are open-ended by design. For these, notes describe what a strong response contains and what a weak one looks like, rather than giving a model answer.
Section A — How muscle grows
A1. Balance = synthesis − breakdown, integrated over time; hypertrophy when positive. Synthesis side: resistance training, protein/leucine intake, adequate energy. Breakdown side: fasting, glucocorticoids, inflammation, disuse. Accept any correct pairing. Strong answers note that both routes produce accretion but may produce different tissue.
A2. Right: amino acid availability is a genuine and necessary input. Missing: mechanical loading, which is the initiating stimulus; energy availability; time. Ranking should place loading first. Watch for students who rank hormones above loading — that is the exact error the chapter is inoculating against.
A3. All three removed the loading term. Balance tips negative; loss is measurable within days.
A4. Hypertrophy = bigger existing fibers; hyperplasia = more fibers. Hypertrophy dominates in adult humans. Matters because the Belgian Blue's phenotype is substantially hyperplasia acquired developmentally, unavailable to an adult given a drug.
A5. † Strong objection: "adjunct" claims are not the same as "substitute" claims, and the chapter may be attacking a strawman — nobody serious claims a peptide replaces training. A strong response concedes this, then observes that (a) the consumer market frequently does make substitute-adjacent claims implicitly, and (b) the adjunct claim is the harder one to demonstrate and is exactly what has not been tested. Weak responses just restate the chapter.
Section B — IGF-1, mecasermin, IGF-1 LR3
B1. GH from pituitary → liver produces IGF-1 → IGF-1 mediates many anabolic effects. Credit mention that muscle also produces IGF-1 locally.
B2. Upward of 99% bound. IGFBPs extend half-life (minutes → hours), restrict distribution out of circulation, and hold IGF-1 inert until released.
B3. Total IGF-1 conflates bound and free; the biologically active pool is free IGF-1 at a specific receptor at a specific moment, governed by a six-protein system the measurement does not capture.
B4. Severe primary IGF-1 deficiency in children (plus GH gene deletion with neutralizing antibodies). Must include: pediatric, growth indication, deficiency state. Deduct for any phrasing that could be read as covering adults or muscle.
B5. IGF-1 is structurally related to proinsulin and engages insulin receptors at sufficient concentration, lowering blood glucose.
B6. † Both paragraphs should be genuinely argued. The "clever solution" paragraph should cite the real pharmacokinetic problem (short free half-life, tight sequestration). The "disabled control system" paragraph should note that buffering, distribution limits, and localization are functions rather than inconveniences. Strong responses recognize both are true simultaneously and that the disagreement is about whether the removed regulation was doing something worth keeping.
B7. † Establishes: the molecule is manufacturable, has been characterized enough for industrial use, and was designed for a non-human purpose. Suggests but does not establish: that nobody performed human-relevant safety characterization. Watch for the genetic fallacy — provenance is suggestive, not disqualifying on its own. The real argument is the absence of the intervening studies, not the origin.
Section C — Myostatin and follistatin
C1. Captures: growth is actively restrained rather than merely permitted; removing restraint increases growth. Obscures: that "the brake" is one member of an overlapping family, that its effect is graded rather than binary, and that release does not necessarily produce functional muscle.
C2. Myostatin → ActRIIB → type I receptor recruitment → Smad2/3 phosphorylation → nuclear translocation → suppression of growth transcription; parallel damping of Akt/mTOR.
C3. Activin A (reproductive signaling, inflammation, fibrosis) and GDF-11 (aging research). Trade-off: narrower blockade = more selective, weaker effect; broader blockade = larger effect, more family silenced.
C4. Mice — engineered, experimental. Cattle — naturally occurring, observational, enriched by selective breeding. Dogs — naturally occurring, observational, with a functional (racing) correlate. Humans — single case report, observational, n = 1.
C5. Activin A, GDF-11, several BMPs (any three). Liability because the off-targets have their own essential functions; broad neutralization is not a precision intervention.
C6. Size puts it in protein territory (Ch 1), so it cannot be made by solid-phase peptide synthesis; glycosylation requires eukaryotic expression machinery. Together they mean an authentic product requires biologic-grade manufacture, which a purchaser cannot verify.
C7. † Must identify the category error: gene therapy is a one-time viral vector producing local expression from within tissue, not systemic injection of a protein. Tone matters — the exercise explicitly asks for non-condescending. Reward responses that concede the underlying biology is interesting before explaining why the citation does not transfer.
Section D — Trials, mass versus function
D1. "Lean mass increased and function did not reliably follow." Insist on both halves.
D2. Mass (DXA lean mass, MRI volume); strength (dynamometry, grip, 1RM); function (gait speed, chair stand, 6MWD, timed function tests); independence (living alone, not falling, staying out of care).
D3. Any two of: DXA lean mass includes water/glycogen/connective tissue; added tissue may be non-contractile; specific force can fall; neural drive limits expressed strength independently of cross-section.
D4. Validated surrogate = evidence that intervention-induced change in the surrogate reliably produces the corresponding outcome change, for that intervention and population. For lean mass: trials showing that drug-induced mass gains produce proportional function gains. The §16.5 trials were effectively that test and were negative.
D5. Because mass has failed as a predictor in this setting; the requirement encodes the memory of that failure so a sponsor cannot market a mass effect as a benefit.
D6. (i) Myostatin expression already low in dystrophic muscle — less brake to release. (ii) Underlying defect is membrane fragility, which extra bulk does not repair. Post hoc = generated after results, so untested — a reason to hold them as hypotheses, not to discount the measured result.
D7. † The exercise is the annotation, not the paragraph. Look for students catching: "increased lean muscle mass" used where readers will hear "stronger"; "trending toward" on functional endpoints; "well tolerated" doing work; the absence of the functional result from the headline. A strong response notices its own rhetorical moves without excusing them.
D8. † Reconstruction should include: training raises every rung simultaneously; the correlation in personal experience is therefore confounded; a drug performs one job only. The resisting example is the interesting half — e.g., someone who gained size on a program and did not get stronger, or an injury period where strength held while size fell. Reward students who can produce a genuine counterexample rather than a compliant one.
Section E — Rating practice
Grade on structure first: population, endpoint, dated, falsifiable. A rating without a named population and endpoint is not a rating.
E1. Note the trap: the chapter rated improves growth, not restores normal adult height. The stated claim exceeds what the approval supports. Expect ⚠️ or a rewritten claim. Full credit for students who notice the endpoint substitution and say so explicitly.
E2. ❌. No human outcome trials; recovery is itself a poorly operationalized endpoint, which is worth flagging.
E3. ⚠️ or ❌ depending on how the student weighs the existing negative data. Both defensible; the reasoning is what is assessed. Strong answers note that "preserves ambulation" is the right kind of endpoint — functional, meaningful — and that it is precisely what did not move.
E4. ❌, and strong answers note the asymmetry: not merely unstudied, but adjacent human evidence points away.
E5. † The system is not broken because the rating tracks the state of the evidence base, not the plausibility of the claim; those come apart in Case C. Proposed improvements usually involve a second axis or a modifier symbol. The important half is the cost: added dimensions reduce legibility, and the rating system's value is partly that it is simple enough to be used. Reward students who name that trade-off.
Section F — Field 6
F1. Endpoints reported; surrogate or outcome; if surrogate, validated?; was the outcome measured too?; direction check; rating consequence.
F2. Grade on correct case assignment and honest handling of "I could not determine what was measured" — which is itself a legitimate Field 6 entry and should be credited, not penalized.
F3. Any real example. Strong entries name the unstated arrow explicitly as a proposition ("changing X reliably changes Y").
F4. † Key distinction: Case A can be resolved by running the study; Case C has already run it. A person facing Case A is deciding under ignorance; a person facing Case C is deciding against evidence. Practical difference: Case A warrants watching the literature; Case C warrants a much higher bar for any new claim in the same family.
Section G — Synthesis
G1. Something equivalent to: "how much tissue you have is not the same as what you can do with it." Reject answers that smuggle the banned terms back in via synonyms like "proxy."
G2. Secretagogue → GH pulse → IGF-1 level → lean mass → strength → function → independence. Six arrows, none validated for this intervention. Accept minor variations in chain length; the count and the reasoning matter more than the exact links.
G3. † The five sentences: mechanism (growth + anti-apoptosis); epidemiology (association with several cancers); confounding (nutrition, body size, GH status, genetics; no causal claim); acromegaly; GH receptor deficiency. The removal exercise is the point — most students remove the confounding sentence, which converts a careful statement into an alarmist one. Have them notice that.
G4. † Assess for: absence of lecturing; at least one question asked before advice given (what are you hoping it does? what has the clinic told you? what else are you taking?); the two or three points usually being (i) the human evidence for the benefit does not exist, (ii) the long-term risk is unquantified rather than known to be small, (iii) product identity and clinician involvement. Penalize moralizing; the book's voice explicitly forbids it.
G5. † Should specify: healthy trained adults, randomized, placebo-controlled, double-blind, standardized supervised training in both arms, minimum several months, co-primary mass and strength or function endpoints, independent analytical verification of drug product. The honest closing — sponsor incentives, cost, regulatory pathway, ethics of enhancement trials, anti-doping status — is the most instructive part.
G6. Ungraded. Collect only if the class has consented to sharing dossier work; several students will have entries about compounds they or people close to them use.
Chapter 17
Worked solutions for the twelve items marked † in exercises.md. Destined for Appendix M.
Exercise 17.2 †
The sentence: As of 2026, there is no completed, peer-reviewed, randomized controlled human trial of BPC-157 for any indication in the published literature.
Why the date is part of the claim. The statement is an empirical report about the contents of a searchable public record at a moment in time, not a permanent property of the compound. It is therefore falsifiable by a single new publication, and a reader who checks in 2029 may correctly obtain a different answer. Date-stamping does three things: it tells the reader the claim is checkable, it tells them how to check (search the same databases), and it tells them the author knows the claim can expire. A rating without a date is presented as a fact about the world rather than a report on the evidence, which is precisely the error the rating system exists to prevent.
Full credit requires the student to notice that "not one" also covers safety studies — there is no published Phase I either — and that "in the published literature" is doing deliberate work, because a registered or completed-but-unpublished study is not a readable result.
Exercise 17.5 †
The five reasons (Ch 5 §5.3):
- Different biology — species differ in metabolism, healing rate, tissue architecture, proteolytic environment, immune response, and lifespan.
- Different disease — a model captures a piece of a condition, deliberately simplified so the experiment is possible. The question is always whether the discarded complexity mattered.
- Different dose — allometric scaling is approximate; route and exposure profile differ; and the laboratory routes used often have no clinical equivalent.
- Different endpoint — animals cannot report pain, function, sleep, or quality of life, so animal endpoints are necessarily surrogates.
- Different publication pressure — animal research is largely unregistered and null results are rarely published.
Least discussed: reason 5. It is least discussed because it is invisible — it concerns the studies you never see, and there is no artifact to point at.
Why it matters especially here. BPC-157's case for efficacy rests heavily on the consistency of a large published animal literature. Consistency is exactly what a publication filter produces. Without a denominator, the published record cannot distinguish "a robust effect" from "a small or absent effect plus ordinary selective publication." The argument that carries the most rhetorical weight is therefore the one most weakened by reason 5, and no additional publication within the same system resolves it. A strong answer notes that this is a structural property, not an accusation, and that the specific remedy is independent replication by unaffiliated laboratories.
Exercise 17.8 †
The Field 3 entry:
MECHANISM — not established. No receptor definitively identified. Proposed pathways (nitric-oxide-related signaling; growth-factor and angiogenic signaling; effects on vasculature) are inferred from downstream observations in animal models, not derived from an identified target. Confidence: low. Source type: preclinical, animal.
Why "acts on nitric oxide pathways" is the wrong entry. It states a hypothesis in the grammatical form of a finding. The observation underlying it is that administering the peptide is followed by changes in markers and tissues associated with those pathways; the inference that the peptide acts on them is a proposed explanation for that correlation, not a measurement of it. Writing the hypothesis into the field permanently launders it into a fact, and every downstream use of the dossier inherits the error.
Credit the student who notes the second-order consequence: a Field 3 that reads like established mechanism invites Chapter 5's rule-3 violation (upgrading a rating with mechanism) at the moment the dossier is next consulted, possibly by the same person months later who no longer remembers how the entry was derived.
Exercise 17.13 †
The two interpretations of a day-14 advantage that has narrowed by day 28:
- Accelerated healing. The treated animals reached a given state sooner; controls caught up. The endpoint of healing is the same, the trajectory to it is faster.
- Better healing, imperfectly measured. The treated tissue really is superior at both timepoints, and the day-28 measurement failed to detect it — because the outcome measure saturates near normal, because variance grew, or because the study was underpowered at that timepoint.
Distinguishing them. Add later timepoints (day 56, day 84) and a measure that does not saturate. If the curves converge and stay converged, interpretation 1. If a difference re-emerges or persists on a non-saturating measure, interpretation 2. Increase group size at the later timepoints specifically, since the expected difference is smaller there and detecting it requires more power, not less. Add a functional measure rather than relying on terminal mechanical testing.
Why it matters: the two interpretations have different clinical implications, and the distinction is routinely lost in summary. "Accelerated" is a real and valuable claim for an acute injury and may be close to irrelevant for a chronic one.
Exercise 17.16 †
Why the count is uninterpretable. A count of published studies is a numerator. To interpret it you need the denominator — how many studies of this kind were performed. Clinical trials supply an approximate denominator through mandatory pre-registration; preclinical research does not register, so no denominator exists, and no one can construct one retrospectively. Additional problems: reviews get counted as studies, overlapping reports from a single experimental program get counted separately or jointly depending on the counter, and abstracts get mixed with full papers.
What follows. Twenty concordant published studies drawn from twenty performed is a strong signal. Twenty concordant published studies drawn from a hundred performed is close to no signal. The published record looks identical in both worlds.
The high-information study: an adequately powered, pre-registered replication conducted by a laboratory with no prior involvement in the compound, ideally with a published analysis plan and a commitment to publish regardless of outcome. This carries far more information than another study from a group that has already published many, because it is drawn from a different (and declared) denominator.
Why requesting it is not an accusation. Independent replication is the ordinary currency of science and is requested of every finding that matters. A literature concentrated in a small number of groups shares methods, reagents, animal sources, and assumptions — not through misconduct but because that is what happens when a small number of people find a question interesting. Asking for independent replication addresses the shared-method risk without implying anything about anyone's integrity, and a strong answer says so explicitly.
Exercise 17.18 †
Three general reasons a rodent mg/kg figure cannot be converted:
- Allometry. Metabolic rate, clearance, and volume of distribution scale non-linearly with body mass. The standard conversions used in early development are approximations that get revised the moment real human pharmacokinetic data arrives.
- Route. Laboratory routes (intraperitoneal in particular) have no clinical equivalent and produce exposure profiles that a subcutaneous injection does not reproduce.
- Species pharmacokinetics. Absorption, protein binding, tissue distribution, proteolytic clearance, and renal handling all differ. A dose is only a proxy for exposure, and exposure is what the tissue sees.
The BPC-157-specific reason: there is no published human pharmacokinetic study. For an ordinary development candidate, the animal-to-human conversion is a starting hypothesis that Phase I immediately tests and corrects. Here the correction step has never occurred, so any human figure in circulation is an unverified extrapolation resting on assumptions that no one has been able to check.
Full credit requires the student to state the implication rather than just the fact: anyone quoting a human dose is quoting an extrapolation, and if they have not said so, they either do not know it or have chosen not to mention it.
Exercise 17.24 †
The five barriers (Ch 4):
- Stomach — hydrochloric acid and pepsin.
- Small intestine — pancreatic proteases (trypsin, chymotrypsin, elastase, carboxypeptidases).
- Brush border — peptidases on the intestinal epithelial surface.
- The epithelium itself — a peptide of ~1,419 Da carrying net negative charge cannot passively diffuse across, and no established transporter carries it.
- The liver — hepatic first-pass metabolism of anything absorbed into the portal circulation.
A stability argument addresses barriers 1–3. Protease resistance is a claim about surviving degradation.
It cannot touch barrier 4. Permeability is an independent physical property determined by size, charge, and lipophilicity. A perfectly indestructible molecule that cannot cross the epithelium has zero systemic bioavailability. Barrier 5 is also untouched, though it is the least decisive of the five here.
The one-line version students should be able to produce: stability answers "does it survive?"; bioavailability answers "does it arrive?" — and only the second is the question.
Exercise 17.29 †
"There are hundreds of peer-reviewed studies on this compound going back decades. Calling it unproven is just ignorance of the literature."
What is true. A great deal. The literature is real, peer-reviewed, indexed, spans multiple model families, and comes from more than one group. Someone who says "there is no evidence" is wrong and will be correctly identified as not having read anything. The speaker is better informed than most of the people arguing with them.
Where it fails. Two places. (a) Level, not volume. Evidence does not accumulate across rungs; studies on the animal rung increase confidence within that rung and never produce a human finding by accumulation. A thousand rat studies and one rat study make the same claim about humans. (b) The count is not verifiable in the way the claim implies — no preclinical registry, reviews counted as studies, overlapping reports counted variously. And "unproven" was never a claim about volume.
Replacement (two sentences): There is a real and substantial animal literature here, and anyone who calls it "no evidence" hasn't read it. The open question is a different one — how many of those studies are in humans — and that has a checkable answer that takes ten minutes to obtain.
Note for graders: the replacement must not concede that the compound works, and must not sneer. The model answer converts an argument into a search.
Exercise 17.31 †
"Thousands of people have used it for years and there are no reports of serious side effects. That's a better safety record than most prescription drugs."
What is true. No large, obvious, acute harm signal has emerged publicly. That is not nothing — a compound producing frequent dramatic harm would likely have generated visible reports even without a formal system. And the concern behind the claim is legitimate: people want to know whether they are in danger.
Where it fails. Three places. (a) "No reports" describes a reporting system, and there isn't one. Attribution is unlikely for nonspecific symptoms, disclosure is discouraged, pharmacovigilance frameworks are built around approved products with marketing authorization holders, and no one aggregates. (b) There is no denominator. "Thousands of people" is itself an estimate no one can source; without exposure data there is no rate, only anecdotes about anecdotes. (c) The comparison is backwards. Prescription drugs have known safety profiles because they were systematically studied and are systematically monitored; the appearance of a "worse" record is a product of looking. A compound nobody monitors will always appear cleaner than one everybody monitors.
Replacement (two sentences): Nothing dramatic has surfaced, and that's worth something — but there's no reporting system, no denominator, and no published human safety study, so "no reports" and "shown to be safe" aren't the same statement. A single well-conducted Phase I would tell us more than all the accumulated silence.
Exercise 17.34 †
"A registered clinical trial exists — I found it on ClinicalTrials.gov. So the human research is happening; it's just not published yet."
What is true. Registries are the right place to look, and the person did the correct thing by looking. A registration is also genuine information: someone with an institutional affiliation committed a design to a public record.
Where it fails. A registration is a plan, and the five states are distinct: registered → run → completed → results posted → peer-reviewed publication. The claim silently promotes state one to state four. Specific checks the student should name: read the recruitment status field (Not yet recruiting, Withdrawn, Terminated, and "Unknown status" are all common and all mean the study is not producing data); check the enrollment figure and whether it is actual or anticipated; check for a Study Results tab; check the sponsor; check the primary outcome and whether it is a clinical endpoint or a surrogate; and check the last update date.
Replacement (two sentences): A registration means somebody filed a plan, which is genuinely different from somebody having an answer — check the status field and whether results are posted. When a completed trial with a pre-specified functional endpoint appears and is published, that's the event that moves the rating, and it's worth watching for.
Exercise 17.35 †
A model answer (students should not reproduce §17.6's wording):
Claim: BPC-157 reduces gastrointestinal mucosal injury in humans taking long-term NSAIDs. Rating: ❌ Hype outpaces evidence (as of 2026). Why: The supporting evidence is rodent work in chemically induced damage models — the oldest and most internally coherent strand of the preclinical literature, and still entirely preclinical. No completed, peer-reviewed randomized human trial exists for this endpoint, and no human safety data exists to support use in a chronic condition. What would change it: a randomized, placebo- and active-controlled trial in patients requiring long-term NSAID therapy, with an endoscopically confirmed or clinically hard endpoint (ulceration, clinically significant bleeding), not a symptom score.
The two-sentence explanation of the difference. The tendon rating's central difficulty is a pathology mismatch — the animal model may not represent the human disease at all (§17.7). The gastrointestinal rating's central difficulty is different: the model maps onto the human condition comparatively well, but an effective, cheap, well-characterized comparator already exists, so the bar a trial must clear is incremental benefit over proton pump inhibition rather than benefit over nothing.
Connection to rule 6: one molecule, many ratings. Same compound, same evidence base, two claims — and the reasons differ even though the symbols match, which is exactly why the one-sentence reason line is not decoration. A source that gives a molecule a single overall rating has compressed away this entire distinction.
Exercise 17.36 †
This item has no single correct answer; it is assessed on execution. What to look for:
Move 1 is the diagnostic. For the compound the student wanted to work, move 1 will usually be either (a) noticeably longer and more generous than for the neutral compound, or (b) conspicuously brief, because the student is defending against their own attachment by overcorrecting. Both are findings and both are worth naming. Students who report "no difference" have usually not written both entries in full.
Move 2 should be one sentence with a date. If it has grown into a paragraph with hedges, the student is arguing.
Move 4 is the most-skipped. An entry missing "what this rating does not claim" is a verdict wearing a rating's clothes.
Move 5 must be specific. "More research is needed" earns no credit. Look for a named population, a named endpoint, and a stated magnitude.
Move 6 should produce at least one actual rewrite. A student who reports that nothing needed changing has almost certainly not performed the read-back honestly, or has written something so hedged that it could not offend anyone because it does not say anything.
The paragraph of reflection is the point of the exercise. The interesting finding is rarely that the student was wrong; it is discovering which direction they were wrong in, and whether they were consistently more generous toward compounds they hoped would work.
Chapter 18
For instructor use. exercises.md ships without answers by design. Quiz answers are published in the
student-facing quiz.md under a collapsed <details> block; they are reproduced in condensed form at
the end of this file for convenience.
Several exercises have no single correct answer. Those are marked [open] and what follows is a description of what a strong response contains, not a solution.
Section A — The parent and the fragment
A1. Thymosin β4 is a 43-amino-acid endogenous human peptide that binds G-actin. TB-500 is a laboratory code for a marketed short synthetic fragment described as corresponding to that molecule's actin-binding region. Strong answers state the length difference and the endogenous/synthetic distinction. Weak answers say "TB-500 is a version of thymosin β4," which reproduces the error.
A2. Accurate rewrite, roughly: "TB-500 is marketed as a synthetic fragment of thymosin β4. Thymosin β4 — the full 43-residue molecule, not the fragment — has been studied in animal models of tissue repair." The commercial usefulness of the inaccurate version is that it is shorter, sounds authoritative, and transfers a peer-reviewed literature without appearing to make a claim.
A3. Rule 3 forbids upgrading a rating with mechanism. "The actin-binding region is the active part" is a mechanistic story about what should happen; it is a hypothesis, and hypotheses do not move ratings. Evidence that would count: a study directly comparing fragment and parent on the activity in question, and — for a clinical claim — a randomized controlled human trial of the fragment on a clinical endpoint.
A4. [open] Five links: (1) the animal data on the parent is sound; (2) the fragment reproduces the parent's activity; (3) the vial contains the fragment; (4) at the stated concentration and purity; (5) the activity produces a clinical outcome. Checks: (1) read the paper; (2) a direct comparative study — rarely available; (3) and (4) independent analytical testing (Ch. 34) — not ordinarily available to a consumer; (5) a registered trial — none identified. The point students should reach: only link 1 is routinely checkable by a reader, and it is about a different molecule.
A5. Has told you: the compound never entered formal drug development far enough for a naming authority to assign a generic name — a fact about regulatory history. Has not told you: anything about pharmacology, efficacy, or safety. Insisting on both halves is the exercise; students who give only the first half have converted a regulatory observation into a pharmacological verdict.
A6. Ch. 1 §1.5: the peptide/protein boundary near 50 residues is a convention, not a chemical distinction. Nothing in the chapter's argument depends on the label. Students who think something does depend on it should be asked to identify which claim would change.
Section B — Reading the animal literature
B1. Dermal wound healing (rodent skin is loose and heals substantially by contraction; human skin does not); cardiac repair after injury (model infarcts are acute, uniform, and induced, unlike most human presentations); corneal healing (arguably the best transferable case, since corneal biology is comparatively conserved and the endpoint is directly observable — accept this as a strong answer).
B2. Species differences; dose and exposure scaling; model artificiality; endpoint mismatch; publication filtering. The forgotten one is nearly always publication filtering.
B3. [open] Any analogy carrying the "no denominator" idea. A serviceable one: a jar of customer testimonials on a counter tells you nothing about satisfaction, because you cannot see how many customers wrote nothing and how many wrote something the owner threw away — and there is no independent record of how many customers there were.
B4. Reason: any number implies a completeness nobody can verify, since animal research is largely unregistered. Objection worth taking seriously: refusing to quantify makes the literature sound smaller and vaguer than it is, and could itself mislead in the other direction. Good responses engage the objection rather than dismissing it — a fair resolution is to describe the literature's scope (areas, duration, institutional character) without a count.
B5. [open] Defense: a composite teaches design without risking misattribution, and it is labeled. Objection: constructed figures are exactly the device bad actors use, and a reader cannot distinguish a teaching composite from an invented result if the label is missed. Strongest resolution: labeling must be inside the figure, not adjacent to it — which is how Figure 18.1 is built.
B6. Volume and rigor do not convert categories. More and better animal studies produce a better case for running a trial, not a substitute for one.
B7. [open] Expected: rating moves ❌ → ⚠️. The "reason" field must now cite the trial and note its size and single-study status. The "what would change it" field should name replication in an independent population as the route toward ✅, and failure to replicate as the route back.
Section C — Split regulatory status
C1. Approved in a number of countries for defined indications, notably hepatitis B and as an immune adjuvant in some jurisdictions; not FDA-approved in the United States. Both halves in one sentence.
C2. Promotional reading: approved elsewhere, so it works and the FDA is slow or captured. Dismissive reading: the FDA declined, so it does not work and other regulators are lax. Shared premise: one regulator is right and the other has erred.
C3. (a) Convergent approval — evidence cleared every threshold applied; tells you the effect is large and consistent enough for different decision rules to agree; does not tell you the effect size for an individual. (b) No approval anywhere with no dossier submitted — tells you about funding, incentives, and who is willing to pay for evidence; tells you nothing about efficacy. (c) Split — evidence in the genuinely contestable middle; tells you serious people disagreed on the same material; does not tell you who is right.
C4. [open] Both sides defensible. "Feature" argument: approval is a benefit-risk judgment, and risk tolerance rationally varies with the severity and prevalence of the untreated condition and the availability of alternatives. "Flaw" argument: it makes "approved" mean different things in different places while patients read it as a single global signal.
C5. [open] Look for: acknowledging that approvals elsewhere are real institutional acts based on real evidence; correcting the implication that non-approval is mere delay; noting that the approval attaches to an indication; ending with a question the patient can act on. Reject responses that condescend or that resolve the tension by picking a side.
C6. Warning against a rating system where the middle is unused or used as a dumping ground for anything controversial. External test: look at the distribution of a source's ratings and at whether middle ratings carry specific "what would change it" content. A middle category with vague or absent change conditions is functioning as a hedge.
Section D — LL-37
D1. Direct membrane-disrupting antibacterial activity (measurable: growth inhibition at a stated concentration against a stated organism); immunomodulatory activity (measurable only after specifying cell type, mediator, and direction). The first yields testable claims easily because the endpoint is built into the activity.
D2. Rule 5: a rating must be falsifiable and checkable, and a rating issued in a section that does not present the evidence gives the reader nothing to check. Ch. 25 presents the evidence.
D3. "Natural human peptide, therefore safe" fails because the mechanism — membrane disruption — does not perfectly discriminate bacterial membranes from human ones; selectivity is a matter of degree and an unsolved design problem. Echoes Ch. 1's ⚠️ Hype Check on "peptides are natural, so they're safe."
D4. [open] Falsifiable version must name a cell population or mediator, an assay, a direction, a population, and a clinical outcome. The marketing version will typically retain only a positive valence. What is lost: everything that could make it wrong.
Section E — Gut peptides
E1. KPV has been studied under its own name in the preclinical literature; TB-500 inherits citations generated on its parent. That distinction favors KPV. It does not change that KPV's literature is preclinical.
E2. Because its target is at the apical surface of the intestinal epithelium, facing the lumen; systemic absorption would add exposure without therapeutic purpose.
E3. (a) Available — target is local. (b) Not available — tendon is systemic. (c) Available — gut motility can be influenced from the lumen. (d) Not available — muscle is systemic.
E4. "Can it be absorbed?" presupposes absorption is required. Asking it first leads you to treat poor bioavailability as a defect even when it is the specification — and, in the other direction, to accept a "it works locally" defense for a systemic claim. Ch. 1 §1.6's linaclotide mention is a good example.
E5. [open] Case for different ratings: larazotide has human trials, KPV does not; a rating system tracking evidence should distinguish them. Case for the shared rating: 🔬 denotes "too early to rate, proceeding properly," which is true of both, and the tier is about the state of the question rather than the volume of work. Either position is acceptable if defended; strong answers note that 🔬 is the one tier where volume genuinely matters less.
Section F — The unfalsifiable claim
F1. Arm/cell population/mediator; assay; direction; population; connection to a clinical outcome.
F2. (a) Not falsifiable — missing all four. (b) Falsifiable — arm, direction, and population specified; missing only the clinical outcome link. (c) Not falsifiable — missing all four; "balance" is the tell. (d) Falsifiable and outcome-linked — the model answer. (e) Not falsifiable — "supports" and "healthy" carry no direction or endpoint.
F3. Because it explicitly refuses to name a direction and presents the refusal as sophistication. An exaggeration can be corrected by better data; this construction is immune to data by design.
F4. Missing: connection to a clinical outcome. It matters more than it appears to because immune markers shift in response to sleep, exercise, stress, recent meals, and time of day, and their relationship to outcomes patients care about is loose. Ch. 6's surrogate graveyard is the reference.
F5. [open] Shared structure: an umbrella term collapsing a heterogeneous, multi-directional system into a single positively-valenced word, after which any finding within the system can be described in the umbrella's terms. Outside examples students commonly produce: "boosts metabolism," "balances hormones," "supports gut health," "detoxifies," "builds functional strength," "optimizes portfolio risk." Accept any where the direction and endpoint are genuinely absent.
F6. Because the claim is not the kind of thing evidence bears on. It is not a cop-out because the field still does work: it says the remedy is specification rather than data collection, which is actionable and testable in a way "we need more research" is not.
F7. [open] The gap between the two rewrites is what marketing fills. Look for students to notice that the generous falsifiable version is usually narrower and less appealing than the original — which is precisely why the original exists.
Section G — The symmetric risk
G1. Angiogenesis, cell proliferation, cell survival. Also required for tumor growth beyond a minimal size.
G2. (1) These compounds are promoted as promoting angiogenesis, proliferation, and survival. (2) Tumors require angiogenesis, proliferation, and survival. (3) A systemically administered signal does not discriminate by intended tissue. (4) Therefore the promoted mechanism is, mechanistically, also a tumor-relevant mechanism. Step 3 is the empirical claim; steps 1, 2, and 4 are definitional or follow from them.
G3. Asserting a link overclaims and violates Rule 4 (downgrading with distaste). Dismissing the concern treats absence of evidence as evidence of absence and abandons Ch. 5's central distinction.
G4. Because a marketed, monitored drug has been subject to years of pharmacovigilance and often epidemiological study — so "no link demonstrated" reflects looking and not finding. For an unstudied compound it reflects not looking. Same words, different informational content.
G5. [open] Strongest objection: mechanism is either admissible evidence or it is not; using it for risk but not benefit lets the author reach whichever conclusion they prefer. Answer must turn on what each claim asserts: an efficacy claim asserts a benefit occurs, which requires demonstration; a safety concern asserts a question is open, which requires only well-formedness. Reject answers that appeal to caution or to which conclusion feels safer.
G6. [open] Look for: accurate mechanism, no asserted link, explicit statement that the question is unstudied, no false reassurance, and a concrete next step (raise it with the oncologist before, name what to ask). The second half — what the oncologist knows that you cannot — should include surveillance status, current therapy interactions, and the ability to notice what the patient did not think to mention.
Section H — The thin-literature dossier
H1. (A) Studied and disappointing — were trials run? (B) Not studied — has anyone registered one? (C) Studied in a different molecule — what substance does each citation name?
H2. Because it does not feel like an absence: a search returns real, peer-reviewed papers, so the box appears full. The feeling of a completed search is the mechanism of the error.
H3./H4./H5. [open] H4's single most informative line is defensibly either "Human RCTs of THIS compound" or "What the literature IS about." Accept either with a defense; reject "regulatory status" as the top answer, since approval status is downstream of the evidence question. H5's diagnostic value: the trigger sentence is usually far easier to write for the compound the student dismisses, which reveals asymmetric standards.
Section I — Synthesis
I1. Shared: no completed RCTs on the marketed compound; substantial animal literature; sold in overlapping markets; laboratory code rather than generic name; both ❌. Differs: BPC-157's animal literature is on the same molecule that is sold; TB-500's is on a parent. The differences list is shorter but the single difference is large.
I2. [open] Grade against: (i) took the animal work seriously — yes, §18.2; (ii) stated the human situation accurately — yes, both ratings; (iii) explained why both are true — yes, five translation problems; (iv) never sneered — the closest failure point, and students often flag the "message boards" line in §18.1 as tonally risky. That is a fair criticism and should be discussed rather than defended.
I3. [open] Human evidence ranking: thymosin alpha-1 ≫ larazotide > LL-37 (for antimicrobial endpoints, per Ch. 25) > KPV > TB-500. Marketing confidence ranking is roughly flat, or inverted. The correlation is absent or negative, which is the observation to draw out.
I4. [open] Strongest case for the category: shared administration route, shared regulatory position, shared consumer context, and a genuine common mechanistic theme (tissue repair signaling). For it to succeed, membership would have to predict something about evidence or effect — and it does not, which is the chapter's thesis.
I5. [open] Assess on: does the reader end able to run the four checks from Case Study 18.2? Is there any moralizing? Did the writer smuggle in a forbidden term under a synonym?
Quiz answer key (condensed)
1 C · 2 B · 3 B · 4 False · 5 A · 6 B · 7 C · 8 B · 9 short answer · 10 B · 11 C · 12 False · 13 B · 14 C · 15 B · 16 B · 17 short answer · 18 B · 19 False · 20 B · 21 C · 22 short answer
Full reasoning for every item, including the four short-answer model responses, is in the <details>
block at the foot of quiz.md.
Chapter 19
Instructor-facing. The exercises.md file deliberately ships without answers; these are model responses and marking notes for the instructor's answer key. quiz.md already contains its own key in a collapsed block — the notes below supplement it only where an item has proven to generate argument.
Section A — Vocabulary and foundations
A1 (release). Look for three elements: a named qualified person, a comparison of batch test results against a written specification, and the creation of traceability / a recall path. A common wrong answer treats release as physical shipment. Credit answers that add "and it is a legal act with a document behind it."
A2 (purity vs content). Purity = fraction of the peptide-like material that is the intended molecule. Content = fraction of total mass that is peptide at all. The honest-label point: gross weight includes bound water and counterions, so a truthful label can still overstate the peptide present. Watch for students who think content is simply "purity expressed differently."
A3 (sterility vs endotoxin). The target sentence is alive now vs ever alive. Full credit requires both failure directions: sterile-but-high-endotoxin, and non-sterile-but-low-endotoxin.
A4 (deletion sequence). Mechanism: a failed coupling on some fraction of chains; synthesis continues; those chains are permanently one residue short. Difficulty of removal: near-identical mass and chromatographic behavior. Strong answers note that unrelated contaminants are easier precisely because they differ.
A5 (epistemic laundering, generalized). Any market works: supplement claims, financial products, nutrition advice. The definition must contain (a) multiple parties, (b) each individually defensible, (c) an aggregate claim none of them made.
A6 (CoA logical form). Accept any phrasing of a report of tests on a sample is not a property of a container. Best answers name at least one mechanism by which the sample and the container come apart.
Section B — The six failure modes
B7. A matrix answer is fine. The key discriminator is the "what would a person observe" column — most entries should be nothing, or something misattributable. That is the pedagogical point.
B8. Two of nine residues; different physiology. The transfer: sequence similarity is not functional similarity, so "related peptide" describes the chemistry of the failure and says nothing reassuring about its consequences.
B9. Amplification (Ch 2 §2.4) plus threshold/saturation. Least consequential: both values on a flat region (deep saturation, or both below threshold with no effect either way). Most consequential: straddling the steep region, where a twofold error is the difference between nothing and near-maximal response. Strong answers note that in an unregulated setting you know neither the concentration nor your position on the curve.
B10 †. Endotoxin survives autoclaving because it is molecular debris, not an organism; it passes sterilizing filters because it is orders of magnitude smaller than the bacteria those filters retain. The misattribution scenario: fever, chills, aches, malaise hours later, indistinguishable from a viral illness. Supervised difference: someone knows what was administered and when, so "reaction to the preparation" is on the differential and gets investigated. Mark down any answer that drifts into procedure.
B11. Aggregation is a handling failure rather than a manufacturing one, which relocates responsibility to whoever transported and stored it — often the buyer, often unknowingly. Strong answers note the vial looks identical either way.
B12 †. Intuitive model: less active material, proportionally smaller effect, wasteful not dangerous. Correct model: aggregated peptide is substantially more immunogenic — a different kind of exposure, not a weaker one. Practical difference: an uncertain handling history is a risk fact, not merely an efficacy fact, and it cannot be inspected away.
B13. Key contrast: a neutralized drug is a treatment failure; cross-reactive antibodies interfere with the body's own molecule and need not resolve on stopping.
Section C — Documents, claims, and the market
C14. Expect: which lot? tested by whom? which attributes? by what method? at what detection limit? does the sample correspond to this container? when? who kept the chain of custody? was sterility tested? was endotoxin tested? were residuals tested?
C15. Any three of: non-random/convenience sampling; heterogeneous test panels; batch- and time-specificity; no denominator.
C16 †. Paragraph 1: existence proofs are strong; identity/purity/content findings on those samples are direct evidence about those samples. Paragraph 2: the reader must not infer suitability for injection. Almost certainly not measured: sterility and endotoxin.
C17. Most-defensible link is usually argued to be the forum (people describing their own lives) or the vendor (selling a labeled reagent). Either is acceptable; the reasoning must connect individual defensibility to the durability of the chain.
C18 †. The strong version must include: synthesis is not exotic; some material genuinely is what it claims; drug pricing critiques are often fair. The response must locate the error in molecule vs batch, not in motive. Banning "dangerous" forces the better argument.
C19. Establishes: shared physical origin of chemistry. Does not establish: what was tested, what was documented, what was rejected, what specification applied, who released it, whether a recall path exists.
Section D — Compounding and supervision
D20. Differences: patient-specific prescription vs batch production; state board vs FDA registration; exempt from CGMP vs required to comply; not FDA-inspected as a facility vs inspected. Common to both: neither is FDA-approved or reviewed for safety and efficacy.
D21. Target sentence: lawful supply establishes lawfulness of supply, and nothing about efficacy. Chapter 5's rating applies to the claim regardless of paperwork.
D22. Salt forms (different chemical entity; evidence does not transfer automatically) and delivery-device substitution (measurement task moved to the patient). General lessons: evidence attaches to a specific substance, and to a specific format.
D23 †. Any ranking is acceptable if defended. The strongest answers pick the differential diagnosis or the baseline and answer the self-provision objection directly: a person can obtain labs but cannot construct a differential for themselves under conditions of personal investment. Watch for answers that slide into "so get labs yourself" — the chapter explicitly forbids treating self-testing as a substitute.
D24. Desensitization/downregulation under sustained stimulation. Dangerous rather than merely wrong because the corrective action makes both problems worse simultaneously: deeper desensitization and higher exposure.
D25 †. Attribution: a yes/no verdict from prior belief. Differential: a ranked list, investigated. Error toward the compound — stopping and feeling reassured while an unrelated treatable condition progresses. Error away — continuing an exposure that is causing harm. Both examples must be concrete.
Section E — Reporting, legality, evidence
E26. Mandatory manufacturer reporting → none; spontaneous reporting systems → none for the product; signal detection → no accumulated reports; label authority → no label; recall → no distribution records.
E27. Must distinguish harm not occurring from harm not being collected, and ideally note that the two are indistinguishable from outside.
E28 †. The case demonstrates detection, not growth hormone pharmacology. Undocumented counterfactual: scattered late-onset cases, no cluster, no denominator, no registry, no question anyone thought to ask.
E29. "My doctor prescribes it" answers question 5 (prescribing). Leaves open: approval, sale, import, possession, and sport.
Section F — Synthesis and transfer
F30 †. Approved medicine → upper left. 503B compounded → upper-left-ish but flag that it is not approved; accept a placement on the boundary if defended. Genuine research-grade of an untested compound → lower left. Strong-evidence compound via unregulated channel → upper right, and the defense must be that evidence attaches to the tested product, not to the molecular formula.
F31 †. Assess tone as much as content. The banned words force away from prescription. The paragraph must contain the unregulated, not peptide distinction and must not moralize.
F32. Independence: a ❌ concerns the state of evidence for a claim; the quality concerns concern a physical container. A world where one changed and the other did not: a positive trial for BPC-157 publishes, and every gray-market vial is exactly as unverified as before.
F33 †. Rating should be ❌. The fourth line is the assessed element: per-lot testing tied to the container by a verifiable identifier, independent lab, published methods and detection limits, full panel, documented chain of custody. Accept answers noting this amounts to reinventing batch release.
F34 †. Assess for: no instruction, no judgment, at least two specific reasons (baseline; differential diagnosis; interaction checking; two-in-the-morning; no reporting pathway). The final sentence about what was left out is where the real learning shows — the best answers say they left out any recommendation to continue or stop, because that is a clinical decision.
Chapter 20
For the instructor answer key. The student-facing exercises.md deliberately carries no answers; the
quiz answer key is published inside quiz.md in a collapsed block. What follows is (a) model answers
for the exercises, and (b) marking guidance for items where the point is the reasoning rather than a
fact.
Section A — The reframe and the three families
A1. Misleading because it implies the receptors exist for morphine. Better: "the receptors for the body's own opioid peptides, which morphine happens to fit." Full credit requires the direction of causation to be explicit.
A2. A specific, saturable, stereochemically selective, naloxone-blockable binding site is not something evolution produces for a plant alkaloid a species may never encounter. Its existence predicts an endogenous ligand. Award credit for the predictive framing — this is the point, not the history.
A3. POMC → proopiomelanocortin → β-endorphin → mu/delta. PENK → proenkephalin → met- and leu-enkephalin → delta (preferred), mu. PDYN → prodynorphin → dynorphin A/B, neoendorphins → kappa.
A4. Tyr-Gly-Gly-Phe. Message = signals "opioid" to a receptor, is the part morphine imitates, and is essentially invariant. Address = the C-terminal extension, determining receptor preference, half-life, and distribution.
A5. Structural: loss of the phenolic hydroxyl on the position-1 aromatic ring, which is a required element of the opioid pharmacophore. Experimental: nociceptin is not blocked by naloxone, so naloxone experiments do not probe the NOP arm — a real practical consequence for study design.
A6. † Open. Look for: the "not a unit" argument resting on tissue-specific processing and functionally unrelated outputs; the "is a unit" argument resting on shared regulation of transcription, shared convertase machinery, and the POMC-deficiency syndrome presenting all three arms at once. The settling evidence would be whether the products are co-regulated in the same cells under the same conditions. Strong answers notice that the question is empirical and cell-type-specific rather than philosophical.
A7. Any two of: more residues for proteases to remove before activity is lost; higher molecular weight slows renal filtration; secondary structure in a longer chain can shield cleavage sites; the enkephalins sit directly in synaptic space where dedicated ectopeptidases are concentrated. Do not require all four.
Section B — Three receptors, three jobs
B1. Inhibition of adenylyl cyclase; opening of potassium channels (hyperpolarization); closing of voltage-gated calcium channels (reduced transmitter release).
B2. Euphoria — mu. Dysphoria — kappa. Respiratory depression — mu. Constipation — mu. Miosis — mu.
B3. Because mu itself mediates both. Selectivity between receptors cannot separate two effects that live at the same receptor. The separation would have to happen downstream of the receptor, which is the biased-agonism idea.
B4. † Dysphoria; patients decline the drug. The design change: peripheral restriction — difelikefalin is built not to cross the blood-brain barrier, so central dysphoria is avoided. Full credit requires noting the approved indication is pruritus, not pain; a common error is to assume difelikefalin is an approved analgesic.
B5. Quick: euphoria and analgesia. Slow/incomplete: constipation; respiratory depression only partially. Safety problem because the desired effect fades faster than the lethal one, so escalating exposure narrows the margin.
B6. Incorrect because it is a receptor-mediated effect, not an immune response. Consequential because an "allergy" label can be carried in the record for life and may exclude an entire drug class from future care, including in situations where alternatives are worse.
B7. † Principle: that a receptor can signal through more than one downstream route, and a ligand can preferentially engage one — so effects sharing a receptor might be separable after all. §20.3's prediction: because effects are distributed by anatomy as well as by pathway, downstream bias has a ceiling — the same receptor in the brainstem respiratory centers and in the PAG is still the same receptor. Reward answers that treat this as a real but bounded strategy rather than a failure.
Section C — Descending inhibition
C1. PAG → RVM → dorsal horn, top-down. Opioids act at the dorsal horn (and within the PAG and RVM themselves, by disinhibition).
C2. Mechanism required, not slogan: descending projections modulate transmission at the first synapse; the modulation is driven partly by input from context, emotion, and expectation; therefore the signal reaching cortex is not proportional to the peripheral input alone.
C3. Because it supplies the physical route from a belief about relief to a change in pain transmission. Without it, §20.7's result would have no anatomy.
C4. † OIH looks paradoxical only if the descending system is assumed purely inhibitory; with a facilitatory arm, sustained opioid exposure driving facilitation produces increased sensitivity without contradiction. Strong answers note this is a hypothesis with support, not a settled account.
C5. Look for: one sentence describing modulation as real physiology; one sentence explicitly denying that this means the pain is imagined, exaggerated, or the patient's fault. Mark down any answer that hedges toward "psychological."
C6. † Chain: (1) stimulating PAG produces analgesia → PAG participates in analgesia; (2) naloxone blocks opioid receptors and nothing else relevant; (3) naloxone reduces the analgesia → the analgesia requires opioid receptor activation; (4) therefore stimulation causes endogenous opioid release. Weakest link: step (2) — naloxone's specificity, and its own effects on pain sensitivity.
Section D — Evidence, measurement, and study design
D1. A peptide measured where it is easy to measure is not evidence about that peptide where it acts; plasma β-endorphin is largely pituitary and does not freely cross into the brain.
D2. Any three of: does not show central β-endorphin changed; does not show the rise caused any subjective effect; does not distinguish the intervention from a generic stress response; does not establish clinical relevance of the magnitude; does not rule out other mediators.
D3. † Look for: randomized, blinded, placebo-controlled crossover; opioid antagonist versus placebo before immersion; pre-specified primary endpoint (pain rating or threshold); a falsifying result stated in advance. Deduct for designs that measure a blood level as the primary endpoint — that is the exact error the exercise is testing.
D4. A null can be produced by inadequate power, inadequate blockade, wrong timing, or wrong endpoint. A positive result under blockade is harder to produce by accident. The determining features: sample size and power, evidence that the antagonist actually achieved receptor occupancy, and pre-registration of the endpoint.
D5. That "the placebo effect" is a family of responses rather than one mechanism, and that the mechanism recruited depends on how the expectation was installed.
D6. † If naloxone independently increases pain, apparent "reversal" could be additive rather than blockade. Design fixes: a naloxone-alone arm in non-placebo subjects; hidden administration; balanced placebo designs.
D7. Because it shows the ✅ claim is specific rather than a general enthusiasm for expectancy effects — a mechanism that predicts which effects it does and does not cover is a stronger mechanism.
Section E — Rating practice
E1. ✅ — matches chapter. E2. ⚠️ — matches chapter. E3. ❌ — matches chapter. E4. ✅ — matches chapter. E5. 🔬 — the chapter states this in prose; accept ⚠️ with a well-argued case that early human data exist, but require the student to name what data.
E6. † Rule violated: a rating attaches to a claim, with population and endpoint, not to a mechanism or a molecule — and the placebo arm of a trial engages this mechanism too, so it can never distinguish a product from its own control. Strong answers make the second point.
E7. † The claim as written is a mechanistic claim, not a clinical one, and the rating system is built for claims with a population and an endpoint. Acceptable rewrite: "Kappa-selective agonists show lower abuse liability than mu agonists in human abuse-liability studies." Award full credit for recognizing that the system itself had to be applied to reject the question's framing.
Section F — Tolerance, dependence, and language
F1. Tolerance: reduced effect at constant exposure. Physical dependence: withdrawal on cessation. Addiction: compulsive use and continued use despite harm. Enforce the no-circularity constraint strictly; it is the point of the item.
F2. Because it is a receptor-level and cellular adaptation to sustained agonist presence, and adaptation does not consult the prescription's legitimacy.
F3. Three of: undertreated pain from misread dependence; abrupt discontinuation and withdrawal; patients withholding disclosure from clinicians; stigma; obscured clinical picture for people who do have opioid use disorder. Require the student to name who is harmed in each case.
F4. † The feature doing the work is the schedule: phasic, local, seconds-long, terminated by peptidases, versus sustained whole-body occupancy. "Same receptor either way" is not a rebuttal because desensitization and downregulation are responses to sustained occupancy, not to any occupancy.
F5. If it meaningfully increases receptor activation, the adaptations follow, because they respond to occupancy, not provenance. If it does not increase activation, it does not work. The claim requires both halves and they are incompatible.
Section G — Synthesis
G1–G3. Open; mark on structure. G3 should reach: an approved drug with commercial incentive to find an easier route has not found one, so the route is evidence about the barrier.
G4. † Open. The valuable output is the student's sentence about what changed. Watch for students discovering an opposing-effects row they had not considered — that is the intended experience.
G5. † Look for: ratings attach to claims; a system supports many claims; averaging destroys the information that distinguishes them; and the four ratings here are not in tension because they concern different endpoints.
G6. Mark for tone as much as content. Any answer that is condescending toward the friend has missed the chapter's voice, and it is worth saying so.
Marking note for the whole set
Several items invite students toward a moralized reading of §20.6. Do not accept answers that frame opioid use disorder as a failure of will, character, or discipline, regardless of how well the rest of the answer is argued. The chapter is explicit on this and the marking should be too. If it comes up in more than one script, address it with the whole group rather than individually — it is a widespread cultural default rather than a personal failing on the student's part.
Chapter 21
For the instructor's answer key. Covers the 34 items in exercises.md (sections A–F, 11 marked †). The
quiz.md key is published with the quiz itself and is not repeated here except where a quiz item needs
extra marking guidance (see the note at the end).
Exercise items are open-ended by design; what follows is what a strong response contains, plus the errors that recur. † items are annotated with the reason they are hard.
Standing marking rule for this chapter: reject any answer containing invented effect sizes, p-values, sample sizes, or confidence intervals. The chapter reports every trial result directionally and deliberately supplies no numbers. A student who produces numbers has either fabricated them or imported them from an unvetted source, and either way the item is not passed.
Section A — Structure and chemistry
A1. Nine amino acids each; differ at two positions. Position 3: isoleucine (oxytocin) vs. phenylalanine (vasopressin). Position 8: leucine (oxytocin) vs. arginine (vasopressin).
A2. Cys1 and Cys6 form a disulfide bridge, closing residues 1–6 into a six-membered ring with a three-residue tail (7–9). Ch 1 §1.2 calls the disulfide bond "biology's staple."
A3. Position 3 swaps one nonpolar residue for another of different bulk. Position 8 replaces a nonpolar leucine with arginine, which carries a full positive charge at body pH — a change in kind rather than degree, on a tail residue doing receptor-contact work rather than structural work (Ch 1 §1.4). Common error: treating "two substitutions" as inherently minor. The number of changes is not the measure; what the changed residue was doing is.
A4. † Hard because the general statement must be made without the examples that make it obvious. Target: no ligand is selective in the absolute; selectivity holds within a concentration range, and outside that range binding to related receptors becomes significant. Clinical consequence: high-concentration oxytocin infusion engages vasopressin V2 receptors, producing water retention; combined with high fluid volumes this can cause hyponatremia. Listen for: whether the student's general sentence contains a concentration. If the sentence would be equally true with "concentration" deleted, they have not got it.
A5. OXTR (uterus, breast, brain — contraction, milk ejection, neuromodulation); V1a (vascular smooth muscle, brain — vasoconstriction, social/territorial behavior); V1b (anterior pituitary — ACTH release); V2 (renal collecting duct — aquaporin-2 insertion, water reabsorption).
A6. Determined oxytocin's structure and synthesized it — the first synthesis of a peptide hormone. Nobel Prize in Chemistry, 1955. Broader significance: it established that a hormone is a molecule and nothing more (synthetic material carried natural activity, with no vital essence left over), and it is the direct ancestor of the synthesis chemistry in Ch 32.
Section B — The uncontroversial physiology
B1. Production: prolactin, in alveolar secretory cells. Ejection: oxytocin, by contracting the myoepithelial cells wrapped around the alveoli. This confusion is extremely common and worth correcting explicitly rather than in passing.
B2. Myometrial OXTR density rises markedly toward term. Therefore the same circulating oxytocin concentration produces little effect at twenty weeks and a large effect at forty. The tissue decides. This is Ch 2's point that effect depends on receptor availability as much as on ligand concentration.
B3. † Hard because it requires importing Ch 3 rather than recalling Ch 21. Acceptable: receptor desensitization/internalization under sustained agonist exposure; downregulation of receptor expression; loss of the temporal information a pulsed signal carries; suppression of endogenous production by feedback. Any two. Most relevant to prolonged administration: desensitization and downregulation.
B4. Labor induction and augmentation; prevention and treatment of postpartum hemorrhage. WHO Model List of Essential Medicines.
B5. † Hard because it inverts a naming intuition students formed in Ch 1. The "-tocin" stem identifies the family — molecules acting at the oxytocin receptor system. It does not encode direction of action; atosiban antagonizes where oxytocin agonizes. The required contrast is "-relin" (agonist) vs. "-relix" (antagonist), where the stem does encode direction. Lesson: stems are family labels of variable informativeness, and you must check which kind you are holding.
Section C — Voles, trust, and replication
C1. Prairie voles form durable pair bonds with biparental care; montane and meadow voles are promiscuous with minimal paternal care. Accompanying difference: distribution and density of oxytocin and vasopressin V1a receptors across brain regions, notably reward-related structures (nucleus accumbens, ventral pallidum). Credit only answers that locate the difference in the receptors, not the peptides.
C2. A conserved signaling molecule can produce species-specific behavior through species-specific receptor placement. The peptide is not the instruction; the peptide plus the receptor map is.
C3. Any three of the five limits in case study 1: that human attachment uses the same architecture; that administering oxytocin to a human produces affiliation; that vole pair bonding and human love are the same phenomenon; that this concerns monogamy in the sense implied (prairie voles are socially but not sexually exclusive); that "the chemistry of monogamy" is a licensed description.
C4. † Hard because it requires naming all three parts of the move, not just objecting to it. Human claim: "oxytocin administration increases affiliation or trust in humans." Mechanism cited: vole receptor-distribution and manipulation work. What is wrong: mechanism establishes that an effect can occur via a pathway; it is not evidence that it does occur in a different species with a different receptor map, and the vole finding's own content is that receptor placement drives species differences — so the result argues against the extrapolation it is being used to support. Award the top mark only for answers that reach that last clause.
C5. Investor receives an initial sum and chooses how much to transfer to a trustee; the transferred amount is multiplied; the trustee chooses how much to return. Amount transferred is treated as the behavioral measure of trust. Replication: larger and preregistered attempts, including a preregistered multi-site replication published in 2020, did not reproduce the effect; critical reviews of the literature as a whole concluded the evidence for a robust prosocial effect is weak.
C6. † Hard because the ban on numbers forces conceptual rather than formulaic clarity. Target: with a small true effect and a small sample, the measurement is dominated by noise; only runs where noise happened to point toward the hypothesis clear the significance threshold; those are the runs that get published; so the published magnitude equals the true effect plus the noise required to reach significance. Hence systematic upward bias in the literature. Listen for: whether the student explains why this happens without anyone behaving badly. If their account requires dishonesty, they have described a different problem.
Section D — Context-dependence and the nickname
D1. Oxytocin amplifies the salience of social cues rather than producing affiliation, so the direction of its behavioral effect depends on the social context. (Or: the molecule supplies the gain; the situation supplies the sign.)
D2. Exaggeration overstates magnitude and can be corrected with a scaling factor. This gets the sign wrong — in identifiable contexts the effect runs opposite to what the nickname predicts. A description with the wrong sign is not an approximation of the truth; it is a different claim.
D3. Any three of: in-group vs. out-group framing of the social target; attachment style (e.g., attachment anxiety); personality or clinical profile (e.g., borderline features); competitive vs. cooperative task structure.
D4. † Hard because of the final clause. The affiliation model routes oxytocin → increased affiliation → more trust/warmth, always. The salience model routes oxytocin → amplified social cues → outcome determined by context, splitting into in-group (cooperation up) and out-group (defensiveness up). The observation the first model cannot place: increased out-group hostility, or decreased trust in specific individuals. The required precision: the first model has a missing term, not a wrong parameter value. It contains no variable for social context, so no setting of its parameters produces the second column.
D5. † General rule: a hormone named for one of its effects will be reasoned about as though that effect were its purpose; the nickname is a compression, and the errors live in the compression. Third example: cortisol, dopamine, serotonin, leptin ("the satiety hormone"), testosterone ("the aggression hormone"), insulin ("the fat-storage hormone"). Credit depends on naming what the compression loses specifically — e.g., cortisol's roles in waking, glucose regulation, and inflammation suppression — not on merely observing that the nickname is incomplete.
Section E — Delivery, trials, and vasopressin
E1. Oral is impossible (proteolysis, Ch 1 §1.3). Intravenous acts peripherally and is not expected to cross the blood-brain barrier appreciably. Intranasal was reached for as a route that might bypass the barrier.
E2. Olfactory and trigeminal nerve pathways, via perineural and perivascular spaces; bypasses rather than crosses the blood-brain barrier.
E3. Any four of: plasma levels definitely rise, so peripheral mechanisms are not excluded; CSF findings have been mixed, and CSF concentration is not receptor-site concentration in any case; assay validity (extraction vs. non-extraction methods differing by more than an order of magnitude); procedural non-standardization across labs (device, volume, head position, sniffing, timing); no direct measurement of receptor occupancy in living human brain.
E4. † For: doses far above physiological mean even a small fractional delivery could produce meaningful central concentrations — an argument for plausibility. Against: whatever happens centrally happens at exposures physiology never produces, so an observed effect cannot be read back as "this is what endogenous oxytocin does." On weight: no fixed answer required; mark for whether the student notices that the two arguments answer different questions — one about whether an effect is possible, one about what an effect would mean.
E5. † A null is ambiguous between "oxytocin does not produce this effect" and "the oxytocin never arrived" — a biological answer versus a logistical one. A positive is ambiguous between central action, peripheral action relayed centrally via afferent signaling, and expectancy in imperfectly blinded designs. The point students most often miss: an effect produced peripherally is a real effect and is still not evidence about brain oxytocin, which is what the theory concerned. Best answers note that this makes the literature hard to learn from in either direction, which is worse than being wrong.
E6. Outcome: the trial did not demonstrate benefit on its primary outcome. Framing required first: autism is a form of neurodivergence, not a disease to be cured; many autistic people object to "treatment" framing; many of the difficulties involved arise from mismatch with an environment built for non-autistic people. Consequence for outcome measures: scales scoring how closely behavior resembles a non-autistic norm are measuring conformity, and encode a value judgment that was never separately defended — so "benefit" in such a trial is not a neutral construct. Credit also for noting that some autistic people do seek help with self-identified difficulties, which is a different goal with different endpoints that the trial literature has largely not been designed around.
E7. V1a (vascular smooth muscle/brain — vasoconstriction); V1b (anterior pituitary — ACTH release); V2 (renal collecting duct — aquaporin-2 insertion, water reabsorption). Sensors: osmoreceptors respond to very small changes in plasma osmolality, continuously; baroreceptors respond only to substantial drops in volume or pressure, and override osmotic control when they fire. Priority revealed: concentration is defended finely and continuously, volume urgently and only when necessary — which is why a hemorrhaging patient retains water at the cost of diluting sodium.
Section F — Rating, dossier, and construction
F1. Rating rule 6 (one molecule, many ratings) and rule 1 (a rating attaches to a claim with a population and an endpoint, never to a molecule). Different interventions, different populations, different endpoints — intravenous oxytocin for postpartum hemorrhage in an obstetric setting versus intranasal oxytocin for social functioning in healthy adults. A student who writes "oxytocin: ✅" or "oxytocin: ❌" anywhere in the answer has failed the item regardless of the surrounding prose.
F2. † Absent = no completed human trials to evaluate. Present-and-negative = the experiment was run and the answer was no. Stronger epistemically because knowledge exists and the question has been narrowed; worse commercially because a failed trial is harder to sell around than a void. Consequence for the unapproved-compound market: it systematically favors compounds that have never been tested, and absence of negative evidence gets marketed as absence of problems.
F3. Must contain all four lines — claim with population and endpoint / rating / one-sentence reason / what would change it — plus a date stamp. Expected tier: ❌, on the grounds of meta-analytically unimpressive results plus the delivery caveat. Accept ⚠️ only if the student explicitly argues the literature is too heterogeneous to have settled the question; this is a defensible minority position and is worth using to open discussion rather than marking down.
F4. † The ⚠️ version must narrow to a population and endpoint where genuine but unsettled human data exists. The ❌ version must be specific enough to be contradicted by the existing negative literature. The unrateable version must omit a moderator that determines direction (not merely magnitude). The marking target is the explanation: whether the student can articulate what they changed and why the tier moved.
F5. Completion check plus honesty check. Both new lines (CONTEXT DEPENDENCE, METHOD CAVEAT) must be present and non-trivially filled. For the second half: no wrong answer, but look for students who reach the point that "nothing would change my mind" describes a preference rather than a rating — and note that this is the observation Ch 40 will ask them to revisit.
Note on the quiz key
The published key in quiz.md is complete. Two items need extra invigilation:
- Item 15 (name three moderators) is frequently answered with three outcomes rather than three moderators. "Trust, cooperation, generosity" is not a passing answer.
- Item 20 (evidence absent vs. present-and-negative) is the item most predictive of whether a student has absorbed the chapter's argument rather than its conclusions. Mark it generously for reasoning and strictly for the direction of the two asymmetries.
Chapter 22
For instructor use. The quiz answer key is published with the quiz; what follows is the exercise guidance, which is deliberately withheld from the student file. Many exercises have no single correct answer — for those, the notes describe what a strong response contains rather than what it says.
Part A — Co-transmission and the neuromodulator idea (§22.1)
A. Classical transmitters: small clear vesicles at the active zone, released by a single action potential. Neuropeptides: dense-core vesicles farther from the active zone, released by sustained, higher-frequency firing.
B. Because the peptide is only released above a firing-rate threshold, intensity is encoded in which message is sent, not only in how much of one message. Strong answers note that this makes the two conditions chemically distinguishable at the receiving cell.
C. (i) The signal persists longer, because termination is by diffusion and extracellular degradation rather than rapid recovery. (ii) Replenishment requires synthesis in the cell body and axonal transport, so stores are exhaustible over hours and the channel is intrinsically low-frequency.
D. † No reuptake transporter exists, so there is nothing for a reuptake inhibitor to block. A drug must instead act at the receptor (agonist or antagonist), at the peptide itself (antibody), or on the enzymes that degrade it. Strong answers spot that the third option is exactly Chapter 28's neprilysin strategy.
E. Any definition capturing "changes how a cell responds to its other inputs" without asserting direct excitation or inhibition. Watch for students who define it by size or speed instead of by function — that is the most common error and it is worth correcting explicitly.
F. † Three steps: (1) neuropeptides change the character of a response rather than its presence; (2) character changes are harder to measure than presence/absence changes; (3) harder-to-measure effects need larger, cleaner trials to detect. Best applied to the NK1 depression program, where the endpoint was a subjective rating scale in a heterogeneous population. Accept the NPY obesity program as a second-best answer with justification.
Part B — Neuropeptide Y (§22.2–22.3)
G. 36 amino acids. "Y" is the one-letter code for tyrosine, which occupies both the N-terminus and (as an amide) the C-terminus.
H. Different receptor subtypes. NPY acts at Y1 and Y5 to drive feeding; PYY(3–36) acts at Y2, a presynaptic autoreceptor that reduces NPY release. Strong answers cite Chapter 2's rule that the receptor determines the outcome, not the ligand.
I. The NPY/AgRP population of the arcuate nucleus. Leptin inhibits it.
J. † The account: sustained energy deficit elevates NPY/AgRP signaling, which persists and defends lost mass; the subjective correlate is hunger that does not resolve. What it does not license: any claim about a specific individual's behavior, willpower, or prognosis. A strong answer states the second part unprompted. This item is partly an ethics check — mark generously for care and firmly against sweeping inference.
K. No clinically meaningful weight loss; programs abandoned. Explanations: redundancy across overlapping circuits; an acute-versus-chronic gap between short-term feeding and long-term body weight; species differences in how much of the system rests on this node.
L. † The inference confuses "necessary for the drug to work" with "necessary for the physiology." NPY does drive feeding; the antagonists did engage NPY receptors. What failed was the clinical hypothesis that blocking this node would produce durable weight loss in humans. This item is a deliberate rehearsal for §22.5 — students who get L right usually get W and Y right.
M. Observational associations between circulating NPY and better performance, lower dissociative symptoms, and faster recovery in military personnel undergoing demanding survival training. Observational/correlational design.
N. (i) NPY causes resilience; (ii) resilience raises NPY; (iii) a third variable — most obviously overall sympathetic and neuroendocrine response magnitude — drives both; (iv) NPY is a marker of system effort with no causal link. Full credit requires all four; three is a pass.
O. † Magnitude: nothing consumer-available has been shown to raise central NPY measurably, and plasma is not brain. Duration: NPY is released in bursts at specific sites; a diffuse sustained elevation is a different signal (Ch 3). Existence proof: no human intervention trial, because there is no delivery route. Joint 3 fails hardest, and strong answers say so and explain that it fails for a physical rather than an evidentiary reason.
P. † Because the two are separated by the blood-brain barrier; a plasma measurement is not a readout of central signaling. Underlying section: §22.8.
Part C — Substance P and the NK1 story (§22.4–22.5)
Q. Phe-X-Gly-Leu-Met-NH2. Substance P, neurokinin A, neurokinin B.
R. Cell bodies in dorsal root ganglia; central terminals in the superficial layers of the spinal cord dorsal horn. These are small-diameter unmyelinated C fibers carrying slow, burning, poorly localized pain.
S. Any three of: localization in C fibers and superficial dorsal horn with NK1 on second-order neurons; release requiring sustained high-frequency firing; slow prolonged depolarization of dorsal horn neurons; reduced pain behavior after substance P depletion; reduced responses to intense stimuli in NK1-null animals; attenuated hypersensitivity after ablation of NK1-expressing dorsal horn neurons.
T. † The point of the annotation is the ratio. In a faithful reconstruction, most sentences are mechanism or animal evidence, and the sentences that carry the clinical claim are inference. Strong answers notice that the inference sentences are also the confident ones. Do not penalize students whose paragraph is too persuasive — that is the exercise working.
U. Acute post-surgical pain, osteoarthritis, painful diabetic neuropathy, post-herpetic neuralgia, migraine. No clinically meaningful analgesia, across compounds and companies. Bonus credit for noting that active comparators separated from placebo in several of the same trials.
V. Positive Phase II in the late 1990s reporting benefit over placebo and comparable to an SSRI comparator; widely publicized; confirmatory program with larger trials failed to replicate; other companies' NK1 antagonists also failed; psychiatric development discontinued.
W. † Eliminates: insufficient dose, failure to reach the brain, failure to engage NK1, wrong target selectivity. Leaves open: species differences in NK1 pharmacology and behavior, redundancy among tachykinin systems, mismatch between rodent behavioral assays and human depression, heterogeneity of the clinical population, and the possibility that the receptor simply is not causally central to these conditions in humans. Full credit requires at least two from each list.
X. Chemotherapy-induced nausea and vomiting; particularly effective against delayed emesis.
Y. † The distinction is between absent evidence and present-but-negative evidence. Any Part III compound with animal-only data works as the contrasting example — BPC-157 is the canonical choice. What would move it: adequately powered randomized human trials with prespecified endpoints. Strong answers add that the two ❌ ratings warrant different responses: one invites a trial, the other records a settled result.
Z. The area postrema lies outside the blood-brain barrier, so access to the emetic trigger site was never in question. Introduced in Chapter 7.
Part D — Orexin (§22.6)
AA. Two groups reported the peptides independently in 1998: "orexins" (from the Greek for appetite) and "hypocretins" (hypothalamic secretin-like peptides).
BB. Canine narcolepsy from an OX2R mutation (dogs); a cataplexy-like syndrome in orexin-null mice (mice); low or undetectable CSF orexin-A with neuron loss on postmortem examination (humans).
CC. † Parallel: autoimmune destruction; of a small specialized peptide-producing cell population; producing a well-defined clinical syndrome; with a measurable deficiency of the peptide. Divergence: insulin can be replaced because its target tissue is within the bloodstream's reach; orexin cannot, because its targets are behind the blood-brain barrier.
DD. Dual orexin receptor antagonists, for insomnia. They remove a specific wake-promoting signal rather than broadly enhancing GABA-A-mediated inhibition.
EE. † Because the drug suppresses the system whose loss causes narcolepsy, the side effects are the disease's own phenomena in miniature — sleep paralysis, hypnagogic and hypnopompic hallucinations, rarely cataplexy-like weakness. Generalization: neuropeptide drugs shift a state rather than switching a function, so their adverse effects tend to look like exaggerations of the state being shifted.
FF. † Delivery: orexin is a peptide, digested if swallowed and excluded from the brain if injected; its targets are deep in hypothalamus and brainstem. Agonism is harder than antagonism because an antagonist need only occupy the binding site, while an agonist must occupy it and induce the conformational change that activates the receptor.
Part E — CGRP, the barrier, and the comparison (§22.7–22.9)
GG. It describes the origin: the calcitonin gene, alternatively spliced. The other product is calcitonin.
HH. Observation: CGRP rises in cranial venous blood during attacks and normalizes after. Provocation: infusing CGRP triggers migraine-like attacks in migraineurs at far higher rates than in controls. Blockade: blocking CGRP or its receptor reduces migraine frequency and severity in randomized trials.
II. † Establishes: causal sufficiency in humans, in the population with the disease, with the clinical event as the outcome. Cannot establish: that the peptide is the only cause, that every attack is peptide-driven, where in the pathway it acts, or that blockade will help — the last requires separate treatment trials.
JJ. Group 1 (monoclonal antibodies, -mab): erenumab, fremanezumab, galcanezumab, eptinezumab.
Group 2 (small-molecule gepants): ubrogepant, rimegepant, atogepant, zavegepant. Neither group
consists of peptides.
KK. † Explanation should distinguish the molecule acted upon from the molecule administered. Three examples: CGRP blocked by antibodies or gepants (Ch 22); a peptide used as a delivery address for a radioactive payload (Ch 27); endogenous peptides raised by inhibiting the enzyme that degrades them (Ch 28).
LL. Structure: tight junctions between capillary endothelial cells; efflux transporters; degrading enzymes (accept pericytes and astrocyte endfeet as a fourth). Exceptions: circumventricular organs; saturable transport systems for a few specific peptides; targeting the receptor from outside the brain.
MM. † The barrier section explains three of the four stories — the NPY delivery gap, the orexin replacement gap, and CGRP's peripheral workaround — and it supplies Field 3 of the dossier. Its role in the NK1 story is precisely that it can be ruled out, which is what makes that failure diagnostic rather than ambiguous. Strong answers make that inversion explicit.
NN. † Strongest surprise argument: same cell type, same release conditions, same receptor class, same clinical domain, comparable preclinical depth. Answer: indication specificity, human provocation evidence, and target accessibility. Reward answers that concede the surprise is genuine before resolving it.
OO. † No fixed answer. Mark on whether the student separates the four items cleanly and, in particular, whether they report the "could not determine from the source" category honestly rather than guessing. That category is the assignment's real content.
PP. Completion check only. Confirm students used "not applicable" where appropriate rather than inventing a CNS claim.
QQ. † Look for specificity: the student should name measurements — a pharmacokinetic study, a central concentration measurement, a receptor-occupancy demonstration — rather than asking for "evidence." Vague answers here usually mean §22.8 has not landed.
Chapter 23
Instructor-side notes. exercises.md deliberately ships without answers; many items have no
single correct response, and several are calibration exercises whose value depends on the student not
being able to check themselves against a key. What follows is what a strong answer contains, for
grading and for discussion — not a key to distribute.
The quiz key is published in quiz.md inside the <details> block and is not duplicated here.
Grading principle for this chapter
More than any other chapter in Part IV, Chapter 23 is graded on the shape of the reasoning, not the conclusion. A student who concludes "⚠️" and a student who concludes "❌" can both be right or both be wrong. What separates them is whether they:
- attached the rating to a claim with a population and an endpoint,
- distinguished what they assessed from what they could not assess,
- named a specific study-quality problem rather than gesturing at a literature's origin, and
- stated a falsifying condition.
A student who writes a confident negative verdict on the Russian literature has committed the chapter's central error just as surely as one who writes a confident positive verdict. Say so, and say why — this is the single most common failure in the set.
Part A — Taking claims apart
A. Strong answers rewrite each phrase with a population, an endpoint, an instrument, and a timeframe, and flag that all four are absent in the original. For A.2 ("neuroprotective"), the best answers notice that the term additionally hides a comparison — protection against what, relative to what — and that it may refer to prevention, to reduction of injury severity, or to improved recovery, which are three different trials.
B. Watch for students who make every row "easy in impaired populations" — correct as far as it goes, but the interesting answers differentiate. Attention is easy to move in sleep-deprived people, hard in rested people. Learning is hard to move in anyone over short durations. Subjective clarity is easy to move in everyone, which is precisely why it is worthless as a primary endpoint.
C. The one-sentence answer: a term that names a direction of change without naming what is changed, in whom, or by what measure, cannot be contradicted by any observation. Fourth terms students commonly offer: "detox," "boosts metabolism," "supports gut health," "balances hormones," "strengthens the immune system." All acceptable.
D. † The strongest answers enumerate readings that include: studied in a rodent model, studied in humans for a different endpoint, studied in humans and found not to work, one small trial, and mentioned in a review. The collapsing question is essentially always some form of in which population, measuring what?
E. The word falsifiable must be doing work, not decorating. Look for the observation that the smaller claim tells you what result would refute it.
Part B — Semax, Selank, design
F–H. Factual. F: active fragment supplies proposed activity, Pro-Gly-Pro addresses degradation. G: proline's cyclic side chain resists exopeptidase attack; glycine's single-hydrogen side chain supplies flexibility. H: the tail addresses degradation only — not absorption, not distribution, not blood-brain-barrier penetration. Students frequently claim it helps with all four; correct this firmly, it recurs in the exam.
I. Good answers note that most Chapter 22 uncertainties produce unpredictable effects rather than directional bias — with the exception of deposition variability, which mostly adds noise and therefore biases toward the null in small trials.
J. † The two claims: (i) the molecule persists, (ii) an effect the molecule initiated persists. Different evidence: concentration measurements vs. a time-course of the outcome. The trophic mechanism makes both possible simultaneously, and — the point students miss — evidence for one is not evidence for the other.
K. Rule 3. Acceptable rewrite: "A BDNF-mediated mechanism has been proposed and has preclinical support; it describes how an effect could occur and is not evidence that it does."
L. Only (ii) is evidence about the compound. Students who say (i) have missed §23.2's clinical callout entirely.
Part C — The core section
M–N. Factual recall plus synthesis. The common element in N: both failure modes substitute a disposition for an assessment, and both are ways of avoiding an unresolved entry.
O. † This is a personal-honesty exercise and should not be graded on content. Grade on whether the student produced a specific replacement question. "I should have been more open-minded" earns nothing; "I should have asked whether the primary endpoint was pre-specified" earns full marks.
P. Answers: 1 = evidence. 2 = access. 3 = evidence. 4 = access. 5 = evidence only if you established that the report exists and omits it — otherwise access. 6 = access. Item 5 is the good one; expect argument about it, and the argument is the point.
Q. Watch for over-reach in both directions. The oxytocin literature licenses: accessible, well-funded, English-language literatures also produce non-replicating findings. It does not license: therefore all literatures are equally reliable or therefore replication failures are expected and unimportant.
R. † The strongest reconstruction usually includes: regulatory review is a real filter; decades of use bound the size of common harms; absence of Western trials is explained by patent economics; and the parochial dismissal is a genuine epistemic failure. What survives the reply: all four of those, actually. What does not survive: the inference from any of them to efficacy. Students who cannot make the strong version have not understood the chapter.
S. Two examples from elsewhere in the book: off-patent and generic compounds generally (Ch 19), and any compound with no commercial sponsor. Genuinely informative absence: when a well-funded sponsor with a clear commercial interest has had the opportunity and the incentive to run trials and has not published them.
T. † Compare against the seven steps in Case Study 23.1. Required elements: a written output, at least one refusal, and a revisit condition. Students routinely omit the refusal.
Part D — Cerebrolysin
U–V. V is the useful one: the entry cannot be completed, and the shape of the failure is the lesson. Students who invent a molecular weight for Cerebrolysin have made an instructive mistake — show the class.
W. Any ranking is acceptable if defended. Flag rankings that put "large effect overlooked" anywhere but last.
X. † The error: treating uncertainty as symmetric. Mixed evidence after many adequately conducted trials places an upper bound on the plausible effect size; it does not leave "very well" on the table. Strong answers say this explicitly.
Y. Both ratings in four-line format, plus the access-limited/conflict-limited distinction. This is the single most important item in the set and should be graded strictly.
Part E — Preclinical and labels
Z–AC. Standard. For AB †, the rhetorical pressure is that both sides read a concession as a surrender; the one-breath sentence is usually some version of "the animal work is real and the human work does not exist, and neither of those changes the other."
AD–AE. AD: ~320 Da places Noopept in small-molecule territory; oral activity, protease resistance, and plausible CNS penetration all follow. What is obscured: exactly the constraints peptides have. AE is fieldwork; grade on whether the student noticed that sometimes nothing true is lost, which is the honest finding for a genuine peptide drug.
Part F — Cognitive endpoints
AF, AH. Recall and transfer. Controls: practice → control group; expectancy → blinding; regression → control group plus recruitment design; state → standardized conditions; noise → repeated measurement and adequate n.
AG. † The full critique should conclude that the design can establish that scores rose and nothing about why. The single most consequential change is a randomized blinded control group. Estimating the inert-compound expectation: students should reason from practice effects being reliably positive plus regression from a low-point recruitment, and should say the expected direction is up without inventing a percentage. Penalize invented percentages presented as fact; reward explicitly labeled illustrative reasoning.
AI. Must reproduce sensation without the hypothesized mechanism; verification is the blinding-integrity check.
AJ. † With 78% correct allocation guessing, the blind failed and the primary result is uninterpretable as a drug effect. The "mild and comparable" side-effect profile does not rescue it — it makes the unblinding harder to explain, which is itself informative, and it does not restore the blind. Strong answers ask what else besides side effects could have signaled allocation.
AK. The placebo arm is a competing intervention through motivation. The reason "so the placebo is just as good" does not follow: the effect is a property of belief, not of the molecule, and the molecule is what is being sold at a price with risk attached.
AL. † Grade against the §23.9 box. All eight elements must appear. The closing sentence should engage §23.3 — i.e., the trial is worth running because the existing evidence is inaccessible rather than because it is presumed bad.
Part G — Dossier
AM. If a student reports zero question marks, ask what they actually read. This is usually a tell that they assessed summaries rather than sources.
AN. † Ungraded on content, graded on candor and specificity. The third category — my feelings changed — is the one the exercise exists for, and students almost never volunteer it unprompted. If the whole class reports zero instances, discuss why that is statistically implausible.
Chapter 24
Instructor-side notes for exercises.md. The quiz answer key lives in quiz.md itself and is not
duplicated here. Several exercises have more than one defensible answer; where that is the case it is
flagged. † items are the harder ones.
Part A — Receptor map and molecule
a. MC1R — melanocytes (skin, hair); activation shifts pigment production toward eumelanin, darkening skin, hair, and existing moles. MC4R — central nervous system (hypothalamus, brainstem, spinal cord); activation suppresses appetite, raises energy expenditure, and participates in sexual arousal pathways.
b. One gene → one precursor protein → several unrelated hormones, determined by which processing enzymes are present in a given tissue. The lesson: "what a gene does" is not a well-formed question without specifying the tissue. Accept any answer that gets to tissue-dependent processing.
c. Cyclization → raises potency (the receptor does not pay an entropic penalty organizing a floppy ligand) and slows degradation (proteases prefer extended accessible backbones). D-amino acid → proteases are stereospecific and were not built for D-residues. Non-standard residue (norleucine) → tunes activity and stability outside the constraints of the biological twenty; also a reminder that synthetic peptides are not limited to Chapter 1's alphabet.
d. The terminal group shifts the balance of receptor subtype activity. Melanotan II sits further toward MC1R; bremelanotide's effects are attributed principally to central MC4R/MC3R signaling. General lesson: structural similarity predicts nothing reliable about clinical profile. Response to "basically the same molecule": near-identity carries no evidentiary weight either — the difference between them is not chemical distance but the presence of trials, a defined population, and oversight.
e. † Corrected version: "α-MSH is a melanocortin peptide; in skin, where MC1R is expressed, it signals pigment production, and in the hypothalamus, where MC4R is expressed, it participates in appetite regulation." What the original gets wrong: it treats one of the peptide's effects as its identity and the other as an accident, when neither is intrinsic to the molecule — both are properties of the receptor and tissue.
f. Because a peptide-like MC4R agonist generally retains some MC1R activity, and MC1R is on melanocytes. On "side effect": defensible both ways, and the discussion is the point. For: it is unintended and unwanted relative to the indication. Against: it is a direct on-target consequence at a different subtype, entirely predicted by the receptor map, so calling it a "side" effect misdescribes the pharmacology. Best answers note that the term conflates unwanted with unexplained.
g. † Required profile: high MC4R potency, minimal MC1R potency. Why it is hard: the receptors are homologous and share a binding pocket architecture evolved for the same endogenous ligand family, so subtype-selective agonism is a genuine medicinal chemistry problem. Strong answers may also raise biased signaling (selecting downstream pathway rather than receptor) as an alternative approach.
h. ~1,025 Da. Comparable: oxytocin (~1,007 Da), octreotide (~1,019 Da), BPC-157 (~1,419 Da). What it predicts: squarely in the peptide range → digested if swallowed, so the default route is injection; too large for reliable oral absorption without an enabling strategy.
i. † Something like: an uncontrolled single observation can generate a hypothesis but cannot distinguish a causal effect from coincidence, expectation, or regression to the mean, because it has no comparison. Full credit requires that the sentence be domain-free — if it mentions peptides or sexual function, push back.
Part B — Desire and blood flow
j. Check against §24.3's table. Common error: writing "brain" versus "body" without naming the molecular targets (PDE5 enzyme; melanocortin receptors).
k. Because it acts downstream: it prevents the breakdown of a signal that arousal produces. No arousal, no signal, nothing to preserve. Good plain-language versions use a "it holds the door open, it doesn't open the door" structure.
l. Any three of: (1) the complaint is a desire problem, not a flow problem — wrong target; (2) an underlying vascular, metabolic, or hormonal condition is severe enough that the mechanism cannot overcome it; (3) a concurrent medication or condition is contributing; (4) the drug was used without the sexual stimulation the mechanism requires; (5) the person is genuinely a non-responder. Only (5) is "the drug is ineffective," and even that is population-specific.
m. † The three: wrong about mechanism (peripheral vascular versus central neural), wrong about the target complaint (response problem versus desire problem), wrong about magnitude (imports an expectation of a large obvious effect). Ranking is open; the chapter argues the third does the most harm because mis-set expectations in this domain are interpreted by patients as personal failure. Accept a well-defended alternative ranking — several students will argue the second error is worse because it produces wrong prescriptions rather than wrong feelings.
n. Because inadequate genital blood flow is frequently the first visible sign of vascular disease elsewhere. Prescribing symptomatically without evaluation can mean treating the presenting complaint and missing an early cardiovascular warning — harm by omission, not by the drug.
o. † Desire → central/psychological/hormonal. Arousal → vascular/hormonal/mechanical. Orgasm → central/neurological/pharmacological (medication effects are common here). Pain → gynecological, dermatological, musculoskeletal, hormonal, psychological. The chapter's peptides address desire (bremelanotide) and, prospectively and at early stage, desire-adjacent processing (kisspeptin).
p. Because removing or changing a contributing medication can improve desire on its own, so any uncontrolled observation of improvement after starting a new drug may be confounded by concurrent changes — and because the diagnostic criteria explicitly exclude medication-attributable low desire, which means uncontrolled settings routinely include patients the trials excluded.
Part C — Approval and effect size
q. Acquired, generalized hypoactive sexual desire disorder in premenopausal women. Exclusions: postmenopausal women; men; low desire attributable to a coexisting medical or psychiatric condition, to relationship problems, or to the effects of a medication or substance. Headlines omit essentially all of it — most reliably the hormonal-status restriction and the "acquired, generalized" qualification.
r. Two of: event frequency measures circumstance and partner availability at least as much as physiology; a person can have frequent sex with no desire, or high desire with no partner; the behavior is influenced by a partner who is not randomized; the diagnosis is defined by desire and distress, so an event count is not measuring the construct.
s. Favorable: the drug targets desire, not behavior; behavior is downstream of many non-drug factors and would need a longer or larger trial to move; instrument endpoints are the prespecified primaries and they were met. Unfavorable: if desire genuinely increased, behavior is the natural consequence, and its absence suggests the instrument change did not correspond to anything the person acted on. Helpful additional information: MCID estimates for the instruments used; longer-term follow-up; qualitative patient-reported outcomes; whether the trial was powered for the behavioral endpoint at all.
t. † Statistical significance answers "is this difference likely to be real?" MCID answers "is this difference big enough that a patient would notice and care?" A large trial can detect a real difference that is smaller than anyone would notice — significance scales with sample size, meaningful magnitude does not.
u. Daily drug: continuous exposure, adaptation over the first days, and a nausea episode does not coincide with the drug's purpose. As-needed drug: each administration is a fresh decision, the effect lands in exactly the window the drug was taken to improve, and infrequent use means no adaptation curve to wait out.
v. † Open. Strong examples: drowsiness from an antihistamine is trivial at bedtime and serious in a drug taken before driving; a mild tremor is irrelevant in most people and disqualifying in a surgeon or musician; transient dizziness matters far more in an elderly patient at fall risk. Reject examples that are really about severity rather than about context of use.
w. Check against §24.4. Most commonly omitted from "what would change it": the discontinuation / durability point.
x. † Best available ⚠️ argument: the effect sits near or below plausible MCID estimates; the behavioral endpoint did not move; real-world persistence is modest; and a modest instrument change against a large placebo response could reflect measurement rather than benefit. Which rule this violates: arguably none. It is a genuine disagreement about whether the endpoint evidence supports the claim, made on evidential grounds. That is the correct conclusion and the point of the exercise — the rating system is not supposed to make disagreement impossible, only to make it about the right thing. Push back hard on any student who argues ⚠️ on the grounds that the diagnosis is dubious; that is rule 4.
Part D — HSDD
y. The condition must cause marked distress or interpersonal difficulty. Without it, the diagnosis would capture everyone at the low end of a normal distribution regardless of whether they experience their desire as a problem.
z. † / aa. † Assess on quality of steelmanning, not on which side the student picks. Marking
signal: a critics' case that does not mention the distress criterion is incomplete; a proponents' case
that does not mention the exclusions is incomplete. See case-study-02.md for full versions.
ab. Parallel: sleep need varies enormously, complaints are context-sensitive, the category has been commercially promoted, and placebo response is large — yet nobody suggests turning away distressed insomniacs. Where it breaks down (any one): sleep deprivation has objective physiological consequences and measurable correlates that low desire does not; insomnia has objective measurement (actigraphy, polysomnography); and sleep is not primarily interpersonal, so the "problem located in the wrong person" critique does not transfer.
ac. † Disease → responsibility with the body/clinician; help is pharmacological or medical; risk is overtreatment and misattribution. Behavior → responsibility with the individual; help is effort, therapy, or lifestyle change; risk is blame. Environment → responsibility with the relationship, the workload, the medication list, the circumstances; help is contextual change; risk is that a person with no changeable context is left with nothing. Best answers note the frames are not mutually exclusive and that the criteria's exclusions are effectively a mandated environment-first search.
ad. Such a reader must still concede that the trials were run in people meeting the stated criteria, were randomized and placebo-controlled, used validated instruments, and were positive on their primaries — that is the ✅. What they are entitled to argue instead: that the category should not exist, that prevalence has been inflated, that the effect is not clinically meaningful (an evidential argument), and that resources are misallocated. The rule forbids moving the rating on category-skepticism; it forbids nothing else.
Part E — Afamelanotide, melanotan II, kisspeptin
ae. Rare inherited disorder of heme synthesis (usually reduced ferrochelatase activity) → protoporphyrin IX accumulates in erythrocytes and skin → the molecule absorbs visible light and generates reactive species, producing burning pain within minutes, escalating over hours, lasting days. Sunscreen: conventional products are designed against ultraviolet, and the offending wavelengths are visible.
af. Because the narrower the claim, the more likely the trial actually tested it. A broad claim usually contains inference; a narrow one usually contains only what was measured. Also: it makes the rating falsifiable in a specific way.
ag. † The four: indication (severe painful disease versus cosmetic change in well people — different acceptable risk); supervision (specialist centers with built-in skin examination versus no monitoring loop); product (regulated identity, purity, and accountability versus an unverified vial); evidence base (randomized trials plus post-marketing surveillance versus neither). Hardest to fix for unsupervised cosmetic use: indication — the risk-benefit denominator is fixed by the fact that the user is well, and no amount of product quality or monitoring changes it. Accept "supervision" with a good argument.
ah. They establish a signal warranting concern. They do not establish a quantified risk. Insist on both phrases.
ai. † The asymmetry follows from what each inference requires. Efficacy requires a counterfactual — what would have happened otherwise — which an uncontrolled series cannot supply, and reporting is selected for success. Harm detection requires only noticing an unusual, distinctive event in temporal association with an exposure, which is exactly what a clinician writing up an odd case supplies; and the distinctiveness substitutes for the missing control, because the background rate of the event is low.
aj. Standard ❌ = a statement about absence of supporting evidence. This one = absence of evidence plus an affirmative harm signal, which changes the practical posture from "wait for data" to "there is a documented concern that would need to be affirmatively resolved." Rule 2 is not violated: ❌ still describes the evidence, and here the evidence includes harm reports.
ak. † Because necessity is not efficacy. Knowing a molecule is required for a physiological process tells you the process depends on it; it does not tell you that administering it to a patient with a clinical complaint produces benefit. That is rule 3 exactly — mechanism, however elegant, does not upgrade a rating.
al. Consequence: continuous GnRH receptor stimulation suppresses the axis, which is the basis of
the -relin class used in prostate cancer and elsewhere. Implication: a kisspeptin-based therapy that
must reproduce a rhythm faces a delivery problem that flat exposure cannot solve, so long-acting
formulations — the usual peptide fix — may be exactly wrong.
am. † In common: a real, well-mapped receptor and coherent physiology. Separating feature: one is being developed in the open with registered trials, published results, and no product for sale; the other is being sold. The mechanism is common to both; the evidence is not.
Part F — Synthesis
an. † Check against §24.8. The per-row explanation should note in each case that changing the population or the endpoint changes the rating without changing the molecule.
ao. Look for: ✅ means adequately designed trials in a defined population supported the claim and a regulator agreed; it says nothing about how large the change is; bremelanotide's effect was small on questionnaires and did not move the behavioral endpoint. Marking constraint: the phrase "statistically significant" is banned precisely because students lean on it as a synonym for "works."
ap. † Assess on whether the student wrote out the exclusions and produced an honest gap line. The most common failure is writing "approved: yes" for a compound available from a clinic — that is availability, not approval. The second most common is treating off-label prescribing as evidence.
Chapter 25
For instructor use. Covers the exercises in exercises.md, the in-chapter 🔍 Check Your
Understanding prompts, and the Spaced Review. The quiz answer key is published in quiz.md and is
not duplicated here.
Items marked † in the exercises are open-ended; what follows is a marking rubric rather than an answer.
In-chapter Check Your Understanding (§25.5)
1. Identical MICs mean identical antibacterial potency; the selectivity indices differ sixteen-fold, so only the first has a plausible systemic future. A headline about "a powerful new bacteria-killing peptide" would almost certainly omit the selectivity index, because the impressive number is the MIC and the disqualifying number is the ratio.
2. The neutrophil uses compartmentalization: defensins are released into a sealed phagosome containing the bacterium and essentially nothing of the host cell, so a lethal local concentration is achieved in a few cubic micrometers without exposing host membranes. None of the body's advantages — compartmentalization, locality, induced regulation, cooperation with other effectors — is available to an intravenous drug, which goes everywhere at once for the duration of therapy. Full credit requires the student to say none, not "some."
3. Acceptable answers describe the consequences rather than naming the concept: a peptide on the skin reaches a small, defined tissue volume; very little of it crosses into the body; and the worst likely local effect is irritation rather than organ injury. An answer that merely paraphrases "lower bar" without mechanism does not earn full credit.
Part 1 — Getting the facts straight
A. 1.27 million attributable, 2019. Associated: about 4.95 million. Attributable is the modeled excess caused by resistance (the counterfactual: susceptible rather than resistant organism); associated counts all deaths in which resistant infection was present in the causal chain.
B. It is not false — it is the associated figure, which is a real quantity. It is misleading if presented as the death toll caused by resistance. The best answers say the number is right and the label is wrong.
C. Short (12–50 residues): places them in the peptide size range, synthesizable, rapidly cleared. Cationic: supplies electrostatic targeting to the anionic bacterial surface. Amphipathic: permits insertion into the bilayer while the charged face remains in water.
D. α-defensins — neutrophil granules and intestinal Paneth cells. β-defensins — epithelial surfaces (skin, airway, gut, urogenital). The only human cathelicidin is LL-37.
E. Xenopus laevis, the African clawed frog; 1987; the observation that frogs with surgical wounds in non-sterile water did not develop infections.
F. MIC — lowest concentration preventing visible bacterial growth under defined conditions. Hemolysis — lysis of red blood cells, the standard readout of mammalian membrane toxicity. Selectivity index — ratio of host-harming to bacteria-killing concentration.
G. Any four of: colistin (systemic, last-line; also topical), daptomycin (systemic), vancomycin (systemic), polymyxin B (both), bacitracin (topical), gramicidin (topical), nisin (neither — food preservative). Accept teicoplanin/dalbavancin/oritavancin/telavancin (systemic).
H. A family of plasmid-borne genes encoding phosphoethanolamine transferases that modify lipid A, reducing outer-membrane negative charge and weakening colistin's electrostatic attraction. "Plasmid" matters because it makes the resistance horizontally transferable between organisms, including across species, rather than confined to a single lineage.
I. 2003, for complicated skin and skin structure infections (2006 for S. aureus bacteremia including right-sided endocarditis — accept either or both). Not used for pneumonia, because pulmonary surfactant binds and inactivates it.
Part 2 — Mechanism and reasoning
J. A target-specific antibiotic depends on precise contact with one protein or nucleic acid site; one substituted amino acid can abolish binding while the target keeps functioning. A membrane-disrupting peptide attacks a bulk physical structure built by many genes and constrained biophysically, so resistance requires coordinated change rather than a point mutation. The strongest objection: resistance nonetheless occurs — charge modification, efflux, proteolysis, shielding — and mcr shows it can be plasmid-borne and mobile. Students who omit the objection have answered half the question.
K. Acceptable observations include: pore size uniformity and ion selectivity (barrel-stave); evidence of lipid headgroups lining the pore, or induced membrane curvature and lipid flip-flop (toroidal); concentration-threshold behavior with wholesale vesicle solubilization and no discrete conductance events (carpet). Credit reasoning over recall of specific techniques.
L. † Rubric. Must cover: targeting (receptor vs. electrostatic), concentration (picomolar– nanomolar vs. micromolar), amplification (cascade vs. none), specificity (structural vs. quantitative preference), reversibility (binding equilibrium vs. lethal structural damage). Strongest answers identify the absence of amplification and the quantitative nature of the preference as jointly responsible for the selectivity problem: many molecules must be present per target cell, and nothing prevents them from acting on host membranes at that concentration. Accept a well-argued alternative choice.
M. Predictions: it will accumulate preferentially at negatively charged bacterial surfaces; it will insert into lipid bilayers and permeabilize them; it will likely fold on contact with a membrane rather than in solution. Likely liability: hemolysis / poor selectivity index; also acceptable — salt and serum sensitivity.
N. More negatively charged than the outer leaflet of a mammalian plasma membrane. The leaflet matters because mammalian anionic phospholipids are largely sequestered on the inner (cytoplasmic) leaflet, which a peptide arriving from outside never encounters. A student who says "mammalian cells have no anionic phospholipids" has the right conclusion by the wrong route — mark down.
O. Cholesterol increases bilayer ordering and rigidity, making insertion less thermodynamically favorable. Depleting cholesterol should make the mammalian membrane more susceptible, lowering the concentration at which host damage occurs and therefore lowering the measured selectivity index. Watch for students who invert the direction.
P. † Rubric. Good questions include: what is the selectivity index / hemolysis result? What were the assay conditions — cation content, serum, salt, inoculum? Was activity retained in physiological salt and serum? Is there any in vivo data? Disappointing answers imply, respectively: no systemic future; the MIC may not be comparable to published values; potency will not survive the body; the result is a screening hit, not a candidate.
Q. It solves the selectivity problem by containment — a lethal concentration in a sealed vesicle with no host membranes exposed. Other strategies: epithelial locality (acting on the outward face of a barrier, in mucus); inducible, terminable expression; cooperation with other immune effectors so AMPs never act alone.
R. Serum protein binding sequesters cationic peptides and lowers free concentration; physiological salt screens the electrostatic attraction the mechanism depends on; host and bacterial proteases degrade the peptide. Each raises the concentration required in vivo while the tolerated concentration is unchanged — the margin closes from one side.
Part 3 — Evidence ratings
S. Model reformulations:
- "LL-37 kills [named organisms] in culture" — ✅ for the in vitro claim; the human-outcome claim is a different, unrated statement. Best answers notice the item as written has no population or endpoint at all.
- "Colistin causes nephrotoxicity in a substantial proportion of patients receiving it for MDR Gram-negative infection" — ✅. Note this coexists with the efficacy ✅; both are supported.
- Not ratable as written. Requires splitting into systemic (🔬) and topical (⚠️) claims.
- "Resistance to teixobactin was not readily generated in laboratory experiments" — supported. "Teixobactin is resistance-proof" — ❌. "Teixobactin will treat human infection" — 🔬, no completed human trials.
T. ✅ asserts that the evidence supports the claim as stated. It does not assert tolerability, pleasantness, or first-line status. Colistin's toxicity is real, documented, and the reason it was abandoned once; the ✅ attaches to the efficacy claim in a defined last-line population.
U. Rule 1 (a rating attaches to a claim, not a molecule) and rule 6 (one molecule, many ratings). The ❌ addresses an assertion about resistance that is contradicted by direct observation; the 🔬 addresses a therapeutic claim on which no verdict exists. Different claims, different evidence states.
V. † Rubric. Rating should be 🔬 (accept a defended ⚠️ only if the student explicitly cites human randomized data, which the chapter does not supply for the inhaled formulation — the intravenous Phase 3 halt is not evidence for the inhaled claim). Must include: a specific safety endpoint (renal function is the obvious one given the intravenous history); both directions; and recognition that the inhaled route is itself a selectivity strategy.
W. The comparator. Pexiganan's 2016 trials failed in part because response rates were high in the comparator arm, leaving no separation to detect. Any wound-healing claim must beat good standard care, and good standard care for wounds is effective.
X. † Rubric. Each rewrite must be specific (a study design and population, not "more research"), dated, and bidirectional. Credit the reflective step: students usually find the harder case is the one where the source had no underlying position to make specific — which is the point of the exercise.
Y. Model answers: ✅ misused when an approval is treated as covering an off-label claim; ⚠️ misused when preclinical-only evidence is dressed up as "early human data"; ❌ misused as a verdict on a molecule's potential rather than on a claim; 🔬 misused to shelter a confidently marketed consumer product from a deserved ❌.
Part 4 — Economics and structure
Z. Chronic therapy: rising curve, sustained plateau across years of daily use. Novel antibiotic: low, flat, intermittent. The second curve is evidence of a functioning clinical system because stewardship restriction — the thing flattening it — is correct practice that preserves the drug's value.
AA. Cost of goods must be recovered across the units sold over the product's life. A chronic therapy sells many units per patient over decades; an antibiotic sells a handful of units per patient, to few patients, under deliberate restriction. The denominator, not the numerator, is the problem.
AB. † Rubric. Credit is for the two objections and the responses, not for the choice. Strong objections: subscription pricing requires someone to set a price with no market signal; market entry rewards may fund a "me-too" agent that clears a low novelty bar; transferable exclusivity taxes patients taking an unrelated drug; push funding does not solve the post-approval commercial problem. Penalize answers that fail to engage seriously with their own objections.
AC. Approval establishes that the drug works; revenue depends on units sold. Stewardship — the feature of good clinical practice at issue — deliberately minimizes units sold for exactly the drugs that most deserve approval.
AD. † Rubric. Would change: the cost-of-goods obstacle specifically; possibly the viability of topical and veterinary applications; possibly the economics of stockpiling. Would not change: the selectivity problem (§25.5), the market failure (§25.6), the regulatory evidence bar, or the fact that a reserved drug sells few units. Strongest answers explicitly map each claim to the obstacle it addresses and note that cheap manufacture of a nephrotoxic peptide accomplishes nothing.
Part 5 — Reading and writing
AE. † Rubric. Look for: identification of the strong resistance claim; identification of an omitted selectivity discussion; identification of "peptide antibiotics as future technology" framing; and — importantly — fair credit where the piece is accurate. An annotation that finds only faults is not a good annotation.
AF. † Rubric. The constraint is that neither paragraph may contain a false statement. The techniques students should surface in the third paragraph include: selective omission, choice of which number to quote, mechanism-as-evidence, appeal to nature, framing absence of approval as either "not yet" or "never," and comparator choice.
AG. Model: "Peptide antibiotics are approved and in daily use; what has not yet succeeded is the narrower project of turning innate immunity's host defense peptides into new broad-spectrum systemic antibiotics." The original errs by treating a whole compositional category as speculative when several of its members are on hospital formularies; the error is common because popular coverage uses "antimicrobial peptide" to mean specifically the host-defense-peptide research program.
AH. † Rubric. Marking is on the audit, not the compound. The rewritten entry must contain all five components; the reflective sentence is where most of the learning is, and the commonest missing component is the bidirectional resolving observation.
Spaced Review
1. Errors: (i) five million is the associated figure, not the attributable one — the correct statement is roughly 1.27 million directly attributable in 2019; (ii) the resistance claim is false as stated. Defensible version: resistance to membrane-active peptides appears harder to evolve and costlier to maintain than resistance to target-specific antibiotics.
2. (Ch 2 link.) Peptide hormones act through receptors with signal amplification — one bound molecule produces thousands of downstream events — so vanishingly small concentrations suffice. AMPs have no receptor and no amplification; killing requires many peptide molecules per target cell and enough local density to destabilize a bilayer. The link to §25.5: the discrimination between bacterial and host membranes is a quantitative preference, and the concentrations required to kill are high enough that the preference can be overwhelmed.
3. (Ch 5 link.) Both correct under rule 1 (ratings attach to claims) and rule 6 (one molecule, many ratings). Colistin's ✅ attaches to an efficacy claim in a last-line population and is supported regardless of tolerability; the ❌ attaches to a resistance claim contradicted by observation, and does not disparage the molecules. One sentence: ✅ asserts that the evidence supports the claim as stated; it does not assert that the drug is safe, tolerable, or preferred.
4. (Ch 4 link.) Topical: small defined tissue volume, minimal systemic absorption of a large charged molecule, and local irritation rather than organ injury as the risk — so the selectivity bar is much lower and products have reached market. Systemic: whole-body exposure at concentrations set by the least accessible infection site, with no compartmentalization available. Daptomycin and pneumonia: the drug reaches the lung but is bound and inactivated by pulmonary surfactant — an site-of-action failure, not a distribution failure.
5. Rubric. Must include: 🔬; a date; one specific obstacle (selectivity margin, or delivery of an intact peptide through cystic fibrosis airway mucus and to a biofilm, or degradation by the high protease burden of CF sputum — all acceptable); and a bidirectional resolving observation. The final sentence should note that inhaled delivery improves the selectivity argument, by concentrating the peptide at the infection site and limiting systemic exposure — while introducing new local tolerability questions and the sputum/biofilm inactivation problem.
Chapter 26
Internal file. Not published with the chapter. The student-facing exercises.md deliberately
carries no answers; these are for instructors and for the answer-key appendix build.
Many items are open-ended. Where that is true, the entry below gives what a strong answer contains rather than a single correct response, plus the most common wrong turn.
Set A — The mechanism
A1. Antigen = anything the adaptive system can recognize, usually a whole protein. Epitope = the specific part a receptor contacts. Immunogen = an antigen that in practice provokes a response. Key discrimination: every immunogen is an antigen, not every antigen is an immunogen, and an epitope is a part rather than a kind. Strong answers note that the adjuvant is often what converts an antigen into an immunogen.
A2. A T cell has no receptor for whole objects. Accurate replacement: T cells recognize short peptide fragments derived from viral proteins and displayed on MHC/HLA molecules at the surface of infected cells. Common wrong turn: substituting "the immune system recognizes viral proteins," which is still wrong for T cells — it must be fragments, and it must be presented.
A3. Class I groove is closed at both ends, pinning peptide length to roughly 8–10; class II groove is open at both ends, permitting 13–25 residues with overhang. Implication: class II can accommodate longer and more variable peptides, and class I imposes tighter length constraints on epitope prediction.
A4. Expected steps: (1) protein synthesized in cytosol; (2) a fraction degraded by the proteasome; (3) fragments transported into the endoplasmic reticulum; (4) loading into the class I groove; (5) trafficking of the loaded complex to the cell surface; (6) surveillance by CD8 T cells. Four of these suffice.
A5. Anchor residues are the side chains that fit pockets in the HLA groove — typically near position 2 and at the C-terminus — and determine whether the peptide is held at all. Two peptides with identical middles but different anchors may differ in whether they are presented; a peptide that is never presented produces no response at all, which is categorically different from a weak one.
A6. Practical consequence: short peptides are a natural vehicle for T-cell epitopes and a poor one for antibody-directed vaccines, because antibodies read folded surfaces. A peptide vaccine intended to raise neutralizing antibodies frequently raises antibodies that bind the peptide well and the pathogen badly.
A7. Acceptable answers describe the expansion of specific lymphocyte clones and a persisting, larger, faster-responding population positioned to act on re-encounter. Reject answers that reduce to "antibodies in the blood," which omits the T-cell arm and the cellular basis.
A8. † Strongest objection: for vaccines whose protection is antibody-mediated, the operative recognition event is conformational, so calling them "peptide-delivery systems" describes a real but non-load-bearing component. Good answers concede this and narrow the claim to T-cell-directed immunity, which is how the chapter actually states it. Reward students who notice that the chapter's own §26.2 supplies the counterexample.
Set B — HLA and the personalization problem
B1. HLA restriction = a T cell recognizes its peptide only in the context of a particular HLA molecule. Illustration: the same peptide presented by one person's HLA and not displayed at all by another's.
B2. Three explanations: (i) the second participant's HLA cannot present the peptide; (ii) they have no T-cell clone specific for that peptide-HLA combination, or it was tolerized; (iii) the response occurred but was not detected by the assay, or occurred in tissue rather than blood. Distinguishing data: HLA typing; a broader assay panel; tissue sampling.
B3. (a) Internal validity is helped — a more homogeneous population reduces variance and the intervention is matched to the biology. (b) External validity is damaged — the result applies to carriers of that allele, and possibly only to them.
B4. † Defense: restriction was necessary given the state of prediction, reference data, and manufacturing; enrolling patients who cannot present the target would guarantee dilution and an uninterpretable trial. What to do differently now: individualized selection against each patient's own HLA; deliberate expansion of reference data across ancestries; pre-specified reporting of results by HLA background.
B5. A whole protein or organism supplies dozens of candidate peptides and lets each recipient's own grooves select whichever fit. A short-peptide vaccine has made that selection in advance, at the factory, for everyone.
B6. Filters between prediction and response: actual generation by cellular processing; adequate surface density; existence of a non-tolerized T-cell clone; activation rather than tolerization. Strong answers note that binding prediction is good while immunogenicity prediction is substantially weaker.
B7. † Losing a single target protein defeats one epitope; losing an HLA haplotype defeats every epitope presented by that haplotype at once — including epitopes the vaccine has not yet used. It degrades the entire presentation channel rather than a single message.
Set C — Adjuvants and the danger signal
C1. Small; rapidly cleared by proteases; carries no danger signal. Any one of these alone is incomplete; all three should appear.
C2. With costimulation: activation, expansion, memory. Without: anergy, deletion, or conversion to a regulatory phenotype.
C3. Mechanism: presentation without costimulation actively tolerizes. Vaccinating with antigen in a non-inflammatory context can therefore induce unresponsiveness to that antigen — which is worse than having done nothing, because the target is now harder to attack.
C4. Antibody-directed prophylaxis needs class II help and B-cell responses, which aluminum salts support. Cancer vaccines need cytotoxic CD8 responses, which aluminum salts historically generate poorly.
C5. † It is preclinical because it was demonstrated in animal models; establishing it in humans would require correlating injection-site persistence with T-cell trafficking and clinical outcome in a clinical trial, which is difficult and rarely a primary objective. Reasonable weight: sufficient to avoid long-lived depot adjuvants absent a specific reason, not sufficient to attribute past clinical failures to that mechanism alone.
C6. For a colleague: local innate activation, consistent with the adjuvant doing its job; reactogenicity and immunogenicity are linked. For a friend: that soreness is your immune system paying attention, which is the part that makes the vaccine work. Both should avoid promising that the absence of symptoms means failure — it does not.
C7. Prophylactic recipients are healthy, numerous, and receive a benefit that is probabilistic and invisible; therapeutic recipients face a disease with its own risks and a very different alternative. Strong answers frame this as who bears the risk relative to what they stand to gain.
Set D — Neoantigens and individualized vaccines
D1. A neoantigen is a peptide carrying a mutation-derived amino acid change absent from the normal proteome. Because it was never present during development, no central tolerance was established against it, so the high-affinity repertoire remains intact — unlike a tumor-associated self antigen.
D2. Expected: tissue → sequencing (tumor + matched normal, plus RNA) → somatic variant calling → HLA typing → epitope prediction and ranking → selection → manufacture (peptides or mRNA) → administration with adjuvant. Most-likely-failure answers are defensible at several steps; reward justification, not a particular choice. The chapter's own emphasis is on prediction of immunogenicity.
D3. Without matched normal tissue, inherited variants cannot be distinguished from somatic mutations. Targeting an inherited variant means vaccinating a person against a sequence present in every cell they own — no tumor specificity, and a tolerance and safety problem.
D4. Minimal epitopes load directly onto class I on any cell, including cells that cannot costimulate, producing tolerizing presentation (Set C). Long peptides require uptake and processing by professional presenting cells, and additionally supply class II epitopes and CD4 help.
D5. Clonal = present in every tumor cell; cannot be escaped by losing a subpopulation. Subclonal = present in some; targeting it selects for the cells that lack it.
D6. † Advantages of shared targets: no individualized manufacturing, far lower cost, faster deployment, conventional trial design, easier regulatory path. Given up: applies only to patients with both the mutation and a compatible HLA allele; fewer targets per patient; a hotspot mutation in a driver gene may be under selection to be retained but is still one target. Either funding answer is acceptable with justification.
D7. Inference: if response to checkpoint blockade tracks mutation count, the responses being released are plausibly directed at mutation-derived antigens. What it does not establish: that neoantigens are the causal target in any individual patient, that vaccinating against them adds anything, or that mutational burden predicts benefit at the individual level.
D8. † Both positions are defensible. Structural case: few mutations means few candidate targets, and no engineering makes a mutation that is not there. Temporary case: prediction and epitope discovery are improving, non-mutational sources of tumor-specific antigen exist, and low-burden tumors may still carry a small number of excellent targets. Reward students who distinguish "fewer targets" from "no targets."
Set E — Evidence, ratings, and reading the literature
E1. Four lines: claim (with population and endpoint), rating, one-sentence reason, what would change it. Jobs: scope the claim; state the verdict; make the verdict inspectable; make it falsifiable and date-stamped.
E2. Rule 3 — never upgrade with mechanism. The ✅ concerns how recognition works; the 🔬 concerns whether a specific intervention helps patients. The first is a premise of the second, not evidence for it.
E3. ⚠️ requires real randomized human data that does not settle the question — which peptide allergy immunotherapy has, including large late-stage trials. 🔬 is early-stage science proceeding properly where translation is unproven; the neoantigen randomized evidence is younger and narrower and confirmatory trials are still running.
E4. † Expect ✅, on the grounds of large randomized trials and population-level surveillance showing reduced infection, precancerous lesions, and cancer incidence. The reconciliation: the HPV claim is prophylactic, in healthy people, against an infectious cause — a different and far more tractable problem than treating an established tumor (§26.8). Students who rate it ⚠️ because "cancer endpoints take decades" should be pushed on what the completed evidence actually shows.
E5. Expected critique: no comparator, so no efficacy conclusion; the immune result and the clinical result are separate findings and their juxtaposition implies a relationship the design cannot support; 15 patients with one year of follow-up cannot characterize recurrence in most solid tumors.
E6. † Reasonable question list: randomized or not; pre-specified primary analysis or exploratory; what both arms received; endpoint and its definition; effect size with confidence interval; follow-up duration and event count; tumor type and disease setting; how patients were selected. Ranking should put pre-specified primary randomized comparison and what the comparator arm received at the top.
E7. Attribution problem: the checkpoint inhibitor is effective on its own, so improvement in a combination arm cannot be assigned to the vaccine without a comparison. Solution: randomize with the checkpoint inhibitor in both arms and the vaccine in only one.
E8. † Case for greater generalizability: the pipeline adapts to each patient rather than assuming a shared target, so it may transfer across HLA backgrounds and tumor types better than a fixed epitope-restricted product. Case against (the chapter's): what gets validated is a process whose performance depends on inputs — mutation count, prediction accuracy for that HLA background, manufacturing quality — that vary in ways a fixed molecule's do not. Either conclusion is acceptable with reasoning.
Set F — Communication and the dossier
F1. Grade against the template in the chapter's Dossier section. Deduct if the immune readout and the clinical endpoint have been merged into one line.
F2. Must contain: the vaccine is not the thing that acts; it teaches; what persists is a changed population of cells; the effect is triggered later by something the vaccine never touched.
F3. † Grade on tone as much as content. A good reply: acknowledges the real progress; states that the trials are randomized and ongoing; states plainly that nothing is approved and access is through trials; does not lecture; does not use the word "hype" at a person whose relative is in treatment; points to the oncologist. Under 150 words.
F4. Two sentences, e.g.: "There are individualized cancer vaccines in randomized clinical trials right now, including in combination with immunotherapy, and early results have been encouraging enough that larger trials are running. None is approved, so the way to access one is through a trial — I can look into whether any of them fit your situation." Grade for accuracy in both directions.
F5. † Grade the rewrite: two separate sentences, no causal connective, and explicit labeling of which is the immune measurement.
F6. Must include: hepatitis B and HPV vaccination prevent cancers by preventing the causal chronic infection; they are prophylactic, given to healthy people; neither is a peptide vaccine. Why it is forgotten: they are filed mentally under "infectious disease," and their success is invisible because it consists of cases that did not happen.
Quiz answer key
The student-facing key is embedded in quiz.md inside a <details> block. Correct letters, for
grading convenience:
1 B · 2 B · 3 B · 4 B · 5 C · 6 B · 7 C · 8 B · 9 B · 10 B · 11 B · 12 C · 13 B · 14 C · 15 B · 16 B · 17 C · 18 B · 19 C · 20 B · 21 B · 22 C
Distractor notes. Item 8 option A ("weaker immune system") is the intuitive and wrong answer and is worth discussing aloud. Item 15 option C ("permanently alters the genome") is false and should be named as false rather than merely marked wrong. Item 19 is the single best discriminator in the set: students who miss it have not internalized the responder/non-responder confound. Item 22 requires eliminating two options that are false on their face (the mechanism is proven; human data do exist), which is the point.
Chapter 27
For the instructor answer key. Quiz answers are already published in the chapter's quiz.md under a
collapsed <details> block; this file covers the exercises (which are published without answers)
and adds marking guidance.
Exercise items are lettered A–AF. Items marked † in the student file are challenge items and should be graded on reasoning quality, not on reaching a particular conclusion.
Part 1 — Recall and comprehension
A. GnRH agonist: delivers continuous receptor stimulation, causing desensitization and sustained suppression of the axis. GnRH antagonist: blocks the receptor directly, suppressing immediately without a surge. Somatostatin analog: supplies a stabilized version of the body's inhibitory hormone, suppressing hormone secretion and, in NETs, slowing progression. PRRT construct: carries a radionuclide to receptor-expressing tumor cells.
B. Leuprolide (mid-1980s) → octreotide (1988) → degarelix (2008) → lutetium Lu 177 dotatate (2017 EU / 2018 US). The forty-year span is the point of the question.
C. The serum testosterone threshold defining successful androgen deprivation; named because surgical castration was the original way of achieving it. Accept any answer noting the historical origin.
D. Targeting peptide (address) / linker (spacer, controls pharmacokinetics and presentation) / chelator, e.g. DOTA (holds the metal ion) / radionuclide (the therapeutic agent).
E. Flushing, secretory diarrhea, sometimes wheezing, and over years right-sided valvular fibrosis. Caused by serotonin and other vasoactive substances released into the systemic circulation by hormone-secreting neuroendocrine tumors, classically midgut tumors with liver metastases.
F. -relix = releasing-hormone antagonist (degarelix, cetrorelix, ganirelix, abarelix). -relin
= releasing-hormone agonist (leuprorelin, triptorelin, goserelin, sermorelin, tesamorelin). Full
credit requires noting the opposite initial effect, not just the class labels.
G. A molecular cage that grips a metal ion and does not release it. Needed because the radionuclide must stay attached to the targeting molecule in circulation; free metal would distribute by its own chemistry rather than by the peptide's targeting.
H. PFS = time to documented disease growth or death. OS = time to death from any cause. Full credit requires stating that a PFS benefit does not entail an OS benefit.
I. (1) Shut down a hormone axis — leuprolide, goserelin, triptorelin, degarelix. (2) Replace or mimic an inhibitory hormone — octreotide, lanreotide. (3) Deliver to an address — lutetium Lu 177 dotatate.
J. A targeting system usable for both diagnosis and therapy. Only the radionuclide changes.
Part 2 — Applying the mechanism
K. Panel 1: pulses → maintained LH/FSH output → maintained gonadal steroid production. Panel 2: continuous exposure → initial surge → receptor desensitization/internalization (label the inflection around 1–2 weeks) → collapsed LH → castrate testosterone. Look for the initial surge being drawn in panel 2; students who omit it have missed the chapter's core point.
L. Should convey: the drug first switches the system on before switching it off; the "on" phase is brief; a second medicine is given to cover it. Should NOT convey: that the worsening is an allergy, an impurity, or a sign the drug is not working.
M. † Strong counterargument for "side effect": from the patient's and the regulator's point of view, an unwanted clinical consequence of a drug is a side effect regardless of mechanism, and the label treats it that way. Strong response: the classification matters because it determines whether you expect it (you do), whether it can be engineered away within the class (it cannot), and whether it generalizes to other axis drugs (it does). Accept either final position; grade on whether the student distinguishes regulatory/clinical classification from mechanistic classification.
N. Truncation (removes non-essential residues and therefore cleavage sites); cyclization via disulfide (blocks exopeptidase access — no free ends to grip — and resists the extended conformation endopeptidases need; also locks conformation, raising affinity); D-amino acid substitution (proteases are stereospecific, so a D-residue at a cleavage site is not recognized). Bonus: C-terminal alcohol (threoninol) closes carboxypeptidase attack.
O. † Should ask: what is the intrinsic plasma half-life of the linear all-L molecule? A depot controls release rate, not clearance rate — so a molecule cleared in minutes will still be cleared in minutes after leaving the depot, and the steady-state concentration achievable may be too low. The specific property to overcome is systemic proteolytic clearance, which formulation cannot fix. Look for students who recognize that formulation and stabilization solve different problems.
P. Because the therapeutic agent is the radionuclide; the peptide only has to deliver and retain it. Beyond binding it needs retention — internalization and residence time long enough for the isotope to deposit its dose. Accept also: adequate tumor-to-background ratio, acceptable normal-tissue biodistribution.
Q. Advantage: crossfire kills neighboring tumor cells that did not take up the construct, covering heterogeneous receptor expression. Consequence: non-expressing normal cells within that radius are irradiated too; targeting concentrates dose, it does not confine effect.
R. † Prediction should include dry mouth / salivary gland dysfunction, renal effects, and GI effects. Grade on the quality of the reasoning trace and on the student's honesty about misses. The transferable point: expression data predicts where toxicity will appear but not how severe or how reversible, because expression level, radiosensitivity, and functional reserve all differ by tissue.
Part 3 — Evidence and rating
S. Both ✅. Symptom control: reason = decades of use, large rapidly observable effects, patient-reported endpoint, guideline first-line status. Antiproliferative: reason = PROMID and CLARINET, randomized placebo-controlled, PFS endpoints; "what would change it" must flag that neither certifies overall survival. The explanation should invoke Rule 6 (one molecule, many ratings) and the requirement to name population and endpoint.
T. Midgut → excludes pancreatic, lung, other primaries. Well-differentiated → excludes high-grade neuroendocrine carcinoma. SSTR-positive on imaging → excludes receptor-negative patients. Progressed on octreotide → excludes treatment-naive and first-line settings.
U. † Failures: (i) "live longer" substitutes OS for the PFS primary endpoint; (ii) "neuroendocrine cancer" drops midgut, grade, and receptor-status qualifiers; (iii) "patients" drops the prior-therapy requirement; (iv) "radioactive peptide therapy" is fine but invites generalization to PSMA agents that are not peptides. Accurate rewrite should keep the endpoint and at least the receptor-status qualifier. Ask students to name what they gave up — usually precision about line of therapy.
V. Best answer is PRRT or the theranostic loop: the mechanism is exceptionally elegant, and the ✅ rests entirely on NETTER-1. An "upgrade with mechanism" would read: "because the peptide provably reaches receptor-positive cells, it must extend survival" — which is exactly the inference the OS data does not support.
W. A class rating hides variation between constructs (different peptides, linkers, payloads, tumor targets) and can launder a failed molecule's reputation via a successful sibling — or vice versa. It is actively misleading when a single construct is being marketed and the class rating is quoted as though it applied to that construct.
X. † Should convey: approval means a regulator evaluated specific evidence for a specific use; an accelerated approval means the evidence was a surrogate endpoint with a promise to confirm; confirmation can fail. Must NOT slide into "approval is meaningless" — the withdrawal is evidence the system worked.
Part 4 — Reading labels and dossiers
Y. Commonly dropped fields: approval type, line of therapy, and "what it does NOT say." Those three are the ones worth calling out in feedback.
Z. Marker's note: the marketer's sentence almost always drops receptor status and line of therapy first, because those are the clauses that shrink the addressable population most.
AA. † Grade on the contrast paragraph, not the entries. The intended insight: a complete "not approved anywhere" entry is not a gap in the dossier — it is the finding.
AB. Acceptable approach: Field 7 records only label content; off-label uses go in a later field tagged with their evidence source (guideline, trial, case series) and explicitly marked off-label. Reject answers that put guideline-supported off-label use into Field 7.
Part 5 — Synthesis and argument
AC. Look for: the forty-year timeframe, at least two named drug classes, an explanation of the invisibility, and the "medicine adopted the peptides with a job it could verify" framing. Marking down for use of frontier is the point of the constraint — it forces the student to describe old, established medicine in the register it deserves.
AD. † No approved counterexample is expected. Strong answers describe their search strategy (regulatory indication databases, guideline documents) and note that near-misses — supportive care agents, for example — fail the test because their approved indications are narrow and specific, not broad and supportive. Reward students who find something and then correctly disqualify it.
AE. † Core argument: peptides' defining pharmacological liability is that they are digested, and relugolix demonstrates it by contrast — same receptor, same clinical effect, oral route, and the only relevant difference is that it is not a peptide. This is stronger than any positive claim because it is a controlled comparison the field ran for us.
AF. No fixed answer. Grade on: absence of dosing or protocol content, presence of the comparator framing, and honesty in the final self-critique sentence. Students who cannot name a sentence they are unsure of have usually not engaged with the difficulty.
Chapter 28
For the instructor answer key. Covers exercises.md (which ships without answers), the quiz.md
key (reproduced in condensed form with marking guidance), and the case study discussion questions.
General marking principle for this chapter: the recurring error is collapsing distinct claims into one verdict. Reward any answer that separates claim, population, endpoint, and study design, even if a detail is wrong. Penalize confident single verdicts on "natriuretic peptides," however well written.
Exercises
Part A — §28.1–28.2
A. Atrial myocardial extract injected intravenously into rats produced rapid, large sodium and water excretion and a fall in blood pressure. The control that mattered was ventricular tissue extract, which did not. Without it, the result could have been a nonspecific response to injected tissue homogenate.
B. The heart was categorized as a mechanical organ; endocrine function was assigned to glands. Nobody looked for a cardiac hormone because the category said there would not be one. Acceptable further examples: adipose tissue as an endocrine organ (leptin, adiponectin); the gut as the largest endocrine organ; bone as a source of endocrine signals; the kidney (erythropoietin) is a weaker example since it was recognized earlier.
C. Expect any three of: the signal responds within seconds-to-minutes of a mechanical change rather than following a chemical set point; it is graded continuously with load rather than pulsed by a pituitary trophic hormone; it has no classical negative-feedback loop through a hypothalamic-pituitary axis; it reports a physical quantity (wall tension) that no chemical assay of the blood would otherwise reveal; it varies with posture, volume status, and rhythm.
D. ANP: atria / stored preformed / NPR-A / fast circulating volume off-loading. BNP: ventricles / made on demand / NPR-A / sustained signal of ventricular load. CNP: endothelium and chondrocytes / made locally / NPR-B / paracrine vascular tone and growth-plate signaling.
E. Storage permits near-instant release, so ANP tracks acute changes — good physiology. But a value that swings minute to minute is a noisy measurement. BNP's transcriptional response integrates over hours, which is a worse acute signal and a far better readout. Full credit requires naming the integration-over-time point explicitly.
F. They are single-pass membrane receptors whose intracellular domain is itself a guanylyl cyclase; the second messenger is cGMP, acting through protein kinase G. It matters because vericiguat (§28.8) reaches the same second messenger via a different enzyme.
G. NPR-C binds all three peptides, does not signal, and internalizes them for degradation — a clearance receptor. It partly explains increased clearance in natriuretic peptide resistance, and it is one proposed mechanism for the lower levels seen in obesity (adipose expression).
H. † Diagram should show natriuretic peptides reducing and RAAS increasing each of the four variables, with RAAS marked as predominant in established heart failure. Intervention points that should appear: block the RAAS arm (ACE inhibitors, ARBs, MRAs); amplify the natriuretic arm (supply or preserve); act downstream on the shared second messenger. Strong answers notice that the chapter tried all three.
I. † Mechanisms: receptor downregulation under sustained exposure; increased clearance (NPR-C and enzymatic); impaired precursor processing so that circulating immunoreactive material includes less-active forms. The prior-lowering argument: adding ligand to a resistant, downregulated system is not the same as restoring a deficiency. The required second half: a lowered prior is a reason to test, not a substitute for testing; it is exactly the kind of mechanistic reasoning rule 3 forbids using as evidence — in either direction.
Part B — §28.3
J. proBNP (108 aa) → NT-proBNP (76 aa, N-terminal, inactive) + BNP (32 aa, C-terminal, active). The irony: the inactive fragment is more stable in blood and in the specimen, so the useless half is the better measurement.
K. Reference ranges differ by roughly an order of magnitude. Practical consequence: a result is uninterpretable until you know which assay produced it.
L. Rule-out. Patient-facing sentence should convey: a low result makes heart failure an unlikely explanation for these symptoms; a high result is consistent with heart failure but is also raised by several other things and does not by itself establish the diagnosis.
M. Age ↑ (rises with age in people without heart failure). Kidney impairment ↑ (reduced clearance; larger effect on NT-proBNP). Obesity ↓ (proposed: increased clearance-receptor expression in adipose tissue, reduced production). Atrial fibrillation ↑ (fibrillating atria produce natriuretic peptides abundantly).
N. † Obesity. Confounders that raise the level cost specificity, and clinicians are already primed to treat a high value cautiously. A confounder that lowers the level attacks the test's rule-out function — the one thing it is best at. Most costly scenario: a person with severe obesity, breathless, with a "normal" result used to exclude heart failure and redirect the workup.
O. Open. Strong medical examples: PSA; D-dimer; troponin in non-cardiac illness; ejection fraction as a single number. Non-medical examples: credit scores, standardized test scores, performance metrics used as targets. Award credit for naming the general failure mode — a validated input becomes the decision.
P. † Echocardiography answers what is the structure and function of this heart; the blood test answers is this heart currently under load. Blood test more useful: 3 a.m. undifferentiated dyspnea with no echo available, where a rule-out redirects treatment immediately. Echo more useful: establishing ejection fraction to determine which disease-modifying therapy applies — a question no blood test answers.
Part C — §28.4
Q. (a) Does measuring it aid diagnosis? — prospective diagnostic accuracy study — ✅. (b) Does infusing a synthetic version improve outcomes in acute HF? — randomized outcome trial — ❌. (c) Does preventing its breakdown improve outcomes in chronic HF? — a different randomized outcome trial — ✅. Bonus credit for the fourth: biomarker-guided titration — randomized trial of a management strategy — ❌.
R. There is no intervention in it. It measures how well a test tracks a diagnosis; nobody is assigned to be treated differently, so no causal claim about treatment can be derived. Size does not help — a large study of the wrong design answers the wrong question more precisely.
S. Example: "NT-proBNP predicts mortality so strongly that lowering it must lower mortality." Fallacy: treating a prognostic association as an interventional claim; equivalently, the surrogate fallacy of Ch 5 §5.6.
T. The trial was stopped for futility because the treatments the biomarker would push toward are the treatments guideline-directed care already prescribes — the number added no new decision. Implication: biomarker-guided strategies would be expected to help where usual care is not already optimized, or where the marker triggers an action guidelines do not specify.
U. † Argument: a compelling mechanism raises the perceived cost of a negative result, so it biases interpreters toward explaining the negative away (wrong population, wrong dose, wrong endpoint) rather than accepting it. Accept any earlier-chapter example where a mechanistically attractive claim is unsupported; BPC-157 and cosmetic peptides are the expected picks.
V. † Marking: the key insight is that this is a third claim — a randomized trial of a testing strategy with clinical endpoints, not a diagnostic accuracy study and not a drug trial. Strong answers note that the diagnostic ✅ does not transfer, that the GUIDE-IT result is about a different use of the same measurement and so does not settle this one either, and that the honest rating is ⚠️ or unrated pending a trial of the strategy. Do not require a particular verdict; require that the claim be correctly typed.
Part D — §28.5
W. Because the failure had nothing to do with molecular identity. Being the exact human sequence guarantees receptor binding; it guarantees nothing about whether raising the ligand in a resistant, downregulated, already-treated system changes outcomes. Sequence fidelity is a chemistry achievement, not a clinical one.
X. Pulmonary capillary wedge pressure — a surrogate endpoint, objectively measured. Dyspnea at a few hours — a short-term patient-reported symptom endpoint; real, but not a hard endpoint. Neither is death or rehospitalization.
Y. Absence of benefit: the treatment did not change outcomes. Presence of harm: the treatment made outcomes worse. The trial found the former and specifically did not confirm the latter. Conflating them misstates what patients were exposed to and misstates what the ❌ means.
Z. † Two independent negative results, in different trials with different molecules, generalize from a claim about a compound to a claim about a strategy — supplying exogenous natriuretic peptide in acute heart failure. Limit: both tested acute decompensation with short follow-up; neither excludes a different phase of illness, a different population, or a differently engineered molecule.
AA. † "Evidence present and negative" means an adequate trial asked the question and the answer was no; the claim is closed until someone gives a reason to reopen it. "Evidence largely absent" means nobody has asked; the claim is open and unsupported. Both get ❌ because both fail the standard of supported claims, but the reader's correct next action differs — in the first case, stop; in the second, look for or demand a trial. A rating system that cannot express the difference misdirects effort.
Part E — §28.6–28.7
AB. ANP, CNP — helpful. BNP — helpful, but a relatively poor substrate, so a smaller contribution. Adrenomedullin — probably helpful (vasodilation). Bradykinin — harmful (angioedema risk), though partly vasodilatory. Substance P — unclear, contributes to angioedema. Angiotensin II — harmful. Amyloid-beta — unclear; a theoretical long-term concern.
AC. Neprilysin degrades angiotensin II. Inhibiting it alone raises angiotensin II alongside the beneficial peptides, increasing vasoconstriction, aldosterone, and sodium retention — releasing the brake and the accelerator together.
AD. Because bradykinin is degraded by both neprilysin and ACE; blocking both compounds its accumulation and raises angioedema risk. The earlier combined neprilysin–ACE inhibitor produced angioedema substantially more often than an ACE inhibitor comparator and was not approved.
AE. † Identical architecture: inhibit the degrading enzyme, let endogenous ligand persist, rather than administering exogenous ligand. Shared ceiling: you can only preserve what the body actually makes, so a degradation inhibitor cannot exceed physiological receptor occupancy the way an agonist can. Inversion explanations (any one, well defended): timing and disease phase; the Chapter 3 physiological-pattern argument; the multi-substrate/network argument; the difference in comparator and endpoint. The physiological-pattern argument is the most interesting; do not require it.
AF. More impressive because enalapril already reduces mortality in this population, so the measured benefit is incremental to an effective therapy. What it prevents: any estimate of the drug's effect versus no RAAS therapy, and any separation of the neprilysin-inhibition contribution from the angiotensin-blockade contribution.
AG. † Run-in: enriches for tolerance, so trial tolerability overstates real-world tolerability; because both arms were run in, the efficacy comparison is less clearly biased. Early stop: trials truncated for benefit tend as a class to overestimate effect magnitude. Neither threatens the direction of the finding; both bear on magnitude and on generalizability of the safety profile.
AH. † Neprilysin degrades BNP but not NT-proBNP. Inhibiting the enzyme raises measured BNP by reducing clearance, independent of cardiac status, while NT-proBNP continues to track wall stress and falls with improvement. Why it summarizes the chapter: the same gene product occupies two roles — therapeutic target and diagnostic marker — and intervening in one role corrupts the other, with the corruption falling precisely on the fragment the enzyme happens to recognize.
Quiz — condensed key with marking notes
1 B · 2 B (accept saline as a control but require the tissue-specific control for full marks) · 3 False, first isolated from porcine brain, overwhelmingly ventricular · 4 ANP→atrial, BNP→ventricular, CNP→endothelium/chondrocytes · 5 B · 6 C · 7 C · 8 natriuresis/diuresis, vasodilation, renin and aldosterone suppression, restraint of hypertrophy and fibrosis — all opposing RAAS · 9 B · 10 False, ranges differ ~10-fold · 11 B · 12 ↑ ↑ ↓ ↑ · 13 obesity, because it is the only one producing false reassurance and it undermines the rule-out function · 14 B · 15 C · 16 C (deduct if the student asserts harm was demonstrated) · 17 angiotensin II · 18 B · 19 B · 20 CV death or HF hospitalization; HR 0.80 (0.73–0.87); stopped early by the DSMB for overwhelming benefit on CV mortality · 21 neprilysin degrades BNP but not NT-proBNP; use NT-proBNP · 22 see key-takeaways.md; require measure/supply/preserve and the system-versus-strategy sentence.
Common wrong answers worth a comment rather than a mark deduction: students frequently answer 16 with "increased mortality," having absorbed the pooled-analysis narrative rather than the trial result. This is the single most useful error in the quiz — use it.
Case study discussion questions
These have no single correct answers. Notes on what a strong response contains are in
_scratch/instructor/ch28.md under the Discussion Guide heading; do not duplicate them here.
Two answers that are substantially constrained by fact:
Case Study 28.1, Q4. The distinction between absence of benefit and presence of harm is factual, not interpretive: the definitive trial did not confirm the mortality or renal signals. Any answer asserting that nesiritide was shown to kill people is wrong and should be corrected. The second half — why the distinction is often unavailable for Part III compounds — must reach: you can only distinguish "no benefit" from "harm" if an adequately powered trial measured both, and for most Part III compounds no such trial exists.
Case Study 28.2, Q6. The washout explanation must reach bradykinin and angioedema, and must identify the unapproved combined neprilysin–ACE inhibitor as the origin. The biomarker half must reach: neprilysin degrades BNP, not NT-proBNP.
Chapter 29
Instructor-facing. Covers the 34 exercises (which ship without answers) and gives marking notes for
the † items. Letters match exercises.md as shipped (A–AH). Quiz answers ship with the quiz and are
not duplicated here except where a marker needs extra guidance.
Part 1 — The inventory
A. No fixed answer. The pedagogical point is the gap between recall before and after. Most students name 2–4 before and 12–18 after. Push them on the second half of the question: nearly every "missed" drug was one they had heard of.
B. teriparatide — severe osteoporosis / high fracture risk. desmopressin — central diabetes insipidus, primary nocturnal enuresis, certain bleeding disorders. cosyntropin — diagnostic testing of adrenal responsiveness. icatibant — acute hereditary angioedema attacks. linaclotide — IBS-C and chronic idiopathic constipation. terlipressin — hepatorenal syndrome (also variceal bleeding in Europe). glucagon — rescue for severe hypoglycemia. secretin — diagnostic assessment of pancreatic exocrine function.
C. Any three of: vancomycin (filed as "antibiotic"), vasopressin ("pressor"), oxytocin ("labor agent"), cyclosporine ("immunosuppressant"), octreotide ("somatostatin analog"), bacitracin ("topical antibiotic"), cosyntropin ("diagnostic agent"), leuprolide ("hormonal therapy"). Full credit requires naming the functional category, not just the drug.
D. Because the peptide/protein boundary is a convention (Ch 1 §1.5). Three decisions that move the total: counting insulin analogs individually or as a family; including or excluding recombinant proteins under ~100 residues; including or excluding glycopeptide and lipopeptide antibiotics. Also acceptable: whether to count diagnostic agents; whether to count peptide-drug conjugates.
E. † Look for two rules, not two numbers. A low-count rule might be: "synthetic peptides of 50 residues or fewer, with a published sequence, approved as therapeutics by the FDA." A high-count rule might be: "any approved agent whose active moiety is a chain of amino acids joined by peptide bonds, including antibiotics, diagnostics, and recombinant products under 100 residues, in any major jurisdiction." Strong answers note that the second rule roughly doubles or more the first. Do not mark on the specific numbers — mark on whether the rules are coherent and whether the student identifies which inclusion decisions carry the most weight.
F. Would have told you: teriparatide, desmopressin, terlipressin, linaclotide, plecanatide,
octreotide, leuprolide (-tide, -pressin, -relin/-relix families). Would not have: vancomycin,
cyclosporine, bacitracin, daptomycin, polymyxin, cosyntropin, secretin, glucagon, oxytocin,
calcitonin. Point to draw out: the naming conventions post-date most of the antibiotics and natural
products, which is why they do not carry the stems.
Part 2 — Teriparatide
G. Continuous elevation removes bone (hyperparathyroidism). Intermittent once-daily exposure increases bone formation and bone mineral density.
H. Two of: (i) the N-terminal 34 residues carry full receptor-activating capability, so the rest is unnecessary; (ii) at ~40% of the mass the molecule falls in the synthesizable peptide range rather than the recombinant-protein range, with manufacturing and cost consequences; (iii) the removed portion contributes to clearance and other functions that are not wanted therapeutically.
I. Anabolic agents stimulate new bone formation; antiresorptives slow removal of existing bone. Teriparatide is anabolic. Antiresorptives (bisphosphonates) dominate prescribing.
J. † Strong answers say explicitly that Ch 3 treats pulsatility as a warning — a reason that flat exposure to a pulsatile hormone may not reproduce physiology — whereas teriparatide treats it as a lever: choose the pattern, get the effect you want. For the second half, growth hormone is the obvious candidate (Ch 3 and Part IV); credit any well-argued choice. To be convertible you need: a known pattern-dependence of effect, a feasible dosing schedule that reproduces the beneficial pattern, and a measurable endpoint. Weak answers assert that pulsatile dosing would "be more natural" without naming an endpoint.
K. PTHrP is a distinct gene product from PTH that acts at the same receptor (PTH type 1). Abaloparatide is a modified PTHrP(1–34) analog. The receptor accepts both because it evolved to respond to a shared N-terminal recognition motif — Ch 2's point that receptors are promiscuous in ways evolution finds useful.
L. † (two-part item)
Part one, the five steps: (1) lifetime high-exposure rat studies show osteosarcoma above control; (2) boxed warning issued at approval, with cumulative duration-of-use limits; (3) two decades of clinical use in large numbers of patients; (4) post-marketing surveillance does not show the expected human excess; (5) boxed warning removed and duration language relaxed. Decisive evidence arrives at step 4, and it is observational post-marketing surveillance data, not trial data. That distinction is worth drawing out: for a rare outcome, an RCT would never have been adequately powered.
Part two, the Chapter 8 comparison: similarities (any two) — both are rodent carcinogenicity findings; both drove boxed warnings; both involve a tissue whose relevant biology differs meaningfully between rodents and humans; both had the question referred to long-term human data. Difference: the teriparatide question has been substantially resolved in the direction of the label being relaxed, whereas the C-cell question remains more open; also, teriparatide's finding was in the tissue the drug acts on, which is a different inferential situation. Deduct for any answer concluding "so animal findings don't matter." The chapter is explicit that this is the wrong lesson.
M. No canonical answer. The correct response is that the student should say what they would need to look up. Award full credit for a well-formed four-line block that names the population (men with osteoporosis), the endpoint (bone mineral density, not fracture), and states honestly that the student would need to check whether a dedicated trial in men exists and what endpoint it used. Penalize students who confidently issue a ✅ by extrapolating from the postmenopausal data — that is rule 3 (never upgrade with mechanism) in a slightly disguised form.
Part 3 — Desmopressin
N. V1 — vascular smooth muscle — vasoconstriction. V2 — renal collecting duct (and vascular endothelium) — water reabsorption / antidiuresis.
O. Deamination at position 1 and D-arginine substitution at position 8. Together they shift selectivity toward V2 and extend duration. The D-arginine substitution does two jobs: it contributes to the selectivity shift and it makes the molecule a poorer peptidase substrate.
P. (i) V2 in collecting duct → water reabsorption → reduced urine volume → central diabetes insipidus. (ii) same mechanism, applied overnight → reduced overnight urine production → nocturnal enuresis. (iii) V2 on vascular endothelium → release of stored von Willebrand factor and factor VIII → increased circulating factor levels → hemostatic support in mild hemophilia A / type 1 vWD.
Q. † Any rewrite that broadens the claim beyond what the evidence supports. Examples that should score well: "desmopressin cures primary nocturnal enuresis," "desmopressin produces durable dryness after discontinuation," "desmopressin is superior to enuresis alarms for long-term resolution." The explanation must identify that the original claim is bounded to wet nights during treatment, and that relapse on discontinuation is well recognized and therefore already outside the claim.
R. Hyponatremia. Mechanistically direct: the drug's therapeutic action is free water retention; excessive retention dilutes serum sodium. Mark down any answer that drifts into fluid restriction protocols — the chapter states the risk without a management plan and students should follow that.
Part 4 — Calcitonin
S. (i) better-evidenced competitors arrived; (ii) fracture evidence never became strong; (iii) regulatory reviews concluded the benefit-risk balance did not support the osteoporosis indication. An argument can be made that (iii) alone was sufficient in Europe; the stronger answer is that (i) alone was probably sufficient in practice, since prescribing had already shifted before the reviews.
T. Salmon calcitonin is substantially more potent at the human calcitonin receptor than human calcitonin is.
U. Three explanations: an unusual biological dose-response; differential attrition across arms distorting the comparison; chance finding among multiple comparisons. Distinguishing information: complete per-arm dropout data with reasons; prespecification documentation showing whether the positive comparison was primary; and, decisively, an adequately powered replication.
V. † Expect a four-line block along the lines of: claim — calcitonin lowers serum calcium in acute hypercalcemia of malignancy; rating — ✅ or ⚠️ depending on how the student bounds it (both defensible if reasoned; the short-term calcium-lowering effect is well established, while tachyphylaxis limits duration and a student may reasonably reflect that in the claim wording); reason and what-would-change-it as appropriate. The second half is the graded part: issuing a separate rating is required because the population, endpoint, and evidence base all differ — a single molecule-level rating would import the fracture-reduction failure into a use where it does not apply.
W. † Look for: (a) a clear statement that nothing was hidden; (b) identification of the mechanism of revision (better trials, comparative evidence, regulatory reweighing); (c) engagement with the obvious objection, which is usually "the manufacturer just stopped promoting it because the patent expired." A strong answer concedes that commercial factors were present and argues that they cannot explain the direction of the revision, since the fracture evidence and the regulatory reviews were independent of patent status.
X. Hypercalcemia, Paget's disease, or acute vertebral fracture pain. What distinguishes them: different endpoint, generally shorter duration of use, and in the hypercalcemia case an effect that is directly and immediately measurable rather than requiring a multi-year fracture trial.
Part 5 — The peptides nobody calls peptides
Y. Cyclization (defeats exopeptidases); D-amino acid content (defeats protease recognition); N-methylation (defeats protease recognition and increases membrane permeability).
Z. N-methylation obstructs protease recognition and removes a backbone hydrogen-bond donor. The second is the one that matters for membrane crossing: fewer hydrogen-bond donors means the molecule is held less firmly by surrounding water.
AA. Cyclization. It solves the problem of secreting a molecule into a hostile extracellular environment full of competitor proteases and expecting it to remain intact long enough to act.
AB. † The exception is possible because cyclosporine is not built like a typical peptide — the three features above. What it licenses: the conclusion that peptide oral bioavailability is a hard engineering problem rather than a physical impossibility, and that specific structural strategies can solve it. What it does not license: the conclusion that any given oral peptide product has solved it. The correct posture is the one from Ch 1 §1.6 — the burden is on the seller to explain what they did, and "cyclosporine works orally" is not that explanation. Mark down answers that treat the exception as general permission.
AC. Oral vancomycin is not absorbed, so it stays in the gut lumen where C. difficile lives. Intravenous vancomycin is used when the infection is systemic. The same principle appears in §29.9 under linaclotide — the compartment does the targeting.
Part 6 — Modern approvals, diagnostics, and the argument
AD. (two-part item)
Part one, linaclotide: its receptor is on the luminal surface of the intestinal epithelium, so the gut lumen is the target compartment; absorption would add off-target exposure without adding effect. Oral semaglutide's ~1% is a triumph because its targets are systemic — the 99% loss is genuine loss. The same number means opposite things depending on where the target is.
Part two, stimulation tests: a stimulation test provokes the response a functioning gland should produce, which is more informative than an ambiguous baseline that varies with time of day, stress, illness, and medications. Cosyntropin can be ACTH(1–24) because the N-terminal region carries the receptor-activating capability — the same principle as PTH(1–34).
AE. An antagonist binds the receptor and blocks the natural ligand without triggering activation. Harder to design because it requires tight binding while failing to induce the conformational change that binding normally produces.
AF. † Box 1: insulin. Box 2: leuprolide, octreotide, icatibant. Box 3: bacitracin, plecanatide. Box 4: bacitracin (also — legitimately dual). Box 5: secretin. Terlipressin is the interesting one: best placed as a box 1/2 hybrid, or argued as box 2 (opposing a pathological vasodilatory state). Oxytocin is the other hard case — it augments a physiological process rather than correcting a deficit or blocking an excess, and strong answers will notice that and argue it. Reward students who identify bacitracin and oxytocin as genuinely awkward rather than forcing them.
AG. † Strongest form of objection: approval requires a clinical endpoint in a defined population; "optimization" is not such an endpoint; therefore box six is empty by construction and the observation carries no information about biology. The three responses: (1) the commercial incentive to enter that space is enormous and companies have repeatedly tried; (2) adjacent approvable categories (prevention, risk reduction, functional improvement) exist, other drug classes hold them, and peptides are thin there too; (3) the same pattern independently appears in Ch 27's oncology survey and Ch 37's master table. Most students find (1) weakest, since "companies have tried" is hard to evidence precisely. Strengthening it would require naming specific development programs and their fates.
AH. † Mark on discipline, not conclusion. Full credit requires: the claim stated in the claimant's own terms (not a strawman); a box assignment with justification; and a statement of what would move it. Deduct for any answer that concludes the compound does not work — the prior does not license that, and the chapter says so explicitly.
Two model answers worth having ready
For the closing move of Prompt 6 / item AH, if a class stalls:
"The claim is that this compound accelerates recovery in healthy trained athletes. That is box six — there is no identified deficiency, no identified excess, no pathogen, and no localized pathology. To move it into box one, someone would have to identify a measurable deficit that the compound corrects, in a population defined by having that deficit. Notice that this is not a hostile requirement: it is exactly what teriparatide, desmopressin, and glucagon all satisfy."
For students who ask what the "field is too new" reply should sound like (this was an exercise in an earlier draft and still comes up in discussion):
"Peptide medicine isn't new — insulin has been in use since 1922 and cyclosporine since the early 1980s, and there are more than eighty approved peptide drugs today. This is a mature field with established regulatory pathways that gets peptides approved routinely. So the question isn't why the field hasn't caught up; it's what specifically happened with this compound — was a trial run and did it fail, is one running now, or has nobody ever run one?"
Chapter 30
Internal. Not for publication in the student-facing chapter. exercises.md deliberately ships without
answers; this file is the instructor's marking companion. The quiz answer key is published inside
quiz.md in a collapsed <details> block and is reproduced here in short form for convenience.
Quiz — answer letters
| Q | Ans | Q | Ans | Q | Ans | Q | Ans |
|---|---|---|---|---|---|---|---|
| 1 | b | 7 | b | 13 | b | 19 | b |
| 2 | b | 8 | c | 14 | c | 20 | b |
| 3 | c | 9 | c | 15 | b | 21 | b |
| 4 | b | 10 | b | 16 | c | 22 | b |
| 5 | c | 11 | c | 17 | b | ||
| 6 | b | 12 | b | 18 | c |
Items that most often trip students: 3 (they assume every cosmetic peptide is oversized — GHK is not), 12 (they want the answer to be "the mechanism is fake"), 14 and 15 together (the size/efficacy inversion is the chapter's whole point), 19 (they choose sample size), 21 (they think split-face solves blinding).
Exercises — marking notes
Reasoning is the assessed object throughout. A student who reaches a defensible wrong conclusion through disciplined reasoning has done better than one who lands on the chapter's stated view by restating it.
Part 1 — The barrier
A. Look for: dead cells, lipids, a physical rather than chemical obstruction, and the point that the layer is doing its job. Full credit requires the student to convey that this is an evolved function, not a design flaw.
B. Acceptable properties: lipophilicity/partition coefficient, ionization or charge at skin pH, vehicle composition, barrier integrity (broken, inflamed, or occluded skin), degree of hydration, follicular density at the site. Two of the six with correct reasoning is a pass; three with worked examples is strong.
C. Order: GHK (~340) < palmitoyl tripeptide-1 (~570) < palmitoyl pentapeptide-4 (~800) < acetyl hexapeptide-8 (~890) < botulinum toxin (~150,000). Best evidence: botulinum toxin, by a wide margin. The non-paradox is that route, not size, determined it.
D. Strong answers ask, in some order: which layer was the compound detected in? — living skin or excised? — was the intact molecule measured, or a metabolite or label? Award extra credit for asking about the concentration relative to what produces effects in culture.
E. † Assess the honesty of the "strongest true version," not the harshness of the critique. Students who cannot construct a charitable reading have not met the chapter's standard.
F. Key point: tape stripping samples the stratum corneum, which is the layer being shed; a signal peptide's target is in the dermis, two layers down. Detection in the sampled layer is compatible with zero arrival at the target.
G. † Both arguments must be made in good faith. Mark down any answer where one side is a straw man. The strongest "acceptable" case usually rests on the small magnitude of enhancement and the reversibility of barrier disruption; the strongest "not acceptable" case rests on chronic daily use and on the fact that the enhancement is unmeasured in the finished product.
Part 2 — Mechanism versus evidence
H. Must convey that palmitate does not improve target engagement; it improves partitioning.
I. Same: lipid attachment alters how the molecule partitions in a lipid environment. Different: Ch 33's target is enzymatic degradation and renal clearance (half-life); Ch 30's is a physical barrier (penetration).
J. Reasonable testable predictions: dermal procollagen markers should rise in treated skin; a dose-response should exist; the effect should require the intact peptide. Credit any design with a biopsy or validated dermal-marker endpoint and a matched vehicle arm.
K. The two paragraphs must not contradict. Common failure: paragraph two implicitly retracts paragraph one ("but it's all in vitro so it doesn't really show anything"). The correct move is scope-limiting, not retracting.
L. † Either position can earn full marks. What is assessed is whether the student engages rule 6 and the population/endpoint requirement of rule 1. A collapse argument that does not address which claim would be lost is incomplete.
M. Expected: steps 1–4 unmeasured or assumed in living humans; step 5 unmeasured. No step is "known crossed." Students who mark step 1 as "known" from an excised-skin study should be pointed back to §30.2.
N. Second location: any route where absorption into circulation removes the molecule from the intended local target — subcutaneous depot pharmacology, intranasal delivery aimed at the brain, or inhaled agents intended to act locally are all acceptable.
O. Must state that botulinum toxin does not cross skin at all — the comparison is not "large molecule crosses better," it is "large molecule is placed past the barrier."
P. † Key insight: a catalytic mechanism turns over many substrate molecules per enzyme molecule, so very low concentrations suffice; a competitive mechanism requires occupancy comparable to the endogenous protein's abundance. The delivery shortfall therefore matters far more for the competitive mechanism. This is the strongest single answer in the exercise set; flag good ones for discussion.
Part 3 — Regulation and language
Q. (i) cosmetic; (ii) drug; (iii) cosmetic; (iv) drug — "reduces the depth of wrinkles" asserts a physical change to skin structure, unlike "reduces the appearance of"; (v) drug; (vi) borderline and worth arguing — "supports the moisture barrier" is generally treated as cosmetic, and students who notice the borderline should be credited, not corrected.
R. Consequence: packaging language is evidence about legal strategy, not about efficacy. A weak claim may sit on strong evidence, and a strong-sounding claim on none.
S. † Assess the taxonomy, not the totals. The third column — sounds like a claim, is not one — is where the learning happens.
T. Under twenty-five words and not smug. Reject anything that would embarrass the friend.
U. Supports it. If chemistry determined the category, the category could not change at a border.
V. † Look for a workable evidentiary standard and honest engagement with cost and with the sheer number of products. Students almost always underestimate the second.
Part 4 — Read this cosmetic study
W. Did right: split-face, vehicle-controlled, blinded assessors, standardized photographs, adequate duration, real n. What it tells you: the peptide added nothing detectable over its own vehicle — a genuine and useful negative. Why it might not be published: it is commercially valueless to the sponsor, and no registry records that it was run.
X. Genuine strength: biopsy with a dermal marker — an actual measurement past the barrier, which almost nothing in this literature has. Problems with the conclusion: no control arm of any kind (baseline-only comparison), and "proven to rebuild collagen" is both an overreach from a marker to a structural outcome and a drug claim (§30.7).
Y. Almost certainly hydration from the vehicle. The additional arm: identical formulation without the peptide, applied contralaterally, with everyone blinded.
Z. † A defensible answer lands at ⚠️. Information that would move it: independent replication; effect size and whether it exceeds a threshold of perceptibility; whether a delivery study supports dermal arrival; pre-registration status; and the full adverse-event and dropout picture.
AA. Expected ranking: Z > W > X > Y. The determining variable is control and blinding, not n — Y has more participants than X and is less informative.
AB. † The final sentence is the assessed item. It must be printable and non-misleading, which usually forces the student into "appearance of" language, and that realization is the point.
Part 5 — Rating, dossier, and advising
AC. The hydration claim plausibly earns ⚠️ or even ✅ depending on how it is scoped — the vehicle hydrates, and if the claim is about the product rather than the peptide, it is well supported. The best answers notice that the claim as written is ambiguous between product and ingredient, and split it. That is rule 1 being applied without prompting.
AD. The category rating is a prior about marketing, not a finding about molecules. It is legitimate as a default expectation and illegitimate as a substitute for per-claim analysis.
AE. Completion check. The required sentence — evidence of arrival vs. evidence of would-work — is the graded element. "No delivery evidence located" plus a date is a correct answer, not a blank.
AF. † No model answer. Mark for tone as much as content. A technically perfect response that would make the friend feel foolish has failed the exercise. The chapter's stated standard is that a ❌ surviving the concession is worth something; one requiring you to ignore the true part is not.
Suggested subsets
- 60-minute tutorial: A, D, H, M, O, Q, W, Y, AC
- 💄 Cosmetic path core: A, D, F, K, Q, T, W, X, Y, AE, AF
- 🔬 Science path core: B, I, J, M, N, P, Z, AD
- Assessed problem set (graded): C, H, M, Q, W, X, Y, AA, AC
- Daggered set as a take-home: E, G, L, P, S, V, Z, AB, AF
Chapter 31
For the instructor answer key. The chapter's quiz.md carries its own key in a collapsed block; this
file covers the exercises, which are published without answers, plus grading notes for the two
case studies.
Items marked † in exercises.md are the harder ones. For those, a strong answer is characterized
here rather than stated, because most admit more than one defensible position.
Part A — The frame (§31.1)
A. Veterinary approval additionally requires demonstrating human food safety for food-producing species (residue tolerances and withdrawal periods) and an environmental assessment. It exists because the patient may become food, so the regulatory system must protect a third party who never met the animal.
B. Any four of: GnRH agonists (gonadorelin, buserelin, deslorelin) — cattle, dogs, exotics — reproductive management and estrus synchronization; insulin — dogs and cats — diabetes mellitus; desmopressin — dogs — central diabetes insipidus and some bleeding disorders; synthetic ACTH analogs — dogs — adrenal function testing; ghrelin-receptor agonist — dogs and cats — appetite stimulation. The diagnostic agent is the ACTH analog (a probe, not a treatment).
C. Look for a sentence that puts the quality claim and the scope claim in the same breath. Model: "Veterinary evidence can be as strong as any evidence in medicine; what it cannot do is answer a question about a species, indication, or endpoint it was not gathered in." Reject rewrites that only say what veterinary evidence cannot do.
D. † Without the ✅, the chapter's ratings would correlate perfectly with species, and a reader would reasonably conclude the machinery is a species filter rather than a claim-evaluation tool. The relevant rules are rule 1 (a rating attaches to a claim, with population and endpoint, never a molecule) and rule 4 (never downgrade with distaste). Strong answers note that the ✅ functions as a control condition. Excellent answers observe that GnRH analogs also carry human ratings elsewhere (Ch 27) and that this is rule 6 — one molecule, many ratings — in action.
E. Two of: the molecule may differ by species (porcine insulin is identical to canine, not to human); the product may be a veterinary-licensed formulation rather than a human one; monitoring practice differs (glucose curves and fructosamine in animals); species-specific management differences between dogs and cats; route and formulation constraints.
Part B — TB-500, horses, and bans (§31.2)
F. Species error: efficacy in a horse does not establish efficacy in a human — and use in a horse does not establish efficacy even in horses. Molecule error: the marketed compound is a short fragment, not the 43-residue parent thymosin β4 whose literature is being cited (Ch 18).
G. Uncertainty about effects; animal welfare, including injury masking; unapproved regulatory status, often as a class; integrity of competition and public confidence; administrability of class-based prohibitions. None implies efficacy. Credit answers that note the first reason is literally the opposite of an efficacy finding.
H. Model: "It is prohibited, which tells me people use it and that an authority decided the risk of allowing it outweighed the benefit — not that it works." Reject rewrites that smuggle the original implication back in.
I. † Both positions are defensible. The case for "less than nothing" is that the phrase captures a real asymmetry: the listener typically comes away more confident, so the ban has moved belief in the wrong direction. The case against is that "less than nothing" is not a coherent quantity of evidence and invites the same rhetorical looseness the chapter objects to. A good proposed alternative: "a prohibition is uninformative about human efficacy, and is routinely misread as informative, which makes it worse than silence in practice."
J. Any three of: rest; altered training load; different shoeing or surface; concurrent conventional therapy; icing, wrapping, or physiotherapy; simple time and natural history; regression to the mean after a bad period; the observer's expectation and financial stake.
K. † Strongest honest case: equine practitioners see many musculoskeletal injuries, develop pattern recognition, and their impressions can generate hypotheses worth testing; equine imaging and lameness scoring are genuinely objective; horses are large mammals, closer to humans in size and clearance than rodents are. Where it stops: hypothesis generation is not hypothesis testing; objective equine endpoints still measure equine outcomes; and none of it addresses the fragment/parent problem. Look for an explicit statement of the stopping point rather than a general hedge.
Part C — Livestock and endpoints (§31.3)
L. Primary: average daily gain, feed conversion ratio (and dry matter intake). Secondary: carcass measures at slaughter, hormone concentrations, morbidity/mortality within the production phase, residue depletion. Absent: strength, function, injury incidence, tendon or ligament healing, recovery, pain, sleep, mood, wellbeing, quality of life.
M. Definition per glossary. Non-medical examples that work well: a very precise fuel-economy dataset offered as evidence about a car's crash safety; extensive standardized-test score data offered as evidence about teaching quality; detailed website traffic analytics offered as evidence of customer satisfaction.
N. Livestock safety endpoints protect a human food consumer, exposed orally, in trace quantities, in cooked tissue, months later. That is the inverse of an injected bioactive exposure. The research is designed to establish that almost none of the compound remains.
O. † Failures, in order of seriousness: (1) the endpoints do not include human-relevant outcomes at all — total silence, not weak support; (2) the safety data is inverted, protecting against trace oral residue rather than characterizing bioactive exposure; (3) the species is wrong, with five named mechanisms; (4) duration is bounded by the production phase; (5) "rigorous" attaches to how well the study answered its question. Most students rank (3) first; push them toward (1) or (2), since a species-matched study with the same endpoints would still be silent.
P. Power determines how precisely a study answers its own question. It does nothing to change which question was asked. Answers that use the phrase "very precise thermometer" or an equivalent are hitting the chapter's intended analogy.
Part D — Provenance (§31.4)
R. "Used in animals" implies veterinary practice: a licensed clinician, a labeled product, a patient with an owner and a follow-up. "Used in laboratory animals" means administered as an independent variable in an experiment, frequently terminal. The reassurance the listener hears is the first; the fact is usually the second.
S. A compound still known by a laboratory code decades after description almost certainly never completed formal drug development, since generic-name stems are assigned to compounds in serious clinical development. BPC-157 has a code, not a generic name.
T. † Strong answers use at least one of: insulin's development involving dogs; exendin-4 from Gila monster venom; leptin from the ob/ob mouse; venom-derived approved drugs. Excellent answers make the structural point that preclinical work is a necessary stage that earns a compound the right to be tested, and that the chapter's objection is to stopping there rather than to doing it.
Part E — Species differences (§31.5)
U. Receptors (rodent thyroid C-cell GLP-1 receptor abundance, Ch 8; insulin sequence differences, Ch 11) · clearance (metabolic rate scaling) · dose scaling (surface area vs. mass) · model-disease mismatch (transected tendon vs. overuse tendinopathy, Ch 17; ob/ob mouse vs. common obesity, Ch 13) · lifespan (two-year rodent study).
V. Observed: thyroid C-cell tumors in rodents given long-acting GLP-1 receptor agonists, which is the origin of the class boxed warning. May not transfer: rodent C-cells express GLP-1 receptors far more abundantly than human C-cells, so the rodent tissue responds to a signal human tissue barely hears. Warning remains: a mechanistic reason to doubt transfer is not proof of human safety, and long-term human surveillance continues.
W. Definitions per glossary. Expected ranking of what vendor-cited papers demonstrate: face validity most often, construct validity sometimes, predictive validity almost never — because predictive validity can only be established by the human trial whose absence is the whole issue. Credit students who notice that circularity.
X. A transected rat tendon looks like a tendon and heals like a tendon (high face validity) but arises from an acute surgical transection rather than chronic degenerative overuse loading (low construct validity). The signs match; the mechanism does not.
Y. Naive mg/kg conversion overestimates the human-equivalent dose, because smaller animals have higher mass-specific metabolic rates and clear faster, so they require more per kilogram to reach comparable exposure.
Z. † The divisor tracks the ratio of mass-specific metabolic rate between the species; the larger the animal, the closer that rate is to a human's, so less correction is needed. A dog's divisor should be smaller than a rat's — it is roughly 2, versus roughly 6 for rat. Accept any answer that predicts "smaller" with the right reasoning, whether or not the number is right.
AA. † A floor-finding tool has a conservative bias deliberately engineered in, plus an additional safety factor applied downstream, plus a dose-escalation process with monitoring that follows it. Used to target an effect, all of that structure disappears and only the number survives. The error is worse than the arithmetic one because it is invisible: correcting the math does not fix it, and the result looks more authoritative afterward. Strong answers name the disappearance of the escalation-and- monitoring process, not just the conservatism.
AB. Triumph: leptin was discovered by positional cloning in that mouse, establishing adipose tissue as an endocrine organ and reorienting metabolic research. Misleading guide: the mouse cannot make leptin, so it models a rare single-gene deficiency rather than common obesity, in which leptin is typically high and responsiveness reduced.
AC. Long-term: it covers most of the animal's adult life; it is designed to be sensitive to late-emerging effects; regulators treat positive findings as serious for that reason. Short-term: two years is about 2.5% of a human lifespan; it cannot address decades of continuous exposure; the compounds people take chronically have no evidence base of comparable relative duration.
Part F — "Used for years" (§31.6)
AD. In whom? (defeats species transfer) · For what? (defeats indication transfer) · Measured how? (defeats endpoint transfer and all subjective claims) · And collected how? (defeats the existence of a record at all).
AE. † A complete answer runs all four and reaches: in whom — food animals, young, in production; for what — growth promotion and feed efficiency, not safety in a chronically dosing adult human; how measured — gain, FCR, carcass, and residue depletion, with safety framed as trace oral consumer exposure; how collected — systematically, but for production endpoints, over a production phase. The sting: this is the rare case where question 4 is satisfied and the claim still fails, on questions 1 through 3. Use it to show that the four questions are conjunctive.
AF. † Case for keeping it as an addendum: it is a consequence of "was anyone recording," since a control group is a recording decision. Case for promotion: it defeats a different class of claim (anecdote and case series) and is independently useful. Either is acceptable if argued; reject answers that do not engage the dependency.
AG. Look for a non-medical example with the right structure — someone insisting a lucky charm works because they have used it for years, with no count of the times it did not; a manager certain a hiring heuristic works with no record of rejected candidates' outcomes; a driver certain a route is faster who has never timed the alternative.
Part G — Regulation, welfare, two-way street (§§31.7–31.9)
AH. Licensed veterinarian; valid veterinarian-client-patient relationship; the animal's health is threatened or suffering/death may result from failure to treat; no approved animal drug labeled for the use is clinically adequate; for food animals, an established extended withdrawal period and compliance with the prohibited-for-extra-label-use list. An online purchaser satisfies none of them. Credit answers that also note the purchaser is not treating an animal at all.
AI. Definition per glossary. Set from residue depletion studies against a tolerance derived from toxicological data, plus a safety margin; enforced by residue testing in slaughter and milk supply chains. Protects a human food consumer. Human medicine has no equivalent because the human patient does not enter the food supply.
AK. Three-sentence answers should contain: performance animals cannot consent or report; decision-makers frequently have financial interests in performance; therefore the arrangement requires external safeguards rather than good intentions. Deduct for moralizing, for imputing indifference to practitioners, or for drifting into the broader ethics of animal use, which the chapter explicitly brackets.
AL. Animal-to-human: insulin development involving dogs; exendin-4 from Gila monster venom; pit-viper-derived ACE inhibitor lineage; cone snail analgesic; salmon calcitonin; leptin from the ob/ob mouse. Human-to-animal: insulin, desmopressin, analgesics, antibiotics, cardiac and oncology drugs used in companion animals, frequently extra-label.
AM. † Strongest counterargument: for a condition with no treatment and a well-characterized mechanism, demanding a human trial before any conclusion imposes a cost measured in untreated suffering, and preclinical evidence plus mechanistic plausibility is sometimes the best available basis for action. Best answers concede that this is a real argument in rare-disease and compassionate-use contexts, then distinguish those settings — regulated, supervised, documented, with adverse-event capture and an explicit acknowledgment of uncertainty — from an unsupervised purchase where none of that structure exists. The chapter's claim survives because it is about what evidence establishes, not about what may ever justify acting under uncertainty.
Case study grading notes
Case Study 1. The pivot is Q1 (are four of the six gaps really "total"?) and Q5 (defeat a conclusion without disputing any premise). Students who insist all evidence is on a continuum have not understood the categorical gap — no measurement of subjective experience exists in these species at any level of rigor. Q6 is the payoff: the hardest design change is outcome ascertainment, and its difficulty is itself the evidence that the human question has not been asked.
Case Study 2. Q3 is the discriminating question. Students almost always rank the twelvefold arithmetic error as worse; the intended insight is that Error 2 survives the correction, is invisible in the output, and gains authority once the math is right. Q6 is a good end-of-chapter check: a pharmacokinetic study eliminates Errors 1 and 3, reduces Error 4 to a bounded uncertainty, leaves Error 5 entirely untouched, and does nothing about Error 2, since exposure is not efficacy.
Chapter 32
How Peptides Are Made — Solid-Phase Synthesis, Recombinant Production, and Why Peptide Drugs Cost What They Cost
Instructor copy. Not published in the student edition.
The quiz answer key is embedded in quiz.md inside a collapsed <details> block and is not repeated
here. What follows are model answers for exercises.md, plus notes on what a strong answer contains
and where students reliably go wrong. Items marked † in the exercises are extension items; model
answers for those are deliberately shorter, because they are positions to be defended rather than
facts to be recalled.
Set A — Merrifield's idea and the synthesis cycle
A.a In solution-phase synthesis, every coupling leaves the product mixed with unreacted starting material, excess reagent, and byproducts, so the product must be separated by crystallization or chromatography before the next step. Those separations lose material and take days, and they must be repeated after every single residue. Strong answers name the between-steps purification, not the coupling reaction.
A.b A hard, non-porous bead would carry chains only on its outer surface, giving a trivially small loading. Swelling turns the bead into a porous gel so chains are distributed throughout the interior and reagents diffuse in to reach them. Common error: describing swelling as a solubility property of the peptide rather than of the support.
A.c (1) It must render the group reliably inert under the conditions of the coupling. (2) It must be removable later under conditions that leave everything else — including the chain, the linker, and the other protecting groups — intact. A group satisfying only (1) would be a permanent cap; the chain could never be extended past it.
A.d Cycle 1 proceeds normally. At the start of cycle 2 the base intended to remove only the N-terminal Fmoc also strips the side-chain caps on every residue already incorporated. On the coupling step of cycle 2, the activated amino acid can now react at the N-terminus or at any exposed side chain — lysine's amino group especially. By cycle 3 the resin carries a population of branched, heterogeneous polymers rather than a single sequence. Look for the word "branched."
A.e It prevents double incorporation: the chain accepting two (or more) copies of the same residue in one pass. Without it, the newly added residue's free amino group would immediately accept another activated monomer, and the product would be a mixture containing insertion sequences — chains with one or more extra residues.
A.f † The strong answer distinguishes the molecule (defined by its sequence, N-terminus to C-terminus, and any modifications) from the process history. The final molecule is identical; the assembly order is not a property of the product. The interesting concession on the other side is that assembly order does determine which impurities form and where, so two identical molecules can arrive in differently contaminated preparations. Credit answers that reach that concession.
Set B — The arithmetic of stepwise yield
B.a 0.99⁸ = 0.923 → 92%; 0.99¹⁶ = 0.851 → 85%; 0.99²⁴ = 0.786 → 79%.
B.b 0.98⁸ = 0.851 → 85%; 0.98¹⁶ = 0.724 → 72%; 0.98²⁴ = 0.616 → 62%. The gap widens with length: three percentage points at n = 8, seventeen at n = 24. One percentage point of per-step efficiency is worth more the longer the chain.
B.c Solve e⁴⁰ = 0.80 → e = 0.80^(1/40) = e^(ln 0.80 ÷ 40) = e^(−0.2231/40) = e^(−0.005578) = 0.99444, i.e. about 99.4% per step. Sanity check against §32.3: the 99.5% column at n = 40 reads 81.8%, slightly above 80%, so 99.4% is the right neighborhood.
B.d Solve e³⁰ = 0.74 → e = e^(ln 0.74 ÷ 30) = e^(−0.3011/30) = e^(−0.010037) = 0.99001, i.e. 99.0% per step.
B.e §32.2's difficult sequences: aggregation on the resin buries the reactive amino group, so efficiency is sequence-dependent and drops sharply in particular regions. The implication is that deletion content is not uniformly distributed across positions — it concentrates at the one or two hard couplings. That in turn means an analytical method must be shown to resolve the specific deletions a given process generates, not deletions in general.
B.f 0.995¹⁵ = 0.928 → 93%. 0.995⁴⁵ = 0.798 → 80%. Fragment route: each 15-mer at 93%, then two ligations at 70% → 0.928 × 0.70 × 0.70 = 0.455 → 46% overall from the fragments' own yields (and lower still if computed strictly across all three fragments). The instructive result is that at 99.5% per step the linear route wins. Fragment condensation only pays when per-step efficiency is lower or the chain is much longer — which is exactly the regime it exists for. Credit students who notice the answer went against the expected direction and explain why.
B.g Crude purity above 99.9% for 34 couplings requires e³⁴ > 0.999 → e > 0.99997. That is roughly three parts per hundred thousand of failure per step, sustained across 34 steps including any difficult residues. Not physically impossible in the sense of violating a law, but far outside demonstrated process performance and effectively a claim that no coupling ever failed. The correct classification is implausible to the point of being uninterpretable, and the right response is to ask what method produced the number rather than to argue about chemistry.
Set C — Deletion sequences and what purification leaves
C.a A chain missing one or more internal residues, formed when a coupling fails on that chain and the chain is then deprotected on the next cycle and accepts the following residue. It does not stop because nothing has blocked its N-terminus — the failure is an absence of reaction, not the creation of a barrier.
C.b 57 ÷ 3,400 = 1.7%. It matters because a 1.7% mass difference is trivially resolvable by mass spectrometry and essentially invisible in a retention time — so which analytical method you use determines whether you see the impurity at all.
C.c Chromatography separates on overall chemical character — hydrophobicity, charge, shape. A residual solvent differs from a peptide in all of those and elutes in a completely different region. A deletion sequence shares the same termini, nearly the same charge, and nearly the same hydrophobic character, so it elutes on the shoulder of the product peak or underneath it.
C.d It permanently blocks failed chains so they cannot grow further, converting them from 24-residue near-twins into truncated chains of assorted shorter lengths that separate easily. The trade is extra cycle time and reagent, plus no yield benefit, in exchange for a much easier purification and a cleaner impurity profile.
C.e Easiest to hardest: (i) residual solvent — chemically unlike a peptide; (ii) the 12-residue truncate — very different length, mass, and retention; (iii) the 29-residue deletion — near-twin; (iv) the racemized full-length chain — identical mass and composition, differing only in configuration at one center, and often barely resolved by reversed-phase methods at all. Accept (iii)/(iv) swapped if defended on the grounds that some racemized species do separate cleanly; the reasoning matters more than the order.
C.f † Manufacturing case: the change eliminates attachment isomers, which are a purification and identity problem and have no receptor rationale. Pharmacological case: arginine versus lysine changes local charge distribution and could affect receptor interaction or degradation, so the choice is not purely process-driven. What would settle it: evidence about whether the lysine-34 variant retains comparable receptor activity — if it does, the change is manufacturing-driven. Strong answers note that the two are not mutually exclusive and that a design that solves a process problem without costing activity is precisely what good analog design looks like.
Set D — Purity, content, and what a document establishes
D.a Purity by peak area is a ratio among UV-detected peptide species; peptide content is the fraction of the vial's total mass that is peptide. Example: a genuinely 98%-pure preparation whose powder is 15% trifluoroacetate and water contains roughly 83% target peptide by mass. A reader who conflates the numbers and reasons quantitatively from the label overestimates the peptide present by a substantial and unquantified margin.
D.b Any four of: trifluoroacetate (or other) counterion; residual water; residual organic solvents; inorganic salts; non-chromophoric process residues. Note: deletion sequences do NOT belong on this list — they are peptides and do absorb.
D.c At least: What method (column, mobile phase, gradient, detection wavelength)? Has the method been shown to resolve this process's known impurities? Is it area percent or something else? What was the sample preparation? Who performed the analysis, and are they independent of the seller? What is the batch number, and can it be connected to the vial in hand? What is the peptide content, and by what method? What is the counterion and water content? Are individual impurities limited, or only the total? Is there any identity confirmation?
D.d Because area percent describes the proportion of detected signal in one peak; it is silent about what that peak is. A well-purified preparation of an entirely different peptide reports the same number. Identity requires orthogonal measurement — mass spectrometry, sequencing, or comparison against a certified reference standard — which is a different kind of measurement because it interrogates the molecule's structure rather than its behavior on a column.
D.e A claim is checkable when it refers to something outside itself that can be inspected — a document, a standard, a testable limit. "Pharmaceutical grade" refers to nothing; no body defines or confers it. "Manufactured under GMP to monograph X" names a published standard with enumerated tests and limits, and implies a facility subject to inspection. The first cannot be wrong; the second can be, which is what makes it informative.
D.f † Model questions: (1) To which pharmacopeial monograph is this material manufactured and tested? (2) Which laboratory performed the release testing, and is it independent of the manufacturer? (3) What is the peptide content by amino acid analysis or qNMR, and what is the counterion? A seller who means something real answers with a monograph designation, a laboratory name, and two numbers. A seller who does not repeats adjectives, offers the same certificate again, or changes the subject to purity.
Set E — Recombinant production and route selection
E.a Insulin is two chains joined by three disulfide bonds involving six cysteines, which can pair in fifteen ways. Expression produces the chains but guarantees nothing about which pairing forms, and fourteen of the fifteen arrangements are not insulin. So the problem was protein chemistry — folding and correct disulfide formation — rather than gene insertion.
E.b As a single continuous chain, folding is a conformational problem rather than a search problem: the fold brings the correct cysteines into proximity, making the correct pairing overwhelmingly favored. The segment's job is to make folding easy and then leave — it contributes nothing to activity and is excised enzymatically afterward.
E.c (i) synthetic — short, no folding requirement; (ii) recombinant — 191 residues puts synthesis far past the exponential; (iii) requires chemical steps — a non-natural residue cannot be installed ribosomally; (iv) synthetic — very short, and cost per gram at small scale favors synthesis; (v) recombinant — multiple disulfides in a specific pattern is a folding problem.
E.d Aib at position 8 is not one of the twenty amino acids the genetic code specifies. No codon encodes it, so no ribosome can install it. Chemical steps are therefore mandatory regardless of length or cost.
E.e Error one: recombinant is cheaper per gram at large scale, not always — at small scale the fixed costs of a fermentation and downstream process dominate. Error two: cost is not the only criterion, and for molecules with non-natural residues or designed chemical modifications the recombinant route is not merely more expensive but impossible.
E.f † A defensible outline: express the 45-residue backbone recombinantly (length favors it; two disulfides favor a cell), allow folding and disulfide formation in the expression/refolding step, then perform the lipidation chemically as a final, defined step — because doing chemistry on a folded molecule risks the fold, but doing it before folding risks the fatty acid interfering with the fold. Strong answers identify the ordering question as the crux and note that regioselectivity of the acylation is the step to worry about, invoking the semaglutide position-34 lesson.
Set F — Manufacturing scale, cost, and price
F.a Aseptic fill-finish (a validated sterile facility takes years to build and qualify, including media fill runs) and injector-pen assembly (precision device manufacturing to medical-device standards, requiring tooling and its own qualification).
F.b Terminal sterilization uses heat or radiation, both of which damage peptides. The manufacturer must instead sterilize every component separately and combine them aseptically in a controlled environment with continuous monitoring.
F.c Capacity is a regulated physical plant whose qualification timeline is largely fixed by inspection and validation requirements. Money could buy an existing qualified facility, or contract external capacity that already exists. Money could not compress the validation period for a new line or manufacture the regulatory approvals.
F.d Blocks and components: Making the molecule — synthesis/fermentation, purification, drying/isolation. Making it a medicine — formulation, aseptic fill-finish, delivery device, cold chain. Proving it is a medicine — QC and release testing, quality systems and regulatory compliance. Everything before — recovered development cost including failures, market structure. A gray-market vial's price includes essentially only the first block.
F.e Manufacturing cost for two identical units off the same line in the same facility is identical. If the prices differ by multiples, the difference cannot be caused by a quantity that is the same in both cases. This is a straightforward controlled comparison and it is why the manufacturing-cost explanation fails.
F.f Restatement: the low price reflects a product from which many stages have been removed, so the gap measures absence rather than efficiency. What it does not claim: that prescription pricing is justified (Chapter 12 argues otherwise); that gray-market material contains no peptide; that the synthesis was incompetent; or that every removed stage is equally important.
F.g † Look for the students to locate the actual point of disagreement, which is narrow: both arguments can accept that manufacturing cost is real, that it does not explain international price differences, and that a research vial omits most of the cost stack. The disagreement is about how much of the remainder is legitimate recovery of development cost versus rent extracted under exclusivity — an empirical question about a specific market, not a general one about peptides.
Set G — Synthesis across the book
G.a † Accidental: the 50-residue convention long predates any consideration of synthetic yield and was drawn on grounds of folding and function. Not entirely accidental: both the convention and the synthetic limit track the same underlying variable — chain length as a driver of complexity — so a loose correlation is unsurprising. The accidental reading requires fewer assumptions and is the better answer; credit students who resist the temptation to find a causal story.
G.b Contamination → aseptic fill-finish (and the quality systems that verify it). Wrong identity → QC and release testing, specifically identity testing against a reference standard. Wrong concentration → peptide content determination and formulation, plus the release testing that would catch a discrepancy.
G.c † §32.3 explanation: short peptides are cheap and high-yielding to synthesize, and yield falls exponentially with length. Chapter 1 §1.6 explanation: peptides of any appreciable size cannot cross the stratum corneum, so a topical ingredient has a physics problem that gets worse with size. Most students should lead with the second for a non-chemist, because it is about whether the product can work rather than about how it is made.
G.d † No model answer; this is a pre-commitment exercise. Collect the responses and return them when Chapter 34 is taught.
G.e † Step 1: synthesis is automated. Step 2: therefore making peptides is easy. Step 3: therefore peptide drugs should be cheap. Step 1 is true. Step 2 is partly true for short peptides and false as stated — §32.3 and the difficult-sequence problem both bite. Step 3 is the weakest: even granting steps 1 and 2 entirely, §32.7 and §32.8 show that drug substance is a minority of both the cost and the constraint. Strong answers identify step 3 and say why.
G.f Assessed on completeness and honesty rather than correctness. The sentence to look for names something specific — impurity profile, peptide content, sterility — rather than a general expression of doubt.
Chapter 33
For the instructor's answer key compilation. Quiz answers are already published in the chapter's
quiz.md under <details>; what follows is (a) a compact quiz key for the master document and (b)
model responses and marking guidance for exercises.md, which ships without answers.
Quiz — compact key
1 B · 2 A · 3 B · 4 B · 5 C · 6 B · 7 C · 8 B · 9 C · 10 B · 11 C · 12 B · 13 B · 14 B · 15 D · 16 A · 17 B · 18 C · 19 B · 20 A · 21 B · 22 B
Items most often missed, and why. Item 5 (Lys34→Arg) — students reach for a pharmacological purpose because they assume every part of a drug serves the patient. Item 8 (receptor activity unchanged) — this contradicts the prior almost everyone brings. Item 15 (the "not" item) — negation items on a technique students have just been told was displaced invite the answer "it didn't work."
Exercises — model responses and marking notes
Set A
A1. Proteolysis; renal filtration. Closing one leaves the other open, and a protease-proof peptide cleared renally in twenty minutes is still a twenty-minute drug. Mark for the second half — the common error is naming both exits and then not explaining the conjunction.
A2. Resist degradation (Aib, D-amino acids, cyclization); increase effective size (lipidation, PEGylation, fusion); improve selectivity or potency (cyclization, stapling, substitution at contact positions); enable a route (macrocyclization + N-methylation, peptidomimetics, stapling — contested).
A3. Aib is alanine with a second methyl on the alpha carbon. It occupies space a protease active site needs empty; at position 8 it happens not to be a residue the GLP-1 receptor depends on heavily. Full marks require the "happens not to be" — students who write as though this were predictable have missed §33.2's point.
A4. No codon, no tRNA, no ribosome; therefore chemical (or hybrid) synthesis, with attendant cost, analytics, and supply consequences (Ch 32 §32.6).
A5. Glycine at the position equivalent to 8. It suggests "engineered" and "optimal" are independent — evolution had already produced a DPP-4-resistant GLP-1 receptor agonist in a lizard.
A6. † Both enzyme and receptor bind the same molecule and both are selective; selectivity comes from reading distinguishing features, and a short peptide does not have many. The N-terminus in particular is both a common protease target and, for class B GPCRs, the activation-critical segment. Look for the general argument, not the specific example.
A7. Mammalian proteases evolved to process L-peptides and cannot accommodate the inverted geometry. Risk: the side chain now projects to the opposite face, frequently destroying receptor geometry.
A8. Charge (Lys→Arg keeps charge, removes the reactive handle); hydrophobicity (rings, halogens, extra methylenes — affects aggregation, formulation, skin penetration); conformation (proline analogs, alpha-methylation, constrained rings — pre-pays the binding entropy).
A9. † Open-ended. Strong answers reach for cyclization or D-substitution elsewhere in the sequence, lipidation alone with an accepted shorter half-life, an exendin-based scaffold (the exenatide route), or peptidomimetics. Weak answers assume a different position-8 substitution would simply have been found. Credit engagement with the possibility that no acceptable substitution exists.
Set B
B1. Head-to-tail (cyclosporine); side-chain-to-side-chain; disulfide (octreotide, insulin, cone-snail toxins).
B2. A linear peptide samples many conformations; adopting the bound one costs entropy paid out of binding energy. Constraint pre-pays it, so measured affinity improves without new contacts.
B3. Truncation removes unneeded residues and narrows receptor-subtype coverage; D-residues block proteases; cyclization removes free termini and locks conformation.
B4. † Because a molecule's residues do several jobs at once. Truncating for stability removed residues that also mediated broader subtype engagement. The general point: the four-goal map is a teaching device, not a claim that interventions are separable.
B5. Better binding (best supported, same entropic argument as cyclization); protease resistance (well supported); cell penetration (contested — ⚠️).
B6. † Reasons: fixation artifact redistributing endosome-trapped peptide; total cellular fluorescence not distinguishing endosomal from cytosolic; fluorophore altering the peptide's behavior. Better: live-cell imaging without fixation, a quantitative cytosolic assay (split-protein complementation, chloroalkane penetration), or best of all a functional readout that tracks with dose and is abolished by target mutation.
B7. Size (complex retained by the glomerulus); shielding (buried against a large protein); reservoir (reversible binding releases free peptide slowly, flattening the profile). All three, or partial credit.
B8. Chain length; presence of a terminal diacid carboxylate; spacer chemistry. Liraglutide C16 simple palmitate ~13 h; semaglutide C18 diacid plus longer spacer ~1 week; detemir C14 myristate tuned for a flat ~day-long profile.
B9. So the lipid can occupy albumin's binding site without dragging the receptor-binding face along. Without it, the peptide is tied directly against a 66,000 Da protein and receptor access suffers.
B10. Potency loss from steric hindrance; anti-PEG antibodies; poor metabolism and tissue accumulation with chronic dosing.
B11. † Defense: findings published openly, often by harmed parties; drugs still approved and used; displacement driven substantially by a better alternative. Scandal framing would require concealment, withdrawal, and harm as the driver. Credit students who name what evidence would flip the framing.
B12. FcRn binds Fc at endosomal pH, diverts it from the lysosome, and releases it at blood pH. More powerful than size alone because it is active rescue, not passive retention.
B13. † Cell-culture expression; chromatographic purification; glycosylation/aggregation/host-cell protein characterization; biologics regulatory framework; orders-of-magnitude cost of goods; biosimilar rather than generic competition; poorer tissue distribution; junction immunogenicity. Four or more.
Set C
C1. Non-peptide scaffold reproducing the binding surface. Discards: protease liability, poor oral absorption, high manufacturing cost, inability to enter cells or brain.
C2. Class A receptors bind short ligands in a compact transmembrane pocket — good small-molecule territory. Class B receptors grip a long peptide across an extended extracellular surface while its N-terminus reaches into the core.
C3. † For: the arc runs from protecting a peptide to replacing it, and orforglipron is the logical terminus. Against: a small molecule at a peptide receptor is medicinal chemistry, not peptide engineering; the peptide only supplied the target. Both are defensible; mark the reasoning.
C4. Ala8→Aib (proteolysis); Lys34→Arg (manufacturing); C18 diacid at Lys26 via spacer (filtration).
C5. Aib closes proteolysis; the diacid closes filtration; Lys34→Arg closes neither.
C6. Two lysines means two possible attachment sites and a mixture of isomers. Arginine keeps the charge but is not a nucleophile under acylation conditions, leaving one reactive amine.
C7. † Open. Strong answers note that reproducible manufacture is a precondition for a drug existing at all, so the framing "compromise vs. excellent design" may be a false dichotomy.
C8. Unchanged. Lesson: the receptor pharmacology was never the problem — evolution had optimized it. What was engineered was survival.
C9. A meal signal must terminate promptly or the organism loses the fed/fasted distinction.
C10. See the §33.9 table. Fixed vs. adjustable ratio; one vs. two PK profiles; one vs. two stability programs; one vs. two immunogenicity questions; new molecule vs. new formulation.
C11. † It establishes that one product outperformed another product on one endpoint in one population. It does not establish that fixed-ratio architecture is superior — the comparison confounds architecture with the specific ratio, the specific molecules, and the specific doses studied.
Set D
D1. Degradation and clearance (engineerable); desensitization and downregulation (not).
D2. Desensitization and downregulation are provoked by sustained receptor stimulation. A drug designed for continuous week-long occupancy delivers exactly that stimulus pattern; the native pulse is short partly so adaptation never fully engages.
D3. Wrong target (substance P, Ch 22); unfavorable therapeutic window; surrogate endpoint that does not predict outcome (Ch 16); absence of clinical evidence.
D4. † Rule: engineering solves delivery problems and only delivery problems. Best candidate counterexamples: cyclization improving affinity (Goal 3), and Lys34→Arg (manufacturing, not delivery). Strong answers concede the rule is a strong generalization with two known exception classes, and that neither exception is about efficacy at the target.
D5. Check against §33.8 and the Dossier worked demonstration.
D6–D8. Individual work; no fixed answer. For D6, the assessment is whether the student recorded what they could not find.
D7. † Questions should map to the four goals: which exit does this close, what is the measured half-life and where does the number come from, and what does "enhanced affinity" mean relative to what comparator at what exposure. Predicted answers: no data, no number, no comparator.
D9. † §33.3 (stapling, ⚠️ — contested, assay artifacts, endosomal escape) and §33.7 (peptidomimetics, ⚠️ — real but not proven at scale; note this route abandons the peptide rather than smuggling one in). Credit students who notice the two push in opposite directions: one keeps the peptide and tries to get it inside, the other gives up the peptide entirely.
D10. Must contain both halves — achievement and boundary. A summary that only celebrates engineering, or only debunks it, is incomplete.
Chapter 34
For the instructor answer key. The quiz answers are published in quiz.md under a collapsed
answer key; they are restated here in compact form for assembly. The exercises have no published
answers by design — the model responses below are for grading, not for distribution.
Quiz — answer letters
| Q | Ans | Q | Ans | Q | Ans | Q | Ans |
|---|---|---|---|---|---|---|---|
| 1 | b | 7 | b | 13 | c | 19 | b |
| 2 | c | 8 | c | 14 | a | 20 | b |
| 3 | c | 9 | b | 15 | b | 21 | b |
| 4 | b | 10 | c | 16 | b | 22 | c |
| 5 | b | 11 | b | 17 | c | ||
| 6 | c | 12 | b | 18 | b |
Full rationales are in quiz.md. Items most often missed in practice: 3 (students accept that MS
detects "wrong molecules" without noticing that stereochemistry is not a mass property), 6
(students pick a peptide impurity rather than a species with no chromophore at all), 9 (students
choose adulteration and miss that this outcome requires no dishonesty), and 21 (students blame the
test's reliability rather than its destructiveness).
Exercises — model responses
Part 1 — the three questions
A. Identity: is this the molecule the label names. Purity: of the material the analysis detected, what proportion is that molecule. Content: how many milligrams of that molecule are physically present.
B. Identity — MS, MS/MS, sequencing, amino acid analysis (composition only). Purity — HPLC, ideally by orthogonal methods. Content — quantitative assay against a reference standard, amino acid analysis, or nitrogen determination.
C. Which technique and mode; which detector and wavelength; whether specificity was demonstrated; whether related species co-elute; what fraction of the vial mass is peptide at all; identity; sterility; endotoxin; water; counterion; whether the number describes this vial. Credit breadth, especially any answer that separates "things a better purity method would tell you" from "things no purity method can tell you."
D. † Purity is a ratio among detected peaks; content is a mass. Counterion, residual water, and salts contribute mass that a purity calculation never examines. Both statements can be true with no error and no dishonesty anywhere in the chain. Strong answers note that gross powder weight versus peptide weight is a specification question the document may simply not address.
E. Diagram: peptide fraction | counterion | water + salt, with purity examining only the internal composition of the peptide fraction.
F. More basic residues means more positive charges at the relevant pH, requiring more counterions for charge balance, each contributing its own mass.
G. † Error one: "percent" is a percent of detected peak area, not of sample mass. Error two: species the detector does not respond to are absent from both numerator and denominator, so they are not counted as impurities at all.
H. Purity is a ratio internal to one chromatogram, so it needs no external reference. Content is an absolute quantity, and an absolute quantity requires something of known quantity to compare against — either a characterized peptide standard or, via AAA, amino acid standards.
Part 2 — identity
I. MS: mass-to-charge ratio of ions, from which neutral mass is derived. MS/MS: additionally the masses of fragments produced from an isolated ion, from which sequence can be read.
J. ESI produces multiply charged ions, so one peptide appears at several $m/z$ values. Software deconvolutes the charge-state envelope to a single neutral mass.
K. Monoisotopic uses the most abundant isotope of each element; average uses abundance-weighted means. A reader comparing a measured monoisotopic value against a calculated average value (or vice versa) may conclude the mass "doesn't match" when the only mismatch is convention. Larger peptides diverge more.
L. † Mass is a function of atomic composition, and L and D forms have identical composition. Shape changes; mass does not. Detection requires chirality-sensitive methods, which are not part of routine identity testing. Strong answers note this is a structural blindness, not a resolution limit.
M. MS/MS sequencing cannot discriminate Leu from Ile, so an MS/MS "sequence confirmation" is strictly a confirmation up to that ambiguity. Also worth crediting: the same argument applies with less force to near-isobaric pairs at low resolution.
N. Signal intensity depends on ionization efficiency, which varies enormously between compounds. A dominant peak means a dominant signal, not a dominant mass fraction.
O. † (i) < (iii) < (ii) < (iv) is defensible, but so is (i) < (ii) < (iii) < (iv). The ranking is less important than the justifications. Required: a label is not an analytical claim; mass constrains strongly but tolerates isomers; composition rules out many molecules but leaves order free; two orthogonal methods reduce the chance that a single method's blind spot is decisive. Best answers argue explicitly for their placement of (ii) versus (iii) rather than asserting it.
P. Acid hydrolysis destroys tryptophan. Its absence in the report may be an artifact of the method rather than of the molecule.
Q. Acid hydrolysis deamidates Asn to Asp and Gln to Glu. Reports give Asx and Glx — the sums.
R. † For n distinct residues there are n! orderings; composition therefore leaves an astronomically large space of candidate molecules for peptides of realistic length, and a molecule of right composition and wrong order is generally biologically inactive. The "something else entirely" is quantitation: AAA is a primary method for peptide content, because it needs only amino acid standards.
Part 3 — purity
S. Two species with similar affinity for the stationary phase emerge together and are recorded as one peak. Every separation method resolves by some property, so any two species alike in that property will co-migrate; the limitation is intrinsic to separation, not to a technique.
T. Would count: a peptide impurity containing tryptophan, tyrosine, or phenylalanine. Would not count at all: an inorganic salt, most residual solvents, and a peptide impurity lacking aromatic residues.
U. † Deletion sequences differ from the target by one residue out of many, so their hydrophobicity and hence retention are nearly identical; unresolved area is counted inside the target peak, inflating reported purity. The coincidence — that the method's characteristic failure produces the assay's characteristic blind spot — means a high area-percent figure provides less reassurance for a synthetic peptide than the same figure would for a compound whose impurities are chemically dissimilar.
V. The relationship between mass and signal. Because it differs between compounds, equal areas do not represent equal masses, so area percent is not a mass percent.
W. Methods whose separation or detection principles are independent. Co-elution in one is unlikely to be reproduced in another, so orthogonality addresses a limitation that improving a single method cannot — no matter how good one method is, it cannot demonstrate the absence of something invisible to its own principle.
X. † Reversed-phase separates by hydrophobicity, so species differing only in charge — for example a deamidation product — may separate poorly. Ion-exchange separates by charge and is more likely to resolve it. Accept size-exclusion for aggregates as an alternative pairing.
Part 4 — sterility, endotoxin, process
Y. Sterility asks whether anything is alive in the vial; endotoxin asks whether anything ever was. Endotoxin is a heat-stable structural molecule, not an organism, so killing or excluding organisms does not remove it.
Z. They measure different things. Sterility detects viable organisms; endotoxin is cell-wall material that persists after the organism is dead or removed.
AA. † LPS molecules and fragments are orders of magnitude smaller than a bacterium, so a filter sized to retain bacteria passes them. Autoclave temperatures do not destroy LPS. The "worse in one respect" point: lysing bacteria releases endotoxin from the cell envelope into solution.
AB. No visual, olfactory, or solubility signature, and it passes a sterility test — so no inspection detects it. Presentation: fever and rigors within hours of administration, often with headache, malaise, tachycardia, sometimes hypotension; commonly misattributed to coincidental viral illness. Medical evaluation is the appropriate response.
AC. † (i) Determining sterility requires opening the unit and attempting culture, consuming it. (ii) Therefore released units are never the tested units; one tests others and infers. (iii) If contamination is a small scattered fraction, most modest random samples contain none, so the test passes while the batch is not sterile — a correct measurement supporting a wrong inference. Credit any answer noting that increasing sample size does not rescue the approach, because each additional test destroys product.
AD. Batch-to-batch consistency, traceability, freedom from cross-contamination between products on shared equipment. "Property of a process" means the property is constituted by how the material came to be rather than by the material's composition, so no measurement of the material can reveal it.
Part 5 — certificates, programs, standards
AE. Can establish: identity by MS; purity by HPLC with method stated; peptide content by quantitative assay; water content; counterion content; sterility and endotoxin if performed. Cannot establish: correspondence to the vial in hand; handling after testing; anything the method did not look for; its own authenticity.
AF. † Principle: a certificate is a claim about a sample, not a property of a vial. (1) follows because samples and vials are different objects and nothing in the document links them. (2) follows because a claim about a past sample is not updated by later events. (3) follows because a claim reports what was measured and silence is not a finding. (4) follows because a document is an assertion by an interested party and contains no evidence of its own provenance.
AG. A predetermined acceptance criterion the material must meet, set before testing. Without one, a result cannot pass or fail, so the document is a description rather than a release decision.
AH. Suggested: Tested for what — identity, purity, content, sterility, endotoxin? On what sample, obtained by whom? Against what specification, and does the document trace to the container I have?
AI. † The four reasons: non-random sampling; a shifting market; inconsistent definitions and methods; publication bias against clean results. On the objection: an approximate number is better than none only when the error is bounded and roughly unbiased. Here the selection mechanisms are correlated with the outcome and the bias direction is not known, so the "approximation" has neither property. Best answers add that the absence of a rate is not reassuring, because it means exposure cannot be calculated.
AJ. † A monograph defines the specification — what a substance must be and meet. GMP is a system of documented process control establishing that a process reliably produces conforming material. A batch meeting a monograph says nothing about the next batch, because conformance was demonstrated by testing a sample rather than by controlling a process. Same shape as AC: both distinguish a property sampled in a product from a property built into a process, and in both cases sampling cannot create what the process did not.
Chapter 35
Peptide Discovery — From Venom to Artificial Intelligence
For the instructor answer key. The chapter's exercises.md deliberately ships without answers;
these notes are what a marker needs. Many items admit more than one defensible response, and the
notes below say where.
Quiz answers live in the collapsed key inside quiz.md and are summarized here only as a letter
string for fast marking.
Quiz — answer string
1 b · 2 c · 3 b · 4 c · 5 d · 6 c · 7 short · 8 c · 9 c · 10 b · 11 b · 12 c · 13 b · 14 b · 15 c · 16 c · 17 short · 18 b · 19 a · 20 c · 21 b · 22 short
Short-answer items (7, 17, 22) are marked against the model responses in the quiz.md key.
Distractors worth discussing in class:
- Q2 (d) and Q9 (a) both fail on the naturalness fallacy. If students pick these, the problem is Chapter 1, not Chapter 35.
- Q5 (b)/(c) attribute semaglutide's albumin-binding solution to exendin-4. This is the single most common conflation in the chapter and is worth catching explicitly.
- Q20 (c) is reversed on purpose. Students who miss it have absorbed "AI is limited" without absorbing which limitation applies to peptides.
- Q21 (c) is the cynical over-correction. Watch for students who swing from credulity to dismissal; the chapter's whole posture is that the technical claim is ✅ and the framing is ⚠️.
Exercises — marking notes
Part 1 — the four routes
A. Look for: find in nature / modify an endogenous ligand / select from a library / design computationally. Full credit does not require the numbering.
B. insulin → route 1 (isolated from pancreas); semaglutide → route 2; exenatide → route 1; ziconotide → route 1; phage-display peptide → route 3; RFdiffusion binder → route 4. Accept "exenatide = route 1 with route 2 successors" as a stronger answer.
C. † Expected ranking (least to most prior knowledge): 1, 3, 4, 2. Route 2 requires knowing the hormone, the receptor, and a reason to want more signal — and has the best hit rate. The insight to reward: knowledge substitutes for luck, and the routes that demand the most up front are buying down risk, not showing off. A strong answer notices this also caps route 2's upside — it can only reach targets the body already signals through.
D. How do I obtain a molecule that binds this target? / Will binding this target help a sick person?
Part 2 — venom
E. Fast onset, potency at minute quantities, action on vertebrate nervous/cardiovascular/muscular systems. Reward students who state the punchline: that is the design brief of a drug.
F. The metabolic-economics argument. Venom is expensive and slow to regenerate; a broad-spectrum weak binder would need bulk delivery; a single-target picomolar binder does the job cheaply.
G. Multiple disulfide bonds staple the chain into a rigid shape. Proteases generally need to thread an extended stretch of backbone through the active site, and a stapled molecule offers very little extended backbone.
H. Not solved by venom: oral bioavailability (still too large and polar to cross the intestinal wall), renal clearance, blood-brain barrier penetration, immunogenicity. Any two.
I. Shares: roughly half its residues, including the N-terminal region that engages the receptor → activates the human GLP-1 receptor. Does not share: the residue at the DPP-4 cleavage position → resists the enzyme. Do not accept a stated percentage identity — the chapter says "roughly half" deliberately.
J. DPP-4 recognizes a specific side chain at a specific position; receptor activation depends on a broader set of contacts. Structural level: primary structure (sequence), affecting recognition rather than the overall fold. Accept an answer that frames it as sequence-level change with no change in the receptor-engaging conformation.
K. † The rebuttal must contain both clauses: exenatide's duration was inadequate, and its successors were engineered. The reflective half of the question is the real one — the pro-nature paragraph is easier to write because it is a story with an agent and a moral, while the rebuttal is a qualification. Reward students who notice that narrative shape, not truth, predicts circulation.
L. Shortcoming: renal clearance / inadequate duration. Solution: fatty acid attached via linker, binding serum albumin, extending half-life from hours to days.
M. Bradykinin-potentiating peptides (peptide) → teprotide (peptide, injectable) → design work on the enzyme's active site → captopril (not a peptide, oral).
N. † Both diagrams: venom peptide → injectable peptide drug → small-molecule drug. Differences worth credit: captopril's target is an enzyme with a definable active site while the GLP-1 receptor is a class B GPCR with a large peptide-binding interface, which is why the second arc took decades longer; also the second arc's peptide drugs remain first-line rather than being superseded. A strong answer says the differences sharpen rather than undermine the parallel.
O. Target (N-type calcium channels) is in the spinal cord; the molecule does not cross the blood-brain barrier; systemic dosing sufficient to reach it would produce unacceptable systemic calcium-channel effects.
P. † The argument: here is a peptide whose central effects are real, valuable, and clinically established, and the medical system's answer to getting it into the CNS was an implanted pump. If casual BBB penetration were available to peptides, that engineering would be unnecessary. Reward students who note this is an existence argument about cost, not a proof that no peptide ever crosses.
Q. Frogs returned to unsterile water after surgery healed without infection; something in the skin was antimicrobial. Opened the antimicrobial peptide field.
Part 3 — rational design and display
R. Any four of: insulin, GLP-1, somatostatin, GnRH, vasopressin, parathyroid hormone.
S. Any well-developed example. GLP-1 → nausea and delayed gastric emptying; somatostatin → gallbladder motility and glucose effects; GnRH agonists → initial flare. The key sentence: these are on-target effects of the parent signal, not contamination or impurity.
T. † Three sentences, plain language, no "on-target." Model: GLP-1 receptors sit in your gut and in parts of your brain that control appetite and nausea. The drug turns those on everywhere at once, and the appetite effect and the nausea effect come from largely the same switch. You can sometimes reduce the nausea by changing how quickly the drug is introduced, but you cannot remove it without weakening the effect people want. Mark down any answer implying the nausea is a manufacturing or purity problem.
U. The peptide is displayed on the phage surface; the DNA encoding it is inside the same particle. Without naming "sequence": the molecule and its instructions travel together, so a molecule selected by binding can afterwards be identified.
V. Cell-free systems avoid the transformation-efficiency cap on how many distinct variants can be introduced into bacteria.
W. † The line: the researcher designs the search, not the molecule. Nothing in the procedure reasons about which side chain should point where; no model of the target is used; the filter is physical. A strong answer concedes that library design encodes real chemical judgment (constrained scaffolds, macrocycles) and holds the line anyway — design of the search space is not design of the hit.
X. (1) Functional assay or binding only — rules out inert occupancy; (2) counter-selection — rules out plate/tag/linker artifacts; (3) cellular and animal activity — rules out irrelevance in a biological context; (4) stability and route — rules out undeliverability.
Part 4 — structure, prediction, design
Y. Membrane receptors must be pulled from the bilayer, kept folded in detergent, and made to sit still in an ordered lattice; class B GPCRs resisted. They are the targets of a large share of this book's drugs. Cryo-EM produced activated receptor–ligand–G-protein complexes, previously unobtainable.
Z. (1) structure ≠ function; (2) short peptides often have no single structure; (3) predicting a structure ≠ predicting a drug; (4) predictions are predictions, least confident where disorder is highest. Most peptide-specific: (2). Reason: this book's molecules are precisely the flexible, fold-on-binding class.
AA. † Malformed because there is no single answer to predict — the conformation is a property of the molecule plus its partner, not of the molecule alone. The trap: the tool returns something anyway, and a returned structure looks like an answer. Reward students who connect this to §1.4's random-coil discussion.
AB. Correct behavior: a well-calibrated model should report low confidence about a region that has no defined structure. Inconvenient: disordered regions mediate much regulatory interaction and host many modifications, so the map is least reliable where a lot of biology lives.
AC. It searches sequence space by design rather than by inheriting a search (route 1), walking locally from a known point (route 2), or sampling blindly and filtering (route 3).
AD. † Look for a properly formed four-line rating with a population and endpoint, a one-sentence reason, and a falsifiable change condition. The ✅-versus-⚠️ distinction should turn on multiple adequately powered trials with consistent results and a known safety profile, not on enthusiasm or on the number of molecules in development.
Part 5 — the bottleneck
AE. † Three cases. Reward specificity and reward the case where the improved stage was the constraint — students who cannot produce one have turned the rule into a general skepticism about technology, which is the failure mode this exercise exists to catch.
AF. Lack of efficacy in humans (the target hypothesis was wrong, or wrong for that population) and unacceptable toxicity. Neither is a candidate-supply problem.
AG. More shots on goal; economically marginal targets (rare disease, neglected infection, antibacterials) become viable. Not supported: that any individual attempt is more likely to succeed, or that time-to-approval shortens.
AH. † The honest answer: the discovery would have been faster and cheaper; the Phase 2 and Phase 3 results would have been identical. Reward students who state that explicitly — the target hypothesis was the failure, and no discovery method addresses a wrong hypothesis. A very strong answer notes the outcome might have been worse: cheaper candidate generation could have produced more compounds against the same wrong target.
AI. † Marking the specification: all six components present; both thresholds written; the "which component carries the weight" answer defended rather than asserted. For the advertisement half, accept any real example; the graded part is the three-part breakdown (origin claim / implication invited / missing evidence).
Case study discussion — marking guidance
Case 35.1 (Gila monster). The failure mode is a student who lands on either "nature knows best" or "the lizard was luck we no longer need." Full credit requires both clauses of nature supplied the lead; chemistry supplied the drug. Q3 is the sharpest item: the distinction between "venom is a good place to look" (✅) and "a venom origin is evidence about a specific compound" (not a claim the evidence supports at all) is the chapter's dossier point in miniature.
Case 35.2 (AlphaFold). The failure mode is treating the exercise as debunking. Insist that students state the ✅ for the technical claim before they state the ⚠️ for the framing. Q2 (the "would it have been reported the other way?" test applied to CASP) is the item that most reliably distinguishes students who have internalized Chapter 5's method from those who have memorized its vocabulary.
Chapter 36
Scope. Model responses for the daggered (†) items and the odd-numbered items (counting a = 1, b = 2, c = 3, and so on through mm = 39). Several of these ask for judgment rather than recall; where that is the case the answer shows the reasoning and says explicitly what a different defensible answer would look like. Dated to the chapter's frame: as of this writing, in 2026.
Part A — The framing problem (§36.1)
a. (item 1)
Model answer. A rating attaches to a claim about a specified effect, in a specified population, on a specified endpoint, and it is assigned by weighing the existing evidence for that claim; for a compound that has not completed its trials the relevant evidence does not exist yet, so there is nothing for a rating to weigh.
What to notice. The word doing the work is exists, not difficult. Prediction being hard would be a reason for a wide error bar. Evidence being absent by construction is a reason there is no rating at all. Students who write "because the future is uncertain" have given the wrong reason for a right conclusion, and the wrong reason licenses a hedged prediction, which is exactly what the chapter refuses to make.
c. (item 3)
The argument. Thirty-five chapters of hedged, evidence-anchored analysis build credibility. In the future chapter the constraint of having to cite something disappears, and if the author cashes the accumulated credibility in on speculation, the reader reasonably transfers trust: this author was careful everywhere else, so this must be careful too. It is not. And the transfer runs backward as well: abandoning the evidence standard the moment evidence becomes unavailable reveals that the standard was never doing the work — the availability of evidence was. The discipline is exposed as a style rather than a commitment, which retroactively devalues every earlier ✅ and ❌.
One reason to disagree. Readers are not that naive. Most people already read future-facing writing in a different register, the way they read a weather forecast rather than a lab report, and can discount a speculative chapter without discounting the empirical ones. On this view the chapter's worry is a little self-important, and a livelier future chapter would cost nothing and inform more.
A reasonable resolution is that the objection holds for readers who are already calibrated and fails for exactly the readers the book was written for — the ones learning the discipline, who have no independent way to tell which register they are in. The cost of being dull is small; the cost of being wrong about which audience you have is not.
d. † (item 4)
Using retatrutide, since §36.5 supplies the material.
Rateable today. "Retatrutide produces clinically meaningful weight loss in adults with obesity." This names a population (adults with obesity), an endpoint (weight loss), a direction, and an implicit magnitude threshold ("clinically meaningful") that can be argued about and specified. There is substantial late-stage weight-loss data as of 2026, so evidence exists to weigh, and the chapter rates it ⚠️ — promising but preliminary, with an explicit statement of what would move it up or down.
Not rateable today. "Retatrutide will transform obesity treatment." No population, no endpoint, no comparator, no timeframe, and no threshold for "transform." There is nothing an observation could contradict, so there is nothing for evidence to bear on. This is the failure §36.2 rates as a claim form: ❌ — hype outpaces evidence, as usually stated — and note that the ❌ describes the evidence available for the claim as constructed, not a prediction that the drug will disappoint.
A third, instructive case. "Retatrutide reduces cardiovascular events more than currently approved agents." This one is well constructed — population, endpoint, comparator, direction, all present — and is still not rateable ✅ or ❌, because the trial has not read out. It gets 🔬: a serious question under investigation with no answer as of this writing.
So what exactly is the difference? Two distinct things, and separating them is the point of the exercise:
- Construction. A claim must name a population, an endpoint, and (for comparative claims) a comparator, so that some observation could refute it. Claims failing here fail before anyone looks at evidence — rating rule 5, falsifiability.
- Availability of evidence. A well-constructed claim can still have no evidence yet, in which case it is 🔬 rather than ✅ or ❌.
The badly constructed claim can never be rated, at any future date, no matter what happens. The well-constructed one is merely waiting. The instant someone supplies endpoint, population, and comparator, an unrateable sentence becomes a gradeable one — which is why the useful response to "this will be transformative" is not disagreement but a request for the missing three.
e. (item 5)
The two kinds of statement.
- Present-tense status claims. "Retatrutide targets the GLP-1, GIP, and glucagon receptors, and phase 3 is underway as of this writing." Checkable against trial registries, regulatory filings, and the published literature. If it is wrong, that is an error the author made, not a prediction that failed.
- Reading technique. "Ask what phase, what endpoint, what population, what comparator, who is claiming, and what the base rate is." Checkable in a different sense — a reader can apply the procedure to a real announcement and see whether it discriminates good claims from bad ones.
How a reader checks each. For the first: look the compound up in a registry and read the study record. For the second: run the six questions on three announcements you have already formed views about and see whether the procedure changes any of them. If it never changes anything, either you were already doing it or it is not working.
Part B — The six questions (§36.2)
g. (item 7)
- Preclinical. A positive result licenses you to believe the compound does something in cells, tissue, or animals; it does not license any belief about humans, because a great many things work in a mouse (Chapter 5). Preclinical is not "almost in trials"; it is a different category of knowledge.
- Phase 1. A positive result licenses you to believe the compound was tolerated at the doses given in the people studied and that its pharmacokinetics behaved roughly as designed; it does not license a belief that it helps anyone, because the trial was generally not built to answer that.
- Phase 2. A positive result licenses you to believe there is a real efficacy signal worth testing properly; it does not license you to expect the same magnitude in phase 3, because phase 2 results are systematically more favorable than what follows (Chapter 9).
- Phase 3. A positive result licenses you to believe the effect is real and approximately the size reported, in the population studied, on the endpoint measured, against that comparator; it does not license extrapolation to a different population, endpoint, or comparator, and it says nothing about effects too rare for the trial to have counted.
i. † (item 9)
Two reasons the refusal is defensible.
- The number is not stable enough to quote. It varies by therapeutic area, by how you count (per program? per indication? from which phase?), and by which decade you sample. A single figure would create false precision about something the chapter is using only as a direction and a magnitude — most fail, and this holds even in phase 3.
- A quoted number invites the wrong use. Readers convert a percentage into an implied probability for the specific compound in front of them, which is not what a base rate is for. The base rate is a starting position that a claim has to move you off, and a bare "roughly nine in ten fail" gets applied as though it were this compound's individual odds.
Why a reader might find it frustrating. A book that spends the whole chapter demanding specificity declines to be specific at the one place a number would be genuinely useful, and "most" is doing a lot of unexamined work. A reader who wanted to calibrate — to know whether "most" means six in ten or nineteen in twenty — has been told to be precise and then handed a vague word. That is a fair complaint, and the honest reply is that the imprecision is deliberate and the reader can look up the figure for their own area of interest, which is the better version of the exercise anyway.
k. (item 11)
The difference. Beating placebo establishes that the drug does something — that its effect is not attributable to the natural course of the condition, regression to the mean, or the act of being in a trial. Beating standard of care establishes something much stronger and much more decision-relevant: that it is better than what a person could already have. A drug can beat placebo decisively and still be the wrong choice for every patient, because something already available is better.
Why the sacubitril/valsartan result is the one cited. Because it is the clean case of the harder test being passed. The comparator was an active drug already known to work, the endpoint was a hard outcome rather than a surrogate, and the population was large. All three at once is rare, and it is why that trial changed practice rather than generating headlines. Citing a placebo-controlled trial here, however impressive, would have illustrated the weaker standard the chapter is warning about.
Extension worth making: hold that standard against §36.5. A new multi-agonist versus placebo on weight loss and the same compound versus tirzepatide on cardiovascular events will be reported in nearly identical language and are not remotely the same claim.
m. † (item 13)
The strongest objection, stated fairly. Drug development is not a lottery. Programs grounded in well-validated targets, with human genetic support and a clear causal chain, do succeed at higher rates than programs built on a thin rationale. Refusing to update on mechanism throws away real information and would have you treat a well-understood receptor agonist and a compound with a hand-waving story as equivalent bets. That is obviously wrong.
Where the objection has force. It has force at the level of classes and portfolios. Target validation quality does predict success across many programs, and the chapter concedes the substance of this when it lists what does move you: precedent in the same mechanistic class — not the story but the track record of approved drugs working that way. A mechanism with drugs already on the market is evidence. A mechanism with a beautiful diagram is not.
Where it fails. It fails at the level of this compound, in front of you, right now. The reason is selection: a mechanism story is present in essentially every candidate, including all the failures, because nobody funds a phase 1 with no rationale. A feature that is universally present in both the successes and the failures cannot discriminate between them. So the mechanism story you are reading carries close to zero information about whether this particular program succeeds, even though mechanism quality, measured across many programs and against a track record, carries some.
The reconciliation. Distinguish the story from the record. "This is how it would work" is a story and moves nothing. "Four approved drugs act this way and three of them met their primary endpoint" is a record and moves something. Chapter 5's discipline is the boundary: knowing how something would work is not evidence that it does.
Part C — Constructed announcements
o. (item 15)
Working from the item n release. Phrases doing rhetorical rather than informational work:
- "POSITIVE." Invites the inference that the result was good. It licenses only that the primary endpoint was met at whatever threshold was prespecified — which is compatible with an effect too small to matter to anyone.
- "NOVEL." Reads as advanced. It means only "not previously approved," which is true of every failure as well.
- "Met its primary endpoint." Sounds like a verdict. Without the endpoint named, it is content-free: the endpoint could be a biomarker nobody would notice changing.
- "Statistically significant." Widely misread as "large" or "important." It speaks to the compatibility of the data with no effect, not to magnitude or clinical relevance.
- "Versus placebo." Presented as validation; it is the weaker comparator, and its presence quietly answers question 4 in the unfavorable direction.
- "Validate our mechanism." The rhetorical center of the release. It invites the mechanism-to- confidence upgrade that rating rule 3 forbids. A phase 2 signal is not validation of anything, and mechanism does not move the base rate in either direction.
- "Position us to advance rapidly." A statement about corporate intention, not about the compound. Every program intends to advance rapidly.
- "Chief executive." Question 5, visible in the byline. This is the least independent voice available, speaking to an audience with a financial interest.
- "Full results will be presented at a future scientific meeting." Reads as a promise of transparency; functions as a reason you cannot check anything today, at an unspecified date, in a venue that is not peer review.
The through-line: every one of these phrases is defensible in isolation and the release contains no false statement. The distortion is produced by selection and connotation, not by lying, which is why "is it accurate?" is the wrong question to ask a press release.
p. † (item 16)
Question by question.
- Phase. "First-in-human" is phase 1 — the earliest human stage, small, typically healthy volunteers, and years from an efficacy answer even in the best case.
- Endpoint. Tolerability and pharmacokinetics, plus "exploratory biomarkers." No clinical endpoint of any kind appears.
- Population. Unstated, and in first-in-human work usually healthy volunteers — people who do not have the condition and in whom improvement is therefore not measurable.
- Comparator. Absent. Dose cohorts are compared to each other, not to a control.
- Source and timing. "Presented today" — a conference presentation or abstract, not a peer-reviewed publication, with no full dataset available.
- Base rate. Attrition from phase 1 is severe; this is the stage at which the prior is least moved by anything the trial could have shown.
"Trends." A trend is a difference that did not meet the threshold for statistical significance. The word is doing the opposite of what it appears to do: it signals the absence of the finding it seems to report. In a phase 1 with small cohorts and no control arm, a "trend" is close to uninformative.
"Exploratory." Exploratory endpoints are not prespecified as the trial's tests. With many measures examined, some will move favorably by chance alone, and reporting the ones that did is expected rather than suspicious — which is exactly why exploratory results do not support conclusions. The word is an honest label, and most readers hear it as a modest synonym for "early."
"Pharmacokinetics supporting once-monthly dosing." A real engineering finding, and note what it is not: evidence about efficacy, and — per §36.4 — not automatically desirable either, since a monthly interval trades away the ability to stop and the ability to titrate.
Conclusion. The announcement is entirely accurate and supports approximately one belief: the compound was given to humans and nothing alarming was observed at the doses used. That is a real milestone and it is not evidence that the drug works.
q. (item 17)
Question by question.
- Phase. Not stated in words, but a randomized trial of 4,200 adults powered on a hard composite endpoint against an active comparator is phase 3 in everything but name.
- Endpoint. A composite of cardiovascular death, myocardial infarction, and stroke — hard outcomes, not surrogates. Every component is something a person would notice.
- Population. "Adults with established disease" — enriched for event risk, which is appropriate, and the point at which to ask the follow-up questions: which disease, what severity, were older adults included, what were the exclusions.
- Comparator. An active comparator already recommended in guidelines. This is the strongest available answer to question 4.
- Source. Simultaneous peer-reviewed publication — the top rung of the ladder, with methods visible and data tables available.
- Base rate. Largely discharged. The base rate is a prior about compounds that have not yet read out; this one has, at the stage where most surprises happen.
Relative to item n. Not comparable. Item n is a sponsor-selected summary of a mid-stage surrogate result against placebo, with no full data. Item q is a completed hard-outcome trial against an active comparator with the paper available. These are different categories of knowledge, and popular coverage would describe both as "promising new drug data."
The two features doing the most work: the hard composite endpoint and the active comparator. Either alone would be strong; together they answer the question a clinician actually has, which is not "does this do anything?" but "is this better than what I would otherwise give?"
What is still not established — worth saying, so the answer does not overcorrect: long-term safety, performance outside the enrolled population, effect size relative to cost and burden, and anything about subgroups the trial was not powered to resolve.
s. † (item 19)
A compliant [constructed teaching example].
In a randomized, double-blind phase 2b trial, 340 adults aged 40 to 75 with [condition] and baseline [severity marker] in a specified range received either the candidate or [an active comparator currently recommended in guidelines] for 26 weeks. The prespecified primary endpoint was [named surrogate], chosen because the trial was not sized to measure clinical events; the candidate produced a greater change than the comparator, of a magnitude whose clinical relevance is not established. Adverse events leading to discontinuation occurred more often in the candidate arm. Adults over 75, and those with [common comorbidity], were excluded, so the result does not describe them. Full results are submitted for peer review and the protocol and statistical analysis plan are posted. A phase 3 trial powered for clinical events would be required to establish benefit; most compounds at this stage do not reach approval.
Every one of the six questions is answered in the text: phase, endpoint (with its surrogate status named), population (with exclusions), comparator, source, and the base rate stated against the company's own program.
Why real releases rarely read this way — reasons that are not dishonest.
- Regulatory constraint. Sponsors are limited in what they may say about an unapproved product, and detailed efficacy discussion ahead of approval carries genuine legal risk.
- Publication convention. Journals and conferences may object to full disclosure before presentation, so "full results to follow" is often a real obligation rather than a dodge.
- Timing. Topline genuinely means topline — a locked database with the full analysis still weeks or months out. Some detail does not exist yet.
- Audience. The primary audience is investors, who need to know quickly and materially that a program advanced; the format evolved for that job.
Reasons that are dishonest, or at least not innocent.
- Selection. Choosing which numbers appear, and omitting discontinuations, subgroups, and the comparator's performance.
- Connotative loading. "Validates," "positive," "breakthrough," "rapidly" — words that carry inference without carrying content.
- Indication drift. Describing trial activity in one indication while discussing prospects in another.
- Base-rate suppression. No release has ever included the sentence "most compounds at this stage do not reach approval," and every one of them could.
The useful conclusion: the compliant version is not impossible, merely commercially unattractive. The gap between what a release could say and what it does say is itself the measurement.
Part D — Oral delivery and duration (§36.3, §36.4)
u. (item 21)
(i) Tablet content. If roughly one percent of what you swallow reaches circulation, an oral tablet must contain far more drug substance than an injection producing a comparable effect — a large multiple, not a modest one. Higher-dose oral formulations have been developed and studied, which is the expected engineering response and makes this arithmetic more demanding rather than less.
(ii) Manufacturing capacity. Every one of those milligrams is made by solid-phase synthesis (Chapter 32): linear, stepwise, solvent-intensive, with yield losses compounding across cycles. So the oral formulation consumes vastly more manufacturing capacity per patient-year than the injectable does. Capacity cannot be conjured by running a reactor harder; it requires more reactors, more solvent handling, more purification, and years plus regulatory qualification.
(iii) Supply. More synthesis demand per patient means the synthesis bottleneck tightens as oral use grows — while the fill-finish bottleneck (§36.9) loosens, because tablets do not need sterile filling into injection pens. Oral and injectable versions of the same molecule can therefore differ substantially in availability and price for reasons that have nothing to do with efficacy.
Which a reader is most likely to meet: (i), because "you can take it as a pill" is the headline. Which is most consequential: (ii), because it governs whether anyone can actually get it. The chain from a pharmacokinetic parameter to a supply constraint is invisible in coverage and is the whole story of who receives treatment.
w. † (item 23)
A defensible ordering, with the reasoning made explicit, since this is a judgment item.
1. Conventional small-molecule manufacturing cost. The binding constraint on access in low-income settings is almost always price, and a change in the cost of goods propagates to procurement, to national formulary decisions, to donor programs, and to whether generic manufacturers can enter at volume. Nothing else on the list moves as many people.
2. No cold chain. The second-largest effect and in some settings the largest. Refrigerated distribution is the reason many effective products stop at the last well-resourced hospital. Removing it changes not just cost but geography — which facilities can stock the product at all, and whether a person has to travel to receive it. It is ranked second only because a cheap product that still needs refrigeration reaches more people than an unaffordable one that does not.
3. No absorption enhancer. Real but largely subsumed. Its consequences — simpler formulation, fewer administration conditions, less material per dose — mostly express themselves through cost, which is already ranked first, so counting it separately double-counts.
A defensible alternative ordering puts cold chain first, on the argument that price can be addressed by policy, donation, or tiered pricing while refrigeration infrastructure cannot be legislated into existence. The reasoning matters more than the ranking; an answer that states its criterion (people reached? facilities enabled? speed of change?) is a good answer whichever order it produces.
y. † (item 25)
Five questions, at least two on reversibility.
- Can it be stopped, and how fast? If the formulation is a depot or implant, is removal possible, does it require a procedure, and what is the exposure profile after removal? (Reversibility.)
- What happens if something else comes up? Surgery, pregnancy, a new diagnosis, a new medication, an acute illness with vomiting or dehydration. A weekly drug can be paused. A three-month depot has already been given. (Reversibility.)
- How is the dose individualized without titration? Chapter 13's escalation schedules exist because tolerability improves with gradual increase. What replaces that at a three-month interval, and what happens to someone who tolerates the first month badly?
- What adherence problem is being solved, and for whom? If a person reliably takes a weekly injection, the convenience gain is small and the control given up is real. The benefit concentrates in people who do not — which is a genuine and large group, and a different group from the one most likely to be reading the announcement.
- What is the comparator and the endpoint in the trial behind it? Almost certainly a pharmacokinetic or weight endpoint against the same molecule at a shorter interval. Non-inferiority on exposure is not evidence of equivalent real-world outcomes, and the trade-offs in questions 1–3 are exactly the kind that show up after approval rather than before.
A sixth worth adding, though not required: has receptor biology been checked over the longer interval? Sustained occupancy drives desensitization and downregulation (§36.4), and a trial short enough to read out quickly may not be long enough to see it.
Part E — Multi-agonists and the muscle question (§36.5, §36.6)
aa. (item 27)
Why they are different claims. Weight loss is a surrogate; heart attacks are a hard outcome. The first is a measurement believed to predict the second, and the inference from one to the other can fail in at least four ways: the relationship may not be linear, so the first portion of weight loss may buy most of the benefit; mechanisms may act on risk independently of weight, as glucagon receptor agonism plausibly does through energy expenditure and hepatic metabolism; compounds may differ in effects that only an outcome trial captures; and body composition may differ, which is §36.6's whole subject. More weight loss than tirzepatide has not thereby been shown to prevent more heart attacks.
Why the second takes so much longer. Hard outcomes are rare events. Kilograms can be measured in everyone, continuously, from the first week. Heart attacks have to be waited for and counted, which requires thousands of participants followed for years to accumulate enough events for a statistically meaningful comparison — and more still if the comparator is an active drug that also reduces them. That asymmetry guarantees the weight number arrives first, dominates coverage, and holds the field alone for years after the question everyone actually cares about has been asked.
bb. † (item 28)
Financial audience. "CagriSema weight-loss result falls short of company guidance; shares fall."
Clinical audience. "Combination of an amylin analog and semaglutide produces weight loss in a large trial that would have been considered extraordinary before this decade."
Both headlines describe the same number. Neither contains a false statement. They differ only in the comparison class each supplies: the first compares the result to a forecast made to investors, the second compares it to what was pharmacologically achievable a few years ago.
What this demonstrates about question 5. Who is making the claim, and to whom, determines the frame before any number is chosen — and the frame, not the number, is what a general reader retains. The market asked "did this beat expectations?" A clinician asks "does this help people, and how does it compare to alternatives?" Merged, they produce headlines a reader could easily interpret as "the drug did not work," which neither audience's own reading supports. Neither the triumph framing nor the failure framing is a clinical judgment, and identifying the intended audience is the fastest way to strip the frame off the fact.
cc. (item 29)
What is established, in one sentence. Substantial weight loss from any cause — dietary, surgical, illness-driven, or pharmacological — includes loss of lean mass and not only fat mass, a finding documented across every modality studied for decades and not specific to GLP-1 drugs.
The four open questions.
- How much of the lean mass lost is functional muscle? Lean mass includes fluid, connective tissue, organ mass, and glycogen with its associated water; some of what is lost may be appropriate remodeling of the support structure of a larger body.
- Does it matter, and for whom? The concern is function — strength, walking speed, chair-rise, falls, independence — and it is sharpest in older adults, who have less reserve and in whom sarcopenia is already a clinical concern, and who also derive documented benefit from weight loss.
- Do resistance training and adequate protein intake mitigate it? Believed and reasonably supported for conventional weight loss; less established during pharmacologically driven weight loss, where appetite suppression may make adequate protein intake harder.
- Should any pharmacological agent be added to prevent it? The question with money behind it, and where the pipeline lives.
ee. † (item 31)
Why this is the most seductive possible surrogate. Two properties combine. First, it is a surrogate in the ordinary sense — a body-composition reading standing in for function, with several unverified links between them (Chapter 16). Second, and unusually, the intervention acts directly on the yardstick. Blocking signaling through the activin type II receptors increases lean mass; that is the mechanism, and it is not in dispute. So a favorable scan result is close to guaranteed by the pharmacology before the trial begins. A result the mechanism all but promises is a weak test of the mechanism — it confirms the drug does what it was built to do and carries almost no information about whether doing that helps anyone. Meanwhile it will be large, highly significant, and easy to describe, so it will read to every audience as strong evidence. A drug that raises the number you are using to judge it demands more care than usual, and in practice attracts less.
Cases elsewhere in medicine with the same structure. The general pattern is an intervention whose mechanism is defined in terms of the measurement used to evaluate it, so that the measurement moves by construction:
- Bone mineral density as the endpoint for agents that increase mineralization. Density can rise while the relevant outcome — fracture — does not follow proportionally, because bone quality and fall risk are not captured by the scan.
- Antiarrhythmic suppression of ectopic beats after myocardial infarction. The drugs did exactly what they were designed to do to the rhythm measurement; the effect on survival did not follow, and the historical lesson is the canonical one for surrogate endpoints.
- Glycemic control measures for agents defined by their effect on glucose handling. The measurement responds by mechanism; the cardiovascular outcome has to be demonstrated separately, which is precisely why outcome trials became a regulatory expectation in that class.
What all three share with the muscle question: the endpoint is not an independent observer of the drug's benefit. It is downstream of the drug's mechanism by definition, so it cannot arbitrate.
Part F — Conjugates, the barrier, supply, and forecasting (§36.7–§36.10)
gg. (item 33)
The inversion. Everywhere else in the book a peptide's properties are problems to engineer around: short half-life, poor tissue distribution, inability to cross membranes, restriction to cell-surface targets. In a peptide-drug conjugate, the peptide is not the active agent. It is the address — a targeting element whose selective binding delivers a payload where the receptor is and, crucially, delivers less of it everywhere else. Its job is not to have an effect; its job is to arrive somewhere specific.
A property that is a liability everywhere else and an advantage here: rapid clearance. For a hormone analog, rapid clearance means the drug is gone before it has done its work, and the whole of Chapter 33's engineering effort is spent defeating it. For a conjugate, whatever fails to find its target is carrying a cytotoxic agent or a radioisotope through the body, and you want it gone promptly. The same pharmacokinetic fact is a design flaw in one architecture and a safety feature in the other.
Others that qualify: restriction to cell-surface targets, which is a limitation for a signaling drug and a non-issue for a delivery vehicle addressing a surface receptor; and high binding selectivity for a single receptor, which narrows a therapeutic's applicability and is exactly what makes an address useful.
ii. † (item 35)
The two anatomical routes.
- Circumventricular organs. Some brain regions have fenestrated capillaries and no complete barrier, specifically so their neurons can sample the blood directly. The area postrema, in the brainstem adjacent to nuclei involved in nausea and appetite, is the relevant example. A peptide that never crosses an intact barrier anywhere can still act on neurons there.
- Vagal afferent signaling. GLP-1 receptors on vagal afferent terminals in the gut and hepatic portal region transmit signals centrally without the molecule going near the brain. The information crosses; the molecule does not.
Why this makes them different claims. "Acts on the brain" is a statement about where the effect is produced. "Crosses the blood-brain barrier" is a statement about where the molecule goes. The routes above make the first true without the second, so the two can come apart — and both directions of the conflation are common. One reader observes central effects and concludes the drug must be entering the brain in quantity, then infers things that do not follow. Another observes correctly that these are large peptides that do not readily cross, and concludes the central effects must be imagined. Both are wrong. The correct statement is narrow: these drugs produce genuine central effects through anatomically specific routes that do not require bulk penetration of an intact barrier — longer and less quotable than either error, which is roughly why you will not see it in a headline.
On the coverage-assessment half of the exercise. This asks you to find a real article, so no model answer can supply the specimen. What to look for: verbs and prepositions rather than claims. "Acts on the brain's appetite centers" is usually fine. "Enters the brain," "crosses into the brain," "works directly in the brain," and "penetrates the central nervous system" are the phrases that assert something stronger. Note also whether the article distinguishes appetite effects from claims about mood, addiction, or cognition — the further from appetite the claimed effect, the more the article is relying on an unstated assumption about bulk penetration. And record whether the conflation changes the article's conclusion or is merely loose phrasing; both are worth noticing, but only the first is an error that matters.
kk. (item 37)
Changes when patents expire:
- Price. Directly. Generic entry is the mechanism by which a molecule's price falls to a fraction of its originator level.
- Who can afford treatment. Downstream of price, and the reason any of this matters. Coverage calculations invert, and health systems that could not consider population-level provision can begin to (Chapter 12; Chapter 19's economics are downstream of this).
- The gray market's main selling point. Its principal advantage has been cost. A cheap regulated product erodes that considerably — which is a change in the market, not a change in any evidence.
Does not change when patents expire:
- The size of the treatment effect. A molecule does not become more effective when it becomes inexpensive. The effect size was measured in trials that patent status does not touch.
- The strength of the evidence. The trials already ran. Their design, size, endpoints, and results are fixed facts about the published record.
- The population in which the drug was studied. A historical fact about who was enrolled. Nothing about market exclusivity can retroactively broaden or narrow it.
The teaching point. A rating attaches to a claim about a molecule's effect in a population on an endpoint (Chapter 5). The price of the molecule does not appear in that sentence. The error runs in both directions — people who dislike pharmaceutical pricing casting doubt on the clinical evidence, and people defending the clinical evidence dismissing pricing concerns as unserious. Errors of the same shape. Keep two ledgers.
ll. † (item 38)
An open item; no single correct answer. What a good response contains, and what to check your own against.
Structural requirements. Three lists, populated, each item specific enough that a reader in five or ten years could determine whether it happened. Every item date-stamped — the date the list was written, which is what makes later grading possible. Items phrased as events, not as trends: "[specific thing] occurs by [date]" rather than "the field continues to advance."
A quick self-test on each item: could a well-informed stranger, handed this list in 2035, mark it without asking you what you meant? If not, it is a mood, not a forecast.
Common failure modes to correct for.
- List one padded with near-certainties. "Would not surprise me" is not "is guaranteed." If every item is a safe bet, the list costs nothing and grades trivially.
- List two containing things you actually expect. If an item in "would surprise me" would not, in fact, surprise you, it belongs in list one. This is the most common error and it is a self-flattering one, because list two is where a forecaster looks bold.
- List three used as a disclaimer. "Would bet against but cannot rule out" is a probability statement, not a hedge. Each item should name something you would genuinely wager against.
On which list is hardest — the reflective half of the exercise. Most people find list two hardest, and the reason is diagnostic: writing down what would surprise you requires knowing what you currently believe strongly enough to be shocked by its contradiction, and most working beliefs turn out to be vaguer than they felt. That difficulty is the exercise working. A respondent who found list one hardest has usually confused "unsurprising" with "likely." A respondent who found all three easy has probably written trends rather than events.
mm. † (item 39)
An open item. What a strong response looks like.
The test, applied. For each of three sources you rely on, find a forward-looking claim they made and ask: did they name what would prove them wrong? Then ask the harder version: is there any past claim of theirs that has since been graded — checked against what actually happened — and did they acknowledge it?
The typical finding, and the useful reading of it. Most sources fail the first test, and this is not evidence of bad faith. Vague claims are easier to write, safer to publish, and more pleasant to read; specific dated ones expose the writer to being caught. So the result is usually not "my sources are dishonest" but "my sources are engaged in a different activity than I thought" — commentary rather than forecasting. That is a legitimate activity. It is simply not one whose track record can accumulate.
The distinction worth drawing explicitly. "Sources that made specific, dated, checkable claims and were wrong are more trustworthy than sources that made vague claims and can never be graded." A respondent who ranks their sources by how often they were right has missed this; a respondent who ranks them by whether they can be scored at all has got it.
Changes to reading habits that indicate the exercise landed.
- Checking dated claims from two or three years ago against what happened, before extending trust to a new claim from the same source.
- Noticing that a source's silence about a past miss is itself information.
- Asking a source's forward-looking claim "what would make this wrong?" and treating a hostile or evasive answer as the finding.
- Distinguishing the two failure modes: sources who never commit, and sources who commit and never revisit. The second is more common and easier to catch, because the record exists.
Applied to this chapter. The honest version of the answer applies the test to §36.10 as well. The three lists are specific, dated to 2026, and gradeable — which is a structural virtue and says nothing about whether they are right. That is the whole point: a forecast tells you what would prove it wrong, and it can still be wrong, in public, in the forecaster's own words.
Chapter 37
Internal reference; do not publish. exercises.md ships without answers by design. These are
instructor-side sketches: what a good response contains, where students go wrong, and which items
have no single correct answer.
The quiz.md key is published inside the file (<details> block); it is not duplicated here except
where an item needs a teaching note.
Exercises — Part A (reading the table)
A. Ten rows, eight distinct claims, four tiers. ✅ ×4 claims (weight loss — rated twice, Ch 5 and Ch 8; MACE — rated twice, Ch 5 and Ch 8; glycemic control Ch 8; CKD Ch 10). ⚠️ ×2 (MASH; HFpEF, both Ch 10). 🔬 ×1 (Alzheimer's, Ch 10). ❌ ×1 (compounded equivalence, Ch 12).
The gap between ten and eight is the point of the item: two claims are rated independently by two chapters each, and both pairs agree. Students who count eight rows have merged a duplicate pair without noticing; students who count eleven have folded in the class-level 🔬 on addictive behaviors, which is rated for GLP-1 receptor agonists generally rather than for semaglutide. Both errors are worth a minute — the first because merging duplicates destroys the book's only public self-check, the second because it is a rule-1 slip (a class claim is not a molecule claim).
B. The differing phrase is with documented growth hormone deficiency. Principle: a rating attaches to a claim, and the population is part of the claim.
C. Ch 5 and Ch 17 rate the identical tendon/soft-tissue claim; Ch 17 additionally rates the GI protection claim. Keep the duplicate because the table is an index into chapters — merging would hide that two chapters independently reached the same conclusion, which is a consistency check the reader should be able to run.
D. Highest ✅ proportion: approved pharmacopeia, 6 of 7. (Cardiovascular at 4 of 6 and oncology at 6 of 9 are acceptable answers if stated as fractions and defended.) Highest ❌ proportion: claim forms, 11 of 14; among therapeutic blocks, veterinary at 3 of 4, then growth, repair, and performance at 9 of 16. Accept any of these with the fraction stated — the exercise is testing whether they compute rather than eyeball.
E. Ch 41 disclosure — obligation. Ch 43 — inevitable. Ch 44 stigma — the problem is not one word but that two opposed mechanisms determine the direction; accept "reduce" if they explain that the direction, not the magnitude, is unforecastable. Ch 44 policy — primarily.
F. † The two are Ch 31's "used in animals for years, therefore safe in humans" and, arguably, Ch 31's two TB-500/GHRP rows, whose claims are explicitly framed on the basis of animal use. Good third candidates: Ch 18's general "immune modulation," Ch 23's "this peptide is a nootropic," Ch 32's pricing claim. No single right answer — assess the defense.
G. Reasonable families: (i) a certificate or test standing in for a product — purity, third-party tested, independent testing, pharmaceutical grade; (ii) a regulatory or list status standing in for evidence — FDA approved, not FDA approved, approved elsewhere, banned list; (iii) an authority or novelty signal standing in for evidence — doctor prescribed it, doctor never heard of it, next generation, this modification. Other groupings are defensible if each family names an inferential error rather than a topic.
H. The ten: NAD+ precursors (Ch 6), GHK-Cu (Ch 6), insulin analogs (Ch 11), CJC-1295+ipamorelin (Ch 15), intranasal oxytocin for autism (Ch 21), therapeutic cancer vaccines (Ch 26), peptide–drug conjugates (Ch 27), calcitonin (Ch 29), topical GHK-Cu (Ch 30), topical peptides as a category (Ch 30). By kind: versions of the claim — GHK-Cu ×2, NAD+, CJC-1295, topical peptides category; sub-population/endpoint — insulin analogs, peptide–drug conjugates; trajectory — intranasal oxytocin, calcitonin, therapeutic cancer vaccines. Accept reasonable reassignment with argument; several are genuinely two kinds at once.
I. † No fixed answer. Look for correct identification of surrogates: MASH histology, visceral fat on imaging (tesamorelin), body composition (MK-677), apnea–hypopnea index, HbA1c-type glycemic endpoints, bone mineral density. Hard outcomes to reward: MACE reduction, vertebral fracture, cardiovascular death or HF hospitalization, progression-free survival, pain-free light exposure, reversal of severe hypoglycemia.
J. Three sentences without "but." Model: Colistin and daptomycin are approved peptide antibiotics supported by trials, so specific molecules in this class demonstrably treat infection. The class-level claim that antimicrobial peptides cannot generate resistance is contradicted by multiple characterized mechanisms, including transferable plasmid-borne mcr. The further claim that the class will yield a new generation of broad-spectrum systemic antibiotics remains open, with named obstacles and active work, which is what 🔬 records.
Exercises — Part B (the two kinds of ❌)
K. See §37.3. Commonly omitted: that the two differ in volatility, not just in what is known.
L. Six marked rows: leptin in common obesity (Ch 13); intranasal oxytocin for autism (Ch 21); NK1 antagonists for chronic pain/MDD (Ch 22); AMPs cannot generate resistance (Ch 25); therapeutic cancer vaccines (Ch 26); nesiritide and NT-proBNP-guided titration (Ch 28). (Note: that is seven rows across six source situations — Ch 28 supplies two. Accept either count if the rows are right.) What they share: the source chapter names studies that were run and did not support the claim.
M. † Open. Strong candidates: Ch 14 growth hormone in healthy adults; Ch 29 calcitonin; Ch 30 Argireline; Ch 16 follistatin-344. The second half of the question is the point — a student who says "then I would drop the marking" has understood; a student who says "then the chapter is wrong" has not.
N. Kitchen-elephant: absence of evidence is strong evidence of absence when a search would have been likely to succeed. Applied here: twenty years of circulation, real money, and no completed randomized trial is informative — usually about sponsor economics rather than about the molecule.
O. † The third situation is Ch 7's native GLP-1: the molecule does the biological thing, and the claim fails on pharmacokinetics. Constructed examples should have the shape molecule works, claim as stated is not achievable. Good student answers often produce oral-delivery cases, which is the right instinct.
P. Defensible ranking: (i) large trials ran and failed > (iii) one small encouraging trial > (ii) no trials > (iv) mechanism in cells. Accept (iii) above (i) only with an argument about what the reader is trying to decide. Reject any ranking that puts (iv) above (ii) without a serious defense.
Q. † No correct answer; both sides are real. The evidence-absent side: nothing is known, including about safety, and the market exploits the vacuum. The present-and-negative side: the question is settled, so continued sale is less excusable. Assess the commitment, not the direction.
Exercises — Part C (the system on itself)
R. Rules: (1) attaches to a claim; (2) ❌ describes evidence; (3) never upgrade with mechanism; (4) never downgrade with distaste; (5) date-stamped and falsifiable; (6) one molecule, many ratings.
S. Best answers: Ch 2's ✅ mechanistic GLP-1 receptor claim beside Ch 10's 🔬 Alzheimer's row; or Ch 26's ✅ peptide–MHC mechanism beside the 🔬 neoantigen vaccine row; or Ch 25's characterized membrane mechanism beside the 🔬 systemic-antibiotic row. In each case the mechanism is settled and the clinical claim is not.
T. MK-677 is the intended target. Strongest honest case: randomized human trials establish durably raised GH and IGF-1 and a measurable increase in fat-free mass versus placebo — real human outcome data most Part III compounds lack. Students should notice that the ⚠️ is earned by data, not by charity.
U. † Argument for a fifth tier: the four NOT RATED claims are not homogeneous (three values questions, one unforecastable direction), so a tier would let them be distinguished. Rebuttal: a tier implies a position on a scale of evidential support, and none of these claims sits on that scale at all; adding the tier would restore exactly the "machine that always returns a verdict" property that §37.4 warns against.
V. Open, two-sided. Reward students who notice that the second argument proves too much — if Part VIII should not have asked the questions, the eleven rateable social claims in it would also have to go.
W. † Open. Watch for the common failure: a "rateable ✅" rewrite that quietly changes the subject from policy to pharmacology. That is the finding, and it should be named rather than penalized.
Exercises — Part D (distribution and aging)
X. Two sentences: 54 against 50 shows the book is not a debunking exercise and that slightly more claims are well supported than unsupported; it does not show that any individual peptide is a coin flip, and the four-rating margin is far too small to argue from — a different but equally defensible chapter list would move it. One sentence on structure: ✅ concentrates where trials were funded and ❌ where they were not, so the distribution is a map of research funding rather than of biology.
Watch for students who treat the ✅ majority as a verdict on peptides. The margin is noise; the clustering is signal. A student who writes "so peptides mostly work" has made exactly the error the item is designed to catch.
Y. ⚠️ requires that a trial happened. Most consumer-market claims never reach a trial, so they never reach ⚠️. It is a distinction earned by having been studied.
Z. † The recurring situation: a modest instrument-measured effect is supportable and the strong marketing version of the same claim is not. Examples: GHK-Cu, NAD+ precursors, topical peptides as a category, CJC-1295. Collective finding: in eight separate cases the evidence supported a version of the claim and not the marketed version.
AA. / AB. Open. Reward specificity — a named result, not a direction. Strong AA candidates: the evidence-absent Part III rows; the ⚠️ metabolic rows exposed to phase 3 regression; the 🔬 rows. Strong AB candidates: the Ch 27 oncology ✅s on progression-free survival; sacubitril/valsartan; teriparatide; glucagon; the mechanistic ✅s in Ch 2 and Ch 26.
AC. † Most exposed: rows resting on early-phase results with no late-stage data. Least exposed: rows where the source explicitly notes substantial late-stage data. Students should notice that the Ch 36 triple-agonist callout itself says the phase 2 magnitude was large enough to be unlikely to be entirely an optimism artifact — which is the nuance the exercise is fishing for.
AD. Persistent 🔬 most likely indicates that the trials capable of resolving the question are not being run — usually a funding or feasibility problem rather than a scientific one. "Too early to say" becomes an answer once enough time has passed.
Exercises — Part E (dossier and drift)
AE–AH. Assess on completeness and honesty. Do not grade the match count. A student with four matches and a well-populated pile one has done better work than a student with eleven matches and no reason sentences. AH (the Step 5 sentence) is the deliverable; it must be about the student, dated, and specific enough to be falsifiable next year.
AI. † Open by construction. The assessment criterion is whether the four-line callout is well-formed: does the claim name a population and an endpoint, does the "why" cite evidence rather than mechanism or vibes, and is the "what would change it" line genuinely falsifiable? The final sentence — whether the disagreement runs with the rest of their drift — is the item's real target.
AJ. Open. Expect students to get stuck at step 1 (finding the population and endpoint, because marketing copy rarely states either) and step 4 (finding out whether anything was ever registered). Both stumbles are the lesson.
Quiz — teaching notes on selected items
Item 3 is the discriminator for the whole chapter. A student who picks (a) has inverted §37.3 and should be walked through Case Study 37.1 before moving on.
Item 10 requires the final distribution (54 / 26 / 50 / 10 / 10 / 4 = 154). Two superseded tallies
circulated during the build; if a student's edition shows 150 total, or shows ✅ and ❌ level at 50
each, the file is stale — check against _scratch/RATINGS-INDEX.md and see
_scratch/continuity/ch37.md §1.
Item 13 — students who choose (a) or (d) have read splits as hedges. Recover with the GHK-Cu example: the split is more precise than a single glyph.
Item 17 — the ten-versus-eight distinction is the whole item. Ten rows; two claims rated twice; eight distinct claims. Counting to eight also requires including the Ch 12 compounded-product row, which is not about the molecule — deliberate, and worth surfacing: a claim about a product is still a claim, and a reader who dismissed it as irrelevant has missed how most people actually encounter the drug.
Items 18–22 are rubric-scored; the published key states what a good answer contains. For item 21, the sentence to look for is some version of being right for the wrong reason is still not knowing.
Case studies
CS1 — Q1–Q5 have determinate answers drawn from §37.3 and Chapter 28; Q6 is open. In Q5 the expected answer is that a single positive trial moves the row to ⚠️ at most, never to ✅, and only if population and endpoint match the claim.
CS2 — Q1 is the sharpest item: entries 3, 7, and 10 involve splits, and a student who argues that entry 3 should be MATCH (against the ⚠️ arm) has a real case, which is exactly why the audit's matching rule specifies "the arm that matches the version you rated." Q6 has no correct answer; assess whether the student can articulate what the audit is measuring — calibration of tiers, or quality of reasoning — because the scoring choice follows from that.
Chapter 38
Model responses for the daggered (†) items and a selection of the others. Several of these questions have more than one defensible answer; where that is true, the response below shows the reasoning rather than pretending to a single verdict. None of these answers is legal advice, and none of them states what any reader may lawfully do.
Group 1 — What approval means
A. The clause almost everyone drops is "the evidence submitted." People reconstruct the definition as "the regulator judged that benefits outweigh risks," which quietly implies the regulator surveyed everything known. It did not. It read a dossier a sponsor assembled. That single omission is what makes case (1) of §38.4 invisible to most readers.
B. Missing: the indication (and, strictly, the population). Defensible rewrite: "Tirzepatide is approved in the United States for type 2 diabetes and, under a separate approval, for chronic weight management in defined populations." Useful-for-insurance rewrite: "Tirzepatide has separate approvals under separate brand names; coverage usually follows the indication on the prescription, not the molecule — which is why the same drug can be covered for one patient and denied for another" (Chapter 12).
E. † The expected finding: indication wording is far narrower than consumer discussion. Labels name a condition, an age range, often a prior-therapy requirement or a body-mass threshold, and sometimes a restriction to use "as an adjunct to" diet, exercise, or another therapy. Advertising describes a benefit; the label describes a population. The gap between them is where most patient misunderstanding lives.
Group 2 — The forty-amino-acid line
F. "In the United States, as of this writing, an alpha-amino-acid polymer with a specific, defined sequence of more than 40 amino acids is regulated as a biological product; 40 or fewer is regulated as a drug." Second sentence: the terms describe which regulatory framework applies, not how the molecule behaves in a body — a 41-residue peptide and a 39-residue peptide are the same kind of chemical object.
G. Oxytocin 9, BPC-157 15, teriparatide 34, semaglutide 31, tirzepatide 39 — all drug side. Insulin 51 and growth hormone 191 — biologic side. Tirzepatide is closest, with two residues of room.
J. † A model chain: (1) the molecule has n amino acids — fact; (2) n determines NDA or BLA — legal fact; (3) application type determines whether the follow-on route is generic or biosimilar — legal fact; (4) route determines the cost and duration of entry for a competitor — strong empirical regularity; (5) cost of entry determines how many competitors arrive and how fast — economic prediction; (6) number of competitors determines price decline — economic prediction, well supported for small-molecule generics and much less so for biosimilars; (7) price determines access — contested, because insurance design intervenes (Chapter 12).
K. Changed: the application framework, the availability of the biosimilar pathway, and therefore the structure of possible competition. Not changed: the molecule, its manufacturing, its clinical evidence, its labeled indications, or its safety profile. The size of the first list relative to the second is the point of the case study.
L. † At a 30-residue line, semaglutide (31) and teriparatide (34) and tirzepatide (39) would all move to the biologic side. Plausible consequences: a more demanding follow-on pathway for the GLP-1 class, later and fewer competitors, and a slower price decline after exclusivity. Speculative elements: whether biosimilar-style requirements would in practice be lighter for a chemically synthesized peptide than for a cell-culture protein — a good argument that the line's location is doing less work than the line's existence.
Group 3 — The pathway
N. An IND permits human administration; it is a permission to test. "Cleared by the FDA" in a press release most often means an IND was allowed to proceed, which is a milestone at the very start of the clinical process and not a finding about efficacy or safety in the sense a reader will hear it.
Q. † Four questions: (1) Is this trial completed, and are results posted? (2) What phase, and what is the prespecified primary endpoint? (3) What population and what comparator — is it randomized and controlled, or single-arm? (4) Is the intervention the same molecule, route, and formulation as the product being sold? A registry entry documents that a study was planned and registered. It is a commitment device, not a result, and citing it as evidence inverts its purpose.
Group 4 — The five meanings
R. Ranked by strength of negative signal about the molecule: (2) submitted and rejected — strongest and the only real one; then (3) approved elsewhere, not here — ambiguous, and worth almost nothing until you know which story applies; then (1) never submitted, (4) approved but not for this use, and (5) not a drug at all, which carry no signal about the molecule at all. Defensible variation: some readers place (3) below (1) on the grounds that a foreign approval is a mildly positive signal, not a negative one. That is a good argument.
S. No filing history → (1). Face serum ingredient → (5). Registered in one country, absent in another → (3). Approved drug used for an unlisted condition → (4). Application withdrawn mid-review → (2).
U. † The expected outcome for most gray-market compounds is (1), with a "cannot determine" attached to at least one jurisdiction. That is the correct answer and should be recorded as such. Credit for documenting the search — which databases, which agency pages, what search terms — because the record of what you checked is what makes "cannot determine" a finding rather than a shrug.
V. The ❌ rests on the published human literature for the healing claims, which is essentially absent. Regulatory status is a fact about a file. What would change the rating: an adequately powered, randomized, controlled human trial with a prespecified clinical endpoint, published in a peer-reviewed venue. Approval, by itself, would not — though in practice approval would normally be preceded by exactly such trials, which is why the two correlate without being the same.
W. Confusion one: "not approved" and "not banned" are treated as opposite ends of one scale. They are outputs of two different systems — a medicines regulator and, if the person is in tested sport, an anti-doping code — and a compound can sit anywhere on both independently. Confusion two: "gray area" implies an unsettled legal question, whereas for many compounds the applicable rules are perfectly determinate and simply unenforced or unexamined. Correction: name the system, name the jurisdiction, and note that what any individual may lawfully do is not something a book can answer.
Group 5 — Off-label and compounding
AB. † Four differences and their risks: salt or ester form — altered solubility and pharmacokinetics, so exposure differs from the studied product; concentration — the documented failure mode, where a person accustomed to a pen's units miscalculates a vial's volume; delivery device — substituting a syringe for a metered pen transfers the metering job to a human being at home; additional ingredients — excipients or co-formulated substances that were never evaluated with the drug. Underlying all four: the approved product's clinical evidence was generated with the approved product, and does not transfer automatically to a different one (Chapter 19).
AC. Model rewrite: "Prepared by a compounding pharmacy on a prescription from a licensed physician. Compounded preparations are not FDA-approved products and are not reviewed for safety or efficacy before sale." That is equally accurate, no longer implies review, and — note — is the sentence almost no marketing page writes.
Group 6 — Research chemicals, supplements, doping
AD. It does two things for the seller: it creates a documentary basis for arguing the product was not offered for human use, and it shifts the consequences of use onto the buyer. It fails to do anything for the buyer at all — it confers no manufacturing standard, no reviewed label, no adverse-event reporting, and no mechanism for learning what happened to anyone else who bought it.
AG. † Model paragraph: The ❌ is a claim about the published human evidence — there is essentially none for the healing claims. The S0 prohibition is a claim about regulatory status — the compound is not approved by any governmental health authority for human therapeutic use, and S0 prohibits exactly that class. Neither judgment consulted the other. The trap for an athlete is to read the prohibition as a backhanded endorsement — they wouldn't ban it if it didn't work — which inverts the rule: S0 bans compounds because nobody knows whether they work, not because somebody found out that they do.
AH. Because S0's test is "not currently approved by any governmental regulatory health authority for human therapeutic use," approval removes a compound from S0's scope entirely. Whether it is then prohibited depends on the other categories on their own terms. Whether this is sensible is a fair argument: it makes the anti-doping rule dependent on medicines regulators in other jurisdictions, but the alternative — an enumerated list — would always lag novel compounds, which is precisely the gap S0 was written to close.
Group 7 — Patchwork, status, and evidence
AJ. † A strong answer includes: what a registration certificate does establish (a dossier existed and an authority acted on it); what it does not (the size of the dossier, the standard applied, the endpoints, the date, and whether the standard resembles the one the reader is imagining); the specific point that registration standards differ substantially between systems; and the constructive move — ask for the public assessment report, which many agencies publish, because that document says what the approval actually rested on.
AL. The strongest defense of the 2001 approval: dyspnea in acute decompensated heart failure is a patient-meaningful symptom, the population was seriously ill with limited options, and the evidence available supported a benefit on that endpoint. The strongest case for caution: a symptom endpoint over hours was allowed to stand in for outcomes over months, adoption was wide and fast, and it took a decade and roughly seven thousand patients to find out that the outcomes were not there. Both are correct; a well-argued answer says which risk the reader thinks a regulator should prefer to take, and why.
AM. † The steelman: clinical evidence is funded almost entirely by parties who expect to recover the cost, recovery requires exclusivity, exclusivity requires patentability, and a short published sequence is hard to patent — so a whole class of compounds goes unstudied for reasons that have nothing to do with their biology. That argument is largely correct and is a real defect in how medical knowledge is produced. The inference it does not license is that any particular unstudied compound works. The same economics predict that useless compounds go unstudied at exactly the same rate. An empty file is uninformative, which is the one thing marketing built on this argument cannot afford to say.
AN. Model: "Paperwork records what somebody asked an authority to judge, and what that authority said. Whether a molecule helps a person is a separate question, answered by trials, and you have to go and look."
Group 8 — Dossier
AQ. The five questions: what the sales channel claims about itself (Ch. 6); who can get it, at what cost, from whom (Ch. 12); what is actually in the vial (Ch. 19); what a certificate of analysis can and cannot establish (Ch. 34); status with regulators and with sporting authorities (Ch. 38). Most readers find Ch. 19 and Ch. 34's questions hardest, because they are the only two that cannot be answered from documents at all — they require testing.
AP. † / AR. † These are self-audit items with no fixed answer. The pattern to look for: readers consistently soften ratings for compounds that turn out to be approved somewhere, and consistently harden them for compounds that turn out to be prohibited in sport. Both are the same error running in opposite directions — letting a fact about a file move a judgment about evidence — and noticing your own direction is the entire point of the exercise (Chapter 40).
Chapter 39
For the instructor's answer key. The chapter's exercises are deliberately unanswered in the student file. Many items have no single correct response; for those, this file gives what a strong answer contains and the common failure mode rather than a model answer.
A general marking principle for this chapter: any answer that reads as a strategy for obtaining a prescription has missed the chapter, however well argued. The chapter's stated goal is a better conversation, not a won one, and answers that treat the clinician as an obstacle to be routed around should be marked down regardless of their sophistication.
Part 1 — The shape of the room
A. Patient braced for judgment (often from memory, not projection); clinician braced for a request they cannot responsibly grant (from base rate, not from this patient). Both want an accurate decision. Failure mode: locating the blame in one party.
B. (1) Short appointments — a substantial conversation inside a slot designed for a simpler one; (2) no clinician can know every compound — thousands, most with no literature; (3) "I don't know" is professionally expensive in a system that rewards confidence. Strong answers explain each as a property of the system.
C. Protects against: judgment, having it entered in the record, being labeled. Costs: the clinician cannot ask follow-up questions about a person who is not there, cannot record it, and cannot connect it to the patient's own history — so the answer given is generic.
D. † Strong answers describe a mechanism, not an attitude. E.g.: a confident-sounding answer is never checked, so the feedback loop that would correct it never closes; a clinician who says "I don't know" absorbs the disappointment immediately, so the two responses have asymmetric local costs regardless of their accuracy. Note that this can be true with every individual acting honestly.
E. Any rewrite that names the specific claim, states what evidence exists or does not, and indicates what would change the assessment. E.g.: "I'm not aware of human trial data on that. What I know of comes from animal work. Do you want me to look?"
Part 2 — Disclosure
F. Differential diagnosis (wrong candidate list, wrong workup); interactions (cannot be checked against an unknown compound); surgery and emergencies (peri-operative protection cannot be applied to an undeclared exposure — GLP-1 gastric emptying and aspiration risk); monitoring (cannot look for a harm not known about).
G. A person who experiences disclosure as a confession minimizes. Minimization produces exactly the four failures. So a moral framing is self-defeating on its own terms — it makes the outcome it disapproves of more likely.
H. Accept any accurate lay explanation. Must contain: the compound slows how fast the stomach empties; fasting instructions assume a normal rate; so the stomach may not be empty even when the patient did everything right; sedation suppresses the reflexes that keep stomach contents out of the airway.
I. Reasonable candidates: endoscopy, colonoscopy, some dental procedures, cardiac catheterization, some imaging requiring sedation, minor dermatological or orthopedic procedures under sedation, cataract surgery, fertility procedures. Applies wherever sedation suppresses airway reflexes.
J. † The key move: relevance is a clinical judgment, and the patient does not have the information required to make it. In Case Study 39.1 the patient sincerely did not think a weight-management injection was relevant to a bowel examination. Strong answers note that a policy of disclosing only when relevant collapses to disclosing only when the patient already knows the answer.
K. Mark on brevity and on the absence of justification. Three sentences, factual, no defense of the decision. Any hedging ("I know it's probably stupid, but…") should be removed — it invites the lecture the patient is afraid of.
L. † The argument: the value of the relationship is the accuracy of the information moving through it; a relationship in which you filter produces decisions from filtered data; therefore a relationship that cannot carry accurate information cannot deliver its core benefit. Strongest objection: this asks a lot of patients with limited choice of clinician, and access constraints make "reconsider the relationship" unavailable to many people. A good answer acknowledges this and notes that the chapter's guardrails (one bad conversation is not a pattern; do not shop for agreement) already narrow the recommendation considerably.
Part 3 — What your clinician needs
M. Mark against the five headings and the word limit. Concision is the skill being tested.
N. Patient knows the marketing claim, not the compound; clinician knows the class, not the marketing. Example: a clinician asked about "a peptide for tendon healing" may answer about peptide pharmacology generally and never learn it is being taken orally, which is the fact that settles the question.
O. For your own care: it is a data point about response and it belongs in the record with a date. For §39.10: absence of change is systematically underreported, which biases clinician impressions toward success.
P. † "I'm taking a compound sold as BPC-157." Adds: the distinction between what the label claims and what is in the vial, which no clinician can verify (Ch 34, Ch 39 §39.9).
Part 4 — The five questions
Q. Population (Ch 5) / endpoint (Ch 16) / comparator (Ch 5) / falsifiability (Ch 2 §2.8) / stopping rule (Ch 10).
R. Mark on whether the questions were genuinely adapted rather than recited. The portability claim is the point.
S. Population — with vs. without type 2 diabetes. Consequence: smaller average weight change in the population with diabetes. Belongs to question 1.
T. † Must have a magnitude, an observation method, and a date. Mark the honesty of the second part generously; the reluctance to commit to a date is the finding.
U. † Mechanism (Ch 6): once invested — money, identity, having told people, having felt something — the reinterpretation of ambiguous evidence runs automatically and is not defeated by awareness. Stopping rule must contain: a symptom meaning stop now; a finding meaning stop and reassess; a date at which no benefit means stop.
V. It has told you the intervention has no evaluation condition and therefore cannot be evaluated. Reasonable next questions: "Then how would we know if it wasn't doing anything?" or "What would you expect me to notice, and by when?" (i.e., fall back to question 4).
Part 5 — Evidence, uncertainty, both failure directions
W. Population, endpoint, comparator, size.
X. All three rewrites should move from assertion to inquiry, and — importantly — should be truthful framings rather than tactical ones. Mark down anything that is the same assertion wearing a question mark.
Y. 1 = "I don't know." 2 = "we don't know." 3 = "it doesn't work." 4 = ambiguous, see Z. 5 = "I don't know," offered voluntarily and well.
Z. † Y.4 ("there's no reason to think that would do anything") is a mechanistic claim, not an evidentiary one. It asserts implausibility rather than reporting either absence or negative results. Strong answers note that this is the mirror image of the mechanism error the book warns about: just as mechanism does not upgrade a rating, implausibility does not by itself constitute negative evidence — though it is a legitimate reason to assign a low prior. The right follow-up is "is that a mechanistic judgment or has it been looked at?"
AA. Any pair where one names the claim and the evidence and the other rejects the class. Cost: degrades the disclosure channel, which is the §39.2 outcome.
AB. Mark on the framing sentences, not the question. Challenge register: preceded by a counter-assertion, delivered as a trap. Curiosity register: preceded by an acknowledgment, delivered as an actual request. Expected products: a defensive non-answer vs. a specific evidentiary standard.
AC. † A strong answer contains all of: the structure is knowable and is not an accusation; dispensing has legitimate reasons; the recommendation cannot be used as evidence; the check is to ask the five questions and evaluate the answers. Mark down both "so it's a scam" and "so it must be fine since they're a doctor."
AD. Both are individual judgments about individual cases, subject to selection and expectation effects. What distinguishes them: training in evaluating evidence, access to the patient's history, professional accountability, and the ability to answer the five questions — all of which affect the quality of the judgment without converting it into evidence about a molecule.
Part 6 — Limits, dossier, repeated game
AE. Verify a vial; make an unapproved compound safe by supervising it; generate the missing trial; predict the individual outcome; escape institutional and licensing constraints.
AF. † Case for ✅: the mechanisms are uncontroversial and the direction is not in doubt. Case for ❌: no comparative data exist, and the claim is often deployed as reassurance for a decision the evidence does not support. Why ⚠️: the direction is sound and the magnitude is unquantified, and the rating is about what supervision adds, not an endorsement of the underlying use. Look for the direction/magnitude distinction explicitly.
AG–AH. Mark against the four conversion patterns. Common error: producing a question that is the verdict with rising intonation.
AI. † The three failure modes: general impression; distaste for the sellers (rule 4); verdict attached to a molecule rather than a claim with population and endpoint (rule 1). Strong answers identify which one and rewrite accordingly.
AJ. Reading a result against the patient's own baseline; knowing how the patient describes symptoms; knowing what has been tried; having seen the patient well.
AK. † Selection: successes are volunteered, non-responses and quiet stops are not. Chapter 6 describes the online version. Intervention: report the outcome, including nothing, at the next appointment.
AL. Under twenty words, factual, includes the stop and the null result. E.g.: "That shoulder thing — I stopped it in March. No difference either way."
Quiz answer key
Reproduced in quiz.md under <details>. Correct options: 1 c · 2 c · 3 b · 4 b · 5 b · 6 b · 7 c ·
8 c · 9 b · 10 b · 11 b · 12 b · 13 a · 14 b · 15 b · 16 b · 17 c · 18 c · 19 b · 20 d · 21 a · 22 b.
Items most often missed in review: 14 (students treat "we don't know" as the stronger claim), 18 (students want enthusiasm to count for more than refusal), and 22 (students overclaim option (c), treating aggregated patient reports as equivalent to a trial — the chapter is explicit that this is weak evidence).
Chapter 40
The quiz key ships inside quiz.md (collapsed under Answer key) and is not duplicated here.
This file covers the two case studies and the exercises. Most Chapter 40 exercises have no single correct answer, because the correct answer is in the student's own dossier. What follows is marking guidance: what a strong response contains, what a weak one looks like, and — for the audit items — what the instructor should expect to see actually changed in the document.
Case Study 1 — Auditing a Finished Dossier
Q1 — Correct Entry 1's unpopulated row.
Strong answers produce two rows, not one, and justify the split from the entry's own Field 5, which records that the same drug at the same dose produced materially different weight results in adults with overweight or obesity without diabetes versus adults with type 2 diabetes. Acceptable corrected rows name population and endpoint, e.g. weight loss, adults with overweight or obesity and without diabetes, body weight at trial duration — ✅ dated, and a separate row for the population with type 2 diabetes. A single populated row is acceptable only if the student explicitly states which population they mean and notes that a second row is required for the other.
Weak answers add a population but no endpoint, or rewrite the row while leaving it applicable to both populations. Watch for students who argue the fix is unnecessary "because everyone knows what it means" — that is the defect's own rationalization and should be surfaced, not corrected quietly.
Q2 — Why the same-evidence test made the method visible.
The key point: reading Entry 2 alone shows a rating that matches the evidence, so nothing looks wrong. The defect is not in the rating but in the procedure, and a procedure is only observable across two applications. Placing Entry 2 beside Entry 3 shows the same reader producing ❌ and ⚠️ on functionally identical Field 5 content, which means the difference came from somewhere else — and the margins say where.
On the second half: checking against Chapter 37 would have shown Entry 2 as agreement and probably flagged only Entry 3. The reader would have concluded they were mostly calibrated and their one disagreement was a judgment call. The external comparison cannot detect a broken procedure that happens to produce a matching output, which is the argument for why §40.3 exists as a separate step after §37.
Q3 — Entry 3's ⚠️ against the definition.
⚠️ means promising but preliminary — real human data that does not settle the question. Entry 3's Field 5 reads: animal work in several injury models, no completed randomized human trial, replication animal only, conspicuously absent — any human RCT. There is no human data of any kind, so the entry fails the definition's threshold condition, not a matter of degree. The correct rating is ❌, kind: evidence absent, dated — the same rating and kind as Entry 2.
The rule invoked by the fix is rating rule 3, never upgrade with mechanism, since the margin note cites mechanism elegance; strong answers also cite Chapter 6 for the anecdote. Accept students who additionally note rule 1, since the row is otherwise well specified.
Q4 — The two margin notes.
Entry 2's margin ("the marketing is shameless... rated accordingly") breaks rule 4, never downgrade with distaste. Entry 3's margins ("very elegant mechanism", "training partner... recovered well"*) break rule 3, never upgrade with mechanism**, plus the testimonial dynamic.
On why neither is more careful: both are ratings moved by something that is not evidence, which is the single prohibition both rules encode. Strong answers add the §40.4 asymmetry — the harsh error receives no social correction and therefore persists — and note that in this dossier the two errors sort by personal proximity (harsh toward what is sold to strangers, generous toward what a friend used) rather than by generosity in general.
Q5 — Answering the "bureaucracy" objection.
Two consequences from the walkthrough: (a) the rating was reached by distaste, so it will not update when evidence arrives — the reader has not specified what a trial would have to show, and will therefore evaluate it after the fact while holding a stated high-confidence position; (b) across the file, Field 12 was written where the reader felt uncertain and skipped where they felt sure, which inverts its purpose — it is most valuable exactly where it feels least necessary.
What the objection reveals: the objector treats a rating as a conclusion to be scored for correctness rather than as the output of a procedure. Under that view, a right answer needs no audit. Under the book's view, a rating is only as good as the method that produced it, and Field 12 is where the method is recorded.
Q6 — Entry 4 and low investment.
Expected explanation: Entry 4 concerns a serum the reader has no strong hope or hostility about, so neither rule 3 nor rule 4 had anything to act on, and the fields were filled by the procedure alone. Investment supplies the non-evidence input that the checks detect.
Concrete counter-procedures worth full credit include: auditing entries in descending order of how much the student cares; having someone else run the four checks blind; auditing the ratings with the compound names masked; requiring a written Field 12 before a rating may be assigned; or rating each entry twice, separated by enough time to forget the first.
Case Study 2 — Twelve Fields, No Peptide
Q1 — The fourth Field 6 row.
It claims that a specific consumer device, with undisclosed parameters, reproduces the effects reported in the clinical literature for the modality. It is ❌ evidence-absent because nobody has tested that inference: the evidence attaches to studies at specified exposures, and no one has measured whether consumer devices deliver them.
The chapters reproduced are 32 (identically labeled vials may differ materially) and 34 (a certificate describes a sample, not the vial in your hand). Accept Chapter 19 as an additional reference on preparation-quality risk being independent of whether the thing works.
Q2 — Cleared vs. approved.
510(k) clearance is a determination of substantial equivalence to a legally marketed predicate device; some low-risk devices are exempt from even that. Premarket approval is a separate, more demanding pathway for higher-risk devices. Neither phrase means "this device works for what it is being sold for," and where a cleared indication exists it is often narrower than the marketing.
The Chapter 38 parallel: approval is a judgment about a specific claim rather than a certificate about a product, and "not approved" resolves into several distinct meanings that must not be collapsed. For the Field 8 wellness claims, the applicable meaning is no regulator has evaluated this claim and none was required to — which is neither "rejected" nor "approved elsewhere," and carries no negative signal on its own.
Q3 — The peptide-equivalent Field 4 line.
Any correct construction of the form: studied by subcutaneous injection in the trials cited; sold and used orally / as a nasal spray / as a topical; route used does not match route studied. Strong answers note the sharper version — route unknowable, where the product does not disclose what it contains or at what concentration.
Structural identity: in both cases the evidence base attaches to a specified exposure, and the thing being used delivers an unspecified one. Neither situation is a claim that the intervention fails; both are statements that the cited evidence does not apply to the thing in hand, and both are invisible unless studied and used exposures are recorded separately.
Q4 — The "no reports" chain.
For a "no reported problems" statement to carry information, all of the following must hold: someone experiences an effect; attributes it to this intervention rather than to something concurrent; knows a reporting pathway exists; the report is collected by someone; and aggregate data is published.
For an approved drug, all five links exist and several are legal requirements. For a consumer wellness device, links three through five are typically missing — there is no surveillance system comparable to post-marketing drug reporting. Strong answers state the conclusion in the chapter's own form: the sentence describes the reporting system, not the safety profile. Accept students who note that link two is also weak for diffuse endpoints like "recovery."
Q5 — The two overall verdicts.
Marketer's version: "red light therapy is clinically proven." Destroys the distinction between the modality and the specific product (row 4), between claims with preliminary human data and claims with none (rows 1–2 vs. row 3), and between an investigator-rated endpoint and a self-reported one.
Skeptic's version: "red light therapy is pseudoscience." Destroys the same distinctions from the other side, and specifically misdescribes rows 1 and 2, where real human data exists.
Both are forbidden by rating rule 6 — one intervention, many ratings. Strong answers observe that the two verdicts destroy the same information, which is the section's point about the symmetry of the two biases.
Q6 — Predict-then-fill.
There is no fixed answer; mark on the quality of the prediction and the honesty of the comparison. Most students predict 5, 9, and 12 and are right, which reproduces §40.2's finding. The graded part is the final question: students should identify where they personally stop looking — typically at the first plausible-sounding source, at the manufacturer's own page, or at the point where the answer would require reading a methods section. Award full credit for a specific, unflattering answer; withhold it for "I should research more thoroughly."
Exercises — marking guidance
Group 1 (A–K), assembly. These are mostly verifiable by inspection of the student's document. Require the completion map in C to be produced before any gap-filling; students who fill as they go lose the by-field tally in D, which is the item that reproduces §40.2's prediction. In H, reject any Field 9 reading "no known side effects" for a compound without human trial data — the required distinction is between no harms and no reporting pathway. In I, the failure mode is one sentence covering all five sub-questions; require five. J † is the heaviest item in the group and should be assigned with time; check that ❌ entries received Field 12s too, since that is the instruction students most often skip.
Group 2 (L–R), the audit. L requires two counts before fixing — students who fix first destroy the measurement. O † is the centerpiece; if a student reports finding no comparable pair, push them toward entries with no human evidence, where non-evidence reasons operate most freely. In P †, the deliverable is a list of failing Field 12s plus rewrites, not a general statement that they could be sharper. R † is the hardest item in the chapter and many students will produce nothing; a defensible "I cannot tell, and here is why that is the problem" is a full-credit answer.
Group 3 (S–W), own error. S must be presented as a column; individual comparisons defeat the exercise. T † should produce one sentence added to the document — check for it. V is the argumentative item: a strong reply cites rule 4, its mirror in rule 3, the absence of a corrective signal, and the concrete cost of misreading a ⚠️ as ❌. W † frequently produces the session's best discussion; ask which part was hard, evidential or emotional.
Group 4 (X–AC), maintenance. X is not simulated — a dated frozen copy should exist. AA † is the item most likely to produce a genuine realization: students routinely find that confidence moved on news that did not satisfy Field 12. AB tests whether the student can state the selection mechanism without leaning on the word "bias," which forces them to describe the actual filtering process rather than name a category.
Group 5 (AD–AJ), use and transfer. AF and AG † are the transfer items and the ones worth grading most carefully; AG † should produce a defensible row with population, endpoint, date and ❌-kind, plus a Field 12 that survives the reasonable-person test. AI † deliberately breaks: the student should find that Field 5 has no trial arm, no endpoint, and no comparator to record, and should attempt substitutes (natural experiments, before/after populations, cross-jurisdiction comparison, self-reported survey data) while noting what each substitution costs. Keep these answers; they are the pre-test for Part VIII. AJ produces the permanent top line of the dossier and should be checked for specificity, not sentiment.
Spaced Review — guidance
1 (Ch 5). The four: population, endpoint (surrogate or hard), comparator, and size/duration. The one that most often turns out to be the problem is the comparator — either absent, or an inappropriate one. Accept "endpoint" with a strong argument.
2 (Ch 6 + §40.6). The rise is predicted because reaching news is filtered for publicity value and the student's own searching filters again in the same direction; accumulating apparent support is what you would feel either way. The one-minute procedure: open Field 12, read the named readout, ask whether that specific thing happened. A conference abstract, a mechanism paper, and a friend's report satisfy none of it.
3 (Ch 37 + §40.4). Each has discovered the direction of their error, which is the finding a per-compound comparison cannot produce. It is a mistake to call the harsh reader more careful because both moved ratings by something other than evidence; the harsh reader simply receives no signal that they did, and will transmit an aesthetic reaction as an evidence assessment.
4 (§40.9). They weight things the dossier does not contain — tolerance for uncertainty, the cost of the problem, the consequences of being wrong. The dossier succeeded at locating the disagreement in values rather than in facts, which is what makes it resolvable in the right register.
5 (the final exam). No single answer. Mark on: whether all twelve fields were attempted; whether Field 1 identified that "a wearable" is a category rather than a specification (which sensors, which algorithm, validated against what reference standard); whether Field 4 was translated into measurement validity rather than skipped; whether Field 5 asked about the reference standard and the blinding problem for self-reported sleep; whether Field 7 was correctly identified as structurally different for a general-wellness device versus a cleared medical claim; and whether the Field 6 row separated the device measures something accurately from using the device improves an outcome — two claims that are almost never separated and that have entirely different evidence bases. That separation is the item's real target and mirrors Case Study 2's modality/product row.
Chapter 41
Instructor copy. Several items are deliberately open — where a defensible range of answers exists, the solution states the criteria for a good answer rather than a single correct response. Items marked † in the student text are noted here.
Grading note applying throughout: no answer should name a private individual. An otherwise strong response that identifies a person has failed the chapter's central rule and should be returned rather than marked down silently.
Exercise 41.1
| Molecule | Brand | Route | Indication family |
|---|---|---|---|
| semaglutide | Ozempic | injection | type 2 diabetes |
| semaglutide | Wegovy | injection | weight management |
| semaglutide | Rybelsus | oral | type 2 diabetes |
| tirzepatide | Mounjaro | injection | type 2 diabetes |
| tirzepatide | Zepbound | injection | weight management |
The recall question: Ozempic and Wegovy share a molecule with Rybelsus. Students who name only the two injectables have missed the oral product, which is the one most often forgotten and the one whose existence most complicates "Ozempic is the injection."
Exercise 41.2
A generic name is assigned under international convention, names a molecule, and is the same everywhere. A brand name is commercial property, names a product — molecule plus formulation plus route plus approved indication — and one molecule may carry several. Any correct non-semaglutide example is acceptable; insulin analogs and their brands work well and connect forward to Case Study 41.2 Case E.
Exercise 41.3
False sense: they contain the same active molecule, semaglutide, so at the level of chemistry they are not different drugs. True sense: they are different products, with different supplied doses, different devices, and different approved indications — which is what determines coverage and what happens at the pharmacy. Good one-sentence versions say something like: "Same molecule, different products, and the difference is what they're approved to treat."
Exercise 41.4 †
Any genericized brand is acceptable. Grade on three things: (i) genuine usage examples rather than assertions at each rung; (ii) an honest account of where it stopped; (iii) a reason for stopping. The strongest answers notice that most brands stop at rung 3 because rungs 4 and 5 require the category to have carried unusual cultural weight — a brand becomes an era only when the thing it names was already a social event.
Exercise 41.5
Liraglutide and dulaglutide both carry the -glutide stem, which identifies them as GLP-1 receptor agonists (Ch 1 §1.8). Tirzepatide carries only -tide, which tells you it is a peptide and nothing about its receptor target. What to check: the actual pharmacology — tirzepatide acts at more than one incretin receptor, which the name does not encode. The teaching point is that stems are informative but not complete, and a name is a hypothesis to be confirmed, not a conclusion.
Exercise 41.6 †
The argument: a genericized incumbent brand becomes the default label for competitors' products, so competitors' molecules are discussed, praised, blamed, and reported under a rival's name. They lose name recognition they paid for and inherit adverse-event attribution they did not cause. But the largest cost falls on neither company: it falls on patients, who cannot articulate what they take, and on clinicians, who must reconstruct it. Credit answers that reach that conclusion; the exercise is built to make students discover that the commercial framing is the least important one.
Exercise 41.7
(i) Patients cannot say what they are taking — cost borne by the patient, and by the clinician reconstructing the history. (ii) Prescribers field requests for the wrong product — cost borne by the prescriber (appointment time) and the patient (denials, delays). (iii) Coverage conversations cannot converge — cost borne by the patient, financially. (iv) The brand name becomes a container for non-drug products — cost borne by the patient who is harmed by an unverified product, and by the approved product's safety record, which is diluted in both directions.
Exercise 41.8
Five: (1) which molecule; (2) which product; (3) which route; (4) which indication it was prescribed under; (5) whether it is an approved product at all, versus a compounded or gray-market preparation. Reasonable rankings differ, but (5) and (1) should rank high — they change what the clinician can trust about identity and content. Accept "how long they have been taking it" as an additional item; it is genuinely clinically relevant, though not one the brand name obscures.
Exercise 41.9 †
Strong answers converge on something close to the chapter's version: "What's the exact name on the box, and is it a pen or a tablet?" — because it recovers product, molecule, and route in one utterance, and implies the indication. Rejected alternatives worth crediting: "which drug are you on?" (returns the same generic word), "what dose?" (presupposes the product and invites a number the patient may misremember). What it still misses: whether the product came through a regulated channel, and how long they have been taking it. Credit answers that name a miss.
Exercise 41.10
A formulary lists products with criteria attached, because coverage is granted for a product used for an indication — not for a substance. So "is semaglutide covered?" has no answer; "is this product covered for this indication under this plan?" does. The expense follows: a request under the wrong product name can generate a denial, a prior-authorization loop, or an off-label designation, each of which costs the patient time and often money, and none of which is a decision about the medicine.
Exercise 41.11 †
The argument: ambiguity flows toward the most famous name, so events caused by other products, compounded preparations, or gray-market vials get reported under the famous brand — and the famous brand's safety and efficacy record gets lent to those products in the other direction. Both errors run the same way, so the bias is directional, not random. One way it could be wrong: if reporting is overwhelmingly generated by clinicians working from dispensing records rather than by patient self-report, product attribution may be accurate at the point of entry and the bias would not materialize. That is an empirical question and it is testable — which is the right note for students to end on.
Exercise 41.12
Claim: visual inspection of a person's face identifies GLP-1 receptor agonist use. Rating ❌. Reason: no reference standard, no published sensitivity or specificity, many alternative causes of facial volume loss, and base rates that defeat even a good test. Falsifier: a STARD-compliant diagnostic accuracy study in a representative sample with verified prescription records, showing accuracy sufficient to overcome plausible base rates.
Exercise 41.13
The claim must first be made rateable: "everyone" is an intensifier, not a quantity. A rateable version specifies a population, a definition of use, and a period — e.g. "more than X% of adults in [defined population] filled a prescription for a GLP-1 receptor agonist in [year]." Once specified, it is answerable from dispensing data, and the honest rating of the unspecified version is that it is not a claim. Full credit requires the student to notice that the rewriting is the work.
Exercise 41.14 †
Should reproduce the chapter's ⚠️: plausible, partly supported by the direction of coverage changes, confounded by shortage effects and by an unregulated parallel market, and dependent on an undefined term ("appropriate"). Strong answers notice that the claim can be simultaneously true for one population and false for another, and that a single rating therefore compresses away the thing you want to know.
Exercise 41.15
Rateable version: "patient-reported product identity matches dispensing records at a high rate." Current state: unmeasured, so ⚠️ at best and arguably 🔬 — the chapter rates the harm claim ⚠️ on mechanistic grounds plus precedent from other genericized classes. What settles it: a medication-history accuracy study. The good student move is noticing that the skeptic's claim and the chapter's claim have the same falsifier — one study resolves both, in opposite directions.
Exercise 41.16 †
Exercise 41.17 †
Take these two together, as they are designed to be. Both are stigma claims about the same period, pointing opposite ways, and both are currently 🔬 — there is a real research literature on weight stigma with validated instruments, but the population-scale question of what these drugs did to it is not settled. The exercise's purpose is to make students notice that they find one of the two intuitively obvious, and that its opposite is equally available. Credit answers that name a measurement: validated stigma instruments administered in comparable populations across the period, with attention to whether attitudes toward people with obesity and attitudes toward drug users moved differently. Chapter 44 owns this.
Exercise 41.18
The claim is partly true and the partial truth is the trap. The phrase does describe something real (Ch 8 §8.8). But phrases have uses, and the uses in §41.3 stages 3 and 4 are a diagnostic guess and an accusation, neither of which is a description. Good answers distinguish the three uses and note that the phrase's clinical surface makes them indistinguishable in practice — which is the reason the term is a problem even though the physiology is not.
Exercise 41.19
Grade only on strength of both cases and absence of tells. A response where one paragraph is noticeably more vivid, better-argued, or longer has revealed the author's position, which the exercise forbids. The for case should include norm-setting, the willpower frame, and commercial entanglement; the against case should include privacy-as-default, selective enforcement, harassment as the predictable consequence, and the absence of a stopping rule.
Exercise 41.20
The distinction: silence makes no claim; misattribution makes a false one. Clear cases are easy to generate. The genuinely hard case is the one to reward — typically something like a person who attributes results vaguely ("consistency and clean eating") without stating anything technically false. Good answers notice that this sits between the two, that its status depends on whether the statement functions as an explanation, and that adding a commercial interest tips it.
Exercise 41.21 †
The best available principled defense is usually some version of: disclosure is owed for interventions that are (a) not obvious from observation, (b) causally sufficient for the outcome, and (c) unavailable to the audience. It fails on all three adjacent cases — surgery satisfies (a) and (b), a personal chef satisfies (b) and (c), and a metabolic advantage satisfies all three while being undisclosable. Credit students who construct the line honestly and then report that it did not survive; that is the intended outcome, and pretending otherwise should score lower.
Exercise 41.22
NOT RATED because the rating system measures the state of evidence for a claim about the world, and requires a population, an endpoint, and a study that could settle it. "Ought" claims have none. ❌ would assert that evidence fails to support it, which is a category error. Rateable neighbor, e.g.: "disclosure by people with large audiences reduces harmful weight-control behavior among their audiences" — currently 🔬, no study designed to test it.
Exercise 41.23 †
What is scrutinized is the claim, not the person and not the silence. This matters because it is what lets the book refuse a disclosure duty while still evaluating the statement: a person who says nothing has produced nothing to evaluate; a person who asserts a causal story has produced an efficacy claim, and efficacy claims are evaluated regardless of who makes them. Weak answers say "the person's honesty," which reintroduces exactly the duty the chapter declined to impose.
Exercise 41.24
The nine: rapid loss, gradual loss, no change, regain, gaunt face, full face, denial, confirmation, silence. Any two plausible additions are acceptable — a change in exercise habits, a change in clothing, a public statement about health, an appearance at an event. The grading point is whether the student's additions would also be absorbed. If a student proposes an observation that would genuinely disconfirm, that is the best possible answer and should be discussed with the class.
Exercise 41.25
"Just saying it's likely" — a probabilistic claim needs a base rate and a likelihood ratio; neither exists, and a guess with "likely" attached is still a guess. "Everyone knows" — describes how widely a belief circulates, which is silent on truth. "Obvious from the pictures" — reports the viewer's confidence, not the subject's state; without a validated index test and reference standard it is not an observation about the world.
Exercise 41.26 †
"Stronger" means it requires fewer shared premises, so it reaches an audience the moral argument cannot. The premise the interlocutor must still grant is minimal and hard to refuse: that a conclusion no possible observation could disconfirm is not knowledge. Anyone who has accepted the book's treatment of supplement claims has already granted it. Watch for students who argue "stronger = more important"; the chapter explicitly does not say that, and the distinction is the exercise.
Exercise 41.27
A social fact is something true because a group treats it as true, independent of external measurement. Good non-medical examples: which neighborhood is "up-and-coming," which job title outranks another within a company, whether a queue is "long." In each case behavior changes on the basis of shared belief, and the belief is not tracking a measurement.
Exercise 41.28 †
Grade the inventory of losses, not the joke. A complete list: population, endpoint, comparator, effect size, duration, confidence interval, harms, funding, prespecification, and the distinction between the compound and the vial. The follow-up question — which loss they most want back — is diagnostic: students who say "effect size" are still thinking about magnitude; students who say "population" or "comparator" have internalized the method.
Exercise 41.29
Because the socially valid response to "is that true?" is "it's a joke," which ends the exchange without answering. The payload therefore enters the audience unexamined — and, since jokes require the audience to supply the premise, the payload is already inside the listener before any evaluation could occur. The premise was not transmitted so much as confirmed.
Exercise 41.30
Completion exercise. Look for source-type distribution rather than raw counts, and for the student noticing that they cannot easily tell rung 2 from rung 3 in live speech. Do not accept any tally that records who was being discussed.
Exercise 41.31 †
Grade the DIRECTION column above all. A log where every source is marked ↑ has almost certainly not been done carefully; comment threads commonly deflate and news commonly distorts. The reflection paragraph should identify which source came closest to Field 5 and whether it surprised them. The best responses notice that a person without the dossier would have had no way to rank the four.
Exercise 41.32
Completion only. Do not grade content and do not require students to share it. If a pattern emerges across the class — and it usually does, with "decline to speculate" reported hardest — it is worth naming aloud without attributing it to anyone.
Chapter 42
Instructor-only. The student exercises.md deliberately carries no answers. Items marked † are
open-ended; what follows for those is assessment criteria rather than a solution.
One global marking rule. Any answer that concludes "therefore this person is lying" or "therefore the claim is false" from a financial interest alone has committed the genetic fallacy and should lose credit regardless of how well written it is. The chapter's discipline is that incentive predicts which claims get made, not whether a given claim is accurate. This is the single most common failure across the whole set.
Part A — The funnel and its economics
A. Attention (a video or post; platform earns ad inventory, creator earns revenue share or sponsorship) → interest (a claim, usually mechanism or transformation; creator gains distribution, brand gains attention it did not buy) → low-friction consultation (often asynchronous, often free or nominal; the operating entity, sometimes the creator per qualified signup) → prescription (prescriber per encounter reviewed, dispenser per fill, sometimes creator per first fill) → subscription (recurring monthly, everyone downstream, for as long as it survives). Most commonly omitted: step 5. Students who stop at "prescription" are treating the structure as a sale rather than as a subscription, which is the error that hides retention content from them later.
B. Lifetime value accrues at the bottom; a single viewer at the top is worth very little in expectation. Rational allocation therefore puts creative effort where conversion is won (steps 1–2) and minimizes anything that reduces throughput at step 3. Prediction: step 3 is optimized for speed and completion rate, not for the probability of producing a correct "no."
C. Commercial: anything slowing progress toward conversion, measured as loss. Clinical: delay or effort that functions as a check. Protective examples: the scheduled visit that forces a history; the pharmacist's independent review; the follow-up appointment that catches a side effect; a disqualifying screening question. Genuine waste: repeat submission of the same demographic data across systems, fax-based referral, travel required purely for a signature. Credit answers that recognize the metric cannot distinguish the two, because it has no term for a harm a delay prevented.
D. Access: specialist deserts; travel/time/cost barriers; shame as a first-contact barrier; consistency of screening. Pressures: revenue tracks fills; brief asynchronous encounters; no owner of month three; commercial linkage of advertiser, prescriber, dispenser. The paragraph must not resolve — an answer that lands on "so telehealth is bad" or "so it is fine" has failed the item.
E. † Marking: did the student find stated disqualifying criteria, or only inferred ones? Strong answers quantify how much of the public page is devoted to speed and approval versus to contraindications, and notice that "who owns follow-up" is usually unstated. Weak answers review the service's honesty rather than its structure.
F. In conventional care, advertising, prescribing, and dispensing sit in entities with different and partly opposing interests; a transaction requires three parties to agree. Integration removes the disagreement, not the good faith. Non-medical analogues: separation of audit from consulting; restrictions on brokers trading against clients; historic separation of film production from exhibition. Credit any analogue where the rule targets structure, not conduct.
G. Best objection: fee-for-service and procedure-driven medicine carry volume incentives too; singling out remote care is availability bias driven by novelty. Best reply: the claim is about integration and continuity, not about screens, and the specific configuration merges functions that conventional arrangements keep separate. Strong answers name the evidence that would move them (comparative appropriateness data with compensation structure recorded).
H. † Marking: does the design actually cost money? Look for salaried rather than per-encounter clinician pay, a named follow-up owner, a synchronous escalation path, symmetric cancellation, and a published declination rate. The second half of the item — which choices reduce profitability and why — is where the credit is.
Part B — Affiliates, disclosure, and selection
I. Per click → curiosity. Per signup/qualified lead → registration. Per first fill/conversion → purchase. Recurring revenue share → persistence. Flat retainer → none directly (relationship continuation instead). Discount/referral code → purchase, plus attribution. Equity → persistence and enterprise growth.
J. Starting-aligned content emphasizes novelty, transformation, and low barriers; continuing-aligned content emphasizes reassurance, troubleshooting, patience, and community. The second is far harder to recognize as commercial because it is indistinguishable in form from genuine support — which it often also is.
K. (i) Does not change the incentive; (ii) does not undo the selection effect; (iii) does not address transferred trust (and does not survive clipping). Any ranking is defensible; the argument is marked, not the order. The strongest answers rank the selection effect first on the grounds that it operates on the entire population of content simultaneously while the other two operate per item.
L. The distortion resides in content that was never produced, so it leaves no trace in any artifact an auditor can examine. Auditing N videos samples only from the published population; increasing N does not reach the unpublished one. This is the same structure as publication bias in a literature — see Chapter 6 and the Goldacre entry in further reading.
M. † Marking: the estimate itself is worthless; the second sentence is the item. Credit answers stating that the ratio is consistent with a selection effect, with genuinely good products, and with sampling bias in the student's own viewing — and that it distinguishes none of them.
N. A strong rewrite: "I'm paid for this. I've used it and I liked it. But you should know I don't make videos about the things I try that don't work, so you're not seeing a fair sample — and 'it worked for me' is one uncontrolled observation by someone who wanted it to work." What remains unaddressed even then: the incentive is unchanged, and the audience's trust was built elsewhere.
O. † Marking: placement (before the claim, not after), persistence (burned into the frame so it survives clipping), legibility without audio, and specificity about the payment form. Full credit requires the honest admission at the end — that no format touches the selection effect.
Part C — Content and platform economics
P. Expect: subgroup dropped first, then comparator, then population, then duration, then the uncertainty, then the fact of a study. Marking: does the student notice that every cut improves performance and that none requires a false statement?
Q. Acceptable without the forbidden word: the system distributes what people watch longest; qualifiers consume seconds and deliver no payoff; audiences leave slightly sooner; the content is shown to fewer people; the creator learns from the numbers, not from a decision to be less careful.
R. The narrative supplies protagonist, before, after, and duration; the viewer supplies causation. Because no causal sentence was uttered, there is nothing to retract and no claim to substantiate — which is precisely why it is both effective and difficult to regulate.
S. (1) A keyword triggers demonetization or suppression. (2) Creators substitute euphemism, misspelling, or oblique reference. (3) The circulating claim now concerns a vaguer object. (4) It cannot be matched to a literature, checked, or rebutted — so precision falls as a side effect of a policy intended to reduce harm.
T. §42.5's core argument is architectural: evidential quality is not in the feature set, which is a fact about what a ranking system operates on. The diffusion study is empirical, drawn from a predominantly political corpus with diffusion as the endpoint. The architectural claim can be held far more confidently; the empirical one transfers to health content only by structural analogy.
U. † Marking: the two ratings should often differ, and the best answers conclude that "rate the claim, not the molecule" needs a further refinement — rate the claim as stated in the version in front of you — because compression can change the claim's evidential status without changing its topic.
V. Best answers argue common cause: confidence, novelty, and narrative are all properties that make a claim easy to transmit, and transmissibility is orthogonal to warrant. Chapter 5 ranks them low because they are cheap to produce; ranking systems reward them for the same reason. Credit a well-argued "coincidence" answer only if it engages the mechanism.
Part D — Pipelines, regulation, compounding
W. (i) Signaling genuine expertise — real, and usually accurate within scope. (ii) Transferring authority beyond scope — happens automatically, because a general audience cannot see scope and a thumbnail does not display it. (iii) Deceiving — requires intent. Only (iii) requires intent, which is why the chapter treats the credential problem as structural.
X. Speak carefully and lose most individual matchups against confident content; stay silent and cede the field to people with no obligations; build an audience and accept that it becomes a patient pipeline. Marking: has the student named the cost of their own choice?
Y. Manufacturer (regulated product → strictest); compounding pharmacy (a pharmacy service → no label to constrain claims to); supplement seller (a food-category product → structure/function regime); telehealth service (a clinical service plus dispensing → uneven, jurisdictionally complex); creator (attention → general advertising and endorsement law, thin in practice). Thesis: the gap is not an oversight; it follows from what each entity is legally selling.
Z. Promotional rules attach to an approval. The manufacturer's claims are tethered to a label, must carry risk information with fair balance, and may not extend beyond approved uses. A preparation with no approval has no label to be tethered to and falls back to general advertising law, which prohibits deception but says far less about what may be claimed for a medicine.
AA. † Marking: does the chosen category actually contain a conditional regulatory opening, or merely an unregulated market? The template requires an opening that closes. Credit honest assessments of where it breaks.
AB. "Same active ingredient" → may contain the same molecule; identity, purity, concentration, and sterility are separately assured, or not. "Compounded" → prepared by a pharmacy for a specific prescription, outside the approval system. "Generic" → not applicable; there is no equivalence demonstration. "Pharmaceutical/research grade" → not a regulatory category; a marketing phrase. "Personalized" → formulated differently from the approved product, which is what places it in a different regulatory position.
AC. Constraints: list price, inconsistent coverage, exclusion of the patient's indication, unavailability during shortage. A "gullible patient" account predicts that better information would have changed behavior; the constraint account predicts that only price, coverage, or supply would. The second predicts the observed migration and persistence of demand; the first does not.
AD. † Marking: a good test names a public, countable signal (advertising language, category of product marketed, jurisdiction of operation) and specifies a falsifying result in advance. Credit the acknowledgment that channel migration is easy to see and hard to attribute.
AE. Any correct example of a filtered sample: restaurant reviews written only by people who returned; alumni surveys reaching only those who stayed in touch; a repair shop's reputation among the customers whose repairs succeeded. Key sentence: the filter correlates with the thing you were trying to measure.
Part E — The read and the dossier
AF. Five answers, then the sixth: this does not tell me the claim is false. Marking: penalize any answer that treats the read as a verdict.
AG. † Marking: did it actually cost something? An answer reporting that the congenial content came through clean may be honest, but it should say what would have counted as a problem. Answers that report discomfort and specify where are the strongest.
AH. Genetic fallacy: "they're paid, so the molecule doesn't work" — rejects a claim by source. Legitimate selection inference: "they're paid, so I should not treat the absence of negative content in this niche as evidence of an absence of negative outcomes" — a conclusion about the sample, not about the claim. The self-assessment at the end is marked for specificity, not for which answer.
AI. † Marking rubric: (1) claims listed as encountered, not as the student wishes they were stated; (2) roles named, never people; (3) the "who gains if the opposite is true" column completed honestly, including where the answer is "nobody with a budget"; (4) Step 2 genuinely attempted, with at least one result identified that runs against its producer's interest; (5) the time comparison recorded — Step 2 taking much longer than Step 1 is the finding, not a failure.
AJ. † Marking: no naming, no accusation, no instruction to stop a medication, short. The best submissions ask a question rather than deliver a conclusion — most often some version of "do you know who you'd call if something felt wrong in a month?" — which conveys the entire chapter without triggering defensiveness.
Quiz answer key
The full key is embedded in quiz.md inside the <details> block and is not duplicated here. Items
most often missed in review: 7 (paid tag is not evidence of unreliability), 15 (the
less-constrained inversion), and 21 (selection effect identified, falsity not established).
Chapter 43
Instructor copy. Not for inclusion in the student edition.
The exercises file deliberately publishes no answers. Roughly a third of the items have no correct answer, and printing a key beside them would suggest otherwise. What follows is assessment guidance: what a strong response contains, what a weak one looks like, and — for the ethics items — what the common failure modes are.
A general standard for this chapter. Across every item, the mark of a strong answer is that the student separates empirical claims from value claims and says which is which. A student who argues confidently for a position without noticing that half their premises are contestable facts has missed the chapter regardless of which position they took. Conversely, a student who reaches a conclusion you disagree with, having stated the opposing view fairly and flagged their own value commitments, has done the work.
Group 1 — Recall and definition (A–H)
A. Diagnosis; indication; clinician with a duty; individualized risk-benefit judgment made under that duty. Full marks require the qualifier on the fourth — "a risk-benefit judgment" alone is incomplete, since the enhancement user also makes one. The distinguishing feature is who makes it and to whom they owe a duty.
B. Positional: promotion competition, rankings, admissions, roster spots, status. Non-positional: personal health improvement, literacy, recovery from injury. Watch for students who classify income as purely positional — it is partly positional, and the "partly" is the interesting part.
C. ❌ describes the evidence, not the molecule. It means the confident version of a claim is not supported by human data; it does not mean the compound does nothing, and it is not a prediction that future evidence will be negative.
D. † Look for all five steps and, critically, for step 3. The commonly dropped step is 3 — the one where nothing happens and the harm enters. A student who jumps from step 2 to step 4 has reconstructed a story about peer pressure rather than the structural argument, and the difference is worth naming in feedback.
E. Externality = a cost falling on someone other than the decision-maker. The word does real work because it identifies the harm as unintended and unattributable to any act, which is why moral criticism of individual adopters misses it.
F. A mechanism permitting an athlete with a genuine condition to use an otherwise prohibited treatment. It concedes that the same molecule is therapy or enhancement depending on the situation of the recipient — i.e. that the prohibition is not about the substance.
G. (1) safety at relevant exposure — empirical; (2) genuine informed consent — partly informational, partly structural; (3) absence of coercive structure — structural; (4) risk borne by beneficiary — institutional; (5) reversibility — empirical. The key discrimination is that 3 is the odd one out.
H. Reversibility limits the cost of being wrong. Under thin evidence, the probability of being wrong is high, so the recoverability of the decision carries more weight than it would under good evidence. Strong answers note that reversibility is currently believed rather than demonstrated, because demonstrating it needs the same long-term data condition 1 lacks.
Group 2 — Comprehension (I–P)
I. An inventory lists absences; a scorecard grades them. If it were a scorecard, the chapter would be arguing that enhancement is wrong, which it explicitly is not. Absences are not automatically wrongs.
J. † The common cause is incomplete clinical development. Compounds that finished development have indications, prescribers, generic names, monitoring guidance, and pharmacovigilance. Compounds that did not are cheap, unnamed, unstudied, and available. Availability and evidence are close to inversely related because they have the same driver. A student who merely observes the correlation gets partial credit.
K. It has an enforced, adjudicated answer, so the consequences are observable rather than predicted. It can tell us what such an answer costs, which no argument can establish a priori. Strong answers note the enforcement's own class of victims.
L. No representative survey, no defensible estimate; inventing one would violate the book's rules. To produce a defensible figure you would need a defined population, a sampling frame, a validated measure that survives incentives to conceal, and a stated definition of use.
M. The consequences of declining differ when the offeror controls your assignments, evaluation, and advancement — even with a genuine right of refusal, no recorded decision, and complete good faith. Look for at least one of: peer visibility, the indistinguishability of "declined" from "not selected" over time, or the social fact of being seen as unwilling.
N. Because a problem caused by bad motives can be solved by changing motives. This one cannot. Good motives remove the obvious remedy without removing the problem.
O. † Two parts. (1) GH administration in healthy adults increased lean body mass without demonstrated increases in strength or exercise capacity. (2) It is corrosive to testimony because the visible effect is real and the functional absence is invisible to the user — personal experience confirms the effect and cannot distinguish it from the benefit.
P. Safety: the black-market failure modes are documented and supervision plausibly addresses some of them. Fairness: it does not address coercion, transfers risk to non-participants, and equal availability has never been implemented.
Group 3 — Application (Q–X)
Q. † Marked on the quality of the mapping, not the conclusion. Require all four features to be addressed explicitly, including any that are absent. The best answers identify a decisive feature and say why it dominated.
R. It addresses condition 2 partially (autonomy, self-funding) and possibly condition 4 (bearing own cost — though only the financial cost, not the risk asymmetry with the employer). It leaves conditions 1, 3, and 5 entirely untouched. The strongest answers note that "nobody's business but mine" is precisely the claim condition 3 denies.
S. Must contain a population, an endpoint, a timeframe, and a stated falsifier. A rewrite that is merely more specific but still unfalsifiable gets no credit.
T. Candidates: appealing to authority by association; implying prevalence without asserting it; transferring credibility from a high-status setting; establishing a social norm; avoiding a regulated efficacy claim while producing the belief one would produce. Three distinct items required.
U. † Grade against the four-line format strictly: claim, rating, one-sentence reason, and a falsifier that is genuinely observable. The most common failure is a "what would change it" line that restates the reason rather than naming an observation. Any rating from ⚠️ to 🔬 is defensible; ✅ is not, and ❌ requires arguing that the claim has been tested and failed, which it has not.
V. Either answer is defensible. Strong answers identify the discriminating feature: is the appearance outcome being compared against others (positional) or against a private standard? In appearance-adjacent and customer-facing work the case for "strongly" is real; in private contexts it is weak. The point is that the setting, not the compound, decides.
W. Assessed on whether item 5 (attributable how?) was actually answered rather than deferred. The honest answer for most cases is "not attributable," and students who write that plainly have done the exercise correctly.
X. Conditions 1 and 5 satisfied; 2 improved but not satisfied (real information yes, real alternative no); 3 and 4 exactly as before. The intended realization: safety research was never going to touch condition 3. That is the payoff of the item.
Group 4 — Analysis (Y–AB)
Y. † The chapter's argument: steps 3–5 all still occur on a false belief, so the population pays the same cost with an empty benefit column. Strong attacks: an inert compound may carry lower physiological risk, so total harm could be lower even if the benefit-to-cost ratio is worse; or the false belief may be more correctable than a true advantage. Credit the student for distinguishing total harm from harm per unit of benefit — that distinction is the whole item.
Z. The strongest objection is that the four-element clinical structure is idealized: real practice includes rushed consultations, diagnostic uncertainty, incentives that conflict with duty, and prescribing that is barely individualized. Best answers conclude it qualifies rather than damages the argument, because the four elements are still present in kind and entirely absent in enhancement — a degraded version of a thing is not the same as its absence.
AA. † The strongest pro-rating argument: the claim does real work in the debate, refusing to rate it leaves it unchallenged, and a 🔬 with an explicit "no evidence base" note would engage rather than sidestep. The chapter's response: rating an unfalsifiable claim degrades the system's meaning and implies an evidence base that does not exist. Credit either conclusion; require that both sides be stated.
AB. Cost falls on athletes, disproportionately the less-resourced. Foreseeable — Chapter 34's supply-chain problems make it predictable. It counts against this enforcement design (strict liability with limited verification support) more than against having an enforced answer at all; strong answers propose what a better design would trade away.
Group 5 — Ethics dilemmas (AC–AG)
No answers exist for these. Assess the reasoning.
Shared rubric. (i) Was the rejected position stated at its strongest? (ii) Were facts and values separated explicitly? (iii) Was the cost of the student's own conclusion named? (iv) Did they avoid naming individuals or institutions, per the standing rule?
AC. † The baseline has moved. Both answers are available. "Not free" must explain what it is instead — constrained, structurally coerced, a choice among options someone else narrowed. "Free" must do work on the word: freedom as absence of interference (satisfied) versus freedom as availability of acceptable alternatives (not satisfied). The failure mode is asserting one sense of "free" without noticing there are two.
AD. Options should include: do nothing; ask openly; propose a policy; change the metric; change the reward structure; seek institutional guidance. The item's payoff is item-by-item cost attribution, and the expected insight is that doing nothing has costs too, and they fall on the non-adopters who are not in the room.
AE. † Plausible safeguards: decision-making by someone with no authority over the person; non-recording of the decision; genuine anonymity of participation; a real and comparable alternative path; independent medical advice. The honest finding is that an institution can deliver the procedural safeguards and cannot deliver the social ones — peers still observe, and reputation is not administrable. Students who conclude "therefore it cannot be made voluntary" should be pushed on whether that proves too much, since the same is true of many accepted workplace choices.
AF. Look for: the value of continued disclosure outweighing the satisfaction of disapproval; honest statement of the limits of the clinician's own knowledge; no abandonment; documentation; safety-netting without instruction. Any answer that has the clinician supply protocol guidance is a fail — that is outside the book's rules and outside the clinician's evidence base. Chapter 39 is the reference.
AG. † Graded on the second part. Proposals typically address supply, supervision, documentation, and reporting. The unanswerable objection is almost always condition 3 — the proposal increases adoption and therefore accelerates the coercion structure it cannot address — or unequal availability. A student who cannot find an unanswered objection should be sent back to §43.6; the inability to find one is the diagnostic.
Quiz key
The quiz key is published inline in quiz.md inside the <details> block. Correct answers: 1 b, 2 c,
3 c, 4 b, 5 c, 6 c, 7 b, 8 b, 9 b, 10 c, 11 b, 12 c, 13 b, 14 b, 15 b, 16 c, 17 b, 18 c, 19 b, 20 c,
21 c, 22 c.
Items worth reviewing aloud regardless of class performance: 4 (why the reason for the ❌ matters), 6 (step 4), 18 (the non-rating), and 21 (condition 3). Those four carry the chapter.
Chapter 44
Instructor-facing. The chapter's exercises deliberately have no published answers, because most of the questions concern a future nobody has observed. What follows is what a strong answer contains, and — more usefully — the specific wrong turns that recur.
Grading principle for this chapter: reward the student who says "I don't know, and here is what would tell me." Penalize confident answers that could not be checked. This inverts the usual grading instinct and should be stated to students in advance, or they will write what they think you want.
Exercises — Part A (§44.1)
A. Strong answer identifies the unit problem: trials operate on individuals; these questions operate on markets, norms, and institutions. Must reach "there is no control society." Common error: answering "because it would be unethical" — the prompt forbids the word ethics precisely because that is the reflexive and wrong answer.
B. Counterfactual = what would have happened otherwise. For the food claim, the counterfactual is the same country in the same years without the drugs, which cannot exist. Accept answers proposing a comparison country as an approximation, provided the student names the assumption (that the comparison country differs in adoption and not in everything else) and admits it is not a counterfactual.
C. Four alternatives should include at least: reverse causation is implausible here but selection is not (regions with high prescribing may differ in income, health behavior, and food retail mix); a common cause (a public health campaign, an economic shift); measurement artifact (category redefinition, retailer mix change); and secular trend (the category was already declining). Ranking by ruleoutability is the discriminating part — secular trend is easiest, unmeasured common cause is hardest.
D. Any non-medical example works. Watch for students who define only one direction; the reverse direction (inferring aggregate patterns from individual findings) is the one usually missed.
E. Any of the three tools is acceptable. The grading target is the assumption: parallel trends for difference-in-differences, no co-occurring shock for interrupted time series, as-if-random assignment for a natural experiment. Placement above correlation and below RCT should be justified, not asserted.
F. † The point of the exercise is the moment the trail goes cold. Full credit requires naming that moment specifically. A student who reports a fully traceable, well-sourced number has probably found a genuine analysis — ask them to identify its sensitivity assumptions instead.
Exercises — Part B (§44.2)
G. Four causal steps: drug reduces intake → aggregate demand shifts → firms observe sales data → firms reformulate. Deduct for steps that skip the observation stage; students frequently jump from consumption to product design without the mechanism that connects them.
H. Direction is a mechanism claim; magnitude is a measurement claim. Only the second can carry a rating with any content — which is exactly why the chapter's rating is 🔬 rather than ⚠️.
I. The self-observation designs will all be weak, and saying why is the assignment: no counterfactual, no sampling frame, confounded with the student's own purchasing changes, and subject to confirmation bias in what gets noticed.
J. The negligible-effect case should include: adoption fraction constrained by cost and coverage, discontinuation, substitution rather than reduction, and firm adaptation. A student who cannot construct this half is arguing from a conclusion.
K. † Completion credit. The pedagogical work is done at the moment of dating the baseline.
L. Externality (positive). Non-medical examples abound — vaccination herd effects, safety regulation, network effects.
Exercises — Part C (§44.3)
M. Must reach "the problem is created by the drug working." A failed therapy is a small line item.
N. Cost-effectiveness is per-patient value; budget impact is aggregate near-term spend. A therapy can pass one and fail the other, and public argument routinely conflates them.
O. The single-payer modification: the entity paying is the entity capturing, so the misalignment disappears — though the budget-impact problem does not. Students often conclude that single payer solves everything; it solves the churn problem specifically.
P. Structural remedies: longer-horizon accountability, risk pooling that survives member movement, price reduction, unified purchasing. The "cruelty" diagnosis yields shaming, publicity, litigation. The teaching point is that different diagnoses generate different remedies and only one class of remedy addresses the incentive.
Q. † Expect wide variation by country. Grade on the published-vs-anecdotal distinction, which is the actual learning objective.
Exercises — Part D (§44.4)
R. Both mechanisms must be stated in their strong forms. Reject a version of the intensification mechanism written as obviously wrong; that is the failure mode this section exists to prevent.
S. Attribution of controllability. The explanation must specify what is being attributed — the condition, or the failure to treat it.
T. The reconstruction is straightforward; the second half is the graded part. The defensible position is that this is not an argument for narrower access, and a student who concludes otherwise should be pressed on what they are trading away.
U. Wide latitude. Look for a clean identification of the enabling condition: the fix became available and widely known to be available.
V. Non-informative rating launders a coin flip. Students who would force a rating should be asked what they lose; the answer is that a reader cannot distinguish "we assessed this and it leans X" from "we had to write something."
W. † Enforce the naming rule strictly. A student who names an individual has demonstrated the chapter's point about themselves, which is worth saying gently in feedback.
X. Routes: care avoidance, disordered eating, psychological harm, differential clinical treatment.
Exercises — Part E (§44.5–44.6)
Y. Table should be complete on both sides. A one-sided table is the most common submission and should be returned rather than graded.
Z. † The crowding-out mechanism requires no intent because attention is a scarce resource allocated by perceived need. The second half — what evidence would show it actually occurred — is the harder and more valuable part; accept funding time-series, policy-agenda analysis, or documented program terminations.
AA. † The compatibility essay is easy; the self-rebuttal is the assignment. Grade the rebuttal.
AB. Any phase placement is defensible with reasoning. The falsifier is the graded element.
AC. Must reach: the curve intersects an inverse distribution of need, so the gap is structural. Must also reach: structural causation does not reduce the harm.
AD. † Good-analogy features: peptide, chronic, large population, long horizon, biosimilar pathway. Poor-analogy features: different population size and necessity, different competitive landscape, different political salience, different endpoint immediacy. Most students find the second list harder — that asymmetry is itself the lesson and should be named in feedback.
Exercises — Part F (§44.7–44.9)
AE. Prediction is unfalsifiable until too late; conditions are checkable as evidence arrives.
AF. Persistence data (either list's Condition 1) is the defensible answer for soonest resolution. Accept price/competitive entry with justification.
AG. † Both arguments must be constructed in good faith. The strongest version of "it's a dodge" is that refusing to weight is itself a choice that implies 50/50, which the author does not defend. A student who finds that argument gets full marks regardless of their conclusion.
AH. Values question, evidence-informed but not evidence-settled, and recognizing one is part of the method because the alternative is laundering a position through scientific vocabulary.
AI. † The meta-questions are the assignment. The question students most want to skip is almost always #3 (who funded it) or #5 (what am I not being shown).
AJ. † Completion plus the thesis statement. Both halves of the thesis must be present and unsoftened.
Quiz answer key (short form)
1-b · 2-c · 3-b · 4-c · 5-b · 6-b · 7-b · 8-c · 9-d · 10-b · 11-b · 12-b · 13-c · 14-b · 15-c · 16-b · 17-a · 18-b · 19-b · 20-c · 21 no key · 22 no key
Items that discriminate. Q9 (the distractor is the most confident-sounding option), Q12 (students frequently invert the access logic), Q13 (students resist NOT RATED as an answer), and Q17 (shape vs. timeline). If a cohort misses Q12 in bulk, reteach §44.4's access argument before proceeding.
Q21 and Q22 are ungraded by key. Q21 fails if either half of the thesis is softened. Q22 is completion plus a checkable falsifier.