58 min read

> "In the fields of observation, chance favors only the prepared mind."

Prerequisites

  • 33
  • 2
  • 5

Learning Objectives

  • Name the four routes by which peptide drugs are discovered and place any given molecule on one of them
  • Explain why venom is a disproportionately productive source of drug leads, and why venom peptides tend to be selective
  • Tell the exendin-4 story accurately, including what the lizard solved and what chemistry still had to finish
  • Trace the peptide-lead-to-small-molecule arc through captopril and explain why it recurs
  • Describe how display technologies work and state precisely what they do and do not do
  • State what AlphaFold-class structure prediction changed and four specific things it did not
  • Explain why faster discovery has not produced proportionally more approved drugs, and name the actual constraint
  • Apply the general rule that improving one stage of a pipeline improves the pipeline only if that stage was the bottleneck

Chapter 35: Peptide Discovery — From Venom to Artificial Intelligence

"In the fields of observation, chance favors only the prepared mind." — Louis Pasteur, lecture at the University of Lille (1854)

Overview

Here is a fact that ought to be more famous than it is.

The most consequential class of peptide drugs of the last twenty years — the GLP-1 receptor agonists, the drugs behind Chapters 7 and 8, behind the shortages and the magazine covers and the largest shift in obesity medicine in a century — begins with a lizard. Not metaphorically. The first drug in the class was a synthetic copy of a peptide found in the venom of the Gila monster, a slow, heavy, bead-skinned animal that lives in the Sonoran Desert and bites almost nobody.

That is a good story, and this chapter tells it properly, because it is usually told badly. It is also a trap, and the trap is the reason this chapter exists. The lizard story invites a conclusion — nature knows best, the answers are already out there — that the rest of the same story flatly contradicts. The lizard peptide became a drug that worked and was displaced within a decade by molecules that were engineered, not found. Nature supplied the lead. Chemistry supplied the drug. Both halves are load-bearing, and people who love the first half usually skip the second.

This chapter is about where peptide drugs come from. There are four routes, they are roughly chronological, and all four are in heavy use right now — the newest has not retired the oldest. We spend real time on the oldest route, because it is the most charming and most misunderstood, and on the newest, because that is where the hype is densest.

And then, in the last section, we ask what all this acceleration actually bought. Every technique here makes it faster and cheaper to find a molecule that binds a chosen target. That is a genuine achievement, and it improves a stage of the process that was never the problem. Chapter 22 already showed you a class of compounds that bound their target beautifully and did nothing for patients. Faster generation of those would not have helped.

In this chapter, you will learn to:

  • Name the four discovery routes and identify which one produced any drug you encounter
  • Explain why venoms are pre-optimized pharmacological libraries, and why that follows from evolution
  • Tell the exendin-4 story accurately, including the irony at its center and the correction to it
  • Trace the arc from venom peptide to injectable peptide to oral small molecule, twice
  • Describe what display technologies select for, and what they conspicuously do not do
  • State precisely what AlphaFold changed and four things it did not
  • Explain why faster discovery has not produced proportionally more approved medicines
  • Apply a general rule about bottlenecks that will outlive every technology named here

Learning Paths

💊 GLP-1 — §35.3 is the chapter for you and possibly the chapter for the whole book. The drug your clinician prescribes has a lizard in its family tree, and understanding exactly what the lizard did and did not contribute is the difference between a fun fact and an actual understanding of the class. §35.5 explains why the nausea is not a design flaw. 🏋️ Performance — §35.5 and §35.10. Almost every compound in Part III arrived by route 2 — somebody found a fragment of a human protein and asked what it might do. §35.10 explains why "there are hundreds of promising peptides in the literature" is a statement about the width of the funnel, not about what comes out the bottom. 🔬 Science — read straight through; this is your chapter. §35.6 through §35.9 are the technical core, and §35.8's four limitations are the ones you will find yourself explaining to other people for years. 💄 Cosmetic — §35.5 and §35.9. Cosmetic peptides are overwhelmingly short fragments of human structural proteins, which is route 2 done cheaply; §35.9 is what route 4 would look like if anyone applied it seriously to skin, and mostly nobody has. 🏥 Clinical — §35.4's ziconotide discussion and §35.10. The first is the most vivid delivery problem in clinical medicine; the second is the answer to every patient who asks why, with all this technology, their disease still has no treatment.


35.1 Four ways to find a peptide

Every peptide medicine in this book arrived by one of four routes.

THE FOUR ROUTES TO A PEPTIDE DRUG
                                                            roughly
                                                          chronological
  1. FIND IT IN NATURE                                          │
     Some organism already makes a molecule that does           │  ancient →
     something. Isolate it, identify it, copy it.               │  1920s–
     → insulin (from pancreas), exendin-4 (from venom),         │
       magainins (from frog skin), conotoxins (from snails)     ▼

  2. START FROM AN ENDOGENOUS LIGAND AND MODIFY IT
     The body already makes the signal. Change it so it
     lasts longer, binds harder, or resists an enzyme.          1970s–
     → insulin analogs, semaglutide, octreotide, desmopressin
       (this is the entire subject of Chapter 33)

  3. SCREEN AN ENORMOUS LIBRARY AND SELECT WHAT BINDS
     Make billions of variants. Wash them over the target.
     Keep whatever sticks. Sequence the winners.                1985–
     → phage display, ribosome display, mRNA display

  4. DESIGN IT COMPUTATIONALLY
     Predict or model the target's structure, then build a
     molecule to fit it — or invent one from nothing.           2000s–
     → structure-based design, de novo design                        2020s

  Note the arrow, and then note what it does NOT mean. All four routes are in
  active, well-funded use today. Route 4 has not retired route 1. Somebody is
  milking a cone snail this week.

The routes are not sealed off, and the most interesting drugs involve several. Exenatide is route 1 chased by route 2. Captopril is route 1 converted by route 4 as it existed in the 1970s. Semaglutide is route 2 informed by route 4. The routes ask where did the first useful molecule come from, not what a finished drug is made of.

Two things are worth noticing.

The routes differ in what they require you to know in advance. Route 1 requires almost nothing — you can find a peptide in venom without knowing what it binds. Route 2 requires the physiology cold: the hormone, the receptor, and a reason to want more of the signal. Route 3 requires the target but nothing about the molecule. Route 4 requires the target's three-dimensional structure, which until recently was the hardest thing on this list to get.

And — this is the thread that runs to the end of the chapter — every one of these routes solves the same problem. They all answer how do I obtain a molecule that binds this target? None answers will binding this target help a sick person? Hold that distinction. It is the whole of §35.10.

🔍 Check Your Understanding

  1. Which route requires the most prior knowledge of physiology, and why does that cut both ways?
  2. A company announces it has "discovered" a peptide that binds a receptor implicated in a disease. Which of the four routes could have produced that announcement, and which question does the announcement leave completely unaddressed?

35.2 Venom: the richest source the field has

Start with the route that has produced the most surprising wins per unit of effort.

Venoms are pre-optimized pharmacological libraries. That phrase is worth unpacking slowly, because it explains a great deal of what follows.

An animal that uses venom to subdue prey or deter predators is under relentless selective pressure to produce molecules that do a very specific set of things. The venom must act fast — a snake that immobilizes a rodent in an hour has lost the rodent. It must act at very low doses, because the animal delivers a small volume and producing venom is metabolically expensive. And it must act on vertebrate physiology — on nervous transmission, on blood pressure and coagulation, on muscle contraction — because those are the systems whose failure incapacitates an animal quickly.

Now read that list again. Fast onset. Potency at minute quantities. Action on nervous, cardiovascular, and muscular targets in vertebrates. That is precisely the design brief of a drug. A company writing a target product profile for a new analgesic or antihypertensive is, without meaning to, describing the specification that evolution has been optimizing against in venomous lineages for tens of millions of years. Venom research is therefore not a novelty beat. It is a systematic search of a chemical space that has already been searched, by a process with an enormous head start, against criteria that substantially overlap with our own.

Why venom peptides are so often selective

Here the intuition runs the wrong way. Venom sounds like a blunt instrument; it is generally the opposite, and the reason is economic. Venom is expensive to make and slow to regenerate — a snake that has emptied its glands is temporarily disarmed. Under that constraint, a molecule binding forty targets weakly is a poor investment, while one binding a single target with picomolar affinity — say one subtype of voltage-gated sodium channel essential for nerve conduction in the prey species — does the job at a fraction of the material cost.

Specificity is cheap and effective; indiscriminate disruption is wasteful. A venom that disabled everything it touched would have to be delivered in bulk to disable anything in particular. Selection therefore pushes venom components toward exactly the property pharmacologists spend their careers trying to engineer: hitting one thing hard and leaving the rest alone. So venoms are not soups of general poisons but collections of precision tools — a single snake or cone snail venom may hold dozens to hundreds of distinct components acting on different targets, a combination therapy evolution arrived at independently many times.

The structural bonus

There is a second gift, and Chapter 4 makes you appreciate it.

Many venom peptides are compact, heavily cross-linked structures — short chains stapled into a rigid shape by multiple disulfide bonds (§1.2 for the refresher) — and several distinct families of this kind recur across unrelated venomous lineages, a strong hint of convergence rather than inheritance.

That architecture is not decorative. A tightly cross-linked peptide is much harder for a protease to chew, because proteases need to thread an extended stretch of backbone through their active site and a stapled molecule offers very little. Venom has to survive contact with prey tissue long enough to work. So venom peptides frequently arrive pre-solved for the stability problem Chapter 33 spends an entire engineering program attacking.

🧬 The Molecule — what a venom peptide typically looks like

Sketch the general shape rather than any one molecule.

```text A DISULFIDE-STAPLED VENOM PEPTIDE (schematic)

   loop        loop
    ___         ___
   /   \       /   \

N---| |=====| |---C === = disulfide bonds _/ ‖ _/ ‖ = a third bridge threading | ‖ | through the ring formed by └══════════─┘ the other two

Typically 10–40 residues. Two, three, or four disulfide bonds. The bridges convert a floppy chain into a rigid, protease-resistant scaffold with a small number of exposed residues doing the binding. ```

Two properties follow, and both matter therapeutically.

It is stable. Rigid, cross-linked peptides resist proteolysis and heat far better than a linear chain of the same length.

It is modular. The scaffold holds the shape while a handful of surface residues do the recognizing, so chemists can sometimes keep the frame and swap the binding face — a natural venom scaffold used as a chassis for a new function. It is a neat illustration of Chapter 1's principle: sequence is what you control, shape is what the receptor sees, and the engineering happens in the gap.

One caution. Stability is not drug-likeness. A protease-resistant peptide is still too large and too polar to be swallowed, still cleared by the kidney, still unable to reach the brain. Venom hands you one hard problem solved, not the rest.

The obvious caveat, stated once

Venom peptides are optimized to incapacitate an animal; drugs are meant to help one. Those overlap in mechanism and diverge completely in intent. A molecule selected for its ability to stop nerve conduction is, by construction, dangerous, and converting one into a medicine is largely the work of finding a route and a target subtype where the useful effect and the lethal effect separate — the therapeutic window of Chapter 6. Most never separate. The examples in this chapter are survivors of a process that discards nearly everything, which is a sentence you should expect about every route here.


35.3 The Gila monster, told properly

Now the centerpiece.

In the early 1990s, work associated with John Eng identified a peptide in the venom of the Gila monster (Heloderma suspectum). The search was not random: there were prior reports that Heloderma venom produced effects on the pancreas in laboratory animals, and Eng had access to sensitive immunoassay methods for detecting hormone-like peptides. The peptide he characterized was named exendin-4.

Exendin-4 shares roughly half its residues with human GLP-1, and that number sits in exactly the right range to be interesting. Much higher and it would be a trivial species variant of the same hormone. Much lower and it would not activate the human receptor at all. At roughly half, you get a molecule similar enough to switch on the human GLP-1 receptor and different enough that human enzymes do not recognize it as a familiar substrate.

That second half is the point of the whole story.

  THE FIRST TEN RESIDUES, ALIGNED

  human GLP-1     H   A   E   G   T   F   T   S   D   V
  position         7   8   9  10  11  12  13  14  15  16    (GLP-1 numbering)
                   |   ✗   |   |   |   |   |   |   |   ✗
  exendin-4       H   G   E   G   T   F   T   S   D   L
  position         1   2   3   4   5   6   7   8   9  10    (exendin-4 numbering)

  Eight of the first ten are identical. The mismatch that matters is the second one.

  DPP-4 — the enzyme of Chapter 33 — cleaves human GLP-1 immediately after the
  alanine at position 8, removing two residues and destroying the activity.
  Exendin-4 carries a different residue at the equivalent position. The enzyme
  arrives, finds the wrong side chain, and does not cut.

Read that diagram and then read Chapter 33 again in your head.

Chapter 33 describes an engineering program of real sophistication: substitute the residue at the cleavage site so the protease cannot act, attach a fatty acid to a lysine so the molecule binds serum albumin and hides from renal filtration, tune the linker, tune the ratio, iterate. Years of medicinal chemistry, across multiple companies, aimed squarely at one problem — native GLP-1 is destroyed within about two minutes by DPP-4 and cleared almost immediately, and a two-minute drug is not a drug.

A lizard had already solved it.

Not approximately. Not partially. The exact modification Chapter 33's chemists arrived at by design — change the residue the enzyme recognizes — was already sitting in Heloderma venom, and had been for however long that lineage has been making the peptide. Nobody designed exendin-4. Nobody derived it from a model of the enzyme's active site. It was found.

Synthetic exendin-4 was developed as a drug and became exenatide, approved in 2005 as the first GLP-1 receptor agonist to reach the market. The most consequential peptide drug class of the twenty-first century — the class that changed how obesity and type 2 diabetes are treated, the class Chapters 7 and 8 are about — opened with a molecule that was discovered rather than designed.

That is the irony, and it is the best teaching moment in Part VI, so sit with it before we spoil it. An entire discipline of rational peptide engineering exists to produce, at great expense, a property that evolution handed over for free to anyone willing to look in an unusual place.

Now the counterweight, because otherwise this becomes mysticism

Here is where most retellings stop, and stopping there produces a conclusion that is comfortable, widely repeated, and wrong: nature already has the answers; we should be searching, not designing.

Watch what happened next.

Exenatide worked. It lowered blood glucose, produced modest weight loss, and validated the target — before exenatide, GLP-1 receptor agonism was a hypothesis; after it, a mechanism with a marketed drug behind it. That is not a small contribution.

But its duration of action was still inadequate for many patients. Resisting DPP-4 is one of two problems; the other is renal clearance, and exendin-4 does nothing about that. The pharmacokinetics did not match the way people live.

The drugs that displaced it — liraglutide, then semaglutide — did not come from another animal. They came from the laboratory. They took the human sequence, substituted the cleavage-site residue (the lizard's trick, now applied deliberately), and then added what the lizard had not supplied: a fatty acid chain that binds albumin, extending the half-life from hours to days. That second step is pure engineering. There is no organism it was copied from.

Nature supplied the lead. Chemistry supplied the drug.

Both clauses are load-bearing. A purely naturalistic reading says look harder in the desert; a purely rationalist reading says the lizard was a lucky shortcut we no longer need. Both are wrong. Discovery and optimization are different activities with different tools, and a field good at only one of them keeps producing molecules that are interesting and not quite medicines.

📊 Evidence Rating

Claim: Venom-derived peptides are a productive source of drug leads. Rating: ✅ Strong evidence Reason: Exenatide, ziconotide, and captopril's documented lineage are approved medicines traceable to venom peptides, across three unrelated venomous lineages and three unrelated therapeutic areas. What would change it: Nothing plausible; this is a historical claim about approvals that have already happened. Note carefully what it does not claim — it says venom is a good place to look, not that any particular venom peptide will become a drug, and not that a venom origin is evidence for anything about a specific compound. (Rated as of 2026.)

💊 In the Clinic — what the exenatide story means for a patient asking about it

Patients who learn the lizard fact usually ask one of two questions, and both have short answers.

"So is my drug made from venom?" No. Exenatide is chemically synthesized; nothing is extracted from an animal. And if you are on semaglutide, tirzepatide, or another current agent, your drug is not exendin-4 at all — it is a modified version of your own hormone. The lizard is a story about the class, not about your prescription.

"If it came from nature, is it safer?" No, and Chapter 1 already dismantled this. What makes the GLP-1 agonists well-characterized is the outcome trials of Chapters 7 and 8 — tens of thousands of participants, hard endpoints, years of follow-up. They would be exactly as well-characterized if designed on a whiteboard, and exactly as poorly characterized without those trials.

The lizard is a wonderful fact and a terrible argument. The coda to this chapter's dossier section exists mostly to make that distinction impossible to forget.


35.4 Frog skin, cone snails, and the sea

Exendin-4 is the most consequential example of route 1, but it is not the strangest, and it is not even the most commercially important. That title goes to a drug that is not a peptide at all.

Captopril: the peptide that became a pill

In the 1960s, researchers studying the venom of the Brazilian pit viper (Bothrops jararaca) noticed that it contained factors which potentiated bradykinin, a peptide that dilates blood vessels — part of how the venom drops a prey animal's blood pressure catastrophically. These bradykinin-potentiating peptides worked by inhibiting an enzyme, and the enzyme turned out to be angiotensin-converting enzyme (ACE), which sits at the center of blood pressure regulation.

One of these peptides was developed as a drug called teprotide. It worked — it lowered blood pressure in humans and demonstrated that ACE inhibition was a real therapeutic strategy rather than a physiological curiosity.

And it had to be injected.

You know why by now. Teprotide is a peptide; peptides are food (§1.3); an injected antihypertensive is not a treatment for a chronic, asymptomatic, lifelong condition managed by millions of people. The pharmacology was validated and the molecule was unusable.

So Ondetti, Cushman, and colleagues did something novel for the time: they used what was known about the enzyme's active site and the peptide's key contacts to design a small molecule reproducing the essential interactions without being a peptide. The result was captopril, approved in the early 1980s, one of the most commercially successful cardiovascular drugs ever launched, and founder of a class still in first-line use for hypertension and heart failure.

Captopril is not a peptide. That is the whole point of including it. It is the earliest and most commercially important peptidomimetic in medicine: a small molecule mimicking the critical features of a peptide's interaction with a target while shedding the properties that made the peptide undeliverable.

And you have read this arc before, one chapter ago. Chapter 33 described orforglipron and the push toward orally available small molecules that activate the GLP-1 receptor — the same receptor the injected peptides hit. The shape is identical:

  THE ARC, TWICE

  1960s–80s   venom peptide  →  injectable peptide drug  →  small-molecule drug
              BPPs from          teprotide                  captopril
              B. jararaca        (worked, impractical)      (oral, class-founding)

  1990s–2020s venom peptide  →  injectable peptide drugs →  small-molecule drugs
              exendin-4          exenatide, then the         oral GLP-1 receptor
              from H. suspectum  engineered analogs          agonists (Chapter 33)

  Different decades, different targets, different chemistry. Same three moves:
  nature validates the target, peptide chemistry makes a usable drug, small-molecule
  chemistry makes a convenient one.

That recurrence is not coincidence. It reflects a permanent asymmetry: peptides are excellent at proving a target is druggable, because their specificity and potency let you interrogate one receptor cleanly. Small molecules are excellent at being taken every morning with coffee. The field keeps making the same trip because both facts keep being true.

Ziconotide: the molecule that made the route surgical

Cone snails are marine predators that hunt using a harpoon-like modified tooth and a venom of extraordinary complexity. A single Conus species may produce hundreds of distinct peptide toxins — conotoxins — most short, disulfide-stapled, and aimed at specific ion channels and receptors in the nervous system.

From Conus magus, the magician's cone, comes ω-conotoxin MVIIA. It blocks N-type voltage-gated calcium channels, which sit on presynaptic terminals and control neurotransmitter release — including the transmitters carrying pain signals in the spinal cord. Block those channels in the right place and you block pain transmission at its entry point.

The synthetic version, ziconotide, is an approved drug for severe chronic pain, genuinely potent and not an opioid — which matters enormously for patients who have exhausted opioid options or cannot tolerate them.

And it can only be given by intrathecal administration: delivered directly into the cerebrospinal fluid, through a catheter, usually from an implanted pump.

Sit with that. Chapter 4 taught you that peptides do not cross the blood-brain barrier. Ziconotide does not cross it either, and its target is in the spinal cord. Injecting it into a vein would not put useful amounts where they are needed, and the amounts required to try would produce systemic calcium-channel effects nobody would accept.

So the route changed. If the molecule cannot reach the compartment, put the molecule in the compartment. The delivery problem was not solved chemically; it was solved surgically.

Keep this as a reference point. When someone tells you a peptide "crosses the blood-brain barrier," remember ziconotide: a peptide whose central effects are real, valuable, and clinically established, and the medical system's answer to getting it into the central nervous system was an implanted pump.

🩺 Safety and Risk — potency is not the same as safety, and ziconotide is the proof

Ziconotide is instructive precisely because it is both a real drug and a demanding one.

The therapeutic window is narrow, the neuropsychiatric adverse effects are well described in its labeling, and the delivery system carries the risks any implanted intrathecal device carries — infection, catheter problems, pump malfunction. Managing a patient on it is specialist work.

None of that makes it a bad drug. It makes it a drug for a specific population — people with severe, refractory pain for whom the alternatives have failed — where the balance tips. That is Chapter 6's therapeutic window doing its job.

The transferable lesson: a molecule's potency tells you nothing about its safety, and its natural origin tells you less than nothing. ω-conotoxin MVIIA is exquisitely selective, extremely potent, entirely natural, and requires an implanted pump and a pain specialist. All four at once.

As always, decisions about pain management belong with the clinicians who know the case, not with a book.

Magainins: the frog that opened a field

In the late 1980s, Michael Zasloff — then working on African clawed frogs (Xenopus laevis) for unrelated reasons — noticed something about the animals' recovery from surgery. Frogs returned to tanks of unsterile water healed without infection. Something in the skin was doing antimicrobial work.

What he isolated were the magainins: short, positively charged, helix-forming peptides that disrupt bacterial membranes. They were not the first antimicrobial peptides described, but their characterization is generally credited with opening the modern antimicrobial peptide (AMP) field, now understood to be an ancient and near-universal component of innate immunity in animals and plants alike.

Chapter 25 tells you what happened next, and it is a more complicated story than the discovery deserves: enormous promise, a plausible mechanism that bacteria should struggle to evade, decades of work, and a clinical record far harder to establish than anyone expected. The discovery was real. The translation has been brutal. That gap is, again, §35.10.


35.5 Rational design from what the body already makes

Route 2 is the least romantic and by far the most productive.

The idea is simple: start with a molecule the body already produces, and change it. You are not searching for a new signal. You are taking a known signal and adjusting its properties — usually its lifetime, sometimes its potency, occasionally its receptor selectivity.

This is the origin of most of the peptide drugs in this book. Look at the list:

Parent hormone Engineered descendants What was changed
Insulin insulin lispro, aspart, glargine, detemir, degludec absorption rate and duration (Chapter 11)
GLP-1 liraglutide, semaglutide, dulaglutide protease resistance, albumin binding, half-life (Chapters 8, 33)
Somatostatin octreotide, lanreotide, pasireotide stability, receptor subtype selectivity (Chapter 27)
GnRH leuprorelin, goserelin (agonists); cetrorelix, degarelix (antagonists) duration, and in the antagonists, a flipped mechanism
Vasopressin desmopressin, terlipressin selectivity between vasopressin receptor subtypes
Parathyroid hormone teriparatide, abaloparatide fragment selection and stability

Six families, dozens of approved medicines, an enormous fraction of peptide therapeutics by prescription volume. Route 2 is where the field earns its living.

The advantage is that you begin knowing almost everything you need. The target and receptor are known. The physiology has been studied for decades — you know what happens when the signal goes up because people with tumors that secrete it exist, and what happens when it goes down because people with deficiency states exist. That is a mountain of free information, and it explains why route 2 has a better hit rate than anything else in this chapter.

The disadvantage is real and under-discussed: you inherit the endogenous ligand's selectivity profile, including its off-target activity.

The body's signaling molecules were not designed to be drugs. They evolved to do several things at once, in several tissues, under tight temporal control (Chapter 3). Take such a molecule, make it last a hundred times longer, and administer it at a flat sustained level, and you get all of its actions — not just the one you wanted.

This is why so many peptide drugs share their parent hormone's side effects, and why those side effects are so often described as "on-target." Consider:

  • GLP-1 receptor agonists cause nausea and delay gastric emptying. That is not a defect in the drug. GLP-1 receptors are present in the gut and in brain regions involved in nausea and satiety, and slowing gastric emptying is part of what native GLP-1 does. The nausea and the weight loss are substantially the same mechanism, viewed from two angles (Chapter 8).
  • Somatostatin analogs affect gallbladder motility and glucose regulation. Native somatostatin inhibits the secretion of many things — growth hormone, insulin, glucagon, gut hormones. A drug built on it inhibits many things too.
  • GnRH agonists produce an initial hormonal flare before suppression. That is the natural receptor's natural response to sustained stimulation, which is exactly what Chapter 3 predicts.

The general statement: route 2 gives you a molecule whose primary activity is guaranteed and whose side effect profile is largely inherited rather than chosen. You can trim it — pasireotide's altered receptor-subtype preference and the vasopressin analogs' selectivity are real successes — but you are trimming, not starting fresh.

One version of route 2 deserves a flag, because Part III is full of it. A great many gray-market compounds are fragments of human proteins: someone took a stretch of sequence from a larger endogenous protein, synthesized it, and asked what it does. That is formally route 2, and it has produced legitimate drugs — teriparatide is a fragment of parathyroid hormone and an approved osteoporosis treatment. It has also produced a long list of compounds whose entire case rests on the parent protein being important. A fragment of an important protein is not thereby important, and the evidentiary bar for the fragment is exactly the same as for anything else. Chapter 18 works through several examples where it was not met.

🔍 Check Your Understanding

  1. Why does route 2 have a better hit rate than route 1, and what does it give up in exchange?
  2. A patient asks why the nausea from a GLP-1 agonist cannot simply be engineered out. Using this section and Chapter 8, give the honest two-sentence answer.
  3. Someone argues that a synthetic fragment of a human protein "must be safe, since your body already contains that exact sequence." Name two distinct errors, drawing on Chapters 1 and 3.

35.6 Display technologies: selection, industrialized

Route 3 changed the scale of the field, and its central idea is elegant enough to state before any of the mechanics.

The problem: you have a target and you want a peptide that binds it, and you have no idea what that peptide's sequence should be. The space of possible sequences is unimaginably large — even at twenty residues drawn from the twenty natural amino acids, you cannot test them one at a time.

The solution: don't. Make billions at once, wash them all over the target simultaneously, discard everything that does not stick, and then — this is the trick — ask the winners what they are.

That last step requires an invention, because a billion peptides in a tube are anonymous. Each molecule needs to carry its own identification.

Phage display

George Smith, in the mid-1980s, demonstrated the solution: a physical link between a displayed peptide and the DNA that encodes it.

Bacteriophage are viruses that infect bacteria, built from coat proteins with their genome inside the coat. Insert a stretch of DNA encoding a peptide into a phage coat-protein gene, and the phage assembles with that peptide displayed on its own surface while the DNA encoding it sits inside the same particle. The molecule and its instructions travel together.

Now the whole scheme falls out:

  THE PANNING CYCLE

  1. BUILD          Construct a library: billions of phage particles,
                    each displaying a different peptide variant on its
                    coat, each carrying the matching DNA inside.
                          │
                          ▼
  2. BIND           Pour the library over the immobilized target.
                    Most phage do nothing. A few stick.
                          │
                          ▼
  3. WASH           Wash away everything that is not bound. This is the
                    selection step, and washing harder selects for
                    tighter binders.
                          │
                          ▼
  4. ELUTE          Recover the phage that stayed.
                          │
                          ▼
  5. AMPLIFY        Infect bacteria with the survivors. The phage
                    replicate. You now have a new library, enriched
                    for binders.
                          │
                          ▼
  6. REPEAT         Three to five rounds. Each round enriches further.
                          │
                          ▼
  7. SEQUENCE       Read the DNA inside the winning phage. It tells you
                    the sequence of the peptide on the outside.
                    ← THIS is the step the physical link makes possible.

Smith's insight, along with Greg Winter's development of the method for antibody engineering, was recognized with a share of the 2018 Nobel Prize in Chemistry — shared with Frances Arnold for directed evolution of enzymes, the same philosophy applied to a different problem.

Ribosome and mRNA display

Phage display has one structural limit: you have to get the library into bacteria, and transformation efficiency caps how many distinct variants you can realize — typically somewhere in the billions, which is a very large number and still a tiny sample of sequence space.

Ribosome display and mRNA display extend the same logic to cell-free systems. In ribosome display, translation is stalled so the finished peptide, the ribosome, and the mRNA encoding it stay associated as one complex — molecule and instructions physically joined without ever entering a cell. In mRNA display, the peptide is covalently linked to its own mRNA through a chemical adaptor, an even more robust connection.

Because nothing has to be transformed into a living organism, library sizes can be orders of magnitude larger. These platforms also permit non-natural amino acids and chemically constrained architectures — including macrocyclic peptides, a serious current focus in the search for orally available peptides (Chapter 4).

The conceptual point, which is the reason this section exists

State plainly what these methods are, because the marketing language around them is confusing.

Display technologies do not design anything.

There is no model of the target. There is no reasoning about which side chain should point where. Nothing in the procedure knows what a hydrogen bond is. The method makes an enormous number of variants, applies a physical filter, keeps what survives, and repeats. It is selection, industrialized — evolution stripped of everything except variation and differential survival, run on a bench, on a timescale of a week.

That is a compliment. Selection is astonishingly effective and does not require you to understand your target, which is exactly the situation you are in when the receptor's structure has never been solved. Display technologies have produced approved medicines and a great many high-quality tool compounds.

But it means the output of a display campaign is a sequence that binds, with no explanation, no guarantee of selectivity against anything you did not counter-screen for, and no information about what happens in an organism. It answers exactly one question, extremely well.

🔬 Read the Study — how to read a display-derived claim

Papers reporting display-derived peptides follow a recognizable structure, and knowing it makes them fast to evaluate.

What the paper will show, and should: library design and size; number of selection rounds; enrichment across rounds; sequences of the top clones; a measured binding affinity for the best ones; and — in a good paper — counter-selection against related targets.

What to check next, in order:

Is there a functional assay, or only binding? Binding a receptor is not the same as activating or blocking it; a peptide can occupy a site and do nothing useful. Affinity and nothing else is chemistry, not pharmacology.

Was specificity tested against anything? Enrichment against one immobilized target selects for binding to the target and to the plate, the blocking agent, the tag, and the linker. These artifacts are common enough to have names in the field, and a paper without counter-selection has not excluded them.

Does anything happen in a cell? In an animal? Each is a large step down in success rate.

Is the peptide stable, and how would it be given? A linear 12-residue peptide with beautiful affinity has all of Chapter 4's problems still ahead of it.

The honest summary of a typical display paper: we found a sequence that sticks to the thing we wanted it to stick to. A real and publishable result — roughly the first mile of a marathon, which the press release will describe as the finish line.


35.7 Structure-based design and the cryo-EM era

Route 4 begins with a simple premise: if you know the three-dimensional structure of a target, you can design a molecule to fit it.

Fit is not a metaphor here. Receptor binding is a matter of complementary shape and complementary chemistry — a hydrophobic pocket that wants a greasy side chain, a charged residue that wants an opposite one, a groove of a particular depth. Given an accurate structure of the binding site, a chemist can reason about which modifications should improve affinity and which should destroy it, instead of making a hundred analogs and finding out.

The bottleneck, for decades, was getting the structure.

Why the receptors this book cares about were so hard

X-ray crystallography, which produced most of the structural biology of the twentieth century, requires a crystal. Crystallizing a soluble globular protein is difficult. Crystallizing a membrane receptor is far worse: you must extract the protein from the lipid bilayer it evolved to live in, keep it folded in detergent, and persuade a flexible signaling machine to sit still in an ordered lattice.

Class B GPCRs — the family that includes the GLP-1, GIP, glucagon, PTH, and calcitonin receptors — were historically very hard to crystallize. That is not incidental to this book. The family is the target of a large share of every peptide drug in Parts II and V. For a long time, the most therapeutically important peptide receptors were also among the least structurally characterized, which meant Chapter 33's engineering was being done substantially blind.

What cryo-EM changed

Cryo-electron microscopy takes a different approach: freeze many copies of the molecule in a thin film of vitreous ice in random orientations, image them individually, and computationally combine hundreds of thousands of noisy two-dimensional images into a three-dimensional reconstruction. No crystal required. Successive improvements in detectors and image processing pushed the achievable resolution from "a blob with suggestive lumps" into the range where individual side chains can be placed — transformative for membrane proteins.

The specific payoff for peptide pharmacology was the ability to solve activated complexes — receptor plus bound peptide plus the intracellular G protein it signals through, captured together in the active conformation. These were essentially unobtainable before. They show not merely where a peptide sits, but which of its residues contact which parts of the receptor, how the receptor's helices move on activation, and how the intracellular face is reorganized to engage the signaling machinery.

This is what made the multi-agonist ratio problem of §33.9 a tractable engineering question rather than a matter of trial and error. When you are designing a molecule to activate the GLP-1 receptor and the GIP receptor at some deliberate ratio — the actual design problem behind tirzepatide and its successors — the difference between having structures of both activated complexes and not having them is the difference between engineering and guessing. You can see which contacts drive activity at each receptor and reason about which substitutions will shift the balance.

That does not make it easy. It makes it a problem with a feedback loop.

The caveat that structural biologists themselves raise

A structure is a snapshot, and receptors are not statues.

A GPCR is a dynamic machine that samples multiple conformations, and its signaling output depends on which conformations a given ligand stabilizes and for how long. Two ligands that appear to occupy the same pocket in two static structures can produce different downstream signaling — biased agonism, which Chapter 2 introduced and which turns out to matter clinically for several peptide classes. A structure tells you a great deal about geometry and comparatively little about kinetics and dynamics.

Keep that in mind for the next section, because it is the same limitation in a more dramatic costume.


35.8 AlphaFold, and what protein structure prediction actually changed

Here is where the hype concentrates, so here is where we will be most careful.

What happened

For roughly fifty years, the protein folding problem — predicting a protein's three-dimensional structure from its amino acid sequence alone — was one of the defining open problems in molecular biology. The information is clearly there in the sequence, since proteins fold reliably on their own, but extracting it computationally defeated an enormous amount of very serious effort.

The field measured its progress through CASP (Critical Assessment of Structure Prediction), a biennial blind assessment: organizers take proteins whose structures have been solved experimentally but not published, teams submit predictions, and predictions are scored against the withheld answer. It is about as clean a test as computational biology has.

At CASP14, in 2020, AlphaFold2, from DeepMind, produced a decisive result — accuracy that was, for a large fraction of targets, competitive with experimental structure determination and far beyond anything previously achieved. The organizers described the folding problem as substantially solved for single domains, an extraordinary thing for the referees of a fifty-year problem to say. The work led to a share of the 2024 Nobel Prize in Chemistry for Demis Hassabis and John Jumper, shared with David Baker for computational protein design (§35.9).

What changed: predicting a protein's three-dimensional structure from its sequence went from a hard, slow, sometimes career-length problem to something that often takes minutes on accessible hardware. And the predictions were not hoarded — hundreds of millions of predicted structures were released publicly, covering essentially the known protein universe.

It is difficult to overstate this as a scientific result. A graduate student in 2015 could spend three years failing to crystallize one protein; a graduate student now downloads a prediction over lunch. Whole categories of question that were not worth asking, because answering them required an unaffordable structural campaign, became askable.

That is real. Now the four things it did not change.

What did not change

(1) Structure is not function.

Knowing a protein's shape does not tell you what it does, when, in which cells, in response to what signal, or what happens when it is disrupted. Structure constrains function and often suggests it — a fold resembling a known enzyme family is a genuine hint — but the inference is loose and frequently wrong. The predicted structure of an uncharacterized protein mostly tells you what an uncharacterized protein looks like.

For drug discovery specifically, this matters because target selection dominates outcomes, and target selection is a question about biology, not geometry. Nothing in a structure tells you that inhibiting this protein will help a patient with this disease. Chapter 22's substance P story is the permanent monument: the target's biology was understood, the pharmacology was excellent, and the hypothesis about the disease was wrong.

(2) Short peptides frequently have no single stable structure to predict.

This one is specific to this book, and it is the limitation most often left out.

Chapter 1 (§1.4) told you that many short peptides are largely random coil in solution and adopt a defined conformation only on binding a receptor or inserting into a membrane. Their shape is not a property of the molecule; it is a property of the molecule plus its partner.

For such a molecule, "predict the structure" is close to a malformed question — there is no single answer to predict. Asked for the conformation of a flexible 15-residue peptide, a prediction tool will return something, because models generally do, and that output may correspond to no state the molecule meaningfully occupies.

This is exactly the class of molecule this book is about. GLP-1, BPC-157, the melanocortin peptides, the cosmetic pentapeptides, most of Part III — short, flexible, folding on binding. The structure-prediction revolution landed most heavily on stably folded globular proteins and most lightly on the molecules in this book's title. That is not the technology's fault; it is a statement about which problem was solved.

(3) Predicting a structure is not predicting a drug.

Suppose you have a perfect structure of your target and of a candidate binder. You still do not know whether the compound is potent enough at achievable concentrations; whether it is selective or also hits four related receptors you did not model; its pharmacokinetics (Chapter 4); its toxicity, frequently mediated by things nobody modeled; whether it can be manufactured at scale, purely and affordably (Chapter 32); or, above everything, whether engaging that target changes a clinical outcome in humans.

Every one of those is downstream of structure, and every one kills programs routinely.

(4) The predictions are predictions.

They come with per-residue confidence scores, those scores vary substantially, and — the part that gets dropped in summaries — confidence is systematically lower precisely for the flexible and disordered regions, because those regions have no well-defined structure to be confident about.

That is not a bug; a model reporting low confidence about a genuinely disordered loop is doing the right thing. But it means the map is least reliable exactly where a great deal of interesting biology lives: disordered regions mediate a large share of regulatory protein-protein interactions, host many post-translational modifications, and are structurally what most short peptides are. A confident core plus a low-confidence floppy tail is an accurate representation of reality. It is not the same thing as knowing the molecule.

⚠️ Hype Check — "AI has revolutionized drug discovery"

The claim, in its usual form:

"AI has solved protein folding. We can now design drugs on a computer. AI will find cures for diseases that have resisted treatment for decades, and it will do it in years instead of decades."

What's true in it, and it is a lot. The structure-prediction advance is genuine, enormous, and was assessed under blind conditions by people whose job is to be skeptical. Adoption has been extraordinarily broad, and the public databases are a permanent addition to the scientific commons. It won a Nobel Prize on a very short timeline for good reason.

Where it fails. The claim quietly swaps two different things. Finding candidate molecules and producing approved medicines are separate activities connected by a decade-long pipeline with a very narrow exit. The advance is overwhelmingly in the first; the attrition is overwhelmingly in the second — and the dominant causes of that attrition, per Chapters 9 and 10, are lack of efficacy in humans and unacceptable toxicity. Neither is a shortage-of-candidates problem.

The second failure sits in "in years instead of decades." The rate-limiting step for a chronic-disease drug is frequently a multi-year outcomes trial with a hard endpoint, because you cannot know whether people live longer without waiting to see whether people live longer. That duration is imposed by biology and by what counts as evidence, not by computation.

How to test it, and this is the useful part. The claim is falsifiable. Computationally derived candidates have been entering trials for several years; within a reasonable window there will be enough to compute a Phase 1-to-approval rate and compare it against the historical base rate. If the cohort clears materially higher, the strong claim is vindicated. If it clears at about the base rate with more candidates entering, the technology increased throughput without changing the odds — genuinely valuable, and a different claim.

Verdict: the technical achievement is ✅. The revolution-in-drug-discovery framing is ⚠️ as usually stated, because it conflates the stage that got faster with the stage that decides outcomes. Ask which stage anyone's number refers to.

📊 Evidence Rating

Claim form — "AI has revolutionized drug discovery." Rating: ⚠️ Promising but preliminary, as the claim is usually stated Reason: The structure-prediction advance is real and enormous, but the claim conflates finding candidate molecules with producing approved drugs, and clinical attrition is governed by efficacy and toxicity in humans rather than by candidate supply. What would change it: A cohort of computationally designed or AI-derived peptides completing Phase 3 with approval rates materially above historical base rates. (Rated as of 2026.)

📊 Evidence Rating

Claim: AlphaFold-class structure prediction is a transformative research tool. Rating: ✅ Strong evidence Reason: A technical claim, demonstrated decisively under blind assessment at CASP and confirmed by extremely broad adoption across structural biology. What would change it: Systematic evidence that the predictions are unreliable for the uses researchers actually make of them, which would show up as a wave of retracted structural conclusions and has not. Note the scope: this rates a research tool, not a therapeutic claim, and says nothing about drug output. (Rated as of 2026.)


35.9 De novo design: molecules that never existed

Route 4 has a second, newer, and more radical form.

Structure-based design, as described in §35.7, starts from something that exists — a natural ligand, a known scaffold, a hit from a screen — and improves it. De novo design starts from nothing. Specify a target surface you want bound. Have a computational method generate a backbone that would complement it. Have a second method choose an amino acid sequence that would fold into that backbone. Synthesize the gene, express the protein, and test whether the molecule you invented does what you specified.

David Baker's group and others have been doing exactly this, with methods including ProteinMPNN (which solves the sequence-design half: given a desired backbone, what sequence will fold into it?) and RFdiffusion (which generates plausible backbones, including backbones shaped to bind a specified target). Baker's share of the 2024 Nobel Prize in Chemistry was for this line of work.

The published results include designed proteins and designed peptide binders with no natural counterpart — sequences that do not appear in any organism and never have. Some of these binders hit their intended targets with high affinity on the first pass, without the rounds of laboratory optimization that every previous route required.

This is genuinely new, and it is worth being precise about why.

Consider what each route does to sequence space. Route 1 inherits a search that nature already ran, over millions of years, against its own criteria — we did not search, we went shopping. Route 2 starts from one known-good point and steps outward a few residues at a time: a local walk. Route 3 samples an enormous random subset and keeps what sticks: a huge but blind sample, filtered by selection. Route 4 specifies the property wanted and constructs a sequence to have it.

For the first time, the space of possible peptides is being searched by design rather than by selection or by luck. That is a change in kind, not just in speed. Routes 1 through 3 all depend on something already existing that can be found or filtered. Route 4 does not.

What is not yet true

Now the discipline.

Almost nothing from this route is in late-stage clinical use. The published successes are overwhelmingly in vitro binding results, some cell-based function, and a growing but still modest set of animal studies. The field is young, the results are real and are being replicated independently, and the trajectory is genuinely steep — and its clinical output does not yet exist in quantity.

Designed molecules also bring problems the older routes do not. Immunogenicity is an open question: a sequence appearing nowhere in nature is, from your immune system's point of view, one it has never been tolerized to, and natural or near-natural peptides benefit from a lifetime of immune education that designed ones do not. Manufacturability is separate: a computationally elegant design may be expensive to synthesize, prone to aggregation, or unstable in formulation (Chapter 32). And every downstream question remains untouched — §35.8's third limitation applies in full.

📊 Evidence Rating

Claim: De novo designed peptide binders are effective therapeutics. Rating: 🔬 Frontier Reason: The design methods are real, published, independently reproduced, and prospectively promising, but almost nothing from this route is in late-stage clinical use as of 2026. What would change it: Randomized human trials of de novo designed peptide binders reporting efficacy on clinical endpoints, plus immunogenicity data across a meaningful number of designed sequences. Note that 🔬 here means too early to rate, not doubtful — this is the rating a serious, well-conducted, early field is supposed to have. (Rated as of 2026.)


35.10 Why faster discovery has not produced more drugs

This section is the chapter's point, and everything before it has been setup.

Look back at what §35.1 through §35.9 described. Venom systematically mined. Endogenous ligands modified with a mature toolkit. Display platforms sampling libraries with more members than the galaxy has stars. Cryo-EM structures unobtainable fifteen years ago. Structure prediction that turned a career-length problem into a lunch break. Design methods producing binders to specification.

Every one of those accelerates the same thing: finding a molecule that binds a chosen target.

But binding was never the bottleneck.

What actually kills drugs

You already have the evidence, distributed across four earlier chapters.

Chapter 9 established Phase 2 optimism: the systematic tendency for early-phase results to look better than what follows — small samples, selected populations, flexible endpoints, and the simple statistics of advancing whichever compounds happened to look best.

Chapter 10 established the base rates: most compounds entering human trials never reach approval, and the failures concentrate in the expensive late phases rather than the cheap early ones.

Chapter 16 established that a surrogate endpoint may not predict an outcome — moving a biomarker in the right direction is not the same as helping a patient, and medicine contains multiple cases where a drug improved the number and worsened the outcome.

Chapter 22 provided the cleanest single demonstration in this book. Substance P antagonists bound their target beautifully. The receptor pharmacology was excellent, the preclinical models were encouraging, and target engagement in humans was verified. They did not treat pain. The molecule did exactly what it was designed to do and the hypothesis about the disease was wrong.

The dominant causes of clinical failure are lack of efficacy in humans and unacceptable toxicity. Neither is fixed by generating candidates faster.

  THE PIPELINE, DRAWN TO SHOW THE CONSTRAINT

  DISCOVERY          PRECLINICAL       PHASE 1     PHASE 2      PHASE 3
  ═══════════════    ═══════════       ═══════     ═══════      ══════════════
  find a molecule    animal work       safety      does it      does it change
  that binds                           in humans   do anything  an OUTCOME?
                                                   at all?
  ┌───────────────┐
  │ ALL of this   │  ┌──────────┐   ┌──────┐   ┌─────────┐   ┌──────────────┐
  │ chapter's     │  │ months–  │   │ ~1 yr│   │ 1–2 yrs │   │  2–5 YEARS   │
  │ technology    │  │ years    │   │      │   │         │   │  huge N      │
  │ acts HERE     │  └──────────┘   └──────┘   └─────────┘   │  hard cost   │
  └───────────────┘                                          └──────────────┘
        ↑                                            ↑              ↑
    got 100× faster                          most failures    THE CONSTRAINT
    and much cheaper                         happen in        (and it is made
                                             these two        of calendar time
                                             stages           and human biology)

  Speeding up the leftmost box does not shorten the rightmost one.

That diagram is the argument. A pipeline whose narrow point is a three-year outcomes trial does not run faster because the first step got quicker. You cannot compress a five-year survival endpoint by improving your molecular docking. The trial takes as long as the disease takes.

And the failures in Phase 2 and Phase 3 are not failures of molecular quality. They are failures of hypothesis — the target was not causal, or was causal only in a subpopulation nobody could identify in advance, or was causal at an earlier stage of disease than the patients enrolled, or was causal and compensatory biology closed the gap. A better binder to a wrong target is a better binder to a wrong target.

Now be fair to the technology

The chapter must not end as a sneer, because the sneer is also wrong. There are two real gains.

First: more shots on goal. If the per-program probability of success is roughly fixed and the cost of a program falls, more programs get run, and more programs at a fixed success rate produce more successes. That is arithmetic and it is a genuine benefit — a throughput argument rather than an odds argument, and not less valuable for that.

Second, and more important: economically marginal targets become viable.

A pharmaceutical program has a cost floor, and if the expected return sits below it the program does not happen, regardless of how much good it would do. That is why rare diseases, neglected tropical infections, and antibacterials (where the commercially rational move is to hold a new antibiotic in reserve, destroying its revenue) have been so chronically underserved. The science was never the obstacle. The arithmetic was.

Lowering the cost of discovery moves that floor. Targets that could not clear the bar begin to clear it. Academic groups and small foundations can run campaigns that previously required a large company's resources, and publicly released structural data puts within reach of anyone with an internet connection something that used to require a synchrotron and a budget. That is a real and significant gain, and it may prove the largest practical consequence of everything in this chapter.

It is simply not the same claim as "AI will cure disease faster."

And the difference between those two claims — this makes more attempts possible versus this makes attempts more likely to succeed — is the discipline this entire book teaches. They sound alike in a headline. They are supported by completely different evidence, they predict different futures, and only one of them has been demonstrated.

The general rule

Extract the principle, because it will outlive every technology named here.

A technology that improves one stage of a pipeline improves the whole pipeline only if that stage was the constraint.

That is a fact about systems, not about biology, and once you have it you will find it everywhere. A faster kitchen does not seat more diners if the constraint is tables. A hundredfold increase in candidate molecules does not produce a hundredfold increase in medicines if the constraint is a three-year trial and a base rate governed by whether the target was the right one.

So when you meet the next announcement — and there will be one, about something not yet invented as this is written — the question is not is this technology real? It usually is. The question is: which stage does it improve, and was that stage the constraint? If the answer is no, the technology may still be excellent and worth funding. It is just not the thing that will change how many people get better.

🔍 Check Your Understanding

  1. State, in one sentence each, what Chapters 9, 10, 16, and 22 each contribute to the argument in this section.
  2. A company announces it has used AI to identify 400 candidate peptides for a disease with no approved treatment. Applying the general rule, what have they demonstrated and what have they not?
  3. Give the strongest honest case for the value of faster discovery, without overstating it.

📋 Your Evidence Dossier

This chapter sharpens Field 12.

Field 12 has been in your dossier since early in the book, and it asks: what would change my mind? Most people fill it in with a sentiment. More research. A big trial. Better data. Those answers feel responsible and commit you to nothing, which is precisely why they are comfortable.

This chapter upgrades Field 12 from a sentiment into a specification.

The exercise: take any entry currently rated ⚠️ — promising but preliminary — the tier where the evidence is real and does not settle the question, and therefore the only tier where "what would resolve this?" has a meaningful answer. Then write down, in enough detail that someone could go and run it, the study that would resolve it.

Why this chapter, of all chapters

Because you have just spent nine sections watching how molecules get found — venom mined, hormones modified, libraries of billions panned in a week, receptor structures that were unobtainable fifteen years ago, structure prediction reduced to a lunch break, binders designed to specification.

Generating a candidate is the easy part. That is §35.10 from the other direction, and once you have absorbed it, the natural next question is not where would I find a molecule? but what would I have to do to find out whether it works? — a question with an answer that has a shape, a cost, a duration, and a sample size.

Here is the sharper version, and the reason this field exists:

A reader who cannot describe the study that would settle their question does not yet know what they are uncertain about.

That sounds harsh. Try it once and you will find it merely true. "I'm not sure whether it works" collapses, under the pressure of specifying a trial, into much more specific uncertainties: works for whom? measured how? compared to what? over what period? by how much? Usually one of those is the question you actually care about, and you did not know it until you wrote the others down.

The specification

Six components. All six, in writing, for each ⚠️ entry.

FIELD 12 — WHAT WOULD CHANGE THIS RATING
  Current rating        ⚠️  and the date you assigned it
  The claim, restated   population + endpoint, in one sentence (rating rule 1)

  THE STUDY THAT WOULD SETTLE IT
   1. POPULATION        who, specifically? "Adults" is not a population
   2. ENDPOINT          what is measured, by whom — and is it an outcome or a
                        surrogate? (Chapter 16: say which, out loud)
   3. COMPARATOR        against what? this is where most weak designs fail
   4. DURATION          how long, and why that long? tie it to the biology
   5. SIZE              enough to detect an effect worth having — 40 people
                        or 4,000? no power calculation needed, just the order
   6. BLINDING/RANDOM   randomized? blindable? if not, what does that cost?

  THE TWO THRESHOLDS
   → UPGRADE to ✅ if:   ___________________________________________
   → DOWNGRADE to ❌ if: ___________________________________________
      (both required. a specification with only one direction is a wish)

  Has anything like this been registered or started?   yes / no / cannot tell

The last two lines are the ones people skip, and they are the ones that make this an evidentiary exercise rather than an exercise in optimism. If you can only describe the result that would vindicate the compound, you have not written a specification. You have written a hope. State the result that would sink it with the same precision, before you know which one arrives.

Worked demonstration — a compound

FIELD 12 — ORAL COLLAGEN PEPTIDES FOR SKIN APPEARANCE   [worked demonstration]
  Current rating   ⚠️ (assigned 2026)
  Claim restated   In adults aged 40–65 without dermatologic disease, oral
                   collagen peptide supplementation improves objectively
                   measured skin elasticity and wrinkle depth.

  THE STUDY THAT WOULD SETTLE IT
   1. POPULATION   Adults 40–65, no dermatologic disease, no concurrent
                   retinoid or procedural treatment, stratified by sun-exposure
                   history — a variable that could swamp the effect
   2. ENDPOINT     Instrumented cutometry and photographic wrinkle grading by
                   blinded assessors. A SURROGATE for what people want, which
                   is to look better to other people — so pair it with a
                   validated participant-reported measure, and say which is
                   primary
   3. COMPARATOR   Placebo matched for protein content, NOT an inert capsule.
                   The crux: it separates "collagen peptides do something
                   specific" from "extra dietary protein does something," and
                   almost no existing trial makes the distinction
   4. DURATION     At least 6 months. Dermal collagen turnover is slow; a
                   12-week study measures hydration, not structural change
   5. SIZE         Several hundred per arm. Small effect, noisy measurement —
                   and small trials of small effects are the machinery behind
                   Chapter 9's Phase 2 optimism
   6. BLINDING     Randomized, double-blind, assessor-blinded — all feasible,
                   so a trial that is not all three has chosen not to be

  → UPGRADE to ✅ if: two independent trials of this design show consistent,
                     perceptible improvement over the protein-matched arm
  → DOWNGRADE to ❌ if: an adequately powered protein-matched trial shows no
                     difference — i.e. the effect, if any, is dietary protein
  Registered?     Look. Then write down what you found, including "nothing
                  matching this design," which is itself a finding.

Notice what fell out of that. Writing the specification located the actual point of uncertainty in a single line — item 3, the comparator — and it was not obvious beforehand. The question was never "does collagen do anything." It was "does collagen do anything that protein does not." A reader who had written more research needed would have read the next positive trial without checking the one thing that determines whether it means anything. The specification usually has one line that carries the weight, and finding which line is the work.

Worked demonstration — a claim that is not about a molecule

Field 12 works on any ⚠️, including the one this chapter issued. Take "AI has revolutionized drug discovery," rated ⚠️ in §35.8, and run the same six components in prose.

The claim restated: among candidate drugs whose discovery used computational design or AI-derived structure prediction, progression from first-in-human to approval materially exceeds the historical base rate. Population: every compound entering Phase 1 in a defined window, with a pre-specified, auditable definition of "AI-derived" — the hardest component, because the term is currently applied to everything from de novo design to routine database lookup. Endpoint: approval, an outcome rather than a surrogate. Comparator: the base rate for matched therapeutic areas, because a cohort weighted toward oncology looks worse and one weighted toward metabolic disease better regardless of method. Duration: ten to fifteen years from cohort entry — §35.10's constraint applied to §35.10's own claim. Size: dozens of compounds at minimum. Design: prospective, with the cohort definition registered before outcomes are known, or "AI-derived" will quietly drift toward the successes.

Upgrade to ✅ if the cohort's approval rate materially exceeds the matched base rate. Downgrade to ❌ if it matches or falls below it, meaning throughput rose without the odds changing. Note that the downgrade condition is not a disaster: it describes a world in which the technology is real, useful, and was sold with the wrong headline, which is the most common way a ⚠️ resolves.

One more thing you already have

Look back at every 📊 Evidence Rating callout in this book. Each has four lines, and the fourth is what would change it — a compressed Field 12 specification, written for you. Field 12 is where you start writing them yourself, which is roughly the difference between recognizing a chord and playing it.

Coda — and while you are in there, note where it came from

A short second pass, because this chapter is the natural place for it. For each dossier compound, add a line recording its origin: found in an organism, derived from a human hormone, selected from a library, or designed. Exenatide — venom, Heloderma suspectum. Semaglutide — engineered from human GLP-1. Ziconotide — a cone snail. A sentence each, and genuinely interesting.

Then write the important part: origin carries no evidentiary weight whatsoever.

Exenatide came from a lizard and is ✅ for glycemic control. Several ❌-rated compounds in this book are exact fragments of human proteins, sequences your own body contains right now. Captopril came from a pit viper and is not a peptide; insulin came from a pancreas and is the best-evidenced peptide medicine in existence; magainins came from a frog and the antimicrobial peptide field has spent decades failing to convert an equally charming origin into approvals. Sorted by origin, the compounds in this book scatter across all four tiers in no pattern at all.

And yet a compound's origin story is the most common substitute for evidence in peptide marketingit comes from nature, so it is safe; it is a sequence your body already makes, so it is safe; it was designed by AI, so it is sophisticated. Three origin claims, three flattering implications, zero evidence. Chapter 31 made the identical argument about veterinary use: a fact about a compound's history offered in place of a fact about its effects. It is one of the least informative facts about a molecule and one of the most frequently cited.

Record it, enjoy it, give it no weight. Then go back to Field 12, which is the field that actually decides anything.


Conclusion

Peptide drugs arrive by four routes. You can find them in an organism, modify something the body already makes, select them out of an enormous library, or design them from scratch. The routes are roughly chronological and all four are running right now.

Venom is the richest natural source, and for a reason: evolution has spent millions of years optimizing venom peptides for speed, potency, and specificity against vertebrate nervous, cardiovascular, and muscular targets, which is the design brief of a drug. The specificity is not an accident — an indiscriminate venom would be a metabolically ruinous way to disable an animal.

The Gila monster's exendin-4 is the field's best story and its best lesson. A peptide sharing roughly half its residues with human GLP-1, activating the human receptor, uncleaved by human DPP-4, became exenatide in 2005 and opened the most consequential peptide drug class of the century. The entire engineering program of Chapter 33 exists to solve a problem a lizard had already solved. And then exenatide was displaced by liraglutide and semaglutide, which were designed. Nature supplied the lead; chemistry supplied the drug. Both clauses.

Display technologies industrialized selection. Cryo-EM opened up the receptors that mattered most and could not be crystallized. AlphaFold turned a fifty-year problem into a routine calculation and released the results to everyone. De novo design is, for the first time, searching sequence space by intention rather than by luck.

And the drugs have not arrived proportionally.

That is not a reason for cynicism. It is a reason for accuracy. Every technique here improves the same stage — finding a molecule that binds — and that stage was not the constraint. The constraint is whether engaging a target changes what happens to a person, and the only instrument that measures it is a trial that takes years and usually says no. Faster discovery gives us more attempts and makes previously uneconomical diseases worth attempting, which is real and matters. It does not change the odds of any single attempt, and the two claims are not interchangeable no matter how alike they look in a headline.

A technology that improves one stage of a pipeline improves the whole pipeline only if that stage was the constraint. Carry that sentence out of this chapter. It will still be true about whatever gets announced next.

Chapter 36 turns from how peptides are found to how they are made at scale — which is where a different set of constraints, entirely physical and entirely unglamorous, decides which molecules ever reach anyone.


Key Terms

Venom peptide — a peptide component of an animal venom: typically short, often disulfide- stabilized, and usually selective for one ion channel or receptor.

Exendin-4 — a peptide from Gila monster (Heloderma suspectum) venom that shares roughly half its residues with human GLP-1, activates the human GLP-1 receptor, and is not cleaved by DPP-4.

Exenatide — synthetic exendin-4, approved in 2005 as the first GLP-1 receptor agonist.

Peptidomimetic — a non-peptide molecule designed to reproduce a peptide's essential interactions with its target while shedding the peptide's delivery liabilities. Captopril is the foundational example.

Teprotide — a bradykinin-potentiating peptide from Bothrops jararaca venom, developed as an injectable ACE inhibitor; the direct precursor to captopril.

Captopril — an oral small-molecule ACE inhibitor approved in the early 1980s, designed from the interactions of venom-derived peptides. Not a peptide.

Conotoxin — one of the short, disulfide-rich peptide toxins produced by marine cone snails, usually targeting a specific ion channel or receptor.

Ziconotide — synthetic ω-conotoxin MVIIA from Conus magus; an N-type calcium channel blocker approved for severe chronic pain, given intrathecally because it does not cross the blood-brain barrier.

Magainins — antimicrobial peptides identified in Xenopus laevis skin by Michael Zasloff in the late 1980s; their characterization opened the modern antimicrobial peptide field.

Rational design — designing a molecule from knowledge of the target, physiology, or structure rather than finding one by search or selection.

Endogenous ligand — the molecule the body produces to activate a given receptor; the starting point for route 2.

Phage display — a selection technology in which peptide variants are displayed on bacteriophage carrying the encoding DNA inside, so binders can be selected and then sequenced.

Panning — the selection step in a display experiment: washing a library over an immobilized target and recovering what stays. Iterated cycles of variation and selection of this kind are directed evolution, recognized alongside phage display in the 2018 Nobel Prize in Chemistry.

Ribosome display / mRNA display — cell-free selection technologies that keep a peptide physically joined to its own genetic message, permitting libraries far larger than can be transformed into cells.

Structure-based design — designing or optimizing a ligand using the three-dimensional structure of its target's binding site.

Cryo-electron microscopy (cryo-EM) — a structural method reconstructing three-dimensional structures from many images of individual molecules frozen in vitreous ice. No crystals required.

Class B GPCR — the receptor family including GLP-1, GIP, glucagon, PTH, and calcitonin receptors; historically very difficult to crystallize.

AlphaFold — DeepMind's method for predicting protein structure from sequence; AlphaFold2 was decisive at CASP14 in 2020 and the work led to a share of the 2024 Nobel Prize in Chemistry.

CASP — Critical Assessment of Structure Prediction, the biennial blind assessment scoring predictions against unpublished experimental structures.

De novo design — computational design of proteins or peptides with no natural counterpart, using methods such as ProteinMPNN and RFdiffusion.

Intrinsically disordered — a protein or region with no single stable structure; common in short peptides, and where structure prediction is least confident.

Clinical attrition — the loss of candidate drugs during human trials, dominated by lack of efficacy and unacceptable toxicity rather than shortage of candidates.

Rate-limiting step — the stage that determines a process's overall speed; improving any other stage does not speed up the whole.


Spaced Review

  1. Exendin-4 resists DPP-4 because of a difference at the residue corresponding to GLP-1's position
  2. Semaglutide resists DPP-4 because of a deliberate substitution at that same position (Chapter 33). Explain what is the same about these two facts and what is different — and then say what semaglutide has that exendin-4 does not, and why that turned out to matter commercially.

  3. Chapter 2 established that a peptide's effect depends on which receptor it engages and what that receptor does downstream. Using that, explain why a display campaign that produces a peptide with picomolar affinity for a receptor has not yet told you whether the peptide is an agonist, an antagonist, or pharmacologically inert — and why that distinction can invert a compound's entire clinical prospect.

  4. Chapter 5 gave you the method for evaluating a claim. Apply it to a technology claim rather than a drug claim: "Our AI platform has identified more novel peptide binders in eighteen months than the entire field produced in the previous decade." What is the claim's population, what is its endpoint, is it falsifiable, and what would you need to see before it implied anything about patients?

  5. Three compounds: one isolated from snake venom, one that is a fragment of a human protein, and one designed computationally with no natural counterpart. Rank them by how much their origins tell you about their likely efficacy in humans. Defend your ranking in one sentence.

  6. A colleague argues that because structure prediction has become fast and free, the pharmaceutical industry should be producing far more approved drugs than it did a decade ago, and its failure to do so proves the technology is overhyped. Using §35.10, identify what is right and what is wrong in that argument, and state the general rule that resolves it.