Appendix I — Peptide Nomenclature and How to Read a Sequence
Most of the confusion in popular writing about peptides is not chemical. It is notational.
A reader encounters GLP-1(7-36)amide and does not know what the parenthetical means, so it becomes
noise. A reader is told that semaglutide changes "position 8" and counts to the eighth residue of the
sequence in front of them and lands somewhere unrelated. A reader sees TB-500 and tirzepatide
presented in the same sentence and has no way to know that one of those names came from a regulatory
naming committee and the other came from a catalog. A reader sees a compound name ending in -mab
in a list of "peptides" and has no reason to suspect that it is not a peptide at all.
None of that is a failure of intelligence. It is a failure of convention transfer. The conventions are real, mostly consistent, and learnable in an hour — and once learned they do a surprising amount of work. Someone who knows the stem system can read an unfamiliar drug name and know roughly what class it belongs to. Someone who knows how peptide numbering inherits from precursor proteins can read a modification claim without miscounting.
This appendix is that hour: how sequences are written, how they are numbered, how modifications are marked, how compounds get their names, and what those names tell you about where a compound sits in the world. Appendix B holds the amino acid code tables themselves; this appendix assumes them and covers everything around them.
One promise up front, because it is the discipline the whole appendix serves: nothing here tells you whether a molecule works. Notation is a description of structure. Efficacy is an empirical question about people. Knowing the notation makes you a more precise reader of claims, which is worth a great deal — and is not the same as knowing which claims are true.
I.1 Direction, and the basic conventions
A peptide has two ends and they are chemically different. One terminates in a free amino group, the other in a free carboxyl group. These are the N-terminus and the C-terminus, and the difference is not cosmetic — it determines which enzymes attack where, which end can be acetylated, which end can be amidated, and which end a chemist can build from.
A peptide sequence is written N-terminus on the left, C-terminus on the right. Always. There is no dialect in which this is reversed. If you see a sequence and do not know its orientation, you do know its orientation.
This convention is not arbitrary. It matches the direction the ribosome works: translation begins at the N-terminus and adds residues toward the C-terminus, one at a time, so a written sequence is a transcript of the order in which the cell actually built the chain. It also matches the direction in which genes are read, so a nucleic acid sequence and its protein product line up when printed above one another.
Here is the part worth naming out loud, because it catches people who have read Chapter 32 and are trying to reconcile it with everything else:
BIOLOGY (ribosome) N ──────────────────────────────► C
builds this way
WRITING (every textbook) N ──────────────────────────────► C
reads this way
CHEMISTRY (solid phase) N ◄────────────────────────────── C
builds this way — backward
Solid-phase peptide synthesis builds in the opposite direction from biology. The first residue is anchored to the resin by its C-terminus, and each new residue is coupled onto the growing chain's free N-terminus. The chain therefore grows C-to-N, which means a chemist assembling a peptide works through the written sequence from right to left.
This is a genuine source of confusion, and it is worth holding clearly rather than half-remembering. It has practical consequences that Chapter 32 developed: the C-terminal residue is the one attached to the solid support and therefore the one whose chemistry the resin choice is optimized around, and synthesis errors accumulate toward the N-terminal end of the molecule because those couplings happen last, on the longest and most sterically crowded chain. When a synthetic peptide preparation contains deletion sequences, the missing residues are disproportionately near the N-terminus. The notation does not change; the direction of work does.
Three-letter codes and one-letter codes
Two code systems coexist, and the choice between them is about readability rather than meaning.
Three-letter codes separated by hyphens — Gly-His-Lys, Ala-Aib-Glu — are used for short
peptides and anywhere a human being needs to read the residues rather than scan the string. They are
unambiguous, they accommodate non-standard residues without special pleading (Aib needs no
explanation; there is no single letter for it), and they are what a chemistry paper uses when
describing a small molecule in detail.
One-letter strings — GHK, HAEGTFTSDVSSYLEGQAAKEFIAWLVRGRG — are used for anything longer
than a handful of residues and for everything computational. A thirty-residue peptide in hyphenated
three-letter code occupies most of a line and resists visual comparison; as a one-letter string it
can be aligned against another sequence and scanned in a second. Every sequence database, alignment
tool, and structure-prediction input uses one-letter code.
The tables mapping between them are in Appendix B §B.1, along with §B.2's account of why the one-letter assignments look arbitrary and mostly are not. The two systems say the same thing, and a document mixing them is normal rather than sloppy — short fragments in three-letter code inside prose, full sequences in one-letter code in figures and listings.
A small habit that pays off: when you see a short peptide named in three-letter form, check whether
the name is the sequence. Sometimes it is. GHK is glycyl-histidyl-lysine — the name is a spelling
of the molecule, not a code assigned by anyone (§I.4).
I.2 Numbering, and why it is often not 1
Residue numbering runs from the N-terminus, starting at 1. Position 1 is the leftmost residue, position 2 the next, and so on to the C-terminus. That is the default and it holds for most molecules you will meet.
The exception is large, common, and responsible for a specific recurring mistake: peptides derived from larger precursors keep their parent numbering.
Many peptide hormones are not synthesized as themselves. They are cut out of a longer precursor protein by processing enzymes — Chapter 7 traced this for the proglucagon system, where a single gene product is cleaved differently in different tissues to yield different active hormones. When the field named and numbered these fragments, it kept the coordinates of the precursor rather than renumbering each from 1, for a practical reason: a literature that renumbers every fragment loses the ability to say where a fragment came from, or to compare two fragments that overlap.
The consequence is that the active form of GLP-1 is written GLP-1(7-36)amide. Unpacked, that notation says four things:
- GLP-1 — glucagon-like peptide-1, the identity of the parent sequence.
- (7-36) — this molecule consists of residues 7 through 36 of that parent, inclusive. Residues 1 through 6 exist in the precursor and are absent from the active hormone.
- amide — the C-terminus is amidated rather than a free acid (§I.3).
- Implicitly: this is thirty residues long, because 36 − 7 + 1 = 30. The residue count and the highest number do not match, and are not supposed to.
A closely related form, GLP-1(7-37), is thirty-one residues and ends in a glycine. Both are real, both are active, and they are not interchangeable in a sequence listing. Appendix B §B.4 explains why that terminal glycine is there: it is the signal that directs the amidation machinery, and the (7-36)amide form is what remains after that machinery has acted.
The practical consequence: "position 8"
Here is where the notation stops being trivia and starts mattering.
Semaglutide's most-discussed structural feature is the substitution of Aib for the native alanine at position 8. Chapter 33 built an argument around it and Appendix B §B.7 explains the chemistry: alanine at that position is the recognition feature DPP-4 uses, and replacing it with a residue the enzyme cannot process removes the cleavage site without disturbing receptor binding.
Now count. Semaglutide's backbone is GLP-1(7-37) — thirty-one residues. A reader who takes "position 8" at face value and counts to the eighth residue of that thirty-one-residue chain lands on glutamate, which is residue 9 in parent numbering, and which has nothing to do with DPP-4.
The correct reading is:
parent numbering 7 8 9 10 11 12 13 14 ...
the actual molecule 1st 2nd 3rd 4th 5th 6th 7th 8th ...
native GLP-1 His Ala Glu Gly Thr Phe Thr Ser ...
semaglutide His Aib Glu Gly Thr Phe Thr Ser ...
▲
│
"position 8" — the SECOND residue of the molecule
Position 8 is the second residue of the actual molecule. It is called position 8 because the numbering was inherited from the precursor, and the field kept it so that everyone discussing GLP-1 biology, GLP-1 receptor pharmacology, and GLP-1-derived drugs would be using one coordinate system.
The same logic explains why DPP-4 is described as cleaving "after position 8." DPP-4 is a dipeptidyl peptidase — it removes two residues from the N-terminus. Two residues from the N-terminus of GLP-1(7-36)amide is positions 7 and 8, His-Ala. The enzyme's specificity is for alanine or proline in the second position of the substrate, which in parent numbering is position 8. Both descriptions are correct; they use different origins.
None of this is a trap laid for the unwary. It is an inheritance, and it is consistent. The rule to carry: when a peptide's name contains a parenthetical range, the numbering in every claim about that peptide refers to the parent, not to the molecule in your hand. If a position will not land on a sensible residue, check your origin before concluding the source is wrong.
I.3 Modification notation
A one-letter sequence string describes the covalent backbone and the side chains. It describes nothing else. Everything else — and "everything else" frequently includes the features that make the molecule a drug — is carried in notation around the string, and that notation is easy to drop when a sequence is copied from a paper into a forum post into a product listing.
C-terminal amidation: -NH2
-NH2 written at the right-hand end means the C-terminal carboxylate has been converted to an
amide. Instead of –COOH the chain ends in –CONH₂.
This is extremely common in peptide hormones and frequently required rather than optional. Amidation removes a negative charge, changes the local hydrogen-bonding pattern, and in many hormone families is read directly by the receptor. Oxytocin, vasopressin, and GLP-1 in its active form are all amidated. Appendix B §B.4 covers the biology: a C-terminal glycine in the precursor signals the amidating enzyme, which consumes the glycine and leaves the amide behind.
The consequence for a reader is blunt. A compound advertised as a given hormone but ending in a free acid is a different molecule with the same sequence listing, and depending on the receptor family, it may be substantially or entirely less active. The bare string does not tell you which you are looking at.
N-terminal acetylation: Ac-
Ac- written at the left-hand end means the N-terminal amine has been acetylated — capped with
an acetyl group. This removes the free amine's positive charge and, importantly, removes the
recognition feature that aminopeptidases use to begin chewing a peptide from its N-terminal end.
It is a common stability strategy in both natural and designed peptides — and, on the page, a two-character prefix that is trivially lost in transcription.
Disulfide bonds
Disulfide bonds are the one structural feature that a linear sequence string genuinely cannot express, because they connect two residues that may be far apart in the chain. Notation varies:
- A bracket or bar drawn above or beside the sequence connecting the two cysteines.
- A note in text: "disulfide bridge Cys1-Cys6," or "disulfide 1-6."
- In structured formats, a separate connectivity record listing the pairs.
Appendix B §B.4 explains why this matters more than it looks. Oxytocin and vasopressin are nine residues with a single disulfide creating a six-residue ring and a three-residue tail; a reduced, linear oxytocin has the identical sequence and is not oxytocin. Octreotide is cyclized through a disulfide, which is part of why it survives long enough to be a drug (Chapter 27).
And when a peptide contains more than two cysteines, the pattern matters, not just the count. Three disulfides among six cysteines can be formed fifteen different ways, only one of which is correct. A sequence listing that names six cysteines and does not specify the pairing has left the most important structural fact unstated.
D-amino acids
The standard twenty are the L enantiomers. Their mirror images, the D forms, appear routinely in peptide drugs because mammalian proteases largely cannot process them.
Two notations are in use:
D-Phe— an explicit prefix in three-letter code. Unambiguous, and the safer form.- A lowercase letter —
ffor D-phenylalanine,wfor D-tryptophan — in one-letter strings.
The lowercase convention is compact and dangerous, because case survives poorly through copy-paste, case-normalizing software, and retyping from memory. A D-substitution changes no mass at all (Appendix B §B.7), so a sequence that lost its case information is not recoverable by measurement. Octreotide contains D-residues; so does cyclosporine, whose oral bioavailability owes something to them (Chapter 29).
Non-standard residues
Anything the ribosome cannot incorporate has no one-letter code and is written out: Aib for
2-aminoisobutyric acid, Orn for ornithine, Nle for norleucine, and so on. N-methylated
residues are usually written N-Me-Ala or MeAla.
This is why a sequence containing a non-standard residue essentially cannot be written as a pure one-letter string. Semaglutide and tirzepatide both contain Aib; any accurate linear representation of either has to break out of the alphabet somewhere. When you see a "sequence" for one of these molecules presented as an unbroken run of capital letters, it has been flattened, and something was lost in the flattening.
The general point
Every modification above changes the molecule's behavior, sometimes completely. Amidation can be required for binding. Acetylation can multiply half-life. A disulfide can be the difference between a folded hormone and an inert string. A D-residue can be the difference between a drug and a substrate. A single Aib can be the difference between a two-minute half-life and a week-long one.
And not one of them is visible in a bare one-letter string.
This is the reason to distrust a sequence quoted without its modifications. It is not that the quoter is lying. It is that the notation makes it easy to quote the part that fits in a plain-text field and drop the part that does not, and the part that does not fit is disproportionately the part that matters. A sequence without its modifications is an incomplete specification, and an incomplete specification of a molecule is a description of a family of molecules, not of one.
I.4 Fragment names, and what they tell you
Many of the compounds discussed in gray-market contexts are not named as drugs. They are named as fragments, as codes, or as descriptions — and the form of the name is itself information.
TB-500 is commonly described as corresponding to a short fragment of thymosin beta-4, the actin-binding protein Chapter 18 examined. The name is a code, not a nonproprietary name, and the relationship between code and parent protein is a description supplied by sellers and secondary sources rather than a designation assigned by a naming authority. Chapter 18 was careful about this because it governs what you can conclude: research on thymosin beta-4 is not automatically research on a fragment of it, and §18.5 walked through why that inference fails more often than it holds.
BPC-157 is a fifteen-residue sequence — a pentadecapeptide, from the Greek for fifteen — described as derived from a gastric protein (Chapter 17). "Pentadecapeptide" is ordinary chemical vocabulary meaning exactly "a peptide of fifteen residues" and nothing more. The family is worth knowing: di- (2), tri- (3), tetra- (4), penta- (5), deca- (10), pentadeca- (15). These are counts, not endorsements, and a compound described with a precise-sounding Greek numeral prefix has told you its length and nothing else.
GHK-Cu is the cleanest case, because the name is the molecule. GHK is glycyl-histidyl-lysine
— the actual three-residue sequence, spelled out in one-letter code — and -Cu denotes complexation
with copper. Chapter 30 discussed it in the cosmetic context. There is no code and no assigned
generic name because the compound is simply named by its composition.
What a name shape tells you
Here is the reading skill this section exists to build.
A compound whose common name is a research code or a description, rather than an International Nonproprietary Name, is telling you that no regulator or naming authority has ever assigned it one. Nonproprietary names — INNs internationally, USANs in the United States — are assigned through a formal process that compounds enter when a sponsor is developing them toward approval. A molecule that has been discussed publicly for years and still goes by a hyphenated letter-number code has, in almost every case, never been through that process.
Two cautions about how far to push this.
First, it is a fact about regulatory history, not about chemistry. A compound with a code name is not thereby impure, dangerous, or inert. Molecules do not become better by being named, and every approved peptide drug spent years under a research code first.
Second, it is nonetheless worth registering. The naming process is a downstream consequence of a sponsor moving a compound through development. Its absence, sustained over many years of public discussion, is weak evidence that no such program exists — precisely the observation Chapter 38 §38.4 built on in distinguishing approved drugs from compounds circulating under "research use only" labeling. The name is not the argument. It is a prompt to go and check for the argument, and Appendix A is where the evidence ratings live.
I.5 The drug-name stem table
This is the most useful section in the appendix, and it is the one to memorize.
Generic drug names are not chosen freely. They are assigned through a formal process, and they are built around a stem — a fixed syllable, almost always at the end of the name, that encodes the compound's class. The prefix that comes before the stem is chosen to be distinctive and pronounceable and carries no information; the stem carries the meaning.
The result is that a generic drug name is partly self-describing. Once you know the stems, you can read a name you have never encountered and know roughly what kind of molecule it is and roughly what it does. This is not a party trick. It is the fastest available way to sort a list of compounds into categories, and it will repeatedly tell you that something being discussed as a peptide is not one.
| Stem | Class | Examples from this book |
|---|---|---|
| -tide | peptide | semaglutide, tirzepatide, liraglutide, exenatide, octreotide, teriparatide, bremelanotide, ziconotide |
| -glutide | GLP-1 related peptides | semaglutide, liraglutide, dulaglutide |
| -relin | peptides stimulating pituitary hormone release | sermorelin, tesamorelin, ipamorelin |
| -relix | peptides inhibiting pituitary hormone release (GnRH antagonists) | cetrorelix, degarelix, ganirelix |
| -pressin | vasopressin analogs | desmopressin, terlipressin |
| -tocin | oxytocin analogs | carbetocin |
| -mab | monoclonal antibody — not a peptide | (contrast only) |
| -cept | receptor fusion protein — not a peptide | (contrast only) |
| -ase | enzyme | (contrast only) |
Read the first two rows together, because they show how the system nests. -tide is the general
peptide stem, covering an enormous range: a GLP-1 receptor agonist, a somatostatin analog, a
parathyroid hormone fragment, a melanocortin agonist, and a cone snail toxin all end in -tide
because all of them are peptides. -glutide is a sub-stem within it, marking the GLP-1-related
family. A reader who meets an unfamiliar name ending in -glutide knows without looking anything up
that it is a peptide acting on the GLP-1 system.
The sub-stem structure is general. -tide names frequently carry an internal syllable identifying
the family, and the pattern is worth noticing even where you do not know the specific convention: if
three drugs you already know to be related share a chunk of name, that chunk is probably doing work.
Teaching point one: -relin versus -relix
One letter separates agonist from antagonist at the same hormonal axis.
-relin names compounds that stimulate pituitary hormone release. Sermorelin and tesamorelin
act on the growth hormone axis; ipamorelin is discussed in the same context. These are agonists —
they push the axis.
-relix names compounds that inhibit it. Cetrorelix, degarelix, and ganirelix are GnRH
antagonists — they block the receptor.
Chapter 27 used compounds from both groups in oncology, and the distinction is not academic there. It determines whether a patient experiences an initial hormonal flare. A GnRH agonist first stimulates the axis before continuous stimulation downregulates it, producing a transient surge in hormone levels before the intended suppression arrives — a surge that, in hormone-sensitive disease, is a clinical problem requiring management. An antagonist blocks the receptor from the first dose and suppresses without the surge. Same axis, opposite mechanism, opposite initial clinical experience, and the whole distinction rides on an n versus an x.
That is worth pausing on as a general lesson. The stem system compresses real pharmacology into name fragments, and the compression is lossless enough that misreading one letter misreads the mechanism. Read drug names carefully, particularly at the end.
Teaching point two: -mab and -cept are not peptides
This is the more important of the two, because it catches a category error that a great deal of popular coverage makes routinely.
A name ending in -mab is a monoclonal antibody. A name ending in -cept is a receptor fusion
protein — a receptor domain stitched to an antibody fragment. Neither is a peptide, and the
differences are not matters of degree:
- Size. Chapter 38 §38.2 drew the regulatory line at forty amino acids: below it, a molecule is generally handled as a peptide drug; above it, as a biologic. A monoclonal antibody is on the order of a thousand residues plus glycosylation — not near the line, not adjacent to the line, but larger than the line by more than an order of magnitude.
- Manufacture. Peptides in this book are made by chemical synthesis (Chapter 32), residue by residue, on a resin. Antibodies and fusion proteins are made recombinantly, in living cells, and then purified — a completely different industrial process with completely different failure modes, cost structures, and quality-control demands.
- Regulation. They are regulated as biologics, under a different pathway with different requirements and different follow-on competition rules.
- Behavior. They are injected, and their long half-lives derive from antibody recycling biology rather than from anything a peptide chemist would recognize.
Compounds ending in -mab and -cept are routinely discussed alongside peptides in both clinical
and fitness contexts — sometimes because they treat the same conditions, sometimes because a list was
assembled by someone who did not know the difference. A reader who can spot a -mab has caught a
category error before it propagates, and has learned something about how carefully that source was
assembled.
-ase is included for the same reason at lower stakes: it marks an enzyme, which is a protein,
and again not a peptide in the sense this book uses the word.
How much to trust a stem
The stems are strong hints. They are not guarantees, and it is worth being precise about why.
Stems are conventions applied by naming committees, not laws of chemistry. They reflect how a naming body classified a compound at the time it was named. Classifications shift; the names do not.
Older drugs predate the system, or predate the particular stem now used for their class. Insulin
does not end in -tide. Oxytocin does not either — it is the parent of the -tocin stem rather than
a member of it. Absence of a stem is therefore not evidence that a compound is not a peptide.
The system has exceptions and near-collisions. Some names carry a stem-like syllable by accident, and some compounds sit at genuine boundaries between classes. Peptide-mimetic small molecules — captopril is the historic example (Appendix J) — mimic peptide function without being peptides, and are named as small molecules.
The right posture: a stem tells you what to expect and where to check, not what is true. Use it to form a fast hypothesis and to notice when something in a list does not belong. Then confirm.
I.6 Trade name, generic name, research code
A single molecule commonly carries three names, and each one encodes something different about who is speaking and why.
| Name type | Assigned by | What it tells you |
|---|---|---|
| Trade name | The manufacturer | Marketing. Tied to a specific indication, formulation, and often dose. |
| Generic (nonproprietary) name | A naming authority (INN, USAN) | The molecule's identity. Assigned once, used globally, encodes the class through its stem. |
| Research code | The originating lab or company | That no nonproprietary name has been assigned yet. |
Trade names are marketing, and they are indication-specific
This is the source of a specific, extremely common confusion: the same molecule can carry different trade names for different approved uses at different doses.
There is nothing improper about this. Different indications mean different labels, different prescribing information, different dosing, and different trial packages. Marketing them under distinct names follows from their being distinct products in a regulatory sense, even though the active molecule is identical.
The consequence for a reader is that popular coverage regularly discusses two drugs when the molecule is one drug. Comparative claims built on that footing — that one "works better," that one has a side effect the other lacks — are usually comparing doses, indications, or trial populations rather than molecules. When two trade names are being contrasted, the first question is whether the underlying generic name is the same. If it is, you are reading about dose and indication, not about chemistry.
Generic names are the molecule's real identity
The nonproprietary name is the one to anchor on. It is assigned once, it is intended to be used worldwide, it is not owned by anyone, and — through its stem — it carries class information. When you want to search a literature, search the generic name. When you want to compare two compounds, compare generic names. When a source discusses a compound and never gives its generic name, notice.
"Global-ish" is the honest qualifier. There are national variations, and a handful of molecules carry different names in different countries for historical reasons. But the ambition is a single global identifier per molecule, and for the compounds in this book it mostly delivers.
Research codes, and the information in their persistence
A research code is a letter-number designation — a company or laboratory abbreviation followed by a serial number. It means the compound has an identity within an organization and has not been assigned a nonproprietary name.
For a compound in early development, this is unremarkable and expected. Every approved drug spent time under one.
What is informative is persistence. A compound still going by a research code years after it entered public discussion has usually not progressed. Nonproprietary names are requested by sponsors moving compounds toward the clinic; the request is a small administrative step that happens naturally as a program advances. A program that has advanced generally acquires a name. A code that has been in circulation for a long time without one is weak but real evidence that the program behind it stalled, was discontinued, or never existed in the form implied.
This connects directly to Chapter 36's pipeline questions, which ask you to distinguish compounds in active development from compounds perpetually described as promising: is there a registered trial, is there a sponsor, has anything been published, has the compound advanced a phase in the last several years. Name form is a fast zeroth check before any of those — it costs nothing and is available from the name alone.
I.7 Reading an unfamiliar sequence: a short worked procedure
You have a sequence in front of you and no context. Here is a six-step pass that extracts most of what a sequence can honestly tell you.
Step 1 — Identify the length. Count the residues; length is the single most orienting fact. Below about ten, conformation is probably flexible and half-life probably very short unless the molecule is modified or cyclized. Ten to fifty is the classic peptide-drug range, where Chapter 32's chemical synthesis is the natural manufacturing route. Above roughly forty you are approaching the regulatory boundary Chapter 38 §38.2 described, and above a hundred you are almost certainly looking at a protein made recombinantly, whatever it is being called.
Step 2 — Note the termini and any amidation or acetylation. Look at both ends. Is there a
-NH2? An Ac-? If the answer is "neither is marked," ask whether that means the molecule genuinely
has free termini or whether the notation was truncated somewhere upstream. For a hormone whose
natural form is amidated, an unmarked C-terminus is more likely a transcription loss than a design
choice.
Step 3 — Count charged residues for a rough net charge. Lysine and arginine positive, aspartate and glutamate negative, histidine roughly a tenth of a charge at physiological pH, plus one negative for a free C-terminus and one positive for a free N-terminus. Appendix B §B.3 gives the method. The result tells you a surprising amount: a strongly cationic peptide is a candidate for membrane interaction and for the antimicrobial mechanism of Chapter 25; a roughly neutral one is behaving like most hormones. Net charge also predicts solubility problems and adsorption to container surfaces.
Step 4 — Look for cysteines, and ask about disulfides. Count them. Zero cysteines means no disulfide structure and a probably-linear molecule. Two means ask immediately whether they are bonded, because a cyclized nine-residue peptide and a linear one are different molecules. Four or more means the pattern is load-bearing and its absence from the description is a serious omission. Appendix B §B.4 covers what disulfides do and §B.5 covers how they fail.
Step 5 — Look for non-standard residues. Aib, a lowercase letter, D- anything, an
N-methylation, an ornithine. Their presence is proof of chemical synthesis. A ribosome cannot
make them, which means no cell made this molecule, which means it came off a resin (Chapter 32).
That is a genuine deduction from the sequence alone, and it is one of the few available.
Step 6 — Ask whether the modifications are stated at all. This is the step people skip and the one that decides whether the previous five were worth doing. A bare run of capital letters, with no termini marked, no disulfides noted, and no case distinctions preserved, is not a specification of a molecule. It is a specification of a backbone, consistent with many molecules of very different behavior. If the modifications are absent, the honest conclusion is that you do not know what the molecule is — not that it has none.
1. LENGTH ──► what class of thing is this?
2. TERMINI ──► amidated? acetylated? or unstated?
3. NET CHARGE ──► cationic? soluble? membrane-active?
4. CYSTEINES ──► disulfides? how many? pattern given?
5. ODD RESIDUES ──► present = chemically synthesized, necessarily
6. IS IT COMPLETE? ──► if not, stop concluding things
Running that pass converts a string of letters into a short list of real statements: this is a thirty-residue amidated peptide, roughly neutral, with no cysteines, containing one non-standard residue and therefore chemically synthesized, with one lysine available for conjugation. Every one of those is checkable and none of them requires expertise beyond this appendix and Appendix B.
The discipline
And then stop, because the sequence has told you everything it can, and it has not told you the thing you actually want to know.
A sequence tells you what a molecule is. It tells you nothing about whether it works.
Appendix B ended on this, and it bears repeating here because notation is seductive in a specific way: a precisely formatted sequence with correct modification notation looks like evidence. It has the texture of rigor. It is easy to read a well-specified structure and feel that it implies a well-supported compound.
It does not. Exenatide came from the venom of a lizard and is a well-supported drug (Chapter 35). Several compounds this book rates poorly are exact fragments of human proteins, specified precisely, with impeccable notation and no completed human outcome trial. Structural precision and clinical evidence are independent axes.
What this appendix buys you is the ability to read a claim accurately: to see that a modification
sits at the DPP-4 site rather than somewhere irrelevant, that a name ending in -mab does not belong
in a list of peptides, that a compound still carrying a research code after fifteen years of
enthusiasm has probably not moved, that a sequence quoted without its amide is quoted incompletely.
Those are the gains notation can give you.
The other question — whether anyone has tested this in humans, and what happened — is answered somewhere else entirely, and Appendix A is the index to it.
Related: Chapter 1 (bonds and sequence) · Chapter 7 (proglucagon processing) · Chapter 17 (BPC-157) · Chapter 18 (thymosin beta-4) · Chapter 27 (somatostatin analogs, GnRH) · Chapter 30 (GHK-Cu) · Chapter 32 (synthesis) · Chapter 33 (engineering) · Chapter 36 (pipelines) · Chapter 38 (regulation) · Appendix A (evidence ratings) · Appendix B (amino acid reference) · Appendix J (timeline) · Appendix K (glossary)