Appendix B — The Twenty Amino Acids: A Reference
Chapter 1 made the argument that a peptide is not a substance but a sequence, and that everything a peptide does — what it binds, how long it survives, whether it can be manufactured, whether it dissolves — is a consequence of which side chains sit in which order. This appendix is the lookup table behind that claim.
It is organized for a reader of this book rather than for a biochemistry examination. Every entry answers three questions: what the residue is, what it does chemically, and where it shows up in a molecule this book actually discusses. The third column is the reason to read it. A textbook can tell you that methionine oxidizes; it is more useful to know that this is why some peptide formulations are packaged to exclude air.
Nothing here is dosing information, and nothing here is a synthesis procedure. It is a reference for reading sequences and understanding claims about them.
B.1 The master table
The twenty amino acids specified by the standard genetic code — the proteinogenic twenty, the ones a ribosome can incorporate directly. Molecular weights are for the free amino acid, averaged over natural isotopes, in daltons (Da). Residue mass — the mass a residue contributes to a peptide chain — is the free amino acid mass minus 18.02 Da, because forming each peptide bond releases one molecule of water (Chapter 1, §1.3).
| Amino acid | 3-letter | 1-letter | Side-chain class | MW (Da) | Side-chain pKa |
|---|---|---|---|---|---|
| Glycine | Gly | G | Nonpolar (hydrogen only) | 75.07 | — |
| Alanine | Ala | A | Nonpolar aliphatic | 89.09 | — |
| Serine | Ser | S | Polar uncharged (hydroxyl) | 105.09 | — |
| Proline | Pro | P | Cyclic, imino | 115.13 | — |
| Valine | Val | V | Nonpolar aliphatic, branched | 117.15 | — |
| Threonine | Thr | T | Polar uncharged (hydroxyl) | 119.12 | — |
| Cysteine | Cys | C | Polar (thiol) | 121.16 | ~8.3 |
| Leucine | Leu | L | Nonpolar aliphatic | 131.17 | — |
| Isoleucine | Ile | I | Nonpolar aliphatic, branched | 131.17 | — |
| Asparagine | Asn | N | Polar uncharged (amide) | 132.12 | — |
| Aspartic acid | Asp | D | Negatively charged (acidic) | 133.10 | ~3.9 |
| Glutamine | Gln | Q | Polar uncharged (amide) | 146.15 | — |
| Lysine | Lys | K | Positively charged (basic) | 146.19 | ~10.5 |
| Glutamic acid | Glu | E | Negatively charged (acidic) | 147.13 | ~4.2 |
| Methionine | Met | M | Nonpolar, sulfur-containing | 149.21 | — |
| Histidine | His | H | Positively charged (imidazole) | 155.15 | ~6.0 |
| Phenylalanine | Phe | F | Aromatic, nonpolar | 165.19 | — |
| Arginine | Arg | R | Positively charged (guanidinium) | 174.20 | ~12.5 |
| Tyrosine | Tyr | Y | Aromatic, polar (phenol) | 181.19 | ~10.5 |
| Tryptophan | Trp | W | Aromatic, nonpolar (indole) | 204.23 | — |
Side-chain pKa values are approximate and context-dependent. The figures above are the standard free-amino-acid values; inside a folded structure, a nearby charge or a buried environment can shift a pKa by two units or more. Treat them as orientation, not measurement.
Two entries in that table deserve a flag before you go further. Leucine and isoleucine have identical masses — they are structural isomers, and no mass measurement can distinguish them. Chapter 34 made this point about mass spectrometry's limits; here is where it comes from. Glutamine (146.15) and lysine (146.19) differ by less than a tenth of a dalton, which is resolvable by a high-accuracy instrument and not by a low-resolution one. Analytical claims about identity depend on which instrument was used, and a certificate that does not say is telling you less than it appears to.
B.2 The one-letter code, and why it looks arbitrary
The one-letter code is not arbitrary, but it is the product of a committee working under a constraint, and it shows.
Eleven residues get the obvious letter: Ala, Cys, Gly, His, Ile, Leu, Met, Pro, Ser, Thr, Val. The remaining nine had collisions to resolve:
- F — Phenylalanine. F for the ph sound.
- W — Tryptophan. The indole ring is a double ring; W is a double V.
- Y — Tyrosine. T was taken by threonine.
- R — Arginine. A was taken by alanine.
- K — Lysine. L was taken by leucine; K is the nearest unused letter alphabetically.
- D — Aspartic acid. D for "acid," and the shorter of the two acidic residues.
- N — Asparagine. The amide partner of aspartate, one letter along.
- E — Glutamic acid. E follows D, as glutamate follows aspartate in chain length.
- Q — Glutamine. Q follows... nothing sensible. It was simply available near E and N.
The pairs are worth internalizing, because they recur: D/N are aspartate and its amide, E/Q are glutamate and its amide, and each amide converts to its acid by deamidation (§B.5), which is one of the most common ways a peptide drug degrades on a shelf.
B.3 Grouped by what the side chain actually does
Classifications differ between textbooks, and the boundaries are genuinely fuzzy — glycine is formally nonpolar but behaves like nothing else, and tyrosine is aromatic and polar and ionizable. This grouping is chosen for usefulness.
NONPOLAR / HYDROPHOBIC POLAR, UNCHARGED
Gly Ala Val Leu Ile Ser Thr Asn Gln Cys* Tyr*
Met Pro Phe Trp*
* Cys and Tyr ionize at high pH;
* Trp is aromatic and largely Tyr is also aromatic
buried, but weakly polar
POSITIVELY CHARGED (pH 7.4) NEGATIVELY CHARGED (pH 7.4)
Lys Arg His (partly) Asp Glu
SPECIAL BEHAVIOR
Cys — forms disulfide bonds Gly — maximum backbone flexibility
Pro — breaks helices, adds rigidity His — buffers near physiological pH
The charged residues are the ones to count first when you read a sequence. A peptide's net charge at physiological pH determines whether it is soluble, whether it sticks to a glass vial, whether it crosses a membrane (Chapter 4), and — for the antimicrobial peptides of Chapter 25 — whether it works at all, since their selectivity for bacterial membranes rests on a net positive charge meeting a net negative bacterial surface.
A rough net-charge estimate at pH 7.4: count Lys + Arg, subtract Asp + Glu, and treat His as roughly a tenth of a charge. Add one negative for the C-terminal carboxylate and one positive for the N-terminal amine, unless the peptide is amidated or acetylated. This will not give you a crystallographer's answer, but it will tell you within a unit or so whether a molecule is cationic, and that is usually the question.
B.4 The residues that matter most in this book
Cysteine — the residue that builds structure
Two cysteine thiols can oxidize to form a disulfide bond, a genuine covalent crosslink. This is the only side-chain-to-side-chain covalent bond the standard twenty can form without help, and it does an enormous amount of work.
- Oxytocin and vasopressin (Chapter 21) are both nine residues with a single disulfide creating a six-residue ring and a three-residue tail. The ring is not decoration; a reduced, linear oxytocin is not oxytocin.
- Insulin (Chapter 11) is held together by three disulfides — two between the A and B chains and one within the A chain. This is why insulin is two chains rather than one, and why producing it correctly is a folding problem rather than only a synthesis problem.
- Somatostatin and octreotide (Chapter 27) are cyclized through a disulfide, and Chapter 33 explained why: cyclization removes the free termini that exopeptidases attack.
The practical consequence, and it is a quality-control consequence: a peptide containing two or more cysteines can form the wrong disulfide pattern, or can crosslink to a second molecule and dimerize. Both produce material with the right mass and the wrong structure. This is one of several reasons Chapter 34 insisted that mass alone does not establish identity.
Proline — the residue that breaks things
Proline's side chain loops back and bonds to its own backbone nitrogen, which has two consequences: the backbone at a proline is conformationally restricted, and the nitrogen has no hydrogen to donate to a hydrogen bond. Proline therefore interrupts alpha helices and beta sheets, and it turns up wherever a chain needs to change direction.
It also matters enzymatically. DPP-4 — dipeptidyl peptidase-4, the enzyme at the center of Chapters 7, 8, and 33 — cleaves after a proline or alanine in the second position. Native GLP-1 has an alanine there, which is exactly why it survives one to two minutes. Semaglutide's substitution of Aib for that alanine (§B.7) removes the recognition feature and the enzyme walks past. An entire drug class turns on the identity of one residue at one position.
BPC-157 (Chapter 17) is proline-rich, which is sometimes offered in marketing as evidence of stability. Proline content does raise resistance to some proteases. It is not, on its own, evidence that a compound survives oral administration intact, reaches a target tissue, or does anything there — and Chapter 17's rating rests on the absence of human outcome data, not on the sequence.
Glycine — the residue that is barely there
Glycine's side chain is a single hydrogen atom. It is the smallest residue, the only achiral one, and by a wide margin the most flexible: the backbone can adopt conformations at a glycine that are sterically forbidden everywhere else. Chains need glycine wherever they must turn sharply or pack tightly.
Glycine has a second role that surprises people. A C-terminal glycine is the signal for amidation, the post-translational step that converts the terminal carboxylate to an amide. Many peptide hormones — oxytocin and vasopressin among them — are amidated, and the modification is frequently required for receptor binding. A synthetic peptide made without it is a different molecule with the same sequence listing.
Lysine — the attachment point
Lysine's side chain ends in a primary amine, which is the most convenient handle in a peptide for attaching something else. Both of the major half-life extension strategies in Chapter 33 use it: lipidation attaches a fatty acid to a lysine, and PEGylation attached polymer chains the same way.
Convenience creates a problem. If a peptide contains two lysines, a conjugation reaction produces a mixture — some molecules modified at the first, some at the second, some at both. This is why semaglutide carries a Lys→Arg substitution at position 34. Arginine is also basic and also positively charged, so the substitution is close to invisible pharmacologically; what it does is leave exactly one lysine, so that fatty-acid attachment yields one defined product. Chapter 33 called this the clearest link in the book between molecular design and quality control: a modification with no patient benefit whatsoever, made so that the manufacturer can know what is in the vial.
Histidine — the residue that changes its mind
Histidine's imidazole side chain has a pKa near 6, which is close enough to physiological pH that small pH shifts change its charge state. No other standard residue does this in the biological range. It makes histidine the workhorse of enzyme active sites, an effective buffer, and a common metal-coordinating residue — zinc coordination by histidines is what allows insulin to form the hexamers that give some formulations their duration profile (Chapter 11).
Methionine and tryptophan — the residues that spoil
Methionine's sulfur oxidizes readily; tryptophan's indole ring is susceptible to oxidation and to photodegradation. Both are formulation liabilities rather than design features. A peptide containing them may need protection from oxygen and light, and a degraded preparation can have reduced potency with no visible change and no change large enough to be obvious on a routine purity assay. This is one of the specific things Chapter 34 meant by saying that a certificate describes a sample at a moment, not a vial after shipping.
Tyrosine — the residue you can label
Tyrosine's phenol ring accepts radioactive iodine easily, which made radioiodination at tyrosine the standard method for producing labeled peptides for receptor binding studies for decades. A great deal of what Chapter 2 says about receptor affinity was originally measured this way. It is a small piece of history that explains why so many classic binding papers used tyrosine-containing analogs.
B.5 How residues fail: the four common degradation routes
A peptide drug does not usually fail by falling apart at random. It fails at specific residues by specific chemistry, and knowing which residues are vulnerable tells you what a stability program has to watch.
| Route | Residues at risk | What happens |
|---|---|---|
| Deamidation | Asn, Gln — especially Asn-Gly | The amide converts to an acid: Asn→Asp, Gln→Glu. Adds ~1 Da and one negative charge. |
| Oxidation | Met, Trp, Cys | Adds oxygen. Met→methionine sulfoxide adds 16 Da. |
| Isomerization / aspartimide | Asp, especially Asp-Gly | The backbone rearranges through a cyclic intermediate, producing isomers with identical mass. |
| Disulfide scrambling | Cys | Correct bonds break and re-form incorrectly, or between molecules. Mass may be unchanged. |
Read the right-hand column carefully. Deamidation and oxidation change the mass and are therefore detectable by mass spectrometry. Isomerization and disulfide scrambling may not change the mass at all. A degradation product that is invisible to the primary identity assay is exactly the kind of problem that regulated stability testing exists to catch — and exactly the kind that nothing in an unregulated supply chain is looking for. Chapter 34 §34.9 argued that testing a sample cannot substitute for process control; this table is the chemistry underneath that argument.
B.6 Alanine scanning, and why alanine is the null residue
If you want to know which residues in a peptide matter, the standard experiment is an alanine scan: synthesize a series of analogs, each with one position replaced by alanine, and measure what each substitution costs.
Alanine is chosen because it is the closest thing to a neutral placeholder. It has a side chain — so it is not glycine, which would add flexibility and confound the result — but that side chain is a single methyl group, contributing almost nothing in the way of charge, hydrogen bonding, or bulk. Substituting alanine therefore asks a clean question: what does this position contribute beyond simply being occupied?
Alanine scans are how the field learned which residues of GLP-1 are essential for receptor activation and which are free to be modified — which is the map Chapter 33's engineering program was drawn on. You cannot decide where to attach a fatty acid until you know which positions can tolerate one.
B.7 Beyond the twenty
The genetic code specifies twenty. Chemistry is not limited to twenty, and several molecules in this book depend on that.
Selenocysteine and pyrrolysine are sometimes called the twenty-first and twenty-second amino acids. Both are genuinely ribosomally incorporated, through special recoding of stop codons, in specific proteins and specific organisms. They are curiosities here rather than drug components, but they make a real point: the twenty are a convention of the code, not a law of chemistry.
Aib (2-aminoisobutyric acid) matters far more to this book. Aib is alanine with a second methyl group on the alpha carbon. It cannot be incorporated by a ribosome, which means any peptide containing it must be made chemically (Chapter 32). Its two methyls sterically restrict the backbone, favoring helical conformations, and — the reason it appears in semaglutide and tirzepatide — it abolishes the DPP-4 recognition site at position 8 without disturbing receptor binding. Chapter 33 stressed that this outcome is unusual: the features an enzyme recognizes and the features a receptor recognizes usually overlap, and a substitution that defeats one normally damages the other.
D-amino acids are the mirror images of the standard L-forms. Mammalian proteases evolved to process L-peptides and largely cannot handle D-residues, so D-substitution is a direct route to protease resistance. Octreotide (Chapter 27) contains D-amino acids; so does cyclosporine, which is orally bioavailable in part for this reason (Chapter 29).
A D-substitution changes no mass whatsoever. L- and D-alanine weigh the same. This means that a peptide synthesized with an accidental D-residue — epimerization is a known side reaction of solid-phase synthesis — is chemically wrong, potentially inactive, and completely invisible to mass spectrometry. Detecting it requires chiral analysis, which is not a routine test. It is the sharpest single example in this appendix of why "the mass is right" and "the molecule is right" are different statements.
Other non-proteinogenic residues used in peptide drugs include ornithine, norleucine, homoarginine, and N-methylated versions of standard residues, which remove a backbone hydrogen bond donor and can substantially improve membrane permeability.
B.8 Reading a sequence: a worked example
Native human GLP-1 in its active 7-36 amide form:
H A E G T F T S D V S S Y L E G Q A A K E F I A W L V R G R-NH2
7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36
position 8 = A (alanine) <- DPP-4 cleaves after this residue
position 26 = K (lysine) <- semaglutide's fatty acid attaches here
position 34 = K (lysine) <- semaglutide substitutes R to leave one attachment site
C-terminus = amide <- not a free acid
Four observations, in the order a working reader would make them.
One: count the charges. Three lysines and two arginines against three glutamates and one aspartate — this is a roughly balanced peptide, slightly cationic, with no extreme charge character. It is not an antimicrobial peptide and does not behave like one.
Two: find the vulnerabilities. The alanine at position 8 is the DPP-4 site and the reason the native hormone's half-life is measured in minutes. There is a tryptophan at 31 — an oxidation liability for any formulation. There is an Asp-Val at 15-16 rather than an Asp-Gly, which is the lower-risk arrangement for isomerization.
Three: find the handles. Two lysines, at 26 and 34. Any conjugation chemistry has to deal with both, and §B.4 explained how semaglutide does.
Four: note the terminus. The C-terminal amide is part of the molecule. A peptide advertised as "GLP-1" with a free acid terminus is a different compound, and the sequence listing alone will not tell you which you have.
Everything Chapter 33 describes as engineering is visible in those four observations. Position 8 becomes Aib and the enzyme loses its site. Position 34 becomes arginine and the chemistry gains a single defined product. Position 26 takes a C18 diacid and the molecule gains albumin binding, a week-long half-life, and — in the sense that matters to the person taking it — its entire clinical identity. The receptor-binding surface is left alone, because evolution had already optimized it.
That is the argument of this book compressed into one sequence: the molecule was never the problem. Survival was.
B.9 Quick lookups
Smallest to largest by mass: Gly, Ala, Ser, Pro, Val, Thr, Cys, Leu/Ile, Asn, Asp, Gln, Lys, Glu, Met, His, Phe, Arg, Tyr, Trp.
Identical or near-identical masses: Leu = Ile (isomers, indistinguishable by mass). Gln ≈ Lys (0.04 Da apart). Any L/D pair (identical).
Ionizable side chains, low pKa to high: Asp (~3.9), Glu (~4.2), His (~6.0), Cys (~8.3), Tyr (~10.5), Lys (~10.5), Arg (~12.5).
Aromatic: Phe, Tyr, Trp. These absorb ultraviolet light near 280 nm, which is how peptide concentration is most commonly estimated spectrophotometrically — and why a peptide containing none of them can be nearly invisible to a standard UV detector, one of the HPLC limitations Chapter 34 §34.3 listed.
Sulfur-containing: Cys, Met.
Cannot be made by a ribosome: Aib, D-amino acids, ornithine, norleucine, N-methylated residues, and every other non-proteinogenic residue. Their presence in a compound is therefore proof of chemical synthesis.
Achiral: Glycine alone.
B.10 What this appendix does not tell you
It does not tell you whether a peptide works.
A sequence is a complete description of a molecule's covalent structure and a nearly useless predictor of its clinical value. Chapter 35 made the point in its own terms: exenatide came from a lizard and is well supported; several compounds this book rates ❌ are exact fragments of human proteins. Sequence is chemistry. Efficacy is an empirical question about people, and the only instrument that answers it is a trial.
Use this appendix to read a claim more precisely — to notice that a substitution is at the DPP-4 site, that a compound has two lysines and therefore a conjugation problem, that a "99% pure" figure says nothing about whether a D-residue crept in. Then go and find out whether anyone has actually tested it in humans. That question is what Appendix A is for.
Related: Chapter 1 (bonds and sequence) · Chapter 4 (delivery) · Chapter 32 (synthesis) · Chapter 33 (engineering) · Chapter 34 (analysis) · Appendix I (nomenclature) · Appendix K (glossary)