Case Study 1 — Sanger and the Sequencing of Insulin
How the world found out that a protein has an exact sequence
Type: Real, public, historical · Tier 1 facts · Relevance: §1.2, §1.3, §1.4, §1.5
Background: the question nobody could answer
By the late 1940s, insulin had been saving lives for a quarter of a century. It was extracted from animal pancreases, purified, standardized, and injected into hundreds of thousands of people. It worked reliably enough that type 1 diabetes had gone from a death sentence to a manageable condition.
And nobody knew what it was.
Not in the sense that matters here. Chemists knew insulin was made of amino acids. They could hydrolyze a sample and measure how much of each amino acid it contained. What they could not do — what nobody could do for any protein — was say what order those amino acids were in.
This gap is hard to appreciate now, because the answer seems obvious in retrospect. But the prevailing view was genuinely uncertain about whether proteins even had a defined sequence. A serious body of opinion held that proteins were statistical objects: mixtures of similar molecules, with amino acids arranged in patterns that were regular in some average sense rather than exact in every copy. On that view, asking for "the sequence of insulin" would be like asking for the sequence of a snowdrift.
Frederick Sanger, working in Cambridge, set out to settle it.
The operating problem
The obstacle was that there was no method. Sanger had to invent one, and the invention is a case study in how a hard problem becomes tractable when you stop trying to solve it all at once.
His approach, in outline:
Step 1 — Label one end. Sanger found a reagent that would attach specifically to the free amino group at the N-terminus of a chain, and would stay attached through the harsh conditions used to break the chain apart. So after breaking up the molecule, you could identify which amino acid had been at the front, because it was the one wearing the tag.
Step 2 — Break the chain into pieces, not into individual residues. Complete hydrolysis destroys all sequence information — you get a soup of free amino acids and learn only the composition. Sanger used partial hydrolysis, deliberately incomplete, producing a mixture of short overlapping fragments.
Step 3 — Separate and identify the fragments. Using chromatography and electrophoresis, he could separate that mixture and determine the composition and end-residue of each small piece.
Step 4 — Reassemble by overlap. This is the conceptual heart of it. If one fragment reads
Gly-Ile-Val and another reads Ile-Val-Glu, they must overlap, and the original sequence must
contain Gly-Ile-Val-Glu. Enough overlapping fragments, and the whole chain can be reconstructed —
in exactly the way a shredded document can be reassembled if the shreds overlap.
It took roughly a decade.
What he found
Insulin turned out to be two chains, which Sanger designated A and B, held together by disulfide bonds between cysteine residues. The A chain has 21 residues; the B chain has 30. There is also an internal disulfide within the A chain.
And critically: the sequence was exact. Every molecule of insulin from the same species had the same residues in the same order. It was not a statistical object. It was a defined chemical structure, as specific as any small molecule, just much larger.
Sanger published the B chain sequence in the early 1950s and completed the full structure by 1955. He received the Nobel Prize in Chemistry in 1958. (He received a second one in 1980, for methods of sequencing DNA — a distinction only a handful of people hold.)
🔬 Read the Study — the insulin sequencing work
text FIGURE 1.CS1 — "Ten years to read fifty-one letters" [real published work] THE STUDY Sequential chemical determination of the amino acid order in bovine insulin. Frederick Sanger and colleagues, Cambridge, roughly 1945–1955. Methods: N-terminal labeling, partial hydrolysis, chromatographic separation, and reconstruction by fragment overlap. THE QUESTION Do proteins have a defined amino acid sequence, and if so, what is insulin's? WHAT IT SHOWS Insulin has an exact, reproducible sequence: two chains of 21 and 30 residues, joined by disulfide bonds. Proteins are defined chemical structures, not statistical mixtures. The method generalizes: any protein can, in principle, be sequenced. WHAT IT DOESN'T It does not show what insulin's three-dimensional shape is (that came later, from X-ray crystallography), how it binds its receptor, or how the body makes it. Sequence is primary structure only — the information, not the machine. THE VERDICT Foundational. This is where the concept "a peptide has a sequence" became an established fact rather than a hypothesis. THE LESSON A question that looks unanswerable is often a question with no method yet. The bottleneck in science is frequently technique rather than insight — and the person who builds the technique changes what everyone else can ask.
Why this mattered far beyond insulin
It made §1.4 real. Primary structure stopped being a concept and became a measurable property. That is the precondition for everything in this book: you cannot engineer a molecule whose structure you cannot specify.
It made synthesis conceivable. Once you know the exact sequence, building it becomes an engineering problem rather than a mystery. Within two decades, du Vigneaud had synthesized oxytocin, Merrifield had automated the process, and by 1982 recombinant human insulin was on the market. None of that happens without knowing what to build.
It made the genetic code answerable. If proteins have exact sequences, something must specify them. That framing helped drive the work that established the DNA-to-protein relationship in the following decade.
And it made "the same molecule" a checkable claim. This is the part most relevant to a reader of this book. When Chapter 34 asks whether the contents of a vial are what the label says, that question is only meaningful because molecules have exact, determinable structures. Sanger established that they do. Everything about identity verification — mass spectrometry, sequencing, certificates of analysis — rests on it.
Discussion questions
-
Before Sanger's work, a serious body of opinion held that proteins might be statistical mixtures rather than exact structures. What would peptide pharmacology look like if that view had turned out to be correct? Could there be a "peptide drug" at all?
-
Sanger's method depended on partial hydrolysis producing overlapping fragments. Complete hydrolysis would have destroyed the information he needed. Describe another situation, in science or elsewhere, where deliberately doing something incompletely preserves information that doing it thoroughly would destroy.
-
The work took roughly ten years to determine fifty-one residues. Modern methods can sequence a comparable peptide in hours. Does that speed-up change what a sequence means, or only how quickly it can be obtained? Argue a position.
-
Insulin had been used clinically for twenty-five years before anyone knew its structure. What does that say about the relationship between mechanistic understanding and clinical usefulness? Can you think of a modern parallel — a treatment that works while the mechanism is contested?
-
Sanger's contribution was a method, not a discovery about a specific molecule. In the peptide field today, which do you think is more limiting: the absence of methods, or the absence of evidence generated by existing methods? Chapter 35 will return to this.
-
This case study contains no evidence rating. Explain why — and identify what kind of claim would need one.