Case Study 32.1 — Recombinant Human Insulin, 1982: When the Hard Part Was Not the Part Everyone Expected

Chapter 32 · How Peptides Are Made


Why this case

Because it is the founding case of an entire industry, and because the popular retelling of it is wrong in an instructive way.

The story as usually told is a molecular biology story: scientists learned to put a human gene into a bacterium, the bacterium made a human protein, and the age of biotechnology began. Every part of that sentence is true. It is also not where the difficulty was, and mislocating the difficulty leads people to a conclusion this book keeps having to correct — that once you can make a molecule, you have solved the problem of having the molecule.

This case study is about the gap between those two things.


The situation before 1982

Insulin had been a therapy since 1922 (Chapter 11). For sixty years, therapeutic insulin came from the pancreases of slaughtered cattle and pigs, extracted and purified at industrial scale.

It worked, and it is worth being explicit that it worked well enough to save an enormous number of lives. But it carried three structural problems.

Sequence differences. Bovine insulin differs from human insulin at three positions; porcine insulin at one. Chapter 1 §1.4 established that a peptide's activity survives some substitutions and not others, and these substitutions preserved activity. What they did not always preserve was immunological invisibility. A meaningful minority of patients developed antibodies to animal insulin, which could alter its time course and, in some patients, cause local or systemic reactions.

Supply coupling. The quantity of insulin available was a function of how many animals were slaughtered for meat. For a medicine whose interruption is measured in days to serious harm, that is an uncomfortable dependency, and projections at the time suggested the coupling would tighten.

Purification burden. Extracting a peptide hormone from a mash of animal pancreas is a purification problem of a very different character from purifying a synthetic product. What you are separating from is not a handful of known process impurities but the entire proteome of another species.

A route to human-sequence insulin, made under controlled conditions, in a quantity decoupled from agriculture, was worth a great deal.


What was actually done

The demonstration was published in 1979 by a group at Genentech working with the City of Hope (Goeddel and colleagues, PNAS), and the resulting product reached approval in 1982 — the first recombinant DNA drug approved anywhere.

The approach had three notable features, and each one is a lesson.

The genes were built, not found. Insulin's chains are short — 21 and 30 residues. Rather than isolate the human gene from tissue, the team chemically synthesized DNA encoding each chain. This was possible precisely because the chains are short, and it sidestepped a great deal of difficulty. It is worth noticing that a chemical capability enabled the biological route.

The chains were expressed as fusions. Neither chain was expressed alone. Each was expressed joined to a bacterial protein, then released. The reason is general and recurs constantly in this field: a small foreign peptide expressed by itself in a bacterium is frequently degraded by the host's own machinery, or simply made in quantities too small to matter. Attaching it to something the cell is already content to make in bulk gets around that. The cost is that the fusion partner then has to be removed, which requires a cleavage step that cuts in exactly one place — a constraint that limits which sequences the strategy suits.

The chains were combined afterwards. Having produced the A chain and the B chain separately, the team combined them and drove the disulfide bonds to form.

That last step is where the case study earns its title.


The actual difficulty: three bonds out of fifteen possible pairings

Insulin is not one chain. It is two, held together by disulfide bonds, and there are three of them: two joining the A chain to the B chain, and one internal to the A chain.

INSULIN'S DISULFIDE ARCHITECTURE

   A chain (21 residues) ─┬──────S──S──────┬─  (intrachain bridge within A)
                          │                │
                          S                S
                          │                │
   B chain (30 residues) ─┴────────────────┴─

   Six cysteines total.
   Six cysteines can pair with one another in FIFTEEN distinct ways.
   Exactly ONE of those fifteen arrangements is insulin.

   The other fourteen are molecules of identical elemental composition
   and identical mass, and they are not insulin, and no amount of
   sequencing will tell them apart.

Chapter 1 §1.2 called the disulfide bond biology's staple, and Chapter 1 §1.4 noted that tertiary structure is what a receptor actually recognizes. Both of those facts converge here. Insulin's activity is not carried by its sequence alone. It is carried by the sequence plus a specific three-dimensional arrangement, and that arrangement is enforced by where the staples go.

So the achievement of expression was, in an important sense, the achievement of making the ingredients. Combining two purified chains in solution and hoping they find each other in the right register — rather than forming the fourteen wrong arrangements, or bonding to another copy of themselves, or aggregating — is a genuinely inefficient process. It worked. It produced a licensed medicine. It was not how you would want to make a drug forever.


Biology's own answer, and how industry adopted it

A human pancreatic beta cell does not make insulin as two chains and hope.

It makes proinsulin: a single continuous chain in which the future B chain and the future A chain are joined by a connecting segment, the C-peptide. Because it is one chain, folding is a conformational problem rather than a search problem. The chain folds back on itself, the correct cysteines are brought into proximity by the geometry of the fold, and the correct pairing becomes overwhelmingly favored. Only after the disulfides are set is the connecting segment excised enzymatically.

The connecting segment's entire function is to make folding easy and then leave. It contributes nothing to insulin's activity. It is a manufacturing aid that biology evolved.

Industrial processes eventually converged on the same trick, by more than one route:

  • Bacterial precursor routes. Express a single-chain proinsulin-like precursor in E. coli. When a bacterium overexpresses a foreign protein it commonly deposits it in inclusion bodies — dense intracellular aggregates of misfolded material. These are actually convenient in one respect (they are easy to isolate and they protect the product from bacterial proteases) and inconvenient in another: the material must be solubilized and refolded before the connecting segment can be cut. Refolding at scale, reproducibly, is a process-engineering discipline in its own right.

  • Yeast secretion routes. Express a precursor in yeast and have the organism secrete it already correctly folded, avoiding the refolding step entirely — at the cost of a different set of complexities in expression level, secretion efficiency, and downstream processing.

Different manufacturers made different choices. What they share is the recognition that folding, not expression, was the problem worth designing the process around.


What this case establishes for the rest of the chapter

Making a molecule and having a usable molecule are different achievements. This is the case study's whole point, and it recurs in every direction. It recurs in §32.3, where a competent synthesis still yields a mixture. It recurs in §32.4, where a purified peptide is still not a characterized one. It recurs in §32.7, where a manufactured drug substance is still not a filled, sterile, deliverable product.

Sequence does not determine everything. Chapter 1 was careful about this and it is worth restating in manufacturing terms: two preparations can have the same sequence and not be the same substance. Fourteen wrong disulfide arrangements share insulin's sequence.

The first demonstration and the industrial process are usually different processes. The 1979 work established feasibility. The routes in use since bear only a family resemblance to it. When you read that something "has been produced recombinantly," you have learned that it is possible, not that it is economical, and the gap between those is often a decade of process work.

Regulatory categories follow manufacturing reality. Insulins are regulated as biologics rather than as small-molecule generics, which is why competitors must demonstrate biosimilarity through their own programs rather than matching a structure on paper. That is a direct consequence of everything above: for a molecule whose identity includes a fold, "same structure" is not a claim a chemical formula can carry. Chapter 12 discusses the pricing consequences, which are real and are not fully explained by this fact.


Discussion Questions

1. The chapter argues that expression was the tractable half and folding the hard half. Suppose you had been advising the 1979 team and had argued the opposite — that expression would be the obstacle. What would you have been reasoning from, and what specific feature of insulin should have changed your mind?

2. Proinsulin's connecting segment contributes nothing to insulin's activity and is removed before the molecule works. Evolution nonetheless conserved it. What does that tell you about the relationship between a molecule's function and its production requirements? Name one other example from this book where a feature exists for production reasons rather than functional ones.

3. Fourteen of the fifteen possible disulfide arrangements are not insulin, and all fifteen share a mass and a sequence. Using Chapter 34's concerns as you understand them so far, what kind of analytical method would be required to distinguish them, and why would sequencing and mass spectrometry both fail?

4. The fusion-protein strategy solved a real problem (degradation of small foreign peptides) and created a new one (the fusion partner has to be removed cleanly). Argue that this is a good trade. Then identify the class of target peptides for which it would be a bad trade, and say why.

5. Two routes to modern insulin are sketched above: a bacterial route requiring refolding from inclusion bodies, and a yeast route that secretes correctly folded precursor. Neither is obviously superior. List the factors a manufacturer would weigh in choosing between them, and say which factor you think dominates and why.

6. This case study concerns a molecule with sixty years of clinical use, a well-characterized structure, and a fully documented manufacturing history. Take a compound from Part III of this book that has none of those things. Write down what the equivalent of this case study would have to contain for that compound — and then note which of those things you can actually find. What does the gap tell you?