Case Study 34.2 — Why You Cannot Test Sterility Into a Product

The idea this case study is about

There is one argument in Chapter 34 that does more work than the rest of the chapter combined, and it is not about peptides at all. It is about the relationship between a process and an endpoint test, and once you have it, a large amount of confusion in this field resolves at once.

The argument is easiest to see with sterility, which is why sterility is the vehicle here. But the shape generalizes, and the generalization is the point. Some properties belong to how a thing was made. Those properties can be sampled in the finished article, weakly. They cannot be created by sampling it. No amount of measuring afterward puts them there.

This case study builds the argument slowly, because people accept it in the abstract and then immediately reason as though it were false.


Part 1 — The intuition that has to go

Here is how most people, reasonably, imagine quality control works:

THE INTUITIVE MODEL (which is wrong for sterility)

     make the thing  ──→  test the thing  ──→  if it passes, it's good

     Testing is the gate. Quality is a property you CHECK FOR at the end.

This model is approximately right for a great many things. Is the mass correct? Weigh it. Is the purity acceptable? Run it on a column. Is the identity right? Take a spectrum. In each case the test interrogates the actual article, the answer is reliable, and a passing result is good evidence.

Sterility looks like it should belong to this family. "Is anything alive in here?" sounds exactly like "is the mass right?" — a factual question about a physical object, answerable by an appropriate measurement.

It is not, and the reason is two features of the test that seem like technicalities and are not.


Part 2 — The two features that break the model

Feature one: the test destroys what it tests

To determine whether a sealed unit is sterile, you have to open it and attempt to culture its contents. The unit is consumed. There is no non-destructive sterility test that examines the sealed article and leaves it intact and saleable.

Therefore you can never test the units you release. You test other units, from the same batch, and infer.

This is already a significant departure. A purity result describes material chemically continuous with what ships. A sterility result describes different units entirely, and the inference from tested units to shipped units rests on an assumption — that the batch is homogeneous with respect to contamination — which is precisely the assumption most likely to fail, because contamination events are usually sporadic rather than uniform.

Feature two: contamination is rare, and rare things hide in samples

Now the statistics, and this is where the argument becomes decisive rather than merely uncomfortable.

Imagine a batch in which a small fraction of units are contaminated. Sample some units at random and test them. What is the probability that at least one contaminated unit lands in the sample?

You do not need the formula to see the shape. If contaminated units are a small fraction of the batch, and you test a modest number of units, then most random samples contain none of them. The test comes back clean. It came back clean because you did not happen to select a contaminated unit, which is not the same thing as there being none.

THE SAMPLING PROBLEM, DRAWN

  A batch. Nearly all units fine; a small scattered fraction contaminated (▓).

  ░░░░░░░░░░▓░░░░░░░░░░░░░░░░░░░░░░░░░░░▓░░░░░░░░░░░░░░░░░▓░░░░░░░░░░░░░░░

  Draw a modest random sample for destructive testing:
            ↑         ↑              ↑        ↑         ↑
  Most draws of this size miss the contaminated units entirely.

        TEST RESULT ................. PASS
        THE BATCH IS ................ NOT STERILE
        THE MEASUREMENT WAS ......... CORRECT
        THE INFERENCE WAS ........... WRONG

  The test did not fail. It answered the question it was asked — "were these
  particular units sterile?" — accurately. The error was in the question the
  reader wanted answered: "is the batch sterile?"

Two things follow, and they should be held together.

The test has low power against exactly the contamination levels that matter. A grossly contaminated batch would be caught. A batch with a low, sporadic contamination rate — the realistic failure mode — would very likely pass.

Increasing the sample size does not rescue the approach. Every additional unit tested is a unit destroyed. To achieve high confidence against low-level contamination by sampling alone, you would have to destroy an impractical share of the batch — and even then you would be inferring about the units you did not test. There is a hard ceiling on what sampling can deliver, and it is set by the destructive nature of the test rather than by anyone's diligence.


Part 3 — So how does anyone make a sterile product?

By changing the question. Not is this batch sterile but is this process incapable of producing a non-sterile batch.

Aseptic processing is the answer, and it is a system rather than a step:

  • The solution is sterilized or sterile-filtered.
  • Containers and closures are sterilized, and depyrogenated under conditions well beyond ordinary sterilization (§34.5 — endotoxin does not care about sterility).
  • Filling happens inside a controlled environment with classified air, engineered to make contamination physically improbable rather than merely unlikely.
  • Personnel are qualified, gowned, and trained, because people are the dominant contamination source in an aseptic area.
  • The environment is monitored continuously, so that a drift in conditions is detected as it happens rather than inferred afterward from a failed batch.
  • Equipment is qualified — demonstrated to do what it is supposed to do, reproducibly.
  • The process itself is validated by process simulations, in which the entire filling operation is run using growth medium in place of product and every filled unit is incubated. Because medium supports growth, contaminated units announce themselves, and because every unit is examined, the sampling problem disappears. This is how a process demonstrates that it reliably yields sterile units.
  • Every batch generates a record of what was actually done, retained and auditable.
THE PROCESS MODEL (which is right for sterility)

  design a process that CANNOT readily produce contamination
        │
        ├── validate that claim about the PROCESS (simulations, environmental data,
        │   equipment qualification, personnel qualification)
        │
        ├── run the process under continuous monitoring, recording what happened
        │
        └── test the finished product as a CONFIRMATION that nothing went wrong
                                        ↑
                    end-product testing lives HERE — as a check on a
                    controlled process, not as a substitute for one

  Sterility is not detected at the end. It is BUILT IN, and then confirmed.

Note carefully where end-product sterility testing sits in that diagram. It has not been abandoned. It is performed, and a failure is taken extremely seriously. But its role is to catch a process that went wrong, not to establish that a process was right. It is a smoke detector, not a fire code.


Part 4 — The generalization

Now the reason this case study exists, because sterility is only the clearest instance.

Ask of any property: is this a property of the material, or a property of how the material came to be?

Property Detectable in a sample? Actually a property of
Molecular identity yes the material
Purity, as detected by a given method yes the material
Peptide content yes the material
Water and counterion content yes the material
Sterility weakly the process
Batch-to-batch consistency no the process
Freedom from cross-contamination between products weakly the process
Traceability of this unit to a documented history no the system
Accountability for a defect no the institution
A route by which harm is detected and acted on no the institution

Everything in the top block can be measured. Everything in the bottom block can, at best, be sampled, and most of it cannot be sampled at all — because it is not a fact about molecules.

This table is Chapter 34 in one page. Sections 34.2 through 34.7 addressed the top block. Section 34.9 is about the bottom one. And the reason "get it tested" feels like a solution is that people generalize from the top block, where testing genuinely settles things, to the bottom block, where it structurally cannot.

Which is the sentence to carry away: you cannot test sterility into a product, and you cannot test consistency, traceability, or accountability into one either. These are made, or they are absent. Analysis afterward does not manufacture them, and no document reports them, because they were never properties of the material for a document to report.


Discussion Questions

  1. State in your own words why the destructive nature of the sterility test is not a mere inconvenience but the thing that breaks the intuitive quality-control model. What would change if a non-destructive sterility test existed?

  2. "The measurement was correct and the inference was wrong." Explain how both can be true at once for a passing sterility test on a contaminated batch. Then find one other example from Chapter 34 where a correct measurement supports a wrong inference, and describe what the two cases have in common.

  3. Process simulations fill every unit with growth medium and incubate all of them. Explain precisely which part of the sampling problem this solves, and why the same approach cannot simply be applied to actual product.

  4. End-product sterility testing is described here as "a smoke detector, not a fire code." Defend the analogy, then identify where it breaks down. Is there any sense in which end-product testing does more than confirm?

  5. Work down the table in Part 4 and, for each item in the bottom block, explain in one sentence why it is not a property a measurement could reveal. Which of them do you find hardest to accept as unmeasurable, and why do you think that is?

  6. This case study argues that "get it tested" appeals because people generalize from properties where testing works to properties where it cannot. That is a reasoning error, but it is a very natural one. What makes it natural? And given that it is natural, what — if anything — should someone who understands the argument say to someone who does not, keeping in mind that the honest answer offers no alternative route?