Case Study 2 — The Vehicle Problem
Why this case
The most common design flaw in cosmetic research is also the least dramatic. It is not fraud, it is not a rigged statistical analysis, and it is not a suppressed adverse event. It is a missing arm.
A cosmetic study that compares an active product against no treatment is not a weak study of the active ingredient. It is a strong study of moisturizer, mislabeled. And because the mislabeling is almost never deliberate, it survives peer review, gets cited, and ends up on a box.
This case study works the problem all the way through, because once you can see it you will see it everywhere — in supplement trials, in device trials, in most of the performance-product literature of Part III. The transferable question is not "did they have a placebo?" It is "was the comparison matched on everything except the thing being tested?"
The physical facts underneath the problem
A peptide is never applied alone. It arrives dissolved in a vehicle: water, humectants such as glycerin and hyaluronic acid, emollients, occlusives, preservatives, sometimes silicones or film-formers, usually a fragrance system.
That vehicle is a moisturizer. Not something like one. It is one, and frequently a good one, because the same formulation science serves both purposes.
And moisturizer does the following, reliably, within hours, in essentially everybody:
- Hydrates the stratum corneum, which swells. Swollen corneocytes push the walls of fine lines toward each other, and fine lines become visibly shallower.
- Smooths the surface optically. Hydrated, evenly coated skin scatters light more uniformly, which reads to an observer as radiance or glow.
- Raises corneometer readings, because capacitance of the stratum corneum is exactly what a corneometer measures.
- Improves profilometry, because the measured surface really is smoother.
- Shifts cutometer readings, because hydrated stratum corneum has different mechanical properties under suction than dry stratum corneum.
- Reduces transepidermal water loss if the formulation contains occlusives.
- Feels good, which affects self-report, and affects it in the direction of the hypothesis.
Every one of those effects is real. Not placebo, not imagined — physically real, measurable, and reproducible.
None of them requires an active ingredient of any kind. Glycerin in water produces a substantial share of the list.
So when a study reports that an active serum improved hydration, smoothness, elasticity, and subject satisfaction versus no treatment, it has reported the expected behavior of the base. The active ingredient has not been tested. It has been carried.
A worked illustration
[constructed teaching example] — the following describes no real study, product, or company. The numbers are invented for teaching and should not be cited as findings.
The study as it would be reported
A 12-week study enrolled 40 women aged 40–60 with visible periorbital fine lines. Participants applied a peptide serum to the face twice daily. Assessment at baseline and week 12 used silicone- replica profilometry (mean line depth) and investigator grading on a 5-point wrinkle scale. Mean line depth improved. Investigator grades improved. Most participants reported visible improvement. The authors concluded that the peptide "significantly improves the appearance of periorbital wrinkles."
Nothing in that paragraph is a lie. Every measurement described could have been made competently.
The study as it would have to be run to answer the question
Now imagine the same investigators ran a second arm: the identical formulation with the peptide left out, applied to the other side of each face, with participants and assessors blinded to allocation.
Here is a constructed illustration of what such a study can look like — again, invented numbers, for teaching only:
[constructed teaching example] — invented numbers, illustrative only
MEAN LINE DEPTH, arbitrary instrument units, lower = smoother
─────────────────────────────────────────────────────────────────────
baseline week 12 change from baseline
Peptide side 100 88 −12
Vehicle-only side 100 89 −11
Untreated (no product) 100 99 −1
─────────────────────────────────────────────────────────────────────
READ IT THREE WAYS:
1. Against NO TREATMENT, the peptide side improved by 11 units.
→ the headline in the version without a vehicle arm
→ "significant improvement in wrinkle depth"
2. Against the VEHICLE, the peptide side improved by 1 unit.
→ the finding that actually concerns the peptide
→ almost certainly within measurement noise at n = 40
3. The vehicle alone captured roughly ELEVEN TWELFTHS of the total effect.
→ the honest summary: the moisturizer worked, and the peptide was along
for the ride
Both readings come from the same participants, the same instrument, the same twelve weeks. The difference is entirely which comparison you make — and if the vehicle arm was never run, reading 1 is the only comparison available, so it becomes the result by default rather than by argument.
This is why §30.8 calls the missing vehicle control close to disqualifying. It does not make a study wrong. It makes a study unable to address its own conclusion.
Why the flawed design persists
It is worth understanding why this keeps happening, because the reasons are not all cynical.
It is cheaper and faster. A vehicle arm means manufacturing a matched placebo formulation, randomizing, blinding, and often doubling the analysis. For a product that will be sold as a cosmetic and requires no efficacy proof at all (§30.7), the return on that investment is unclear.
The result is guaranteed to be positive. A no-treatment comparison essentially cannot fail, because moisturizer works. A design that cannot fail is not a test.
The regulatory environment does not require better. Since a cosmetic may not make structure/function claims anyway, and "reduces the appearance of fine lines" is satisfied by hydration, the study is in one sense fit for purpose. It substantiates the claim that will actually be printed. The trouble is that consumers read the claim as being about the peptide.
Negative results are business documents. A vehicle-controlled study that shows no peptide effect does not get published. It gets filed. Nobody registers cosmetic trials in advance, so the studies that vanish leave no trace — which means the published literature is not a sample of the research, it is a selection from it.
And genuinely: some of these studies are run in good faith by people who want an answer. Study design is hard, the vehicle's contribution is easy to underestimate if you have not thought about stratum corneum hydration, and "we compared to untreated skin" sounds rigorous. Not every bad design is a motivated one.
The generalization
Strip the skincare specifics and the principle is this:
A control arm must differ from the treatment arm in exactly one thing — the thing you are testing. Everything else must match: the base, the ritual, the frequency, the sensory experience, the attention paid to the participant, and the expectations set at enrollment.
When the control differs in several ways at once, the study measures the bundle. That is fine if you are selling the bundle and say so. It is not fine if you are attributing the bundle's effect to one component and charging for that component.
Applications outside this chapter:
- A supplement trial where the active arm also received dietary counseling
- A device trial where the active arm attended more clinic visits
- A peptide injection study where the comparator was no injection rather than a saline injection
- Any trial where the placebo is distinguishable by taste, smell, texture, or sensation
The question is always the same, and it is short: matched on everything except the variable?
Discussion questions
1. Using the constructed illustration above, explain to someone with no research training why "improved by 11 units versus no treatment" and "improved by 1 unit versus vehicle" are both accurate descriptions of the same data. Then say which one belongs on the box, and why.
2. A company argues: "Our vehicle contains the peptide. There is no meaningful way to remove one ingredient without changing the formulation, so a matched vehicle control is impossible." Assess this argument. Is it ever true? What would you ask to find out whether it is true in a specific case?
3. The case study lists five reasons the flawed design persists, only some of them cynical. Rank them by how much you think each contributes, and defend your ranking. Would fixing the regulatory reason fix the others?
4. Design a vehicle-controlled study of a signal peptide that you would personally find convincing. Specify: duration, sample size and how you chose it, the primary endpoint, who is blinded to what, how you would standardize photography, and one thing you would pre-register that companies usually do not.
5. Split-face designs are described in §30.8 as a real help but not a fix. Identify two ways a split-face vehicle-controlled study could still mislead, and propose a practical mitigation for each.
6. Suppose a well-run, six-month, vehicle-controlled, blinded-photograph trial of a peptide serum found a small but real effect over vehicle. Write the 📊 Evidence Rating you would issue, in the four-line format. Then write the sentence a marketing department would want to print, and the sentence you would be willing to defend — and explain the distance between them.