Case Study 23.2 — Practice Effects and the Uncontrolled Cognitive Trial
Section 23.7 made a strong claim: an uncontrolled before-and-after cognitive study produces approximately no evidence, because the result you expect if the compound does nothing is the same as the result you expect if it works.
That is a claim people nod at and then fail to act on. It is easy to agree in principle that practice effects exist, and still find yourself impressed by a study that reports a 15% improvement over eight weeks. This case study exists to close that gap by working an example all the way through, so the "approximately no evidence" verdict stops being a slogan and becomes something you can see.
Everything numerical below is a constructed teaching example. No real study, compound, or dataset is described. The numbers are chosen to be plausible and to illustrate the arithmetic; they are not data and must not be cited as such.
The study
[constructed teaching example]
A wellness clinic wants to evaluate a cognitive-enhancement product. They design what sounds like a sensible study:
- Participants: 40 adults, aged 30–55, recruited from the clinic's mailing list, who report "brain fog" or difficulty concentrating
- Design: cognitive battery at baseline, then eight weeks on the product, then the same battery again
- Control group: none — "each participant serves as their own control"
- Battery: six tasks producing eleven scores in total
- Also collected: a weekly self-rated focus score, 1–10
- Analysis: paired comparison of each score, baseline versus week 8
Result as reported: mean improvement across the battery of 12%, with statistically significant improvements on four of the eleven scores. Mean self-rated focus rose from 4.6 to 6.8. The clinic writes: "Participants showed significant cognitive improvement after eight weeks."
Read that and notice how reasonable it sounds. Forty people is not trivial. Eight weeks is a real duration. The battery is validated. The statistics are correctly computed. Nothing in the description is dishonest.
And the study cannot distinguish a genuinely effective compound from an inert one. Here is why, term by term.
Taking the result apart
Recall the six-term decomposition from §23.7:
Observed change = practice effect
+ expectancy
+ regression to the mean
+ state differences
+ measurement noise
+ TRUE DRUG EFFECT
Every term except the last is systematically positive in this design. Work through them.
Practice effect
Participants took the identical battery eight weeks earlier. They have seen the tasks, learned the instructions, developed strategies, and lost the first-time anxiety of being tested. Retest gains on cognitive instruments are among the most reliable findings in the measurement literature, they vary by task, and they can persist for months.
This term alone is expected to be positive and non-trivial for essentially every participant. It is not noise that averages out — it points one way for everyone.
Expectancy
Every participant knows they are taking an active product. There is no blind to break because there is no blind. They were recruited on the basis of a complaint they want resolved, and many of them paid for the product or received it from a clinic they trust.
Expectancy acts on this study through two channels: it inflates the self-rated focus score directly, and it increases effort on the objective tasks. Trying harder on a cognitive test improves scores. So "expectancy" is not confined to the subjective measure — it leaks into the objective ones through motivation.
Regression to the mean
Look at the recruitment criterion: adults who report brain fog or difficulty concentrating. They enrolled because they were having a bad stretch. Cognition fluctuates with sleep, stress, illness, workload, mood, and season. Selecting people at a low point in a fluctuating measure guarantees that the average follow-up measurement will be higher, for reasons that have nothing to do with any intervention.
This is the term most often missed, and this particular recruitment strategy maximizes it.
State differences
Baseline testing happened whenever each participant could come in. Follow-up happened whenever they could come back. Time of day, sleep the previous night, caffeine, food, and whether the appointment followed a stressful morning were all uncontrolled and all vary between the two sessions.
Uncontrolled state differences usually add noise rather than a systematic direction — with one exception worth noting. If participants approach the follow-up session more invested in the outcome, they may arrive better rested and more caffeinated. Small effect, and it points up.
Measurement noise
Any single administration of a cognitive test estimates ability with error. With eleven scores and no correction for multiple comparisons, some will move up by chance alone. This is the term that produced "four of eleven significant."
That phrase deserves its own examination.
The multiplicity problem, concretely
[constructed teaching example, continued]
Eleven scores were tested. Four reached significance. Reported that way, four out of eleven sounds like a partial success — the compound "worked on some measures."
Now ask the §5.10 question: what was the primary endpoint?
There wasn't one. Eleven scores were collected, eleven were tested, and the four that moved were written up as the finding. No pre-specification, no correction, no denominator in the abstract.
Under those conditions, "four of eleven moved" is not a result that needs a compound to explain it. Practice effects push all eleven up to varying degrees; noise scatters them; the ones that land past a threshold get reported. The write-up did not lie about anything. It simply never told you that seven other scores were also tested and did not move, and that no one said in advance which score would count.
The self-report
Focus rose from 4.6 to 6.8 — the largest and most impressive-looking change in the whole study, and the least informative. It is an unblinded self-report of a subjective state, collected weekly from people who are actively hoping to feel better, with no comparison condition of any kind. §23.8 covers why this is the worst-case endpoint. Here it is worth noting only that it will move the same amount whether the product is active or inert, and that its size relative to the objective measures should raise a question rather than settle one.
What the study would have looked like if the compound did nothing
Here is the exercise that makes the point land. Suppose the product is completely inert.
[constructed teaching example]
Practice effects: participants improve on the battery, more on some tasks than others. Expectancy: they try harder at follow-up and report feeling sharper. Regression to the mean: enrolled at a low point, they drift back toward their own average. Noise across eleven scores: a few land past the significance threshold. Self-report: rises substantially, because everyone knows what they took and why.
Expected write-up: "Participants showed significant cognitive improvement after eight weeks, with significant gains on several measures and a large improvement in self-rated focus."
That is the same sentence the clinic actually wrote.
This is what "approximately no evidence" means. It is not that the study is weak, or that it needs more participants, or that the effect is uncertain. It is that the observed outcome is what you would expect under both hypotheses, so observing it does not distinguish between them. A larger version of this study — 400 participants instead of 40 — would produce tighter confidence intervals around a quantity that still does not answer the question. The sample size was never the problem.
Repairing it
What changes would help, in order of how much they buy?
A control group is the whole fix. A randomized, blinded control group receiving an inert preparation experiences the same practice effects, the same regression to the mean, the same state variation, and — if blinding holds — much of the same expectancy. Subtracting one group from the other removes five of the six terms simultaneously. Nothing else in this list comes close. Everything below is a refinement of a design that already has one.
One pre-specified primary endpoint. Name it before enrollment, with its instrument. The other ten scores become exploratory and are labeled as such. This restores the denominator.
Practice-effect mitigation. Alternate test forms reduce item-specific learning. A practice session before baseline pushes participants past the steepest part of the curve. Neither eliminates the term — they shrink it, which matters mostly for how large a trial you need.
Standardized testing conditions. Same time of day, recorded sleep and caffeine, consistent task order. This reduces noise rather than removing bias, so it buys you power rather than validity.
A blinding-integrity check. Ask at the end which arm each participant believes they were in, and report it. Costs one question, and tells the reader whether the expectancy term was actually controlled or only nominally controlled.
An active comparator, if the product produces any noticeable sensation. §23.8 explains what this buys and why it is rare.
Notice the shape of that list. The first item is categorical — it changes what the study can establish. The rest are quantitative — they change how efficiently it establishes it. A study without the first item cannot be rescued by any quantity of the others, which is why "we'll just add more participants" is the wrong response to this critique.
The self-experiment version
Everything above applies with more force to an individual testing themselves, and the differences all run the wrong way.
A single person has a sample size of one, no control condition, complete knowledge of what they took, maximal expectancy, no standardization of testing conditions, and — this is the part people underestimate — usually no record of how much their performance varies week to week when nothing at all has changed.
That last omission is the fatal one, and it also suggests the only genuinely useful thing a careful self-experimenter can do: measure your own baseline variability first. Take the test repeatedly, under varied conditions, changing nothing, until you know how much your score moves on its own. Most people who do this are startled. Only after that does a change during an intervention carry any information at all — and even then, it carries information about you, on that instrument, under those conditions, with expectancy fully intact.
That is not nothing. It is also not a finding about a molecule, and the distinction is the whole subject of §23.8.
Discussion Questions
-
The clinic's sentence — "Participants showed significant cognitive improvement after eight weeks" — is defended as literally true. Is it? Identify precisely which word or words carry the misleading implication, and rewrite the sentence so it is both true and non-misleading. How does the rewrite read to a prospective customer?
-
Five of the six terms in the decomposition are systematically positive in this design. Construct a design in which one of them would point downward, and explain what that would do to how the result should be read.
-
The case study argues that increasing the sample size from 40 to 400 would not fix the problem. Explain why in your own words. Then identify a situation in which a larger uncontrolled study is more informative than a smaller one, and say what makes that situation different.
-
"Four of eleven measures improved significantly" and "the pre-specified primary endpoint improved significantly" can describe datasets that are numerically identical. What exactly is different between them, and why does the difference live in the study's history rather than in its data?
-
The repair list separates one categorical fix from several quantitative ones. Suppose a researcher tells you a control group is impossible for practical reasons but they will implement every other item on the list. What can the resulting study establish? Write the strongest honest claim its results could support.
-
The self-experiment section suggests measuring your own baseline variability before testing any intervention. Almost nobody does this. Why not? And is there a version of the suggestion that someone might actually follow — or is the recommendation only useful as a demonstration of how uninformative the usual approach is?