Practical Evaluation Exam — Open Book
Sixty points. Two hours. Open book.
You may use: the full text, Appendix A (master evidence table), Appendix C (the twelve fields), Appendix D (reading a clinical trial), Appendix F (red flags), Appendix H (worked evaluations), Appendix K (glossary), and your own dossier.
You will not need to recall anything. Every fact required is in front of you. This exam tests whether you can run the method on material you have never seen, which is the only outcome the course actually claims.
⚠️ Read this before you begin
All three sources below are constructed. They are labeled [constructed teaching example]. None of
them is real.
- COMPOUND A, COMPOUND B, and COMPOUND C do not exist.
- CONSTRUCTED BIOSCIENCES is not a company.
- The trials, figures, percentages, and p-values below were written for this exam and describe no actual study.
- Nothing on this page may be quoted, cited, repeated, or shared as though it described a real product, company, study, or result.
The sources are built to be realistic in structure, because that is what you are learning to read. They are entirely fictional in content.
Course policy applies to your answers: no dosing, no protocols, no sourcing, no vendors — including in quotations. If a source contains such material, elide it. Your analysis does not need it.
Source A — a marketing page (18 points)
[constructed teaching example]— invented for this exam. COMPOUND A does not exist.
RENEWAL COMPLEX — with COMPOUND A
Clinically proven to reduce the appearance of fine lines.
In our study, 92% of users reported smoother-looking skin after four weeks.
COMPOUND A is a fragment of a structural protein your body already makes. Applied topically, it signals your skin to produce more of its own collagen — working with your biology, not against it.
Most peptides are too large to get where they need to go. Our proprietary delivery system solves that.
Over 40,000 five-star reviews. This product is not intended to diagnose, treat, cure, or prevent any disease.
A1. (6 points) Produce the four-line rating for the page's central claim.
CLAIM:
EVIDENCE:
RATING: [✅ / ⚠️ / ❌ / 🔬] + [date] + [if ❌: absent / present-and-negative]
WOULD CHANGE IF:
A2. (4 points) The page reports one number. Identify what was measured, who measured it, and what it was compared against. State which of the three the page does not tell you.
A3. (4 points) The page makes a delivery claim. Using Chapter 30 and dossier field 4, explain why route is the central question for this class of product, and state what the page would need to provide for its delivery claim to be checkable.
A4. (2 points) Identify the origin appeal and the mechanism appeal, and name the rule each would violate if you let it move your rating.
A5. (2 points) The final line — "not intended to diagnose, treat, cure, or prevent any disease" — tells you something about the product's regulatory category. State what it does and does not tell you about the evidence.
Source B — a press release (18 points)
[constructed teaching example]— invented for this exam. CONSTRUCTED BIOSCIENCES is not a company and COMPOUND B does not exist. Every figure below was written for this exam.
CONSTRUCTED BIOSCIENCES ANNOUNCES POSITIVE TOPLINE RESULTS FOR COMPOUND B
Constructed Biosciences today announced positive topline results from its open-label Phase 2 study of COMPOUND B in adults with chronic tendon pain.
Forty participants completed twelve weeks of treatment. Participants showed a 38% reduction from baseline in a serum marker of tissue turnover (p < 0.01).
In an exploratory comparison against historical records from a previous observational cohort, treated participants experienced 50% fewer symptom flares.
A pre-specified subgroup of participants under 45 showed the largest response.
"These results validate our mechanism and represent a major step toward bringing COMPOUND B to patients," said the company's chief medical officer.
The company intends to advance COMPOUND B into a registrational program.
B1. (6 points) Produce the four-line rating for the claim "COMPOUND B reduces symptom flares in adults with chronic tendon pain."
B2. (4 points) Name the three design features that most limit what this study can establish. For each, say what it prevents you from concluding.
B3. (3 points) The release reports a 50% reduction in flares. State every piece of information you would need before that figure could be interpreted, and say which of them the release provides.
B4. (3 points) "p < 0.01" appears once. Explain what it does and does not establish here, and why its presence does not repair the study's central limitation.
B5. (2 points) The chief medical officer says the results "validate our mechanism." Assume that is true. State precisely what a validated mechanism licenses under the rules — and what it does not.
Source C — an abstract (18 points)
[constructed teaching example]— invented for this exam. COMPOUND C does not exist and no such trial was conducted. All figures below were written for this exam.
Effect of COMPOUND C on physical function in adults with chronic musculoskeletal pain: a randomized, double-blind, placebo-controlled trial
Background. COMPOUND C is a synthetic peptide under investigation for musculoskeletal indications.
Methods. Adults aged 40–70 with chronic musculoskeletal pain of at least six months' duration were randomized 1:1 to COMPOUND C or matching placebo for 24 weeks. Participants with inflammatory arthritis, prior joint replacement, or diabetes were excluded. The primary endpoint was change from baseline in a validated physical function score at 24 weeks. Secondary endpoints included change in a serum marker of tissue turnover and patient-reported pain.
Results. 280 participants were randomized (140 per arm); 238 completed. The primary endpoint did not differ significantly between groups (mean difference 1.8 points, 95% CI −1.0 to 4.6, p = 0.21). The serum marker declined significantly more in the COMPOUND C arm (p = 0.03). In a post hoc analysis restricted to participants under 55, the physical function score favored COMPOUND C (p = 0.04). Adverse events were similar between arms.
Conclusions. COMPOUND C demonstrated favorable effects on tissue turnover and supports further investigation for physical function in chronic musculoskeletal pain, particularly in younger adults.
C1. (6 points) Produce the four-line rating for the claim "COMPOUND C improves physical function in adults with chronic musculoskeletal pain."
C2. (3 points) The abstract reports three p-values. State which endpoint each belongs to and rank the three findings by how much weight they should carry. Justify the ranking.
C3. (3 points) Quote the phrase in the Conclusions that does the most work, and explain the gap between it and the Results.
C4. (3 points) Describe the population precisely enough that a reader could tell whether they are in it. Name two groups the trial says nothing about, and state what that does to a claim made about "people with chronic pain."
C5. (3 points) "Adverse events were similar between arms." State two things this establishes and two things it does not.
Part D — Across the three (6 points)
D1. (3 points) Rank Sources A, B, and C by how much they establish about their own central claim. Justify the ranking in three sentences.
D2. (3 points) One of these three is by far the best-conducted piece of work and still does not support its own concluding sentence. Identify it, and explain why "better study" and "supported claim" came apart here.
Answer key
**A general instruction to graders.** This is an applied exam and there is no single correct wording. Grade the *moves*: did the student find the missing comparator, did they identify the endpoint's type, did they refuse to let mechanism or origin do evidentiary work, did they date the rating, is the falsification condition a study rather than a wish. A student who reaches a different rating from the key by a defensible route earns full credit — **a well-reasoned ⚠️ that you disagree with outscores a ✅ that matches the key.** --- ### Source A — the marketing page (18 points) **A1 — four-line rating (6 points).** Model answer:CLAIM: COMPOUND A, applied topically, reduces the appearance of fine
lines. Population: unspecified by the source (adult users of
a cosmetic serum). Endpoint: "smoother-looking skin," reported
by the users themselves -- a self-reported appearance measure,
not a clinical or objective one.
EVIDENCE: One in-house study. Design unstated. No comparator, no
blinding described, no sample size, no measurement instrument,
4 weeks. Result is a percentage of users who "reported"
something. Replication: none described. Conspicuously absent:
any controlled comparison, any objective measure, any evidence
that the peptide reaches the layer where the mechanism is
claimed to operate.
RATING: ❌ [date] -- kind: evidence ABSENT
WOULD CHANGE IF: A randomized, controlled, blinded trial in a defined
population, with an objective or validated appearance
endpoint assessed by blinded raters, against a vehicle
control (the same product without COMPOUND A), of adequate
duration -- plus evidence that the peptide penetrates to the
claimed site of action.
| Element | Pts |
|---|---|
| Claim stated with population and endpoint; missing population **named as missing** | 1.5 |
| Endpoint identified as **self-reported appearance**, not a clinical outcome | 1.5 |
| **No comparator** — specifically, no vehicle control — identified | 1.5 |
| Rating ❌, tagged evidence-absent | 1 |
| Falsification condition specified as a trial design, ideally naming the **vehicle control** | 0.5 |
**The vehicle-control point is the sophisticated one on this item.** A serum without the peptide is the
comparator that matters, because everything else in the formulation — occlusion, hydration, the act of
applying something twice daily — plausibly produces the reported effect. Award the full comparator credit
for "no control group"; note vehicle control in feedback for students who missed it.
**A2 — the one number (4 points).**
- **What was measured:** users' *reports* of smoother-looking skin — a self-reported, subjective
appearance measure, not an objective one. (1.5)
- **Who measured it:** the users themselves, in a study run by the seller, with no blinding described.
(1.5)
- **What it was compared against:** **the page does not say.** (1) There is no control arm, no vehicle
comparison, and no before/after standard stated. Full credit requires the student to identify the
comparator as the missing item.
Accept students who additionally note that 92% is a proportion of an unstated denominator, and that
"users" may mean everyone enrolled or only those who completed.
**A3 — the delivery claim (4 points).**
Route is the central question because a topically applied peptide must cross the **stratum corneum**, a
barrier that large, hydrophilic molecules cross poorly (Chapter 30). If it does not reach the layer where
collagen is produced, the mechanism cannot operate regardless of how well established the mechanism is in
a dish. **This is dossier field 4's closing line in cosmetic form: does the route it is used by match the
route in which it was shown to work?** (2)
For the delivery claim to be checkable the page would need to state what the "proprietary delivery system"
is and provide penetration data — evidence that the intact peptide reaches the target layer at a relevant
concentration, in human skin, not in a model. (2)
**A4 — origin and mechanism (2 points, 1 each).**
- **Origin:** "a fragment of a structural protein your body already makes." Field 2 — origin carries no
evidentiary weight. Letting it raise the rating is the field-2 error.
- **Mechanism:** "signals your skin to produce more of its own collagen." **Rule 3 — mechanism never
upgrades a rating.**
Accept "over 40,000 five-star reviews" (popularity/testimonial, Chapter 6) for either point.
**A5 — the disclaimer (2 points).**
It tells you the product is being sold in a **cosmetic** category rather than as a drug, which means it
was not required to demonstrate efficacy for a disease claim to any regulator. (1) It tells you **nothing
about the evidence** — a product in this category may be well studied or not studied at all, and the
disclaimer is silent between those. (1) **Deduct if the student treats the disclaimer as a downgrade
signal**: that is the Rule 4 error and the mirror image of treating approval as a ✅.
---
### Source B — the press release (18 points)
**B1 — four-line rating (6 points).** Model answer:
CLAIM: COMPOUND B reduces symptom flares in adults with chronic
tendon pain. Endpoint: symptom flares -- a clinically
meaningful outcome, but here measured against a
non-concurrent comparison group.
EVIDENCE: Open-label Phase 2, single arm, 40 completers, 12 weeks. No
concurrent control; the flare comparison is against a
historical observational cohort. Primary reported result is a
serum marker (surrogate), 38% from baseline. Flare figure is
exploratory and relative only -- no event counts, no rates, no
absolute difference. Subgroup result reported. Not replicated.
Conspicuously absent: a randomized, blinded, concurrently
controlled trial with flares as a pre-specified endpoint.
RATING: ⚠️ [date] (a defensible ❌ evidence-absent is also
creditable -- see grading note)
WOULD CHANGE IF: A randomized, double-blind, placebo-controlled trial in adults
with chronic tendon pain, with symptom flares as the
pre-specified primary endpoint, adequate duration, and enough
participants to detect a clinically meaningful difference --
reporting absolute event rates in both arms.
| Element | Pts |
|---|---|
| Claim stated with population and endpoint | 1 |
| **Single-arm / no concurrent control** identified | 1.5 |
| **Historical comparison** identified as the flare figure's basis | 1.5 |
| Serum marker identified as a **surrogate**, distinguished from the flare claim | 1 |
| Rating dated, with a study-shaped falsification condition | 1 |
**Grading note on the rating.** Both ⚠️ and ❌ (evidence absent) are defensible here and both earn full
credit **when justified**. The case for ⚠️: a completed Phase 2 in humans with a plausible signal is
meaningfully more than nothing. The case for ❌: the specific claim — that it reduces flares — rests
entirely on a historical comparison, which is not evidence for a causal claim, so the claim as stated is
untested. **What is not creditable is a rating that ignores the single-arm design**, in either direction.
**B2 — three limiting design features (4 points; up to 3 named + 1 for quality of explanation).**
1. **Single-arm and open-label — no concurrent control and no blinding.** Prevents attributing any change
to the compound: natural history, regression to the mean, co-interventions, and expectation effects are
all uncontrolled, and chronic tendon pain fluctuates on its own.
2. **The comparison group is historical.** A cohort observed at a different time, under different care,
selected by different criteria, and measured by different methods is not a comparator. This is
frequently the single most misread feature in press releases.
3. **The primary reported result is a surrogate.** A serum marker of tissue turnover is not pain and is
not function; accepting it as benefit is a bet.
Also creditable: **small sample (40 completers, with no stated enrollment number — dropouts unaccounted
for)**; **short duration relative to a chronic condition**; **a subgroup finding reported alongside
primary results**; **topline announcement — no peer review, no full protocol, no full data**.
**B3 — interpreting the 50% figure (3 points).**
Needed, none of which the release provides (award up to 2.5 for the list, 0.5 for stating that the
release provides essentially none of it):
- **Absolute event rates in both groups** — 50% fewer than what? Two flares versus four is a very
different finding from twenty versus forty.
- **How a "flare" was defined**, and whether it was defined the same way in both groups. In a historical
observational cohort, almost certainly not.
- **How many participants contributed** to each figure, and over what period.
- **Whether the comparison was pre-specified** — the release calls it exploratory, which is itself the
answer to this one.
- **Who assessed flares, and whether they knew the assignment.**
- **Any measure of precision** — a confidence interval.
**B4 — the p-value (3 points).**
It establishes that the observed within-group change in the serum marker is unlikely to be **that large by
chance alone, given the statistical model** — a statement about sampling variability within a single arm.
(1) It does **not** establish that the compound caused the change, because with no control arm there is
nothing to compare against; it does not establish that the change matters clinically; and it does not
apply to the flare claim at all. (1) It does not repair the central limitation because **a p-value
quantifies chance, not confounding** — a single-arm study can produce a very small p-value for a change
that a placebo arm would have matched. (1)
**B5 — "validates our mechanism" (2 points).**
Granting it entirely: a validated mechanism licenses **a hypothesis worth testing** — nothing more.
(1) It does not license a rating upgrade, an efficacy claim, or a change in field 6. **Rule 3.** Chapter
22 is the canonical case: a target engaged exactly as designed and no clinical benefit. (1)
---
### Source C — the abstract (18 points)
**C1 — four-line rating (6 points).** Model answer:
CLAIM: COMPOUND C improves physical function in adults with chronic
musculoskeletal pain. Population as studied: adults 40-70,
pain >= 6 months, EXCLUDING inflammatory arthritis, prior
joint replacement, and diabetes. Endpoint: validated physical
function score at 24 weeks -- a patient-relevant outcome.
EVIDENCE: Randomized, double-blind, placebo-controlled, 280 randomized
(140/arm), 238 completers, 24 weeks. Good design. PRIMARY
ENDPOINT NOT MET: mean difference 1.8 points, 95% CI -1.0 to
4.6, p = 0.21 -- the interval includes no difference. A
secondary surrogate (serum marker) improved, p = 0.03. A POST
HOC subgroup (under 55) favored the compound, p = 0.04.
Replication: none. Conspicuously absent: nothing -- the right
trial was run and it answered the question.
RATING: ❌ [date] -- kind: evidence PRESENT AND NEGATIVE
WOULD CHANGE IF: A second adequately powered randomized trial with physical
function as the pre-specified primary endpoint -- ideally in
the younger population the post hoc analysis suggests, tested
prospectively rather than found retrospectively -- showing a
difference that excludes no effect and is large enough to
matter to a patient.
| Element | Pts |
|---|---|
| Claim stated with the **actual enrolled population**, including exclusions | 1.5 |
| **Primary endpoint not met** identified as the governing result | 2 |
| ❌ tagged **evidence present and negative** | 1.5 |
| Secondary and post hoc findings correctly subordinated | 0.5 |
| Falsification condition specifies a **prospective** test of the subgroup | 0.5 |
**The evidence-present-and-negative tag is the highest-value single mark on this exam.** A student who
rates ❌ but tags it *absent* has missed the entire point of the source — this is the one place in the
three where a real answer exists. Award no credit for the tag in that case, and say so in feedback.
**C2 — the three p-values (3 points).**
- **p = 0.21** — the **primary endpoint**, physical function at 24 weeks. **Carries the most weight**, and
it is negative.
- **p = 0.03** — a **secondary endpoint**, a serum marker (a surrogate). Second in weight, and much less
than its position in the Conclusions implies.
- **p = 0.04** — a **post hoc subgroup**, participants under 55. **Carries the least weight**:
hypothesis-generating only, chosen after the data were seen, with no correction for the multiple
comparisons that were possible.
1 point for correct assignment, 1 for the ranking, 1 for a justification that names **pre-specification**
and **multiplicity** — the two reasons the ranking is what it is. **Deduct if the student ranks by p-value
magnitude**; the smallest p-value here belongs to the least trustworthy finding, which is exactly the trap.
**C3 — the Conclusions gap (3 points).**
Quote: **"supports further investigation for physical function… particularly in younger adults"** — or
**"demonstrated favorable effects on tissue turnover."** Either is creditable.
The gap: the trial was designed to answer whether COMPOUND C improves physical function, and **it
answered no.** The Conclusions foreground a surrogate the trial was not designed around and a subgroup
found after the fact, and use the word "supports" for a result that did not support it. A reader who reads
only the Conclusions comes away believing the trial was encouraging; a reader who reads the Results
learns it was negative on the question it asked. (2 for the gap, 1 for the quotation.)
Award full credit to students who note that *"supports further investigation"* is a defensible sentence
for a sponsor to write and still misleading as a summary — the honest version would begin with the
primary result.
**C4 — the population (3 points).**
Precisely: **adults aged 40–70, with chronic musculoskeletal pain of at least six months, who do not have
inflammatory arthritis, have not had a joint replacement, and do not have diabetes.** (1)
Two groups the trial says nothing about — any two of: **adults under 40; adults over 70; people with
inflammatory arthritis; people with prior joint replacement; people with diabetes; people with pain of
shorter duration.** (1)
Effect on a claim about "people with chronic pain": the phrase covers a far larger and different
population than the one enrolled, so **the claim as stated was not tested.** Rule 1 requires the
population, and silently broadening it is the most common way a correct trial result becomes an incorrect
public claim. (1)
**C5 — "adverse events were similar between arms" (3 points).**
**Establishes** (any two, 0.5 each): that adverse events were collected and compared in a randomized,
blinded setting, which is a far stronger safety observation than an uncontrolled report; and that no large
difference in common events appeared over 24 weeks in this population.
**Does not establish** (any two, 1 each): that the compound is safe — this trial was sized to detect a
difference in *function*, not to detect uncommon harms; that longer-term or rarer harms are absent, since
24 weeks and 280 participants cannot see them; that the excluded groups (diabetes, inflammatory arthritis,
prior joint replacement) would show the same profile; or that "similar" means "identical," since no
numbers, definitions, or severity grading are given.
---
### Part D — Across the three (6 points)
**D1 — ranking (3 points).** The ranking, by how much each establishes about its own central claim:
**C > B > A.**
- **Source C** ran the right study — randomized, blinded, placebo-controlled, adequate duration, a
patient-relevant primary endpoint — and got an answer. It establishes the most **even though the answer
was negative**, and that is the whole lesson.
- **Source B** generated a human signal but cannot attribute it: single-arm, open-label, historical
comparison, surrogate primary. It establishes that something was measured, not that the compound did it.
- **Source A** establishes essentially nothing about its claim: no control, no objective measure, no
blinding, self-reported endpoint, seller-run.
1 point for the ranking, 2 for a justification that turns on **design** rather than on tone, size, or
source type. **A student who ranks C last because its result was negative has made the exam's central
error and should receive 0 for D1** — with the reason stated plainly in feedback.
**D2 — better study, unsupported claim (3 points).**
**Source C.** (1) The gap opened because **the quality of a study and the support for a claim are
different things.** A good trial produces a trustworthy answer; it does not produce a favorable one.
Source C's design is exactly what you would want, which is precisely why its negative primary result is
credible — and the concluding sentence then reaches past that result to a surrogate and a post hoc
subgroup, neither of which the design was built to test. (2)
Full credit for any answer that lands on: **the design earns the trust, and the conclusion spends it
somewhere the design did not go.**
Award full credit to a student who argues Source C is *also* the most useful of the three to a reader,
because it is the only one that changed anyone's state of knowledge.
Related: Chapter 5 (the method) · Chapter 6 (hype cycle) · Chapter 30 (route and cosmetic peptides) · Chapter 34 (what analysis establishes) · Chapter 37 (the table) · Appendix C · Appendix D · Appendix F (red flags) · Appendix H (twenty worked evaluations) · Claim Evaluation Rubric · Midterm Exam · Final Exam