Rubric — The Peptide Evidence Dossier
For the capstone assembled in Chapter 40 from entries built across the term using Appendix C. Five to ten compounds, twelve fields each, plus a drift statement.
State the grading principle to students before the first entry is due, and again when you return the first graded one:
A well-reasoned ⚠️ that the instructor disagrees with outscores a ✅ that happens to match the book.
This is not generosity. It is the only defensible standard in a course whose own ratings are date-stamped and whose stated position is that some of them will be wrong. Reasoning is the graded object. Verdicts are not. Students will not believe this until they see a high mark on an entry you argued with in the margin, so make sure that happens early.
The rubric
| Criterion | Exemplary | Proficient | Developing | Insufficient |
|---|---|---|---|---|
| 1. Claim specificity — every rating attaches to a claim with a population and an endpoint (Rule 1) | Every claim names a population and an endpoint precisely enough that a reader could tell whether a given trial tested it. Populations are the ones actually studied, not silently broadened. | Population and endpoint present on every claim; occasional imprecision ("adults," "improvement") that a reader could still work with. | Population or endpoint missing on some claims; several claims broadened beyond the population studied. | Ratings attached to molecules rather than claims. Field 6 has one row per compound. |
| 2. Evidence by design, not by count — field 5 | Best available evidence described by design, population, endpoint, comparator, size, duration, result, with replication status and what is conspicuously absent filled in and substantive. | All design elements present for most compounds; "conspicuously absent" attempted rather than merely present. | Evidence summarized narratively or by volume ("many studies show"); design elements partly present; absence line blank or generic. | Study counts substituted for design. No comparator or population recorded. Animal and human evidence not distinguished. |
| 3. Rating justified and date-stamped — Rules 3, 5 | Every rating dated. Justification traces to field 5 and nothing else. No rating is upheld by mechanism, origin, popularity, or regulatory status. Multiple ratings per molecule where warranted (Rule 6). | Ratings dated and mostly traceable to evidence; one or two lean on mechanism or status. | Some ratings undated; justification frequently mechanistic; molecule-level rating persists in places. | Undated ratings. Justification is mechanism, marketing copy, or personal conviction. |
| 4. The two kinds of ❌ | Every ❌ tagged evidence absent or evidence present and negative, correctly, with the distinction reflected in the verdict. | Tags present and mostly correct. | Tags inconsistent or several reversed. | Distinction absent; ❌ used to mean "bad." |
| 5. Field 12 — present and specific | Field 12 names the study: population, endpoint, comparator, duration, rough size. Separate up-moving and down-moving conditions. States whether such a study is underway. Falsifiable as written. | Field 12 present for every entry and specific enough to be checkable, though not fully specified as a trial design. | Field 12 present but vague ("better evidence," "more studies"). Down-moving condition often missing. | Field 12 blank or absent on multiple entries. An entry with an empty field 12 is a belief, not a conclusion, and is graded as such. |
| 6. Honest blanks | Blanks used deliberately and marked, with the reason where it is known ("no human data"; "not looked"). Field 9's "unknown because nobody has looked" line is filled in where it applies. Nothing plausible-but-unsourced appears anywhere. | Blanks generally honest; a small number of fields filled with unsourced but harmless plausibility. | Several fields filled with confident material the student cannot source; few blanks in an evidence base that clearly has holes. | Every field filled. A dossier with no blanks on a thinly evidenced compound is a red flag, not an achievement. |
| 7. Fields 7 and 8 kept separate | Approved use recorded precisely (indication, population, jurisdiction, line of therapy); claimed use recorded in the seller's own terms without softening; overlap stated; the gap between them is analyzed. | Both fields present and distinct; overlap noted. | Fields overlap or one is a paraphrase of the other; claimed use edited into something more reasonable. | Collapsed into a single "what it's for." |
| 8. Risk in three categories — field 9 | Studied use, unstudied use, and preparation quality separated and populated. Interactions and contraindications noted. Procedure-relevant risks flagged. | All three categories present; one thin. | Only risks of misuse recorded, or only label risks. | "No known side effects" written for a compound with no human trials. |
| 9. Field 10 not adjusting field 6 | Regulatory, sport, sales, access, and quality kept as five separate questions. No rating anywhere moved because of approval status, in either direction. | Field 10 mostly separated; rating independent. | Field 10 collapsed into one line; at least one rating tracks approval status. | "It's approved, so ✅" or "not approved, so ❌" reasoning present. |
| 10. Internal consistency across entries | The same standard of evidence produces the same rating across compounds. A student who demanded a hard outcome for one compound demanded it for all. Inconsistencies, where present, are named and defended. | Consistent standard with minor unexplained variation. | Standard visibly relaxes for compounds the student favors and tightens for ones they do not. | Ratings track preference. Rule 4 violated: a compound downgraded for sounding like marketing rather than for the state of its evidence. |
| 11. Sourcing and citation honesty | Every non-obvious statement carries its source in the margin. Tier is clear: named-and-checkable, described-but-unnamed, or labeled [constructed teaching example]. Nothing invented is unlabeled. |
Sources present for most substantive claims; tiers mostly distinguishable. | Sourcing sporadic; unclear which statements were read directly. | An unlabeled invented source, statistic, trial, or citation appears. See edge case 5. |
| 12. Drift statement | Names which ratings moved, in which direction, triggered by what, and what that reveals about the student's own bias. Compares current ratings against the field 12 written before investment. | Movement documented and attributed; some self-analysis. | Movement listed without cause or reflection. | Absent, or asserts no drift without evidence. |
Scoring
A twelve-criterion rubric is heavier than most capstones need. Two workable approaches:
Weighted (recommended). Criteria 1, 2, 3, 5, 6, and 10 carry double weight; the rest single. Exemplary 4 / Proficient 3 / Developing 2 / Insufficient 0. Maximum 72.
The insufficient-floor rule. Insufficient on criterion 1, 3, or 11 caps the dossier grade regardless of everything else — those three are the course's non-negotiables. Announce the cap in advance; it is a fair rule and an unfair surprise.
Do not average away criterion 6. A dossier that scores well everywhere by filling every field with confident material is precisely the artifact this course is designed not to produce, and the arithmetic of a twelve-row rubric will reward it if you let it.
Common grading edge cases
1. The student's rating differs from Appendix A. Not an error. Read the justification. If it traces to field 5 and the student engaged the evidence the book cited, it can be Exemplary — including when you think they are wrong. The disagreement memo assignment exists to make this explicit, and the first time a student sees a high mark on an entry you argued with, the whole class recalibrates.
2. A gorgeous dossier on five well-evidenced compounds. Technically excellent, pedagogically thin — the student never had to identify an absence. Grade the entries on their merits and address the selection separately. This is why the compound list is collected in week one and binding. A student who swapped out their ❌ compounds in week ten escaped the harder half of the course, and the drift statement should say so.
3. Everything is ⚠️. Two very different students produce this. One has genuinely found preliminary evidence across a preliminary field — check whether any claim has a hard outcome behind it and whether any ❌ appears; if the evidence supports the ratings, this is fine. The other is hedging to avoid being wrong. The tell is field 12: real ⚠️ ratings have sharp, distinct falsification conditions; defensive ⚠️ ratings have the same vague sentence twelve times.
4. Everything is ❌. Rule 4 violation until proven otherwise. Check whether the ❌ ratings are tagged absent versus present-and-negative — a student who cannot distinguish them has produced a mood, not an assessment. Also check the well-evidenced compounds: a ❌ on a claim with a large outcome trial behind it is either a serious misreading or a serious argument, and the justification will tell you which within a sentence.
5. An invented statistic, trial, or citation. This is the one failure that is categorically different from the others. Everything else on this rubric is a skill in development; this is the failure the book spends forty-four chapters teaching against. Treat an unlabeled fabricated source as academic dishonesty under your institution's policy, and distinguish it clearly from a labeled constructed example, which is a legitimate and encouraged teaching device. Make the distinction explicit in the assignment prompt so the line is impossible to cross innocently.
6. Dosing, protocols, or sourcing in the submission. Against course policy and against the book's own guardrail. Most first instances are innocent — a "typical protocol" line copied into field 4, a supplier name in field 10 offered as evidence of availability. Return it for revision without penalty the first time and with penalty the second. Note in feedback that field 4 asks about route and duration of action, not about regimens.
7. A compound the student personally uses. Grade the entry, not the choice, and be aware that your marginal comments will be read personally in a way they would not be on another entry. Keep every comment attached to a field number and a rule number. If an entry's problem is that the student's verdict outruns their evidence, say exactly that, in the same words you would use for any compound. Never write a comment that would read differently if the student were not using it.
8. Two students with near-identical entries on the same compound. For a compound with one obvious literature, convergence is expected and is not evidence of collusion. Look at fields 11 and 12: verdicts are written for the student's own situation and falsification conditions reflect individual standards. Identical field 11 and 12 text is the signal; identical field 3 is not.
9. The late-term compound swap. A student replaces a compound in the final weeks and submits a complete entry for it. Almost always an escape from a compound with no evidence — which was the more instructive entry. Ask for both. Grade the original.
10. Excellent entries, no assembly. Twelve good fields per compound and no synthesis, no cross-compound comparison, no drift statement. Chapter 40 is an assembly chapter, and criteria 10 and 12 are where the assembly is graded. A dossier that is ten strong entries stapled together is Proficient, not Exemplary, and the feedback should name what is missing: the patterns that only appear across compounds.
Related: Chapter 5 · Chapter 6 · Chapter 37 · Chapter 40 (capstone) · Appendix A · Appendix C (the workbook this rubric grades) · Appendix H · How to Teach This Book · Where Students Get Stuck · Claim Evaluation Rubric · Participation Rubric