Case File Project Rubrics
The student-facing description of the project is in The Systems Integration Case File, and the four-level scale sketched there is deliberately coarse. This page turns it into an analytic instrument you can defend at a grade appeal, and — more importantly — into one you can actually apply to twenty-eight entries from sixty students without losing your term.
The rationale, stated for a skeptical colleague
The objection is reasonable and usually goes like this: "This is a writing assignment in a science course. My students already cannot name the cranial nerves. Why am I spending grading hours on six sentences of prose?"
Three answers.
First, it assesses the thing we claim to teach and never test. Read your own last exam. Almost every item is answerable within one chapter. That is not a criticism of your exam — it is what item banks and chapter structure produce. But every downstream course, and every clinical encounter, is between-system. Students are being graded on a skill adjacent to the one they will need, and they respond rationally by studying the adjacent skill. The Case File is the only component of most A&P courses that puts points on integration.
Second, it is a writing assignment only in the sense that a proof is. Six sentences is not an essay; it is a constrained argument with a required form — name the system, name two prior systems, state the direction of causation, use the numbers, make a falsifiable prediction. It is closer to a chemistry problem than to a reflection paper, and it is graded like one.
Third, the marginal cost is small and front-loaded. The 90-second rubric pass below, with the spot-check schedule, costs about twenty minutes a week for a section of sixty. That is less than the time you currently spend regrading a badly written exam item.
The honest counterargument: if your course is a 10-week accelerated module with 120 students and no teaching assistant, do the checkpoint-only version described under Grading efficiency and do not apologize for it. Four graded entries out of twenty-eight still produces most of the effect, because the effect comes from students writing the entries, not from you reading them.
The analytic rubric
Twenty points per entry. Six dimensions, weighted toward the two that actually distinguish integrated thinking: the number of correct connections, and whether the direction of causation is stated.
| Dimension | 0 — absent | 1–2 — emerging | 3–4 — proficient | 5 — integrated | Max |
|---|---|---|---|---|---|
| Mechanism accuracy — is the physiology of the new system correct? | Wrong or absent | Partly correct; one substantive error | Correct, appropriately detailed | — | 4 |
| Cross-system connections — how many prior systems, correctly linked? | None | One system named | Two systems, both mechanistically linked | Three or more, all correct, none padded | 5 |
| Direction of causation — is A causes B stated, not A relates to B? | Association only | One direction stated, others vague | All connections directed | — | 4 |
| Use of the patient's data — are Amara's or the family's actual numbers used? | No numbers | A number cited without interpretation | Numbers used to support the mechanism | — | 3 |
| Feedback loop — is at least one loop identified and labeled? | None | A loop implied but not named | A loop named, with its sign (negative or positive) | — | 2 |
| Prediction — is it specific, mechanistic, and falsifiable? | None or "I predict problems" | Specific but not justified | Specific, justified, and checkable against a named later chapter | — | 2 |
Level bands. 18–20 = Level 4 · Integrated. 14–17 = Level 3 · Connected. 9–13 = Level 2 · Adjacent. 0–8 = Level 1 · Isolated.
Two scoring rules that prevent most disputes.
- A connection that is asserted but not mechanistically linked scores zero on dimension 2 and cannot earn points on dimension 3. "The kidney is related to the heart" is not a connection. Say this on the first day and enforce it on the first entry.
- Padding is penalized, not rewarded. Three correct connections score 5; four connections of which two are hand-waved score 3. Students discover quickly that the instrument rewards correctness per claim rather than claims per paragraph, which is the habit you want.
A worked example: Entry 22, at all four levels
The Chapter 26 entry gives students these findings on Amara, eighteen months after her infarction:
| Finding | Value | Reference |
|---|---|---|
| Serum creatinine | 1.4 mg/dL (was 0.9 two years ago) | 0.6–1.1 |
| Estimated GFR | 48 mL/min/1.73 m² | > 90 |
| Urine albumin : creatinine | 180 mg/g | < 30 |
| Serum potassium | 3.4 mEq/L | 3.5–5.0 |
| Hemoglobin | 10.6 g/dL | 12.0–15.5 |
| Blood pressure | 138/84 on lisinopril and furosemide | < 130/80 |
| Ejection fraction | 58% (HFpEF) | 55–70% |
Level 1 · Isolated — 6/20
"The kidney filters blood through the glomerulus and reabsorbs what the body needs in the tubules. Amara's kidneys are not working as well as they should. Her creatinine is 1.4, which is high, and her GFR is 48, which means stage 3 chronic kidney disease. She will need to be monitored."
Scores. Mechanism 3, connections 0, direction 0, data 2, loop 0, prediction 1.
What is here: correct renal physiology and correct use of two numbers. What is missing: every other system in the file. This is a competent Chapter 26 summary and a failed Case File entry, and the distinction is worth making explicit to students — the assignment is not "describe the kidney."
Level 2 · Adjacent — 11/20
"Amara's kidney function has declined to a GFR of 48. Kidney disease is related to high blood pressure, and she has had hypertension for years. Her heart problem is also connected to her kidneys, since both are cardiovascular risk factors. Her hemoglobin of 10.6 is low, and anemia is associated with kidney disease. I predict she will have more problems in the future."
Scores. Mechanism 3, connections 2, direction 1, data 2, loop 0, prediction 0.
What moves an entry from Level 2 to Level 3: replace every "is related to" with a verb that has a direction. This student has correctly identified three relevant systems — cardiovascular, renal, hematologic — and has committed to nothing about any of them. "Anemia is associated with kidney disease" is true and empty. The Level 3 version of that same clause is "her failing kidneys secrete less erythropoietin, so marrow red cell production falls, so hemoglobin is 10.6." Same three nouns; one causal arrow inserted; the sentence now makes a claim that could be wrong, which is what makes it worth points.
The single most effective intervention here is mechanical: return the entry with every directionless verb circled and nothing else marked. Students fix it in one iteration.
Level 3 · Connected — 16/20
"Amara's GFR has fallen to 48 mL/min/1.73 m². Two mechanisms from earlier chapters explain it. First, her heart: HFpEF with an EF of 58% means high filling pressures, so renal venous pressure rises and renal perfusion pressure falls, which lowers GFR. Second, years of hypertension have damaged the glomerular capillaries, which is why she now spills albumin at 180 mg/g. Her kidneys also secrete less erythropoietin, so her hemoglobin has fallen to 10.6. Her potassium of 3.4 is from furosemide, which blocks NKCC2 in the thick ascending limb and delivers more sodium distally, where it is exchanged for potassium. I predict her kidney function will keep declining."
Scores. Mechanism 4, connections 5, direction 4, data 3, loop 0, prediction 0.
What moves an entry from Level 3 to Level 4: close a loop, and make the prediction falsifiable. This is genuinely good work. Every connection is real, directed, and supported by a number. What it is missing is the thing the whole project exists to teach: all of these arrows run one way. The heart affects the kidney; hypertension affects the glomerulus; the kidney affects the marrow. Nothing here comes back.
And "I predict her kidney function will keep declining" is unfalsifiable within the course. A prediction earns points when a later chapter can check it.
Level 4 · Integrated — 20/20
"Amara's GFR of 48 is the closing of a loop, not a new problem. Her stiff left ventricle (Ch. 18, EF 58% but high filling pressure) raises central venous pressure, which raises renal venous pressure and lowers the perfusion gradient across the glomerulus, so GFR falls. The kidney reads reduced perfusion as volume depletion and activates RAAS (Ch. 16, 19): angiotensin II constricts the efferent arteriole — which defends GFR briefly at the cost of glomerular hypertension, explaining the albumin of 180 mg/g — and aldosterone retains sodium, which raises blood volume, which raises preload, which raises filling pressure in a ventricle that cannot accommodate it. That is a positive feedback loop: the kidney's correction for low perfusion makes the heart worse, which lowers perfusion further. Her lisinopril interrupts it at ACE, which is why her pressure is 138/84 rather than 168/98. The furosemide interrupts it by volume — and produces the potassium of 3.4, because blocking NKCC2 delivers sodium distally where aldosterone-driven reabsorption exchanges it for K⁺ (Ch. 31). Her hemoglobin of 10.6 is the same failing kidney seen from Chapter 17: less erythropoietin, less marrow stimulation.
Prediction: because anemia reduces oxygen content while her heart cannot raise cardiac output, I expect her exercise tolerance in the Chapter 30 rehabilitation entry to be limited by oxygen delivery rather than by symptoms of angina, and I expect any further diuresis to raise her creatinine transiently before it improves her congestion."
Scores. Mechanism 4, connections 5, direction 4, data 3, loop 2, prediction 2.
The four moves that made this Level 4, in the order students acquire them:
- A loop is named and signed. Not "the heart and kidney affect each other" but a traced circle with the word positive attached and the reason it is positive.
- Drugs are used as evidence about mechanism. Lisinopril and furosemide are not treatment trivia here; each is an experimental interruption of a specific step, and naming the step is what proves the student holds the chain.
- Every number does work. 58%, 180 mg/g, 3.4, 10.6, 138/84 each support a specific claim. Numbers cited without a claim earn dimension-4 points but not dimension-1 points.
- The prediction names a chapter and a mechanism, and can be wrong. "Limited by oxygen delivery rather than by angina, in the Chapter 30 entry" is checkable. That is the whole difference between a prediction and a hope.
Grading efficiency: 28 entries × 60 students
The arithmetic that frightens people is 1,680 artifacts. The arithmetic that matters is that you do not grade all of them.
The 90-second rubric pass. With the six-dimension instrument in front of you, an entry takes 90 seconds because you are not reading for prose. Scan in this fixed order: (1) count the prior systems named — that is dimension 2; (2) scan for directional verbs — dimension 3; (3) scan for numbers — dimension 4; (4) look for the word loop or a traced circle — dimension 5; (5) read the last sentence only — dimension 6; (6) read the middle for accuracy — dimension 1. Sixty entries is ninety minutes, which you will not do weekly, which is why the next section exists.
The spot-check schedule. Grade every entry from every student four times a term; grade a rotating 20% sample otherwise; give completion credit for the rest.
| Entries | What you do | Time for 60 students |
|---|---|---|
| 7, 14, 21, 28 (checkpoints) | Full rubric, all students, written feedback on the weakest dimension only | ~90 min each |
| All others | Random 20% sample, full rubric, no written feedback | ~20 min each |
| All others | Everyone else: completion credit, verified by a two-second glance for the three required headers | ~10 min |
Students do not know in advance which entries are sampled, which is what preserves effort. Say openly that this is how it works — the transparency costs nothing and removes the suspicion that grading is arbitrary.
Peer review, used properly. On the weeks you sample, have students exchange entries in class and score one another on dimension 2 and dimension 3 only — count the systems, circle the directionless verbs. Those two dimensions are objective enough that untrained peers agree, and the exercise teaches the rubric faster than your comments do. Do not have peers score mechanism accuracy; they cannot, and the feedback is worse than none.
Checkpoint grading at 7, 14, 21, and 28. These four are chosen deliberately. Entry 7 is early enough to correct habits before they set. Entry 14 falls at the end of the endocrine chapter, where RAAS first makes a four-system loop available and the level distribution suddenly spreads. Entry 21 catches metabolic integration. Entry 28 is the capstone. Weight them 15% / 20% / 20% / 45% of the project grade and let the sampled entries ride on completion.
Three alternative formats
Concept-map version. Students maintain one growing diagram instead of prose: systems as nodes, mechanisms as labeled, directed arrows. Each chapter they add one node and at least two arrows, and they may not add an arrow without a verb on it. Scoring converts directly — dimension 2 becomes the number of new correct arrows, dimension 3 becomes the fraction of arrows with a direction and a verb, dimension 5 becomes whether any cycle exists in the graph. Best for visual learners and for large sections, because a map is faster to grade than prose. Weakness: it lets students hide vagueness inside a short arrow label, so require the verb.
Oral / presentation version. Four to six students per term present a five-minute "case conference" on one chapter's entry, taking questions from the class. Use the same rubric with one addition: a responsiveness dimension worth 3 points for answering an unscripted "what would happen if…" question. This format is the single best preparation for clinical case presentation and it is the version students remember years later. Weakness: it scales badly, so use it for a subset of entries or a subset of students.
Group version with role assignment. Teams of four, with rotating roles: Physiologist (mechanism of the new system), Connector (links to prior systems, and owns dimension 3), Data analyst (owns every number and its reference range), Predictor (owns the prediction and audits the group's previous ones). Roles rotate every chapter so nobody specializes. Grade the group entry with the standard rubric and add an individual role-performance mark. This is the best option for large sections and the best for teaching the fact that clinical reasoning is a team activity. Weakness: the usual one; the rotating Predictor role and the audit requirement are what keep the free-rider problem manageable, because the audit exposes who wrote what.
Capstone options for Chapter 33
Choose one; all four are scored with the same six dimensions scaled to 60 points.
- The full model, run forwards. Given three new pieces of information about Amara at Chapter 33 — say, a new medication, an intercurrent infection, and a fall — predict the consequences across at least six systems, with directions, and identify two feedback loops. The default, and the closest to the project's stated purpose.
- The audit. Students revisit all 27 of their own predictions, mark each right, wrong, or untested, and write 500 words on what their wrong predictions had in common. This is the metacognitive version and it produces the best writing of the term. It is also the option most resistant to outsourcing, since nobody else has their predictions.
- The teaching artifact. Build a one-page systems map, a five-minute recorded explanation, or a written case for a first-year student, covering the whole arc. Graded on accuracy and on whether an untrained reader could follow the causation.
- The new patient. Given a short vignette of an entirely new patient — a 30-year-old with type 1 diabetes and a foot ulcer, a 55-year-old with COPD and cor pulmonale — build a complete integrated model from scratch in one sitting. The hardest option and the best evidence of transfer; use it as a final exam component rather than as homework.
Academic integrity when a language model can write six sentences
It can, and it will write them well. Pretending otherwise wastes everyone's time, and prohibition enforced by suspicion produces a worse course. Here is the practical position.
Start from what the assignment is for. The Case File exists to make a student build a model in their own head. An outsourced entry produces a correct artifact and no model, which the student discovers on the final exam. Say that plainly, once, in week one, and then design the assignment so that outsourcing is more work than doing it.
The structure already resists it, in three specific ways.
- The PREDICT step is personal and dated. A prediction is yours, written before the evidence exists, and Entry 28's audit requires you to grade your own 27 predictions. A model can generate a plausible prediction; it cannot generate the prediction you made in week four, and a student who outsourced every entry has nothing coherent to audit. Requiring the audit is worth more than any detection tool.
- The patient's numbers accumulate and must stay consistent. By Chapter 26 an entry has to be consistent with a creatinine, an ejection fraction, a potassium, and a medication list established across twenty-one prior entries. Generated text is characteristically fluent and numerically unmoored. Grade dimension 4 strictly and this surfaces on its own.
- Oral checkpoints cost two minutes. At entries 7, 14, 21, and 28, ask three students at random one question each about their own entry: "You wrote that aldosterone raises preload. Walk me through it." Two minutes per student, four times a term. The possibility of being asked does the work; the asking is almost incidental.
A defensible course policy, stated positively. Permit students to use a language model to check their work, to explain a mechanism they did not understand, or to critique a draft entry — and require them to say so in one line at the end of the entry, naming what they asked and what they changed. Prohibit submitting generated text as their own entry. This policy is enforceable, matches what students will actually do in practice, and — this is the part worth noticing — converts the tool into a study aid that rehearses the mechanism rather than replacing it. A student who asks a model to critique their causal chain has done the assignment twice.
What not to do. Do not run entries through an AI-detection tool. They are unreliable at this length, they misfire on non-native English writers at higher rates, and a false accusation costs more than every outsourced entry in the course combined.
Next: Common Misconceptions and How to Break Them — the most practically useful page in this companion.