Case Study 2 — DRG Creep: The Measure That Became a Target
A real, public phenomenon — documented in the peer-reviewed literature, in CMS rulemaking, and in statute — told qualitatively. The term at its center was coined in print before the national system it describes had even launched. No figure in this case study is asserted with precision; the primary sources are in the further reading, and the magnitudes belong to them.
Background
Case Study 1 ended with a warning: when a classification prices the product, describing the product becomes a financial act. This case is what that warning looks like when it comes true — twice, a quarter-century apart — and what it cost the second time.
In 1981 — two years before IPPS existed nationally — a medical informaticist named Donald Simborg published a short piece in the New England Journal of Medicine titled "DRG creep: a new hospital-acquired disease." Watching the New Jersey experiment, he predicted that once payment attached to DRGs, the recorded case mix would drift upward — not primarily through fabrication, but through the accumulation of small, individually defensible choices in documentation, sequencing, and coding, each made in the direction the money pointed. He named the disease before the epidemic.
The issue
The prediction rests on a structural fact this chapter's §33.6 states as bluntly as it can: the case mix index cannot distinguish "our patients got sicker" from "our documentation and coding got better." Both raise it identically. And a third thing raises it identically too: coding pushed past what the record supports. From inside the CMI, legitimate acuity, legitimate capture, and creep are one number.
That would be a curiosity if the CMI were only a statistic. It is a payment multiplier — so the question "how much of the measured case-mix change is real?" is a question about billions of dollars, asked annually, with hospitals and CMS on opposite sides of it.
What happened — the first time
Through the middle and late 1980s, the measured Medicare case mix rose faster than anyone believed the patients were changing. The research literature of the period took the number apart and concluded that a substantial share of the increase reflected changes in documentation and coding behavior rather than changes in patients — hospitals were learning the system, hiring the coders, buying the encoders, and capturing the severity the old cost-based world had given them no reason to record. Some of that was correction of historic under-coding; some was creep; the boundary between the two was and is genuinely hard to draw. CMS (then HCFA) responded with recalibrations and budget-neutrality adjustments — the beginning of a permanent institutional wariness about case-mix change.
What happened — the second time
In 2008, Medicare replaced the old DRGs with the MS-DRGs this chapter teaches — a refinement that tripled the sensitivity of payment to documented severity, since the new three-tier CC/MCC splits paid precisely for what secondary-diagnosis documentation could establish. CMS understood exactly what it was deploying: the agency projected, in advance and in rulemaking, that measured case mix would rise from documentation and coding improvement alone — hospitals responding to the new tiers by recording severity more completely — without any change in actual patient acuity, and it applied prospective documentation and coding adjustments to payment rates to offset the projected effect.
The projection came true. Measured case mix rose after the MS-DRG transition, analyses attributed a substantial share to documentation and coding change, and the aftermath ran for years: disputes over the size of the effect, statutory intervention — the American Taxpayer Relief Act of 2012 directed further recoupment of the documentation-and-coding effect — and a long tail of rate adjustments as the system clawed back payment it judged had come from better description rather than sicker patients.
Around the same machinery, a parallel and harsher thread developed: severity capture as an enforcement subject. The same MCC arithmetic that §33.5 prices at \$1,867.44 per occurrence made certain high-value secondary diagnoses statistically visible, and patterns in their reporting — clusters of specific MCC diagnoses out of line with clinical prevalence, query programs engineered to harvest severity — have appeared repeatedly in OIG work, payer analytics, and False Claims Act matters. The claims-data methods are the ones this book keeps describing: a hospital whose rate of a lucrative diagnosis sits far from its peers has, in Chapter 21 §21.10's phrase, answered a question nobody asked yet.
The limit this case teaches
Here is what makes this a Chapter 33 case study rather than a Chapter 5 war story: most of the behavior in it was legitimate, and the system still had to claw the money back.
Sit with that. The MS-DRG design asked hospitals to document and code severity completely — that is what a severity-adjusted system is for — and hospitals that did exactly that, compliantly, raised the national CMI, and the budget-neutrality machinery took the aggregate difference back anyway. At the level of the individual record, complete capture is an obligation; at the level of the system, its aggregate effect was treated as an artifact to be offset. Both positions are coherent. The lesson is not that anyone was wrong; it is that a measure that carries payment stops being a pure measure — permanently, structurally, no matter how honest the participants are. Goodhart's law is usually quoted about targets; DRG creep is what it looks like in a classification.
What it shows
Simborg was right before the system existed, which means the vulnerability is in the design, not the people. Every severity-adjusted payment system — MS-DRGs, Chapter 34's APCs at their margins, and most acutely Chapter 36's HCC risk adjustment, where the identical drama is running now with diagnosis codes and Medicare Advantage — carries the same structural incentive. A reader who understands this case has pre-read Chapter 36's enforcement landscape.
The defensible line is the record, and only the record. Between under-capture (the §33.3 reversed loss), complete capture, and creep, the only boundary an auditor, a court, or your own conscience can apply is documentation: severity the record supports, all of it, and nothing else. A program that queries in both directions and can show its work sits on the defensible side of a line that dollar outcomes cannot define.
And the honest CMI conversation is a professional obligation. §33.6's split — service mix versus capture rate — is not managerial hygiene; it is the individual-hospital version of the exact analysis CMS runs on the nation, and the hospital that cannot perform it on itself will eventually have it performed on them, by someone with extrapolation authority (Chapter 37 §37.6).
Discussion questions
-
Define DRG creep in one sentence that distinguishes it from both fraud and legitimate documentation improvement — and then explain why that sentence is hard to operationalize in an audit.
-
CMS predicted the post-2008 case-mix rise in advance and offset it prospectively. What does the accuracy of that prediction imply about how well payment engineers understand provider behavior — and what does the years-long recoupment fight imply about the limits of that understanding?
-
A hospital's CDI program can honestly claim its CMI gains reflect previously under-documented severity. The national adjustment machinery treated aggregate gains as an artifact to offset. Can both be right? Whose problem is the contradiction, and where in a hospital's leadership should it be owned?
-
§33.10 says the accurate record and the defensible record are the same record. Use this case to state the organizational version of that sentence: what does a severity-capture program have to be able to show, and to whom, for its CMI gains to survive scrutiny?
-
Chapter 36 will describe risk-adjusted payment for populations, where diagnosis codes alone — with no stay, no procedure, no bill — carry the payment weight. Using this case as the template, predict where that system's creep will concentrate and what its version of the documentation-and-coding adjustment will look like. Then check your prediction against Chapter 36.