Case Study 1 — The Audit of the Chart Review: Medicare Advantage Risk Adjustment Under Enforcement
What this is. A real, public, and still-unfolding regulatory and enforcement story: how Medicare Advantage came to be paid on documented diagnoses, what happened to the diagnoses once that was true, and how the government built an audit around it. Every institution, statute, rule, and program named here is real. No settlement amount, penalty, or audit finding percentage is printed in this case study, because the figures move and because this book does not invent them; where the size of something matters, it is characterized rather than quantified. Verify the current state of every rule and every case before relying on it — this is one of the most actively litigated subjects in American health law.
Background: how a payment system came to read charts
Medicare Advantage — Part C, Chapter 3 §3.6 — pays a private plan a fixed monthly amount for each enrolled beneficiary and makes the plan responsible for that beneficiary's care. The design problem was visible from the start and is §36.2's subject: a fixed payment per member rewards an organization for enrolling healthy members, and the early program was widely understood to be doing exactly that.
Congress addressed it in stages. The Balanced Budget Act of 1997 directed that payments be risk-adjusted. An early method adjusted only for inpatient history, which had an obvious defect — it paid attention only to members sick enough to be hospitalized. The Medicare Prescription Drug, Improvement, and Modernization Act of 2003 directed a comprehensive model, and CMS phased in the CMS-HCC model, which reads diagnoses from inpatient, outpatient, and professional encounters alike, beginning in 2004.
That is the moment the diagnosis code changed jobs. Before it, a diagnosis on a Medicare Advantage encounter was an administrative fact. After it, the diagnoses documented for a member during one year determined what the plan was paid for that member for the next.
The issue
Every participant understood immediately what the new design implied, because the same thing had happened twenty years earlier in the inpatient hospital. Chapter 33's Case Study 2 tells that story: when a classification prices a stay, the classification becomes an object of attention, and "DRG creep" was named in the medical literature before the system had finished launching.
Here the object of attention is the diagnosis, and the pressure has three distinct forms, only one of which is misconduct:
Legitimate improvement. Fee-for-service Medicare had no reason to code completely — a chronic condition that did not justify the day's service was frequently left off the claim, and nothing happened. Medicare Advantage organizations invested in getting the whole picture recorded, and a substantial part of the resulting increase in coded severity was real conditions that had always been there and had never been written down. Chapter 33 §33.6's warning about the case mix index applies word for word: a measure of sickness that is also a measure of documentation capture cannot, by itself, tell you which one moved.
Aggressive but arguable practice. Retrospective chart reviews. In-home health risk assessments performed by clinicians who provide no treatment. Prospective "suspect condition" lists delivered to physicians before visits. Every one of these has a defensible version and an indefensible version, and the difference is frequently a matter of program design rather than of intent.
And unsupported submission. A diagnosis reported for payment that the medical record does not support.
What happened
Four things, over roughly two decades, and they are best read as one system responding to itself.
1. Congress built a discount into the payment before anyone was accused of anything. Because diagnoses are coded more completely in Medicare Advantage than in fee-for-service Medicare, and because the model is calibrated on fee-for-service data, scores in Medicare Advantage run higher for reasons that are partly documentation rather than health. The Affordable Care Act established a statutory minimum coding intensity adjustment — an across-the-board reduction applied to Medicare Advantage risk scores every year. Verify the current percentage; it is set by statute and rule and has been revisited. A payment system with a built-in haircut for coding intensity has conceded the mechanism in advance, which is a useful thing for a coder to know before anyone in a practice describes complete coding as free money.
2. The Medicare Payment Advisory Commission kept measuring the gap. MedPAC's annual Report to the Congress: Medicare Payment Policy has, for many years, estimated that Medicare Advantage risk scores exceed what the same beneficiaries' scores would be in fee-for-service, and has attributed a substantial share of the difference to coding intensity — including specifically to chart reviews and health risk assessments. The reports are free, public, and the single best plain-language source on this subject. Read the current one.
3. The Office of Inspector General audited the practice itself. OIG has published a sustained body of work on Medicare Advantage risk adjustment: contract-level audits of individual organizations testing whether submitted diagnoses were supported by medical records, and thematic reports examining chart reviews and health risk assessments as sources of diagnoses that appeared nowhere else in a beneficiary's record. That last analysis is the detection method §36.8 describes, performed at national scale: it needs no chart at all, only the data, and it asks a question a legitimate program can answer easily.
4. CMS finalized an audit rule with teeth. Risk Adjustment Data Validation (RADV) is CMS's mechanism for testing submitted diagnoses: a sample of enrollees from a contract, medical records requested, and diagnoses that the records do not support removed. The mechanism existed for years without the feature that makes an audit financially consequential — extrapolation.
In January 2023, CMS published its contract-level RADV final rule. Two provisions matter to anybody working in this field:
- Findings will be extrapolated from the audited sample to the contract, beginning with payment year 2018.
- No "fee-for-service adjuster" will be applied. The industry had argued at length that because the risk model is calibrated on fee-for-service diagnoses — which are not themselves audited to a medical-record standard — an audit holding Medicare Advantage to a stricter standard produces a systematically unfair result. CMS declined to build an offset for that argument into the rule.
A legal challenge followed. As with any active rule, check the current status of both the regulation and the litigation before relying on any description of them, including this one.
The outcome, so far
There is no tidy ending, and pretending otherwise would misrepresent a live area of law. What can be said with confidence:
- Risk adjustment is not going away. It is the mechanism that makes any capitated arrangement survivable, and every alternative payment model built since the 1990s uses some version of it.
- The audit apparatus has been strengthened, and extrapolation is the change that matters. Chapter 37 §37.6 owns the arithmetic; the short version is that a finding rate in a sample of a few hundred members, projected across a contract's population and several payment years, is not a small problem.
- False Claims Act enforcement has been extensive and sustained. The government's recurring theory is straightforward: submitting a diagnosis for payment that the medical record does not support is a false claim, and a chart-review program designed to find only the diagnoses that raise payment — while leaving unsupported ones in place — is evidence about the program's purpose. Cases have been brought against plans and against provider groups, frequently initiated by insiders under the qui tam provisions Chapter 5 §5.3 explained. Amounts have been large. This book prints none of them; read the Department of Justice's own announcements for the current record.
- And the overpayment obligation is separate from the submission. CMS's 2014 rule applying the sixty-day overpayment requirement to Parts C and D was itself litigated — vacated by a district court, with the vacatur later reversed on appeal. Verify the current state of that law too. The practical point survives every twist of it: discovering that you have been paid on an unsupported diagnosis starts a clock, and Chapter 31 §31.9 owns what happens next.
What it shows
Payment design is behavioral design, and everyone involved knew it in advance. Nothing in this case study is a surprise that emerged later. The incentive was legible in the design, Congress discounted for it before any enforcement action, MedPAC measured it annually, the OIG audited it, and CMS eventually built an audit that could reach the money. This is what a payment system looks like when it works on itself — noisily, slowly, and in public.
The hardest problem here is not fraud. It is that three very different things produce the same number. A risk score can rise because patients genuinely got sicker, because documentation finally captured conditions that were always there, or because unsupported diagnoses were submitted. The number cannot distinguish them, and neither can an outsider without the charts. That is the same epistemic problem Chapter 33 §33.6 identified in the case mix index and Chapter 29's Case Study 2 identified in a falling denial rate, and this book's answer has been the same every time: a number that improves is a question, not a conclusion.
Detection increasingly happens in the data, before anybody reads a chart. The signature that matters most — a diagnosis that exists in a chart review or a health risk assessment and nowhere else in a beneficiary's record — is visible from claims data alone. An organization can run that query on itself, in an afternoon, and most do not.
And the defense is contemporaneous and bidirectional. Account 31-2245 taught this in Chapter 21 with modifier 59: eleven of forty-two claims were, in fact, separately documented and defensible, and the practice could not prove it after the fact. Here the equivalent is a trace from every submitted diagnosis to a page of a signed, dated, face-to-face encounter note — and a deletion history. A review program that has never removed a code has told an auditor everything about itself before the first record is opened.
The lesson
The accurate record and the defensible record are the same record, and this is the chapter where that sentence stops being a slogan and becomes an audit standard.
Notice what the enforcement landscape does not say. It does not say that reviewing charts is improper — CMS's own rules contemplate that additional supported diagnoses may be submitted. It does not say that a practice should code conservatively; Chapter 5 §5.8 disposed of that, and under-reporting here describes a population as healthier than it is, distorts the quality data computed from the same codes, and leaves the practice unable to answer basic questions about its own patients.
What it says is that the record has to be able to bear the weight of the payment. Every dollar in this payment system rests on a sentence somebody wrote in a chart, and the only defensible position for the coder is the one this book has held for thirty-six chapters: report what the record supports, report all of it, and report only that.
Discussion questions
-
Congress applied a coding intensity adjustment to Medicare Advantage risk scores before any of the enforcement actions described here. What does it mean for a payment system to discount in advance for a behavior it expects? Is that a concession, a control, or both?
-
The fee-for-service adjuster argument holds that the model is calibrated on unaudited fee-for-service diagnoses, so auditing Medicare Advantage to a medical-record standard is comparing two different things. State the strongest version of that argument. Then state the strongest response. Which do you find more persuasive, and what evidence would change your mind?
-
Three causes produce the same rise in a risk score: sicker patients, better documentation, and unsupported submissions. Design a measurement an organization could run on itself that would distinguish the second from the third. What data would you need, and what would you look at first?
-
A chart-review vendor tells you it finds three to five additional conditions per chart and works on contingency. Chapter 36 §36.8 says a contingency arrangement is not per se improper. Do you agree? Write the contract terms you would insist on before signing one.
-
Chapter 33's Case Study 2 and this one describe the same phenomenon in two payment systems twenty years apart. Name three things the second system did that the first did not — and then name the one thing neither system solved.
-
(Chapter 37 §37.6) Extrapolation is what turned RADV from an inconvenience into an existential audit. Explain to a practice administrator, in plain language and without arithmetic, why a finding rate in a sample of two hundred members can produce a demand that threatens a contract.