Case Study 2 — The Number That Gets Read as Fraud: The Improper Payment Rate and Its Contested Meaning

Real and public, and a case about a measurement rather than a claim. The federal improper payment estimate is required by statute, computed by a published methodology, reported annually, and repeatedly explained by the agencies that produce it — and repeatedly reported as something it is not. This case study asserts no percentage, dollar amount, or year-over-year figure. It describes the definition, the mechanism, the contest over its meaning, and what a working revenue cycle professional should take from all three. Every current figure belongs to the primary sources in this chapter's further reading.


Background

Since the early 2000s, federal law has required agencies to estimate and report improper payments in programs susceptible to significant improper payment. The statutory line runs from the Improper Payments Information Act of 2002 through amendments in 2010 and 2012 and into the Payment Integrity Information Act of 2019, which consolidated the framework. The requirement is unglamorous and sensible: an agency that cannot say how much of what it pays out is paid in error cannot manage the program.

For Medicare fee-for-service, the estimate is produced by the program §37.5 introduced: Comprehensive Error Rate Testing (CERT). A random national sample of claims is pulled, records are requested, independent reviewers score each claim against Medicare coverage, coding, and billing rules, and the results are projected to produce a rate. The rate is published each year, prominently, in the department's financial reporting.

Now read the statutory definition of the thing being measured, because the whole case is in it:

An improper payment is any payment that should not have been made, or that was made in an incorrect amount, under statutory, contractual, administrative, or other legally applicable requirements. It includes overpayments and underpayments. It also includes any payment for which insufficient documentation prevents a reviewer from determining whether the payment was proper.

Three things are inside that definition that almost nobody carries out of a headline.

It includes underpayments. A payment that was too small is improper. The rate is not a measure of money owed back to the government.

It includes payments that cannot be evaluated. If the reviewer cannot tell from the record whether the payment was correct, the payment counts as improper. Not "wrong" — unestablished.

And it says nothing whatever about intent. Fraud requires a knowing false statement. The improper payment estimate is computed from records, by reviewers, against rules. It has no instrument for detecting what anybody intended, and it is not designed to have one.


The issue

The rate is reported every year, and every year a substantial share of the public discussion treats it as an estimate of money lost to fraud.

The agencies that produce it say otherwise, consistently and in writing. CMS has stated repeatedly that the improper payment rate is not a fraud rate and is not designed to measure fraud. The Government Accountability Office has made the same point in its own reporting, and has separately noted that improper payments and fraud are distinct concepts requiring distinct measurement. The HHS Office of Inspector General, whose actual job includes fraud, treats them as different objects.

The correction does not travel. The rate is a single, memorable percentage attached to an enormous program, and the two available readings of it — "the government is being defrauded" and "the government cannot document what it bought" — are not equally quotable.

And the misreading has operational consequences, which is why this is a case study and not a complaint. An improper payment rate that is read as a fraud rate produces political pressure for enforcement. Enforcement pressure produces review programs. Review programs produce records requests, documentation denials, and prepayment review — which land on the desks of people who committed no fraud and, in a large share of cases, delivered the service correctly and described it correctly.


What the number actually contains

Here is the part that makes this chapter's whole argument concrete: the largest single contributor to the Medicare fee-for-service improper payment rate has persistently been insufficient documentation. That has been true across many reporting years and is stated plainly in the published error-rate reporting — verify the current composition in the current report, because the categories and their shares move.

Sit with what an "insufficient documentation" finding actually is. A patient was seen. A service was furnished. A code was assigned that describes it. A claim was submitted and paid. Then, months or years later, a reviewer asked for the record — and the record did not contain a signature, or an order, or the physician's documentation of medical necessity, or a note establishing that the service the code describes is the service that occurred.

In an enormous share of those cases nobody did anything dishonest. A signature was illegible. A standing order was never re-signed. The reason for a test lived in a nurse's note the reviewer did not receive. The records went to an old address. Somebody sent the office note and not the order.

That is not fraud. It is the first theme of this book, measured at national scale: if it isn't documented, it didn't happen. And the money involved is real, because a payment the record does not support is a payment the program cannot defend, whatever actually occurred in the room.


The contest, stated fairly from both sides

The reading this book rejects is that the improper payment rate estimates fraud. It does not, its producers say it does not, and the categories inside it say it does not.

The reading this book also rejects is the comfortable inverse — that because the largest category is documentation, the number is a paperwork artifact and the underlying payments were fine.

That second move is more tempting for a revenue cycle audience, and it is wrong for a reason this book has been building since Chapter 4. A coder does not know what the provider did; a coder knows what the provider wrote. A payment supported by a record that does not establish the service is not a payment somebody has proved was correct. It is a payment nobody can evaluate — and the entire discipline this book teaches exists because that distinction has consequences.

Chapter 5 §5.8 said upcoding and downcoding are both errors and refused to let "conservative" stand in for "accurate." The same refusal applies here in a different direction: "it was only a documentation error" is not a defense, it is a diagnosis. It tells you exactly which control failed, and the control that failed is cheap to fix and expensive to leave — which is the whole argument of §37.3's category 4.

So the honest position is a narrow one, and it has to be held on both sides at once. The improper payment rate is not a fraud rate, and saying so is not minimizing. It is also not noise, and saying so is not alarmism. It is a measure of payments the program cannot substantiate from the record, which is a real and serious thing to be unable to do — and which is fixable by exactly the mechanisms this chapter describes.


What it shows

First, a measurement's name determines how it is used, and the name is usually chosen by somebody who is not thinking about how it will be used. "Improper payment" is precise, statutory, and reads to a non-specialist as a synonym for "wrongful." This book has now produced several instances of a number that told a true story that nobody read correctly — Chapter 23's Case Study 2, where a rising collection ratio was produced by underpayments; Chapter 27's Case Study 1, where a denial rate improved because claims were never reaching adjudication; Chapter 28's Case Study 1, where a number with a story attached stopped being a question. This is the same failure at national scale, with the same mechanism: a plausible interpretation, universally repeated, never tested against the definition.

Second, the composition of a number is worth more than the number. A rate is a summary; the categories underneath it are the instruction. A practice that learns its own denial rate has learned almost nothing (Chapter 29 §29.7's denominator discipline). A practice that learns which category is largest has learned what to fix on Monday. The same is true of the national rate, and the national rate's largest category has been telling the profession the same thing for years.

Third, this is the chapter's clearest statement of its own limits. §37.3 insisted that a scoring sheet separate not-supported from wrong-code, and report direction, and keep a category for "supported but fragile." All three of those disciplines exist because a pooled number cannot be acted on. An error rate with no composition is a headline, not a finding — whether it is CERT's or your own.

Fourth, and least comfortable: the documentation category is the one a coder cannot fix alone. A signature, an order, a stated reason for a test, a physician's note establishing medical necessity — none of those is produced by the coding department. What the coding department can do is find them missing before a payer does, which is §37.2's internal audit and §37.7's step 5, and escalate the pattern rather than the instance. Chapter 38's clinical documentation integrity apparatus exists because that gap is structural rather than individual.


The lesson

Never quote a rate without its definition and its composition. Not the national rate, not your denial rate, not your audit's accuracy rate. State what is in the numerator, what is in the denominator, and what the largest category is. A number that arrives without those three things is an argument wearing a statistic's clothes.

And when somebody in your building says "it was only a documentation error," treat that as the beginning of the analysis rather than the end of it. It is the most actionable finding in this chapter — cheap to fix, invisible until somebody asks, and, at national scale, the largest single thing standing between what American healthcare actually delivered and what it can prove.


Discussion questions

  1. State, in two sentences a physician would accept, why the improper payment rate is not a fraud rate. Then state, in two sentences, why "it's only documentation" is not a defense. Notice how hard it is to hold both — and say which one your organization is more likely to get wrong.

  2. The statutory definition of an improper payment includes underpayments. Explain how a rate that includes underpayments can nevertheless drive enforcement pressure in only one direction, and connect your answer to Chapter 28 §28.8.

  3. §37.3 requires that an audit report accuracy with its denominator and its direction. Using this case, write the two-sentence rule you would put at the top of every audit report your department issues.

  4. A payment counts as improper when insufficient documentation prevents a reviewer from determining whether it was proper. Name three specific documents whose absence produces that finding, and for each say which department in a practice actually controls it.

  5. Compare this case with Chapter 27's Case Study 1 and Chapter 28's Case Study 1. All three involve a number that was accurate and understood wrongly. What do the three failures have in common that a control could actually address? Answer using §37.2's favorable-trend discipline.

  6. This chapter argues that an audit is a reading by somebody who was not there. A national improper payment estimate is that reading, performed on a random sample of the entire program. Is the rate therefore the best available description of American medical documentation, or the worst? Argue one side, then argue the other, and say what would change your mind.