Case Study 1 — The Recovery Audit Contractor Program: A Design, Its Consequences, and Its Repair
Real and public. The Recovery Audit Contractor program is created by federal statute, implemented by published rulemaking and contract, and documented in Government Accountability Office reports, HHS Office of Inspector General work, CMS's own program pages, and litigation. This case study asserts no precise statistic, contingency-fee percentage, denial count, or recovery total. Where the size or direction of an effect matters, it is characterized qualitatively and the reader is sent to the primary sources in this chapter's further reading. Chapter 30's Case Study 1 tells the appeals half of this story — the backlog at the Office of Medicare Hearings and Appeals — and this one deliberately does not repeat it.
Background
By the early 2000s, Medicare had a measurement problem it could not act on. The Comprehensive Error Rate Testing (CERT) program told the agency, every year, that a meaningful share of fee-for-service payments were improper — and CERT is a measurement, not a recovery mechanism. It samples, it scores, it publishes a rate. It does not go and get the money.
The Medicare Administrative Contractors, meanwhile, had medical review authority but a competing mandate. Their primary job is to process claims — quickly, at enormous volume, within performance standards that reward throughput. Asking the same organization to both pay claims fast and comb through paid claims looking for its own errors is asking it to prioritize against itself.
The idea Congress reached for was structural: create a contractor whose only job is to find improper payments already made, and pay it out of what it finds.
The Medicare Prescription Drug, Improvement, and Modernization Act of 2003 authorized a demonstration project. Recovery Audit Contractors began work in a small number of states, reviewing Medicare fee-for-service claims after payment, identifying overpayments and underpayments, and being compensated on a contingency fee — a percentage of what they found. The Tax Relief and Health Care Act of 2006 made the program permanent and directed CMS to take it nationwide; the national rollout was completed around 2010.
Three design features are worth stating precisely, because everything downstream follows from them.
Contingency-fee compensation. The RAC is paid a share of the improper payments it identifies. The percentage is set by contract, has been publicly reported in a range, and differs by contractor and claim type — verify the current rates at CMS rather than taking any figure from a textbook.
A postpayment orientation. The program's core is review of claims that have already been paid, through automated review (a determination from claims data alone — a duplicate, a units-versus- descriptor conflict, an edit violation, no records needed), semi-automated review, and complex review (records requested, a clinician or coder reads them).
Both directions, by statute. The program is required to identify underpayments as well as overpayments. This is real and it is systematically under-appreciated by providers, who experience the program entirely as a demand letter.
The issue
The design worked, in the narrow sense that it found improper payments at scale. What it also did was generate an entirely new object: a volume of contested determinations that no other part of the Medicare program had been sized for.
The largest single driver was the medical necessity of a short inpatient admission. A patient arrives at a hospital sick enough to need care and not obviously sick enough to need a bed for several days. The physician admits. The stay lasts a night or two. Under Medicare, the hospital is paid under the inpatient prospective payment system (Chapter 33) rather than the outpatient one (Chapter 34), and the difference in payment is substantial.
Reviewing that decision after the fact is a genuinely hard problem, and it has the shape this whole chapter is about: an audit is a reading of a decision by somebody who was not there. The RAC's reviewer has the record. The record documents the patient's condition, the physician's assessment, and the orders. What the record cannot fully contain is the clinical judgment made at 2 a.m. about a trajectory that had not yet happened. When a contractor concluded, from the same record, that the patient could have been treated as an outpatient, the hospital's payment was recharacterized and the difference recovered.
Two structural consequences followed, and both are teachable.
First, the disputes did not stop at the contractor. Hospitals appealed, at scale, and won a substantial share at the administrative law judge level. Chapter 30's Case Study 1 covers what that did to the appeals system, why the wait itself became a penalty, and how the backlog was eventually litigated and worked down. Read that case beside this one; they are two halves of a single design failure.
Second, the incentive question became unavoidable. A contractor paid on what it finds has a reason to find. That is the point of the design — and it is also the objection to it. Under the program as originally run, the contingency fee attached at the point of recovery, which meant a determination later overturned on appeal had, in the interim, already generated a fee and already taken the money.
Neither of those consequences was hidden, and neither was a scandal. They were the predictable results of a design that had optimized for one thing — finding improper payments — without pricing the disagreement it would produce.
What happened
In 2013 and 2014 the program largely stopped. CMS wound down the existing RAC contracts while it procured the next round, and the procurement was contested, which extended the interruption. For a period, the RACs were not issuing new documentation requests. The pause was not a repudiation of the program; it was a contracting event. But it created something unusual: a natural interval in which everyone involved — CMS, hospitals, the appeals system, and the contractors — could look at what the first phase had produced.
What CMS did next is the part worth studying, because it is a real example of a review program being re-engineered rather than defended. In announcing the next round of contracts, CMS published changes to the program's operating rules. The published set included:
- Additional documentation request limits tied to a provider's own compliance. Rather than a flat cap by provider size, the number of records a RAC may request is adjusted by how the provider has performed on prior review — fewer requests for providers with low denial rates, more for providers with high ones. A review program that reviews the compliant less is a review program that has accepted a piece of feedback.
- A discussion period before the claim goes to the MAC for adjustment. A defined window in which a provider can talk to the RAC about a finding before it becomes a recovery. This is the origin of a step that now appears throughout postpayment review, and §37.8's response letter is written for it.
- Contingency fees withheld until after the second level of appeal. Rather than being paid on recovery and repaying the fee if reversed, the contractor waits. This is a direct answer to the incentive objection above, and it is a small change with a large behavioral implication: a determination that will obviously be overturned is now worth less to the contractor who makes it.
- Accuracy and overturn thresholds. Contractors are held to published performance standards on how often their determinations survive review.
- A shortened look-back for patient-status reviews where the hospital submitted the claim promptly.
Separately, and at the same time, the underlying clinical question was addressed by rule rather than by audit. The FY 2014 inpatient prospective payment system final rule introduced the two-midnight benchmark for inpatient admission — the standard Chapter 16 §16.3 teaches — replacing a case-by-case judgment with a stated expectation about the length of stay. Short-stay medical review then moved, for a period, to a probe-and-educate approach and subsequently to the Beneficiary and Family Centered Care Quality Improvement Organizations rather than the RACs. The program stopped auditing a question it had proved could not be audited consistently, and the question was answered upstream instead.
CMS also offered hospitals administrative settlements for eligible pending inpatient-status appeals — a stated percentage of the net allowable amount in exchange for withdrawing the appeals. Chapter 30's Case Study 1 publishes those percentages and analyzes them; this case does not repeat the figures.
What it shows
First, a review program is a system with feedback, and it will be gamed by its own incentives unless somebody prices the disagreement it produces. The contingency fee did exactly what it was designed to do. What nobody costed in advance was the appeals volume, and the appeals volume is where the money and the credibility actually went. The design constraint on a review program is not how much it finds. It is how much of what it finds survives.
Second, the reforms are a model of what a corrective action plan looks like at national scale — and they map onto §37.10's durability ladder almost line for line. Withholding the contingency fee until after level two is a rung-1 intervention: it changes what the system permits, not what anybody knows. Scaling record requests to prior performance is rung-2 — it makes the reviewer's behavior depend on measured accuracy. The discussion period is rung-3, a place for a finding to be read by a human before it becomes money. Almost none of it is education, which is why almost none of it depended on anybody's good intentions.
Third — and this is the finding a working coder should take personally — the RAC's most productive tool required no medical record at all. Automated review is a query against claims data: a duplicate, a unit that exceeds the descriptor, a code pair subject to an edit, a modifier that does not belong on a code. Every one of those is computable by the provider, on the provider's own data, before the claim goes out. §37.10's outside-in analysis is not a metaphor for what a contractor does; it is literally the same query.
Fourth, the "underpayments as well as overpayments" mandate is real and almost nobody uses it. Providers experience the RAC as a one-way instrument because the demand letter is the only artifact that arrives with a deadline. Chapter 28 §28.8 is the reason this matters: the underpayment does not announce itself in either system.
Fifth, the two-midnight change is the most under-remarked event in the story. After years of auditing a judgment call, the answer was not a better auditor. It was a rule that made the judgment call reviewable — a standard stated in advance, which is exactly what §37.2 requires of an internal audit and exactly what the disputed short-stay reviews had lacked. An audit against an unwritten standard is a disagreement with extra steps.
The outcome
The program continues. Its issues are approved by CMS and published in advance on each RAC's website — a list of exactly what is about to be reviewed in your region, free, and read by almost nobody. Its record-request volumes are governed by published limits. The discussion period, the delayed contingency fee, and the accuracy thresholds are part of the operating design.
Short-stay inpatient status is no longer the program's center of gravity. The appeals backlog it helped create was litigated, scheduled by a federal court, and eventually worked down — Chapter 30's Case Study 1.
None of that means the design question is settled. A contractor paid for finding things will find things, and the honest reading of two decades of this program is that the incentive is neither illegitimate nor free. What changed is that the cost of a wrong determination now falls, partly, on the party making it.
The lesson
Read the approved-issues list. It is published, it is specific to your region and provider type, and it is a forecast of what will be reviewed. Together with the OIG Work Plan (§37.2), it is the cheapest risk-based audit scope a practice will ever get.
Run the automated review on yourself. Every finding a RAC can make without a record, you can make without a record. Duplicates, units against descriptors, edit pairs, modifiers that do not belong on a code — that is a query, not a project.
Use the discussion period. It exists because providers asked for a place to be heard before a finding became money. A finding you can defeat in a conversation is a finding you should not be appealing.
And take the design lesson into your own department. If you build a program that measures how much your auditors find, you have built a smaller version of the same incentive. Measure how much of what they find survives rebuttal. That is the number the federal program eventually learned to watch, and it took a decade and a court order.
Discussion questions
-
Chapter 30's Case Study 1 asks whether the volume of disagreement a review program generates should be a design constraint on that program. Answer it, using this case's reforms as evidence, and say which single reform you think did the most work.
-
Contingency-fee compensation is defended as aligning a contractor's incentives with the program's goal, and attacked as rewarding volume over accuracy. Both are true. Design a compensation structure that keeps the first and blunts the second — then name what your design would cost.
-
The two-midnight benchmark replaced a case-by-case judgment with a stated expectation. Using §37.2's standard stack, explain why the earlier arrangement was structurally unauditable, and name one thing in your own organization that is currently audited against an unwritten standard.
-
The program is required to identify underpayments. In practice, providers rarely see one. Give two explanations that do not require anyone to be acting in bad faith, and say which one Chapter 28 §28.8 supports.
-
Automated review requires no medical record. List five automated-review findings a contractor could make against your own claims data this afternoon, and say for each whether your practice would have detected it first.
-
This case describes a program being re-engineered rather than defended. Compare it with §37.10's durability ladder: where does each reform sit, and what does it tell you that almost none of them was education?