Case Study 2 — The National Guidelines That Were Never Written

A real, public matter, told from the rulemaking record: the decade in which CMS said it intended to publish national guidelines for facility evaluation and management visit levels under the Outpatient Prospective Payment System, the principles it published instead, and the CY2014 decision that solved half the problem by eliminating the measurement rather than standardizing it. The complement to Case Study 1: there, a published rule was not read; here, the rule was never published — and the consequences are still on every emergency department claim in the country. No figures are asserted; the rules themselves are free and are the source.**


Background

When the Outpatient Prospective Payment System (OPPS) began on August 1, 2000, hospitals faced a question CPT had never been written to answer.

CPT's evaluation and management (E/M) codes describe a physician's work: history, examination, and medical decision making, and since 2021 decision making or time (Chapter 15 §15.3). But under OPPS a hospital also reports a visit code — on its own claim, for its own resources: the nursing time, the monitoring, the room, the medications administered, the equipment, the staff who did not write a note. Those are different quantities, and no code set had ever defined the second one.

CMS's answer at the time was interim and honest: until national guidelines exist, each hospital should develop and apply its own internal guidelines for assigning facility visit levels, so long as they reasonably relate the resources used to the level reported and are applied consistently. The agency stated its intention to develop a national standard.

Hold that word: interim. It is now a quarter of a century old.

The issue

Through the 2000s, CMS returned to the question in successive OPPS rulemaking cycles. It described the problem, solicited comment, and considered a series of alternatives — among them proposals from its own advisory panel on outpatient payment, models based on counting interventions, models based on time, and models that would have simply mapped facility levels to the physician's code.

Every one of them ran into the same wall, and the wall is instructive because it is not bureaucratic. A national facility E/M standard has to be:

  • General enough to fit a rural critical access emergency department and an urban level I trauma center;
  • Specific enough to be applied the same way by two people;
  • Cheap enough to apply thousands of times a day without adding documentation whose only purpose is billing;
  • Resource-based, not diagnosis-based, because the whole point is that it measures something the physician's code does not;
  • And stable, because a national leveling standard that changes materially is a system-wide reprogramming exercise.

Those five constraints pull against each other. What CMS published instead of a standard was a list of principles that hospitals' own internal guidelines were expected to satisfy — roughly a dozen of them, and the ones that matter most to this chapter should sound familiar: the guidelines should follow the intent of the CPT descriptor; be based on hospital facility resources rather than on physician work; be written; be applied consistently across patients; not change frequently; be available for review; require only documentation that is clinically necessary; not facilitate upcoding or gaming; and produce coding decisions that other staff and outside reviewers could verify.

That list is where §35.6's five properties come from, and it is the only national standard that exists: a standard for the rule, not a rule.

What happened

In the CY2014 OPPS final rule, CMS did something with the clinic half of the problem that almost nobody predicted. It stopped measuring.

Effective January 1, 2014, hospitals no longer report five levels of E/M for a hospital outpatient clinic visit. They report a single HCPCS Level II code — G0463, a hospital outpatient clinic visit — for a visit of any level. The agency's reasoning was empirical: the facility resources consumed by clinic visits did not vary enough across the five levels to justify pricing them five ways, and a single code removed a leveling problem the agency had been unable to standardize for thirteen years. Chapter 34 §34.8 records the result, and Chapter 20 §20.6 records the mechanism — a G-code exists because CMS needed a code CPT does not supply.

And the emergency department half was left exactly as it was. Five levels, 99281 through 99285, assigned by each hospital's own internal criteria, under the principles above. That is the state of the world §35.6 describes, and it is not an oversight: emergency department resource intensity genuinely does vary enormously across encounters in a way clinic visits apparently do not, so collapsing the ED levels would have thrown away a real distinction to solve an administrative one.

The outcome, and the contest

The contest never resolved, and both sides of it are reasonable.

The case for a national standard, made repeatedly by payers, oversight bodies, and some hospitals themselves: an entire category of Medicare payment is assigned by rules the paying party has never seen, which makes cross-hospital comparison impossible, makes distribution analysis the only available oversight tool, and puts every hospital to the expense of building and defending something that could have been built once.

The case against one, made by hospitals and their associations: departments differ enormously, resource-based criteria that fit everyone would fit no one well, and the internal-guidelines arrangement has the advantage of being adaptable to how a department actually works — which is also, not incidentally, why it produces levels that relate to real resource use rather than to a national average.

The uncomfortable middle, which is this case study's actual subject: a non-decision is a decision, and its costs are real and unevenly distributed. Every hospital pays to build criteria. Every payer audits distributions instead of rules. Emergency department facility level distributions became — and remain — a recurring subject of federal and payer attention, precisely because a distribution is the only comparison available when the underlying rules are private. And a hospital whose criteria are perfectly sound but whose scoring records were never retained cannot prove any of it, which means the organizations most exposed are not the ones leveling wrongly; they are the ones who cannot demonstrate how they leveled at all.

What it shows

An interim arrangement that works becomes permanent, and nobody re-decides it. Chapter 25's Case Study 2 named the shape in a practice: a workaround that mostly works becomes institutional knowledge. This is that pattern at the level of federal payment policy, and the mechanism is identical — the arrangement is tolerable, the alternative is hard, and there is no event that forces a review.

"Write your own rules and apply them consistently" is a heavier obligation than it sounds. It is easy to read as permissive. It is not: it transfers the entire burden of design, documentation, training, reproducibility, retention, and defense onto each organization, and it is enforced not by approving the rules but by testing the results. Consistency is a claim about an organization's behavior over years, and it can only be proved with records nobody thinks to keep on a Tuesday.

Sometimes the correct fix for an unmeasurable measurement is to stop measuring it. G0463 is worth sitting with, because a book about coding accuracy should be honest that the accurate answer is occasionally "this distinction is not worth its cost." The agency looked at thirteen years of data, concluded the levels were not measuring a real difference in clinic resources, and deleted them. That is a legitimate outcome, and a coder who reflexively treats fewer codes as worse precision has mistaken granularity for accuracy.

And the half that was kept tells you the deletion was not laziness. The ED levels survived because the variation there is real. The test was evidence, applied twice, with two different answers.


Discussion questions

  1. Take the five constraints on a national facility leveling standard and try to design one anyway. Which constraint do you sacrifice, what does the sacrifice cost, and who bears it? Then say whether your design would survive Chapter 37's audit — and what a reviewer would ask you first.

  2. CMS published principles for hospitals' internal guidelines rather than the guidelines themselves. Compare that with Case Study 1's published review framework. Both are "a standard for the rule rather than a rule." Which one worked better, and what was different about the two situations?

  3. G0463 replaced five clinic visit levels with one. Argue that this reduced coding accuracy, then argue that it increased it. Which argument is stronger, and what evidence would settle it? What does your answer imply about the book's recurring claim that specificity is a virtue?

  4. A hospital's ED criteria are sound, resource-based, written, and consistently applied — and the hospital retained no record of how any individual encounter was scored. A payer requests documentation for forty encounters. Describe exactly what the hospital can and cannot demonstrate, what it should do next, and what it should have been doing all along. Then connect this to Chapter 26's case of a fact that was detectable from outside before it was detectable from inside.

  5. Emergency department facility level distributions are watched because the underlying rules are private. Is distribution analysis a fair oversight tool in that situation? Argue both sides, and then state what a hospital can do proactively so that a distribution question arrives as a conversation rather than as a records request.

  6. This case study and Case Study 1 fail in opposite directions — a published rule nobody read, and a rule nobody published. Write the one-sentence lesson each teaches, then write the sentence that covers both. What does the combined lesson say about where a new specialty coder should spend the first two hours of §35.10's method?