Case Study 1 — When an Algorithm Decides Coverage: The Medicare Advantage Utilization Management Rules

Real and public. Everything asserted below comes from federal rulemaking published in the Federal Register, from CMS's own published guidance and program pages, and from congressional oversight that is a matter of public record. This case study asserts no denial rate, no settlement figure, no penalty amount, and no company's performance statistic. Where litigation is mentioned it is identified as allegations in publicly filed complaints, which are not findings. Rules in this area are recent, are being implemented on staged compliance dates, and are being litigated and amended — verify the current requirements at CMS before relying on any characterization here, including this one.


Background

Every chapter of this book so far has looked at automation from the provider's side of the claim: a scrubber that appends a modifier, a template that inserts an attestation, a posting rule that classifies an adjustment, an engine that suggests a code. Chapter 37 §37.10 built a control for the whole family — the assertion register — and made a point of saying the mechanism is not the provider's alone. A payer's configuration can make an assertion too, and Chapter 31's Case Study 1 showed one doing exactly that, with the provider still carrying the exposure.

This case is that observation at national scale, in the one place where an automated assertion decides whether a patient gets care.

The setting is Medicare Advantage. Chapter 3 introduced it: a Medicare beneficiary may receive their benefits through a private plan under contract with CMS rather than through original fee-for- service Medicare. Enrollment has grown to a very large share of the Medicare population. The plans are paid a risk-adjusted amount per member per month (Chapter 36), which means — and this is the structural fact everything else follows from — the plan retains the difference between what it is paid and what it spends. Utilization management is therefore not an incidental administrative function. It is where the money is.

And utilization management is exactly the kind of work that automates. It is high volume, it is repetitive, its inputs are structured, and a decision produces an immediate measurable saving. Plans and their vendors built and licensed tools that predict expected lengths of stay, score the appropriateness of a level of care, and flag cases for review. Some of those tools are widely licensed clinical criteria sets that have been used in utilization review for decades. Some are newer predictive models.

None of that is improper on its face, and it is important to say so before anything else, because the argument in this case is not that automation in coverage review is wrong. Chapter 22 established that medical necessity is a coverage concept and that some mechanism has to apply it. A tool that helps a nurse reviewer find the cases most likely to need a physician's attention is doing something useful. The question this case is about is narrower and sharper: what happens when the tool's output becomes the decision rather than an input to one?


The issue

Post-acute care is where the question came to a head, and it is worth understanding why that setting in particular.

A Medicare beneficiary is admitted to a hospital, treated, and discharged to a skilled nursing facility or an inpatient rehabilitation facility to recover. How long they need is a clinical judgment made day by day: how the patient is progressing, what they can do today that they could not do yesterday, whether they can safely go home. Under original Medicare there is a benefit period with defined coverage rules. Under a Medicare Advantage plan, continued coverage typically requires ongoing authorization, and the plan decides.

A predictive tool can produce a number for that situation. Given a diagnosis, an age, a functional score, and a set of comorbidities, it can estimate how long patients like this one typically stay. That estimate is a genuinely useful piece of information for planning.

It is not a coverage determination, and the distinction is the whole case. An expected length of stay describes a population. A coverage determination is about one person, on one day, whose recovery is either tracking that population or is not. When the estimate becomes the target — when coverage terminates on the day the model predicted rather than on the day the patient's own documented condition supports — an assertion about a population has been substituted for a judgment about a patient.

That is precisely the mechanism this book has watched a dozen times, in a new and much more serious place. In every earlier instance a configuration asserted something about a claim. Here a configuration asserts something about a person's care.

Three features made it hard to see from outside:

The reason given was clinical, not algorithmic. A denial or termination notice cites criteria and clinical findings. Nothing on the notice says a model produced a number.

Appeals are individually rational and collectively rare. Chapter 30 taught the appeal machinery and Chapter 30's Case Study 1 taught what happens when the volume of disagreement exceeds what the system was sized for. A patient in a skilled nursing facility, or their family, is not well positioned to appeal on the day coverage ends, and a great many do not.

And the aggregate pattern is invisible to any single provider. A facility sees its own cases. The distribution across a plan's entire post-acute population is visible only to the plan and to a regulator — which is Chapter 37 §37.10's outside-in view again, from the outside.


What happened

Three separate public processes converged, and each produced a different kind of artifact.

1. Rulemaking

For contract year 2024, CMS finalized a Medicare Advantage and Part D rule with substantial utilization-management provisions. Read the rule itself rather than any summary of it — including this one — but the provisions most relevant to this chapter are these:

  • Coverage criteria must comply with Medicare's. An MA plan must provide the basic benefits that original Medicare covers, and it may not apply criteria more restrictive than Medicare's coverage rules. Where Medicare's rules are not fully established, a plan may use internal coverage criteria only in defined circumstances, and those criteria must be based on current evidence in widely used treatment guidelines or clinical literature and made publicly available.
  • Medical necessity determinations must be based on the individual patient's circumstances, including the treating provider's information, rather than on a criteria set applied without regard to the case in front of the reviewer.
  • An approved prior authorization remains valid for the duration of the approved course of treatment, which removes the mechanism by which a course of care could be interrupted by re-review.
  • A Utilization Management Committee must review the plan's utilization management policies annually, ensure consistency with Medicare's rules, and — the part that matters most for this chapter — be accountable for those policies by name and structure rather than diffusely.

Read that list against Chapter 37 §37.10's assertion register and the correspondence is close enough to be startling. Publish what your criteria assert. Have a named body own them. Review them on a schedule. And test the assertion against the individual case rather than against the configuration. A federal rule and an internal control arrived at the same design, from opposite directions.

2. Sub-regulatory clarification

CMS subsequently published guidance addressing artificial intelligence and algorithmic tools directly, in response to questions the rule had raised. The clarification was narrow and precise, and its precision is the teachable part: a plan may use an algorithm or software tool to assist a coverage decision, and the tool may not itself be the basis for denying or terminating coverage. The decision must rest on the individual patient's circumstances as documented.

Notice what that does and does not do. It does not ban the tool. It does not require the tool to be disclosed to the patient. It does not say the tool is inaccurate. It relocates the decision — the same move §38.8 makes about coder-in-the-loop, written as regulation instead of as workflow design. And it leaves exactly the question §38.8 says decides whether a loop is real: is the human review substantive, or is it a click?

3. Oversight and litigation

Congressional oversight examined Medicare Advantage prior authorization, with particular attention to post-acute care denials; the committee reports are public and are the best available source for the scale of the practice, which this book will not characterize numerically. Separately, class action complaints were filed against Medicare Advantage organizations alleging that algorithmic tools drove post-acute coverage terminations. Those are allegations in publicly filed complaints. They are not findings, they are not admissions, and a reader should treat them as what they are: a public record that the question was contested in court, not an answer to it.

Alongside all of this, CMS finalized a separate interoperability and prior authorization rule requiring impacted payers — including Medicare Advantage organizations — to send prior authorization decisions within specified timeframes, to provide a specific reason for a denial, and to publicly report prior authorization metrics, with application programming interfaces phased in on staged compliance dates. The transparency provisions are the ones to watch: a published denial rate is the outside-in analysis of §37.10, run by the regulator, on data the plan is required to hand over.


What it shows

First, and most importantly for a coder: the assertion mechanism is not a provider problem. It is a health care problem, and it runs in both directions across the claim. This book has documented a dozen instances on the provider's side of the transaction — a macro, a template, a mapping, a posting rule, a form a patient signs, a script handed to a person. This case is the same mechanism operating on the payer's side, with a patient's care as the output rather than a claim's content. The register you build after Chapter 37 should have a row for every assertion your processes rely on that is made by somebody else's system, because Chapter 31's Case Study 1 established the principle that being unable to change it does not make it stop being your problem.

Second, the regulatory answer to an algorithmic decision was not a technical standard. It was a governance requirement, and it consists of four moves this book has been making since Chapter 5: publish the criteria, name the owner, review on a schedule, and decide the individual case on the individual record. Nothing in that list is about the algorithm. Every item is about who is accountable for what it asserts — which is the design principle behind the assertion register, and the reason §37.10 defined the register by assertion rather than by system.

Third, "a human reviewed it" is a claim that has to be tested, not a box to be checked. §38.8's three design questions transfer directly to a utilization review desk: does the reviewer see the case before the score, or the score before the case? Is disagreeing with the tool cheap or expensive? Was the reviewer's productivity standard recomputed for review work? A rule that requires a human decision does not, by itself, produce one. This is the same gap the chapter identified between a coder-in-the-loop workflow and a rubber stamp, and it is not solved by regulation any more than it is solved by software.

Fourth, the transparency provisions may end up doing more work than the prohibition. A rule that says a decision may not rest on a tool is enforced case by case, after the fact, by whoever complains. A rule that requires published denial metrics creates a distribution somebody can compute — and Chapter 26's Case Study 1 established the general principle in a very different setting: some errors are detectable from outside before they are detectable from inside. A published rate invites the comparison that a case-by-case standard never produces.

And fifth, the honest complication. Utilization management is not illegitimate and neither is automation in it. Some care is not necessary; some stays are longer than the patient's condition supports; somebody has to apply a coverage standard, and applying it consistently across an enormous population is a real problem that human review alone has never solved well either. The failure mode in this case is not that a model was used. It is that a model's output about a population was allowed to answer a question about a person, and that nobody outside the plan could see it happening. A version of this story in which the tool prioritized cases for a physician's substantive review, and the physician's decision cited the patient's own documented condition, would not be a case study.


The outcome

The rules are in force on their stated effective and compliance dates, with the transparency provisions phasing in later than the utilization management provisions. Plans have stood up utilization management committees, published internal coverage criteria, and revised notice language. Congressional oversight is ongoing. The litigation is ongoing. Whether the practical effect matches the rule's intent is exactly the sort of question that takes several years of published metrics to answer, and the metrics are only now being required.

None of which is a conclusion, and this case study is not going to manufacture one. What can be stated is that a specific automated assertion was identified, characterized in rulemaking, addressed by a governance requirement rather than a technical one, and made measurable by a disclosure requirement — and that the design of the fix is recognizable to anyone who read Chapter 37.


The lesson

For the working revenue cycle professional, four things transfer immediately.

Read the plan's published criteria. They are now required to be publicly available where internal criteria are used. That is a document you can hold against a denial, and Chapter 30 §30.4 taught what it is worth in an appeal: the policy proves the standard the payer owes itself. An organization that never reads a payer's published criteria is choosing to argue from the record alone.

When a denial's reasoning does not match the record, say so specifically. A notice that recites criteria without engaging the patient's documented condition is a notice that can be answered on that ground. Quote the record (Chapter 37 §37.8) and identify what the determination did not address. This is not a rhetorical move; individualized consideration is what the rule requires.

And ask the same three questions about your own loops that you would ask about theirs. Which comes first, the case or the score? Is disagreeing cheap? Was the productivity standard reset? If those questions are fair to ask of a payer's utilization review desk, they are fair to ask of your own coding department, and §38.8 is where you ask them.

Finally: the register has a row for this. Somewhere in your organization a process depends on an assertion made by a payer's system — an eligibility response, an authorization determination, a predicted length of stay, an edit, a remittance code. You cannot change any of them. You are still answerable for the claims and the patient communications that rely on them, and the entry is not the payer's software. It is the fact that a process of yours relies on an assertion made by somebody else's system, with an owner and a test.


Discussion questions

  1. CMS's clarification permits an algorithm to assist a coverage decision and forbids it from being the basis of one. Using §38.8's three design questions, describe how you would test whether a given utilization review desk complies with that in substance rather than in form — and name what evidence you would ask for.

  2. Compare the four governance moves in the rule (publish the criteria, name the owner, review annually, decide the individual case) against Chapter 37 §37.10's five register fields. Which register field has no counterpart in the rule, and does its absence matter?

  3. This case involves a payer's configuration. Chapter 31's Case Study 1 established that a payer's assertion can still be the provider's exposure. Write the register entry for one assertion your organization currently relies on from a payer's system, including the evidence test — and say honestly whether anyone would run it.

  4. The transparency provisions require published prior authorization metrics. Chapter 29 §29.7 established that a denial rate is not one number — it varies by lines or claims, by zero-pay or any non-contractual adjustment, and by adjudicated or submitted. What does that imply about comparing published rates across plans, and what single definitional disclosure would make the comparison meaningful?

  5. Argue the other side, in good faith. Construct the strongest case that predictive tools in utilization management improve outcomes for patients as a group, and then say precisely what safeguard would have to be true for you to accept that argument. Is your safeguard in the rule?

  6. This chapter's §38.4 says a CDI program can corrupt by selection without sending a single non-compliant query. Describe the utilization-management analogue: how could a plan comply with every provision of this rule, in every individual case, and still produce a systematically skewed result? What would detect it?