Case Study 2 — The A1c Below Seven: When Hitting the Numerator Is Bad for the Patient
What this is. The complementary failure to Case Study 1. That case was about risk adjustment overstating how sick a population is. This one is about quality measurement overstating how well it is being cared for — and about a measure that, pursued hard enough, harms the patients it was written to protect.
The clinical evidence and the guideline history below are real and public. The organization in the second half is a clearly labeled composite, built from documented patterns in the pay-for-performance literature. No real practice, plan, or measure steward is described. Verify every current measure specification with its steward and every clinical target with current professional guidance; both have changed and will change again.
Background: the measure everybody agreed on
Hemoglobin A1c reflects average blood glucose over roughly the preceding three months. It is objective, inexpensive, drawn from a tube of blood, and it correlates with the microvascular complications of diabetes — retinopathy, nephropathy, neuropathy — that make the disease expensive and disabling. It is, in other words, an almost perfect quality measure: cheap, standardized, resistant to gaming at the level of the individual test, and pointed at an outcome rather than an activity.
For years the target that entered measure sets, practice dashboards, and physician incentive arrangements was an A1c below 7.0%. CPT Category II code 3044F was created to report exactly that fact on a claim — the most recent A1c below 7.0% — and §36.9 uses it as the book's example of an outcome measure's numerator.
The reasoning was sound. Lower A1c, fewer complications, less cost, better lives. A practice whose diabetic patients were mostly under 7.0% was doing something right, and one whose patients were mostly above it was probably not.
The issue: the evidence turned
Then the large trials reported, and the picture became more complicated in a way that measure sets are structurally bad at absorbing.
The ACCORD trial — Action to Control Cardiovascular Risk in Diabetes — randomized patients with type 2 diabetes at high cardiovascular risk to intensive glycemic control targeting an A1c below 6.0% versus standard control. The intensive arm was stopped early, after a mean of about three and a half years, because of increased mortality in that arm. The result was published in the New England Journal of Medicine in 2008 and it was not what anybody expected. Companion trials of intensive control in similar populations found no mortality benefit, with some microvascular benefit.
The mechanism most often discussed is hypoglycemia. Driving glucose down requires more medication, and more medication — particularly insulin and sulfonylureas — produces more episodes of low blood sugar. In an older adult, hypoglycemia is not a nuisance. It is falls, fractures, confusion, emergency department visits, hospitalizations, and occasionally death. The harm from the treatment scales with the aggressiveness of the target, and it lands hardest on exactly the patients whose remaining life expectancy makes the microvascular benefit least relevant.
Professional guidance moved accordingly, and moved toward a shape that a single-threshold measure cannot represent: individualized targets. Tight control for a newly diagnosed fifty-year-old with decades of complication risk ahead of her. Substantially less stringent targets for an eighty-two-year-old with limited life expectancy, multiple comorbidities, cognitive impairment, or a history of hypoglycemia. The American Geriatrics Society's contribution to the Choosing Wisely campaign advised against using medications to achieve tight glycemic targets in most older adults, and recommended moderate goals instead.
So the correct target became a function of the patient. And a measure is a function of a population.
What happened to the measures
Measure stewards responded, and it took years. Over time the tightest threshold was retired from the major diabetes measure sets, and the surviving glycemic indicators reorganized around a less stringent control level and, more importantly, around poor control above a high threshold — a formulation that flags the patient whose diabetes is genuinely out of control without rewarding anyone for pushing a frail patient's number down. Verify the current specification with the steward; this family of measures has been revised more than once and will be again.
The CPT Category II codes moved too. The A1c reporting family has itself been revised, and the ranges the codes describe have changed. Verify the current descriptors every January — this is precisely the kind of quiet revision that leaves a practice reporting a retired code to a measure that no longer exists.
What happened inside an organization
The following is a composite, constructed from documented patterns in the pay-for-performance literature. It is not a real organization.
A multi-site primary care group enters a commercial value-based contract with a quality gate. Among the measures is diabetes glycemic control, specified at the tight threshold — because the contract was drafted from a measure set the payer had used for years and nobody at either table re-read the specification.
Year one: the rate is poor. The group's diabetic patients skew older than the payer's book average, and a substantial number are over seventy-five.
Year two: three things happen, in this order, and only the first was a decision anybody made deliberately.
One — real work. Care managers call patients. Medication adherence improves. Some patients who had been drifting get genuinely better care and their numbers come down. This is the measure working, and it is the largest single component of the improvement.
Two — the exceptions get used. The measure's exclusion mechanism — in claims-based reporting, the performance measure exclusion modifiers 1P, 2P, 3P, and 8P — exists so that a clinician can record a legitimate reason the standard was not met: a medical contraindication, a patient's decision, a system failure. It is a good mechanism and it is why the measure is fair. It is also the only lever in the measure that a practice can move without changing anybody's health, and the volume of exceptions rises noticeably. Nobody instructs anyone to do this. The exceptions are individually defensible. Whether they would have been recorded in a year with no contract is a question nobody asks.
(The phenomenon has a name and a research literature. The United Kingdom's Quality and Outcomes Framework, a large national pay-for-performance program, built an explicit "exception reporting" mechanism for the same good reason, and the resulting variation between practices has been studied extensively. It is one of the few places where this behavior has been measured rather than speculated about.)
Three — and this is the one that matters — some patients get more medication than they should. Not by anyone's decision. By the accumulated weight of a number on a dashboard, reviewed monthly, attached to a payment, applied to a physician at 4:40 on a Friday looking at an eighty-year-old whose A1c is 7.4%. The clinically correct action for that patient is very often nothing. The action the measure rewards is an increase.
The rate improves. The contract's quality gate is met. The distribution is paid.
And the harm, if it occurred, is invisible to every number in the arrangement, because a hypoglycemic fall arrives as an emergency department claim under a different diagnosis, in a different measure, sometimes in a different year, attached to nothing.
The limit this case teaches
A measure is a proxy, and a proxy pursued as a goal stops being a proxy. Chapter 33's Case Study 2 made this argument about severity coding and a case mix index. Here it is about a laboratory value, and the version is worse for one specific reason: in the coding case, the patient was unaffected. DRG creep moved money without touching care. This moves care.
Three properties made it possible, and all three are general:
The measure was specified once and the evidence moved. No mechanism inside a contract updates its own measure specification. Somebody has to re-read it, and the year in which a rate is improving is the year nobody does.
The population was averaged. The measure's denominator treated a fifty-year-old and an eighty-two-year-old as the same patient with the same right answer. Risk adjustment's entire purpose is to stop that from happening, and it was applied to the payment side of this contract and not to the quality side. That asymmetry is common and it is worth looking for in any arrangement you are asked to work under.
And the exclusion mechanism was both the fairness safeguard and the pressure valve. There is no version of this measure without exceptions — a measure that cannot accommodate the patient for whom the standard is wrong is a worse measure. But the same mechanism is the only one an organization can move without changing health, which means an exception rate is a number a responsible organization monitors on itself, in both directions.
What it shows
The two halves of this chapter are one system, and this case is where they touch. Risk adjustment exists because populations differ and payment must account for it. Quality measurement compares organizations whose populations differ — and where the measure does not account for it, the organization with the older, sicker, more complex panel is punished for having it. That is precisely the incentive risk adjustment was invented to remove, reintroduced through the quality door.
A number improving is a question, not a conclusion — instance five in this book. Chapter 21's clean edit report produced by not billing. Chapter 23's payment that rose for a bad reason. Chapter 27's denial rate that improved because rejected claims never denied. Chapter 28's contractual adjustments that rose and were assigned a plausible untested cause. And now a quality rate that improved partly through better care, partly through exceptions, and partly through treatment that should not have been given. All three components were real; the aggregate number reported one thing.
And the coder is closer to this than it looks. Somebody assigns 3044F. Somebody appends an exclusion modifier. Somebody notices that the practice's exception rate doubled in the year the contract started, or does not. The measure runs on codes, and codes are assigned by people who are in a position to see the pattern before anybody else in the building.
The lesson
Ask what the measure is a proxy for, and then ask whether the proxy still holds for this patient. That is a clinical question the coder does not answer — but it is a question the coder is often the first person in the organization able to raise, because the coder sees the distribution.
And carry the practical version, which is short: when a rate improves, decompose it before you celebrate it. How much came from care, how much from exceptions, how much from a change in who is in the denominator? Chapter 33 §33.6's coding manager gave the model answer about a case mix index, and it transfers here without a word changed: before we take credit, let me split the change into components, because if we call it all our work this year, then the year the gains plateau, this same number will say we failed.
Discussion questions
-
The measure was specified at a threshold the clinical evidence had moved away from. Whose job was it to notice? Name every role in a practice and a plan that could plausibly have caught it, and say which one you would actually assign it to and why.
-
Exception reporting is simultaneously the fairness safeguard that makes a measure defensible and the only lever an organization can move without changing health. Design a monitoring approach that preserves the first property while surfacing abuse of the second. What would you look at, how often, and what would you consider a signal rather than noise?
-
This contract risk-adjusted its payment benchmark and did not risk-adjust its quality measure. Explain what that asymmetry does to a practice with an unusually old panel. Then decide: is this a contracting failure, a measure-design failure, or both?
-
Chapter 33's Case Study 2 and this one are both "the measure became a target." Name the property that makes this one worse, and say whether that property changes what an organization should do differently.
-
A physician tells you that a patient's A1c of 7.4% at age eighty is the right number for that patient and she is not going to increase the medication. The measure will count it as a failure. What, if anything, should the coder do — and what is squarely outside the coder's role?
-
(Chapter 36 §36.9) 3044F reports an outcome; 0001F reports a process. Which of the two is more vulnerable to the failure described in this case study, and why? Does that mean process measures are safer, or only that they fail differently?