83 min read

> "The engine read the note in less than a second and got seven codes right. It took me eleven minutes

Prerequisites

  • 4
  • 15
  • 33
  • 36
  • 37

Learning Objectives

  • Define clinical documentation integrity, state the two things it is not, and explain why a program that only queries in one direction is not a documentation program.
  • Deliver an operational answer to copy-forward and build the case for documented negatives from an actual procedure note.
  • Run a concurrent review: the worklist, the working DRG, the reconciliation, and what concurrency buys that a retrospective review cannot.
  • Write a compliant query — its required elements, its option set, and its retention — and identify the five markers that make a query leading.
  • Name the two CDI metrics that corrupt a program, and explain how a program can corrupt by selection without sending a single non-compliant query.
  • Describe what a computer-assisted coding engine actually does, and place it on Chapter 37's assertion register with an owner and an evidence test.
  • Trace a clinical sentence through natural language processing and identify the four attributes that decide whether a concept is codable.
  • State the four properties that make a domain autonomously codable, and apply the test to your own work without predicting a date.
  • Compute precision and recall on an engine evaluation, explain the cost asymmetry between a missed code and a fabricated one, and name the three drivers of model drift.

Chapter 38: Clinical Documentation Integrity, Computer-Assisted Coding, and AI: The Technology Changing the Work

"The engine read the note in less than a second and got seven codes right. It took me eleven minutes to find the eighth, and the eighth was the one that wasn't there." — constructed

Overview

Thirty-seven chapters have taught you to read a record and produce a claim. This chapter is about the two forces currently reshaping that work from opposite ends.

At one end is clinical documentation integrity — a mature profession, thirty years old, whose entire subject is the gap between what happened to a patient and what the record says happened to them. Chapter 33 §33.10 priced one instance of that gap at \$1,867.44: the difference between a record that says "acute respiratory failure with hypoxia" and one that says "hypoxic," on a patient whose clinical care was identical either way. Every chapter since has been circling the mechanism that closes such a gap without inventing anything, and every one of them has deferred it here.

At the other end is the software. Computer-assisted coding has been in production for two decades. It is genuinely good at some things, reliably bad at others, and it is now being joined by tools whose capabilities move faster than any regulatory apparatus can respond to. Some of the work you are training for is more exposed than the rest of it, and you deserve a straight answer about which parts before you spend a year on a credential.

The practitioner's question that opens the chapter is not "will a machine take my job?" It is narrower and more useful: what is the human in the loop actually for? This chapter answers it with a list rather than a reassurance — four things, each demonstrated on a record you have read since Chapter 4 — and then says plainly what the answer implies about a career.

One discipline runs through both halves. A query and an engine are the same kind of object: each one produces an assertion about a patient that nobody in the room decided in that individual case. Chapter 37 §37.10 built a control for exactly that — the assertion register — and this chapter's job is to apply it rather than to rediscover it.

In this chapter, you will learn to:

  • Say what clinical documentation integrity is, and what it is not
  • Fix copy-forward operationally, and install the documented negatives that decide codes
  • Run a concurrent review and reconcile a working DRG to a final one
  • Write a compliant query, and recognize the leading one at a glance
  • Name the two CDI metrics that corrupt a program, and what to publish instead
  • Describe what a coding engine does, mechanically, and where it belongs on the register
  • Trace a sentence through natural language processing and find where it breaks
  • Apply the autonomy test to a domain without predicting a date
  • Compute precision and recall, and price the asymmetry between the two errors

38.1 What CDI is and what it is not

Clinical documentation integrity (CDI) is the discipline of making the medical record describe the patient accurately, completely, and in language the classification systems can read — reviewed against the clinical evidence in the record itself, and corrected only by the clinician who owns the statement.

Read that definition twice, because three clauses in it are load-bearing and each one excludes something the field is regularly accused of.

"Describe the patient accurately" — not favorably. A CDI program's product is a record that matches the patient. Sometimes that record supports a higher-weighted classification and sometimes it supports a lower one, and a program that only ever moves in one direction is telling you something about itself rather than about its patients.

"In language the classification systems can read" — this is the real problem CDI exists to solve, and it is not a physician's failure. A physician writes for the next clinician. ICD-10-CM reads for a payment system. Those are different audiences with different vocabularies, and the translation loss between them is enormous. "Hypoxic" is perfectly good clinical communication. It is not a diagnosis, and Chapter 33 §33.10 showed what that costs on one admission.

"Corrected only by the clinician who owns the statement" — Chapter 4 §4.7's line has not moved and does not move in this chapter. A coder, a CDI specialist, and an engine may all notice that a sentence is missing. None of them may write it.

What CDI is not

It is not a revenue program. Programs are frequently sold as one, measured as one, and staffed as one, and §38.4 is about what happens when they are. The distinction is not decorative: a program that exists to raise the case mix index has adopted an objective its own method cannot serve honestly, because the method — asking a physician an open question — cannot be pointed at an outcome without becoming the thing Chapter 4 §4.9 forbids.

It is not a second coding department. CDI reviews the record; coding assigns from it. Where the two functions merge, the most common casualty is the reconciliation in §38.2, because nobody is left to disagree with anybody.

And it is not an audit. Chapter 37 taught the reading of finished work against a written standard. CDI reads unfinished work while it can still be completed by the person who made it. The two look similar and their timing makes them entirely different instruments.

What to do about copy-forward

Chapter 4 §4.6 defined cloned documentation and named the three ways it becomes a problem. Chapter 15 §15.12 called copied-forward text "the single most damaging thing in a modern medical record" and sent the remedy here. Chapter 36 §36.7 added the risk-adjustment version: six identical annual assessments of the same condition are one year of documentation and five years of a copy operation. Here is the operational answer, and it starts with a concession.

You will not win a ban, and you should not want one. Copy-forward exists because a physician with a fifteen-minute visit and a problem list of nine items needs the previous note. Carrying a medication list, an allergy list, or a stable surgical history forward is reasonable and safe. Attacking the feature loses the argument and the relationship; attacking the fields wins both.

Five moves, in the order they usually succeed:

1. Decide, in writing, which sections may carry and which may not. The rule that survives committee is short: the sections that record what happened today do not copy forward. In practice that means the assessment, the plan, the examination findings for the problem being treated, any time statement, and any decision-to-proceed statement. Everything else is negotiable.

2. Make the copied text visibly copied. Most modern systems can display provenance — what was carried, from where, and when. Turning that display on for the reviewer's view costs nothing and changes behavior, because the physician sees what the auditor will see.

3. Make the fields that decide codes structurally uncopyable. This is the highest-yield move and it is a build, not a memo. Chapter 16 §16.4 promised this section a specific instance: the difference between a practice that reports a greater-than-thirty-minute discharge day service correctly and one that does not is a template field that requires the clinician to state the total discharge time on that date — blank on arrival, required before signature, never inherited from yesterday. Do the same for the three or four statements in your specialty on which a code actually turns.

4. Measure it on your own data. A note-similarity comparison across a provider's own panel requires no chart review and no vendor: pull a sample of notes per provider, compare each note against that provider's previous note for the same patient and against notes for different patients, and look at the distribution. Chapter 37 §37.10 called this the outside-in view, and this is one of its cheapest applications — a payer's special investigations unit can run the same comparison on the claims and notes it already has.

5. Then have the conversation, and lead the way Chapter 4's 📞 leads. "The note is underselling what you're actually doing" is both true and the version a busy clinician will act on. Leading with audit exposure makes the conversation adversarial; leading with the record understating the work makes it collaborative, and the fix is identical.

The limit, stated in the same breath. None of this reaches a note that is genuinely typed fresh each visit and genuinely says nothing. Copy-forward is a specific failure with a specific remedy; it is not a general theory of bad documentation, and a program that treats every thin note as a cloning problem will spend its credibility on the wrong argument.

The case for documented negatives

Chapter 17 §17.7 made this argument in three sentences and promised it here at length. Here it is at length, because it is the single cheapest documentation improvement in this book.

A documented negative is a sentence whose only purpose is to record that something did not happen. It converts silence, which proves nothing, into evidence, which decides a code.

The March 14 procedure note printed in full at Chapter 4 §4.10 contains two of them:

"No aspirate obtained." · "No imaging guidance used."

Each closes a coding question that would otherwise require a query.

Without the second clause, a coder faces a procedure note that is silent about ultrasound guidance. Silence is not a documented negative — it is an absence, and an absence has to be asked about. With it, the choice between 20610 (major joint injection, no imaging guidance) and 20611 (the same injection with ultrasound guidance and a permanent recording and report) is made by the record rather than by an inference, and the more highly valued code is excluded on the physician's own statement rather than on the coder's assumption. Chapter 37 §37.11 read the same clause from the auditor's chair and reached the identical conclusion from the other direction: an auditor notices absent negatives. This record has them, which is why the line survives an outside reading.

The general principle, which transfers to every code set in this book: the classification is full of pairs distinguished by a single feature — with guidance and without, with contrast and without, complete and limited, bilateral and unilateral, with aspiration and injection only. The note that names the absent feature chooses the code. The note that omits it starts a query.

Three practical consequences.

Install them as fields, not as prose habits. Ask a specialty's coders which three questions they query about most often, and put those three in the procedure template as required selections. The physician's added burden is three clicks; the coder's saved burden is a query, a delay, and a chart that sits in a work queue.

A documented negative that is not true is worse than silence, and this is where the temptation lives. A template that pre-populates "no imaging guidance used" as a default value is not documentation. It is a configuration making an assertion nobody chose — the mechanism this book has now watched produce a dozen separate failures across every part, in claim fields, in posting rules, in a payer's software, in a note over a physician's signature, and in a script handed to a person. Chapter 37 §37.10's register exists for exactly this, and a templated negative belongs on it the day it is built, with an owner and an evidence test run against a real record.

And negatives are not padding. The objection you will hear is that requiring them lengthens notes that are already too long. The answer is that these are the shortest sentences in the record and they are the only ones that close questions. Three clauses in the March 14 note took no longer to dictate than the sentence before them and closed three separate coding questions permanently.

⚠️ Where Claims Die

The CDI program that only queries in one direction.

It rarely looks like misconduct from inside. The worklist is built from records that grouped without a major complication or comorbidity, because that is where the financial materiality is. Every individual query is compliant in form. Nobody is asked to write anything untrue. And after two years the program has never once sent a query that removed a diagnosis, because it never reviewed a record where removal was the likely answer.

Why it matters more than it looks: the direction of a program's findings is fully visible in claims data, and it is exactly what a reviewer computes first. A severity-capture pattern that runs one way is a question the organization will eventually be asked to answer, and "every query was compliant" is a defense of the queries rather than of the program.

What the disciplined organization does: builds at least part of the review population on something other than payment materiality — a clinical condition list, a random sample, a service line, a physician cohort — and reports adds and removals as separate lines on the same dashboard. §38.4 gives the metrics; the design decision is here, because it is made when the worklist is built and never again.


38.2 The concurrent review

A concurrent review is a review of a record while the encounter is still open — during the admission, before the bill drops, while the physician who wrote the note is still on the unit and the patient is still in the bed.

Everything CDI does differently from auditing follows from that one word.

Why timing is the whole method

Chapter 4 §4.5 governs amendments, addenda, and late entries: they are dated when they are made, attributed, and flagged as amendments in the legal health record. That machinery is correct and necessary, and it is also a tax. A clarification obtained on hospital day 2 is an ordinary progress note. The identical clarification obtained three weeks after discharge is an addendum, and every reader downstream can see that it arrived late and can ask why. Chapter 30 §30.11 made the sharpest version of this point: an amendment made after a denial, asserting exactly the disputed element, is the single worst document a file can contain.

Concurrency also buys three things a retrospective review cannot:

The physician's memory is current. Chapter 21 §21.8 stated the limit plainly and sent it here: a query cannot ask a surgeon to remember something eighteen months later. Asking on day 2 is asking about this morning.

The record is still being written. A clarification on day 2 shapes the rest of the stay's documentation. The same clarification after discharge fixes one sentence in a closed file.

And the care itself can improve. This is the part CDI professionals mention least and it is real. A physician asked, in a neutral form, whether a documented finding represents a specific diagnosis is a physician thinking about that patient again.

The mechanics

CONCURRENT vs. RETROSPECTIVE - the same clarification, two calendars
                                                     [constructed]

  ADMIT       DAY 2         DAY 4        DISCHARGE   BILL     PAYMENT
    |-----------|-------------|--------------|---------|---------|
                ^                            ^
                |                            |
      CONCURRENT REVIEW                RETROSPECTIVE REVIEW
      - progress note                  - addendum, dated when made
      - physician on the unit          - physician has moved on
      - working DRG still moving       - coded DRG already assigned
      - no bill exists yet             - bill may already be out;
      - no amendment trail               a correction is a corrected
                                         claim (Ch. 29 §29.6)

  SAME QUESTION. SAME ANSWER. DIFFERENT DOCUMENT, DIFFERENT WEIGHT.

The worklist. Nobody reviews every admission. A program picks, and §38.4 is about what happens when it picks badly. Common and defensible criteria: admissions with a working DRG that has a severity split and no documented complication or comorbidity; specific conditions with known documentation ambiguity (sepsis, respiratory failure, heart failure specificity, malnutrition, renal staging); records where a clinical indicator appears without a corresponding diagnosis; and a random sample, which is the leg that keeps the other three honest.

A working DRG is the Medicare Severity Diagnosis-Related Group a record would group to as it stands right now, mid-stay, on incomplete documentation. It is a review tool, never a bill.

The review itself is a read of the record against the clinical evidence in it: the admission history and physical, the progress notes, the medication administration record, the vitals, the imaging and laboratory results, the nursing documentation, the operative reports. The reviewer is looking for one of three things — a diagnosis that is clinically evidenced and never stated, a stated diagnosis whose clinical evidence is thin or contradicted, and a stated diagnosis that could be specified further.

The reconciliation. At discharge, the coder assigns the final codes and the record groups to a final DRG. Where the coded DRG and the CDI reviewer's working DRG disagree, somebody has to adjudicate — and that process is the best single indicator of whether a CDI program is a coding program with a different name. A healthy reconciliation has a written procedure, a named adjudicator who is neither reviewer, and a log. An unhealthy one has whoever argues hardest.

Outside the hospital

Concurrent review is usually described as an inpatient activity because that is where the money is concentrated, but the professional-services version is real and is growing.

Pre-visit review is Chapter 36 §36.6's worklist: reading the chart before the appointment and surfacing the chronic conditions the visit will probably address, so the assessment can address them in the note rather than in somebody's memory. Post-visit, pre-bill review is a hold on defined encounter types for a documentation read before the claim goes out — which is Chapter 37 §37.10's rung 2, and it costs throughput exactly as that ladder predicts.

Both have a hard boundary and it is worth stating before anyone builds one. Surfacing an unaddressed chronic condition before a visit so the clinician can decide whether to address it is care. Surfacing it after so somebody can decide whether to report it is a different activity with a different name, and Chapter 36 §36.8 is where that line is drawn.

📞 On the Phone

The concurrent conversation, at the nurses' station, with a physician who has four minutes.

What does not work: "Doctor, I need you to document acute respiratory failure on this chart so we can get the MCC." Two fatal errors in one sentence — it supplies the diagnosis and it states the payment consequence — and it is a leading query delivered verbally, which is still a leading query (§38.3).

What works: "I'm looking at the admission on room 412. The record documents an oxygen saturation of 84% on room air, four liters by nasal cannula started in the ED and titrated up overnight, and the assessment reads 'COPD exacerbation, hypoxic on arrival.' Based on your clinical judgment, can the respiratory status be further specified in today's progress note — or is the assessment as written what you intend? Either answer is fine; I just need it in the note rather than in the vitals."

The failure modes. The physician says "you know what I meant." The answer is that you do, and the record does not, and nobody downstream will be allowed to know. The physician says "just put whatever you need." The answer is no — the clinical statement is theirs and cannot be anyone else's. And the physician says "is this a billing thing?" The honest answer is that documentation specificity affects payment, quality measurement, and the next clinician's reading of the chart, that you are not asking for any particular answer, and that "no" is a complete response you will record and move on from.

Every verbal query gets written down — who asked, whom, when, what was asked, what was answered — and the answer only counts when it lands in the record itself. §38.3 is strict about this and the reason is Chapter 4 §4.9's: a documentation change with no visible provenance appears to have arisen from nowhere.


38.3 The compliant query, and the leading query that ends careers

This is the most-promised section in Part VIII. Five chapters sent you here by name, each one leaving a real question unanswered, and this section owes all five:

Who deferred What they left here
Ch. 21 §21.8 A query cannot ask a surgeon to remember something eighteen months later. So what can a retrospective query ask?
Ch. 22 §22.6 The right fact in the wrong section — conservative therapy documented in the history rather than the assessment — is "a documentation improvement opportunity of the most ordinary kind." What is the conversation that fixes it, and is it even a query?
Ch. 33 §33.10 The query is the occasion created by a record that is not wrong, just unwritten. What form does it take, and how does it run in both directions?
Ch. 36 §36.11 The risk-adjustment query on the diabetes and kidney-disease linkage that should have been sent — and the hard limit that a query cannot manufacture an encounter.
Ch. 15 §15.7 You may ask what happened. You may not decide what happened. Where exactly is that line drawn in the wording?

Chapter 4 §4.9 defined the query and gave its four rules. This section gives its anatomy, its option set, its retention, and its failure — the last of which is a compliance act with personal consequences, and the reason the section title says what it says.

The anatomy of a compliant query

A compliant query is a written request to a clinician to clarify documentation that is conflicting, ambiguous, incomplete, illegible, or clinically inconsistent — presenting the relevant clinical evidence from the record, asking an open question, offering clinically reasonable options including a way to decline, naming no code and no payment consequence, and retained with its response.

Seven elements. A query missing any of the last four is defective; a query missing the fourth is leading.

THE SEVEN ELEMENTS OF A COMPLIANT QUERY

  1  IDENTIFICATION   patient, encounter, date of service, who is asking,
                      and how to respond

  2  THE EVIDENCE     the specific clinical indicators FROM THIS RECORD that
                      raise the question - quoted or cited by location, and
                      ONLY those (minimum necessary, Ch. 5)

  3  THE QUESTION     open, answerable in more than one direction, about a
                      clinical fact - never about a code, a level, a DRG,
                      or a dollar

  4  THE OPTIONS      clinically reasonable alternatives, INCLUDING
                      "clinically undetermined" and "other, please specify"
                      - and including at least one that adds nothing

  5  THE DISCLAIMER   a statement that no particular response is intended
                      or preferred

  6  THE LANDING      instruction that the response be documented in the
                      MEDICAL RECORD, not merely on this form

  7  THE RECORD       the query and its response retained per a written
                      policy, retrievable on demand

Element 2 is where most weak queries fail. "The patient appears septic" is not evidence. A temperature, a white count, a lactate, a pressor order, and the sentence in the assessment that raises the ambiguity — those are evidence, they are in the record, and they let the physician answer from clinical facts rather than from a coder's impression.

Element 4 is where compliant queries become leading ones, and the mechanism is subtler than people expect. Chapter 4 §4.9 gave the rule: a query with two options, one of which is obviously what the coder wants, is a leading query with a fig leaf. Here is the version that catches more programs — an option list can be leading even when it has five options, if every clinically reasonable answer on it adds or specifies a condition. The test is not how many options there are. The test is whether the record's most likely answer, including "the assessment as written is what I intend," is on the list with equal weight.

Yes/no queries are permitted in narrow circumstances and are misused constantly. The defensible use is confirming a relationship or a detail already documented somewhere in the record — the diagnosis is stated in a consultant's note and not in the attending's, and the question is whether the attending concurs. The indefensible use is introducing a diagnosis the record has never named. The difference is whether a "yes" tells you something new about the record or something new about the patient.

And "clinically undetermined" is a real answer, not a courtesy. Chapter 33 §33.7 sends its hardest case here: when a condition's timing is genuinely unclear, a query about present-on-admission status must accept "cannot be determined" — which is what the W indicator exists to report — with exactly the weight it gives every other option. An option set that treats honest uncertainty as a non-answer is asking the clinician to guess, and a guess entered as a fact is worse for the record than the ambiguity was.

Verbal queries are legitimate and are the daily reality of concurrent review. They carry one extra obligation: they get documented — who, whom, when, what was asked, what was answered — because otherwise the clarification enters the record with no visible provenance, and Chapter 4 §4.9 named that as the problem.

The four questions the deferring chapters asked

What can a retrospective query ask? Only what the record can still answer. Chapter 21 §21.8's surgeon, eighteen months on, is being asked to remember whether a debridement was in a different anatomic region — and a memory is not documentation. The line: a retrospective query may ask a clinician to clarify what the record shows; it may not ask them to supply what the record never captured. If the clinical evidence supporting the answer is not in the record, there is nothing to clarify, and the honest outcome is the less specific code and a prospective template fix.

The right fact in the wrong section. Chapter 22 §22.6 found six weeks of conservative therapy documented in the history of present illness, where a reviewer scoring medical necessity may never reach it, and it is not in the assessment where a structured review would look. The coder may not relocate it — that is the physician's reasoning and moving it is authoring. A query is available: it presents the history excerpt and asks whether, in the clinician's assessment, conservative therapy was tried and did not adequately control the symptoms. That is an open clinical question with a real "no" available. But be honest about which fix matters: the retrospective query fixes one note and the template fixes every note after it, and this particular gap is a template problem wearing a query's clothes.

Both directions. Chapter 33 §33.10 said a defensible program can show an auditor its queries without flinching, and the reason is that some of them removed things. A query that asks whether a documented diagnosis is supported by the clinical evidence — and produces the answer "no, remove it" — is the same instrument pointed the other way, and it is the single most persuasive artifact a CDI program can produce. If your program has never generated one, that is a finding about the program.

And the hard limit. Chapter 36 §36.11 established it and it governs everything in this section: a query cannot manufacture an encounter. If the kidney disease was not addressed on March 14, no answer to any question makes it reportable for March 14. A query completes a record of work that was done. It cannot create the work, and it cannot change what a past encounter was about.

📋 Read the Chart

text FIGURE 38.1 - "The same gap, twice" [Account 10-4471; both query drafts are constructed] THE DOCUMENT Two drafts of a written query to the family physician about the March 14 encounter, generated from the gap Chapter 22 §22.6 identified. Only one may be sent. THE CONTEXT The assessment never states that conservative therapy failed. Six weeks of intermittent ibuprofen with partial relief IS documented - in the history of present illness. Many payer policies for joint injection expect documented conservative therapy. WHAT IT SHOWS Draft A names the policy, names the consequence, and asks for a specific sentence. Draft B quotes the record's own history, asks a clinical question about the assessment, and offers a real "no." WHAT IT DOESN'T Neither draft tells you the clinical answer. The coder does not have one and must not appear to. Neither changes the code on this claim: M25.561 and 20610-RT are correct either way. THE DECISION Send B, or send nothing and fix the template. Do NOT send A. If A has gone out, tell your compliance officer today rather than hoping. THE LESSON A query asks what the record does not say. It never proposes what it should have said - and the strongest version of this fix is not a query at all.

```text --- DRAFT A - LEADING. DO NOT SEND. --------------------------------- Dr. -, Northfield's policy for joint injection requires documented failure of conservative therapy. Please add to the assessment that conservative measures failed so the injection is not denied. Thanks!


FAULTS: names the payer policy as the reason to document states the financial consequence supplies the sentence wanted offers no alternative and no way to answer "no" asks for a conclusion rather than a clarification

--- DRAFT B - COMPLIANT --------------------------------------------- Dr. -, Clarification is requested on the assessment for the 03/14 encounter.

The history of present illness documents: - six weeks of right knee pain, worsening - intermittent ibuprofen 400 mg as needed, with partial relief

The assessment for the knee reads: "new complaint this visit ... consistent with a degenerative process; no definitive diagnosis established today."

Based on your clinical judgment, does your assessment of this problem include the response to the measures tried to date? ( ) Conservative measures tried and did not adequately control symptoms - please document in the assessment ( ) Conservative measures tried with adequate control; injection elected for another documented reason ( ) Conservative therapy not yet adequately tried ( ) Clinically undetermined ( ) Other: ____

Please document any response in the medical record. No particular response is intended or preferred.


```

The retention question

Where does a query live afterward? This is a real policy question with two defensible answers, and the field has not converged on one.

Model one: the query becomes part of the permanent legal health record. Chapter 4 §4.8 defined the legal health record and the designated record set; under this model the query and its response are inside it, produced on any request for records, and visible to every later reader.

Model two: the query is retained as a business record, separate from the legal health record, in the CDI or coding system rather than the chart.

Your organization must pick one, write it down, and apply it consistently. Professional guidance from the American Health Information Management Association (AHIMA) and the Association of Clinical Documentation Integrity Specialists (ACDIS) addresses the question, state law and payer contracts may speak to it, and your health information management director and compliance officer are the authorities for your organization — this book is not. What this book can tell you is that the choice matters less than the three rules that hold under either model:

1. The response must land in the medical record itself. A code may never be assigned from an answer that exists only on a query form. If the physician's clarification is not documented in a progress note, an addendum, or a late entry under Chapter 4 §4.5's rules, there is nothing to code from — the query answered your question, not the record's.

2. Every query and its response must be retrievable. Chapter 33 §33.6 published the follow-up question an auditor asks when a case mix index moves: "show me the queries behind the shift." That is §38.4's subject and it is unanswerable by a program that cannot produce the query text, the response, the resulting code change, and the direction, keyed to the account.

3. Retention runs at least as long as the audit look-back. A retention schedule shorter than the window in which somebody can ask about a claim is a schedule that destroys your own defense. Match it to your organization's records retention policy and to any payer contract requirement, and never let it be the shortest number in the building.

The leading query, and why the section title says "ends careers"

A leading query is one that suggests the answer, names a code or classification, states or implies a financial consequence, offers no clinically reasonable alternative, or signals which response is preferred.

Five markers. Any one of them is disqualifying. They are easy to state and, under time pressure at 4:40 on a Friday, easy to write anyway.

THE FIVE MARKERS OF A LEADING QUERY

  1  IT SUPPLIES THE ANSWER
     "Can you document acute blood loss anemia?"

  2  IT NAMES THE CLASSIFICATION OR THE CODE
     "...so this will group to the correct DRG"
     "...to support a level 4"

  3  IT STATES OR IMPLIES THE MONEY
     "...so the account is not underpaid"
     "...the payer requires this for coverage"

  4  IT OFFERS NO REAL ALTERNATIVE
     one option; or five options that all add something;
     or no "clinically undetermined"

  5  IT SIGNALS THE PREFERRED RESPONSE
     "Please confirm that..." / "I'm sure you meant..."
     / repeated queries on the same encounter after a "no"

Marker 5's last clause deserves its own sentence. Querying the same encounter repeatedly after receiving an answer you did not want is leading regardless of how each individual query is worded. The record of three queries and one changed answer tells its own story, and the record is what survives.

⚖️ Compliance Check

What a leading query actually exposes, and to whom.

Chapter 4 §4.9 named the exposure in one line and it is worth expanding: a leading query can be characterized as the provider organization manufacturing documentation to support a code. That is a materially different allegation from having coded something incorrectly. A coding error is an error. A record created to justify a payment is a false record, and Chapter 5 §5.1's certification logic reaches it: every claim to a federal health program is a certification, and the False Claims Act's theory of a false record does not require anyone to have lied about the service. Chapter 37 §37.6 supplies the multiplier — a consistent practice extrapolates cleanly and satisfies the precondition for projecting a sample across a universe.

And now the part that is personal, which is why this section is titled the way it is. The claim is the organization's. The query has an individual's name on it. A pattern of leading queries is attributable to the person who wrote them, and the consequences do not require a prosecutor: termination, loss of a credential under the credentialing organization's own code of ethics (Chapter 39), and a professional reputation in a field where the hiring pool is small and everyone eventually works for someone who worked somewhere else. In the enforcement matters that have been brought in this area, the query practice itself has been part of the record.

Requirements change, state law and payer policy vary, and professional query guidance is revised. Verify current practice with your compliance officer, your credentialing organization's published ethics standards, and the primary sources — not with a textbook, including this one.

The protective habit, and it is cheap: query templates reviewed and approved before use, so an individual is never improvising wording under production pressure; a second reader on any query that would move a severity tier; and the retention in this section, so the provenance of every documentation change is visible. You want the file to show what you asked.


38.4 CDI metrics and the ones that corrupt the program

Every CDI program is measured. Chapter 24's Case Study 2 supplied the sentence this book has carried since — an unmeasured function is indefensible — and it applies here as much as anywhere. The problem is not measurement. The problem is that two of the standard measures become corrupting the moment they become targets.

The standard dashboard

THE CDI MEASURES, AND WHAT EACH ONE IS FOR       [structure, not benchmarks]

  REVIEW COVERAGE   share of eligible records reviewed
                    -> capacity and worklist design

  QUERY RATE        queries sent / records reviewed
                    -> ***CORRUPTING AS A TARGET***

  RESPONSE RATE     queries answered / queries sent
                    -> physician engagement; a real problem when low

  TURNAROUND        hours or days to a response
                    -> whether concurrency is actually working

  AGREEMENT RATE    responses that added or specified / responses
                    -> ***CORRUPTING AS A TARGET***

  DIRECTION         adds and specifications vs. removals and
                    de-specifications, reported SEPARATELY
                    -> the measure almost nobody publishes

  DRG CHANGE RATE   records whose group changed after a query
                    -> volume of effect; says nothing about validity

  CMI              average relative weight of discharges
                    -> Ch. 33 §33.6: diagnostic, never a target

Benchmarks for these vary enormously by setting, staffing model, payer mix, and how each numerator is defined, and any specific figure you are quoted is a comparison of definitions at least as much as of performance — which is Chapter 29 §29.7's denominator problem in a new department. Ask what counts as a review, what counts as a query, and what counts as agreement before you compare anything to anyone.

Why query rate corrupts

A query rate target is met by sending queries. That sentence is the whole argument, and it is not cynical — it is what happens to any rate that becomes a quota, in any department, staffed by people who want to keep their jobs.

Concretely: a reviewer who has read a clean record and found nothing to ask has done the work correctly and produced nothing the metric can see. A reviewer who has read the same record and sent a marginal query has produced a countable unit. The marginal query is by construction the weakest one in the queue — the one where the clinical evidence is thinnest and the answer least clear — and it is also the one most likely to be worded suggestively, because a weak clinical case makes a suggestive wording feel necessary.

The measure is still useful diagnostically. A query rate near zero says the program is not looking or is afraid to ask; a query rate that jumps in a month says something changed. Watch it. Do not target it.

Why agreement rate corrupts, and this is the important one

Agreement rate is usually computed as the share of responses that added or specified a condition. Stated that way it sounds like a quality measure. It is not. It is a measure of how often physicians gave the program the answer that produces a code.

Now watch what maximizing it does, because this is the mechanism by which a program corrupts without ever sending a non-compliant query.

To raise agreement rate you do not need to write a leading query. You need to send only the queries you already believe will be answered your way — and drop the ones where the clinical evidence is genuinely ambiguous, where "clinically undetermined" is the likely answer, and where the honest result is a removal. Every query you send is compliant in form. Every query you didn't send is invisible. The selection is the leading act, and no review of the query text can detect it.

This is exactly what Chapter 6 §6.3 warned about, and it is worth quoting the mechanism because that chapter named this section by name. A grouper "what-if" run tells you what a documentation change would be worth; it says nothing about whether the change would be true. A program that lets "will this change the DRG?" drive "is this true?" has become the leading-query problem at institutional scale.

Where the line actually falls is narrower than people expect, and it is worth being precise:

  • Prioritizing which records get reviewed by financial materiality is defensible. Attention is finite and it has to go somewhere.
  • Deciding whether to send a query you have already concluded is clinically warranted, based on what the answer would be worth, is not.
  • Deciding whether to send it based on the likely direction of the answer is the corruption itself.

The first is triage. The third is selection. They can be performed by the same person, on the same day, at the same desk, and only one of them is defensible — which is why the design decision in §38.1's ⚠️ belongs in the worklist build and not in anybody's judgment at 4:40 on a Friday.

🧮 Run the Numbers

Two CDI programs, one month, and the dashboard that hides the finding. [constructed teaching figures — no benchmark is asserted]

```text PROGRAM A PROGRAM B records reviewed 400 400 queries sent 152 88 query rate 38.0% 22.0% responses received 139 79 response rate 91.4% 89.8%

OF THE RESPONSES: added or specified 128 52 clinically undetermined 11 18 removed or de-specified 0 9 ---- ---- 139 79

agreement rate (added/responses)  92.1%            65.8%

DRG changes after query 61 34 upward 61 31 downward 0 3 ```

The arithmetic, checked. Program A: 152 ÷ 400 = 38.0% · 139 ÷ 152 = 91.4% · 128 + 11 + 0 = 139 ✓ · 128 ÷ 139 = 92.1%. Program B: 88 ÷ 400 = 22.0% · 79 ÷ 88 = 89.8% · 52 + 18 + 9 = 79 ✓ · 52 ÷ 79 = 65.8% · 31 + 3 = 34 ✓.

On every published metric, Program A wins. Higher query rate, higher response rate, an agreement rate twenty-six points better, and nearly twice the DRG movement. A vendor would put Program A in a case study.

Program A cannot answer one question: where are the queries that removed something? Zero removals, zero downward changes, and eleven "clinically undetermined" responses out of a hundred thirty-nine — on four hundred real records containing real patients, some of whom were described in the chart as sicker than they were. Either Program A's physicians document with a precision no hospital has ever achieved, or Program A only reviews records where the answer can go one way.

Program B's 65.8% is the number a director gets asked about, and it is the better program. Nine removals is nine records that now describe the patient. Eighteen "clinically undetermined" responses is eighteen physicians who felt free to say so, which tells you the option set is real. Three downward DRG changes is the artifact Chapter 33 §33.10 said a defensible program can show an auditor without flinching.

What to publish instead of an agreement rate: adds and removals as separate lines, "clinically undetermined" as its own line, and the query text available on request. The composite percentage should not exist, because a single number that blends the two directions can only be improved by suppressing one of them.

What the auditor asks

Chapter 33 §33.6 staged the conversation: a case mix index rises, an executive credits the CDI program, and the coding manager insists on splitting the change into service-mix and capture-rate components before anyone takes credit. The auditor's follow-up is the sentence that section promised this one: "show me the queries behind the shift."

A program that can answer it produces, per account: the query text, the date, who sent it, the clinical evidence cited, the response, whether the response landed in the record, the code change, and the direction. A program that cannot produce that has a case mix index it cannot explain, which is the same position as Chapter 28's Case Study 1 — a number with a story attached stops being a question — with an auditor holding the other end.

And one compensation rule, which Chapter 36 §36.8 already stated for chart review and which applies identically here: nobody whose pay, bonus, or performance rating depends on the direction of findings should be reviewing records or writing queries. Not because the people are dishonest. Because the arrangement is indefensible from outside, and defensibility from outside is the entire subject of the chapter before this one.


38.5 Computer-assisted coding: what the engine does

Computer-assisted coding (CAC) is software that reads clinical documentation and proposes codes for a human to review, presenting each suggestion with the text that produced it.

The word doing the work is proposes. A tool that assigns and bills without a human reviewing that chart is a different thing with a different governance problem, and it is §38.7.

CAC is not new and it is not exotic. It has been in production in large facilities for roughly two decades, most heavily in hospital outpatient and emergency department coding, where volume is high and documentation is comparatively uniform. If you take a coding job in a hospital, there is a reasonable chance the software is already on your desktop.

What it actually does, mechanically

THE CAC PIPELINE - what happens between a signed note and a suggested code

  1  INGEST        the record arrives: notes, reports, results.
                   Scanned documents go through optical character
                   recognition first - and OCR errors propagate
                   silently downstream

  2  SEGMENT       identify the document type and its SECTIONS
                   (HPI, ROS, exam, assessment, plan, family history,
                   procedure note, discharge instructions)

  3  EXTRACT       find clinical concepts in the text and normalize
                   them against a clinical terminology

  4  QUALIFY       for each concept, decide: negated? whose? when?
                   how certain?                       <-- §38.6

  5  MAP           terminology concept -> ICD-10-CM / CPT / HCPCS
                   candidate, with the coding rules the vendor has
                   implemented

  6  SCORE         attach a confidence value

  7  PRESENT       show the coder the code, the confidence, and the
                   SOURCE TEXT that produced it, highlighted in place

Step 7 is the most valuable thing on the list and the least discussed. A good engine does not merely produce a code; it produces a code with the sentence that generated it, in place, in the document. That is an audit trail no manual process has ever had, and it is why an assisted workflow can be faster to defend than an unassisted one even when it is not more accurate.

Two technologies, both in production, often in the same product. The older approach is rule-based: dictionaries, patterns, and hand-built logic, written by people, inspectable, predictable, and brittle at the edges. The newer approach is statistical: models trained on coded records that learn associations between text and codes, better at variation and worse at explaining themselves. Most commercial products are hybrids. Neither one understands the note. Both produce correlations between text and codes, and the difference between a correlation and a reading is the entire subject of §38.6.

An engine is an assertion-maker, and it goes on the register

Chapter 37 §37.10 built the assertion register: an inventory of every place an assertion can be made about a claim or an encounter by something other than a person deciding about that claim, each entry naming what it asserts, where it lands, whose voice it speaks in, a named owner, and an evidence test run against a real claim the configuration touched. A coding engine is the largest single entry most organizations will ever make on it, and the register is why this chapter does not need to re-derive the governance from scratch.

THE REGISTER ENTRY FOR A CODING ENGINE            [constructed template]

  WHAT IT IS      the CAC engine, version and rule-set date, plus every
                  local configuration: suppressed codes, auto-accept
                  thresholds, specialty rule packs, crosswalks

  WHAT IT         "this documentation supports this code"
  ASSERTS         and, at any auto-accept threshold:
                  "no person needs to look at this one"

  WHERE IT        the CLAIM (correctable) - unless the tool writes
  LANDS           into the NOTE, in which case the assertion is in the
                  clinical record over a physician's signature, which
                  is Ch. 29's Case Study 2 in a new module

  OWNER           a person, by name. Not "HIM." Not "the vendor."

  EVIDENCE TEST   pull a claim the engine touched - as transmitted -
                  and ask whether the code it asserted is supported by
                  the record it read. Date the check.

Two features of that entry are worth pausing on.

The "where it lands" row decides how expensive the entry is to be wrong about. An engine suggesting a code on a claim is corrected by correcting the claim. A tool that writes into the clinical note — ambient documentation products increasingly do, and so do templates that insert attestations — puts the assertion in the legal health record over a clinician's signature, where Chapter 4 §4.5 governs and a note can be amended but never un-asserted. Chapter 29's Case Study 2 is the worked example: a template that auto-inserted a modifier 25 attestation into two years of signed notes, true most of the time, which is why nobody looked.

And the evidence test is run against output, not against configuration. Chapter 27's Case Study 2 is the proof: an eleven-year-old local mapping produced a malformed element on every claim, the clearinghouse silently normalized it, and the practice's copy and the payer's copy were accurate records of different files. Nothing on a setup screen would have shown it. Read a claim.

What CAC is good at, and what it is reliably bad at

Reliably good Reliably bad
High-volume, template-driven documentation with a narrow code space Anything requiring an argument rather than a match
Finding codes a tired human skips at 4:30 Sequencing and principal-diagnosis selection
Consistency across coders and across days Linkage the record implies but does not state
Suggesting the specific code once the concept is named The absence of a sentence nobody wrote
Producing the source-text audit trail Rules that live outside the note: payer policy, edits, the contract, the calendar
Flagging clinical-indicator patterns for CDI review Judging whether a condition was addressed at this encounter

That right-hand column is not a list of things the technology has not gotten to yet. Four of the six are structurally outside what a text-reading system can do, and §38.6 explains why for each.

⚠️ Where Claims Die

The document the engine never saw.

Step 1 of the pipeline is the one nobody audits, and it produces the most confident wrong answers in the whole workflow. An engine codes from the documents that reached it, and in a real organization the record is assembled from more sources than anyone's diagram shows: notes from the electronic health record, results from a laboratory interface, an outside consultant's letter that arrived by fax, an operative report dictated into a different system, a pathology report that posts two days after the encounter, a scanned intake form.

What goes wrong is not that a document is missing. It is that nothing reports it as missing. The engine reads what it has, finds clean text, extracts clean concepts, and returns high-confidence suggestions from a partial record. A confidence score is a statement about the text the engine read. It is not a statement about the record. Chapter 4 §4.1 established what a coder is actually reading, and the answer has never been "one document."

Where the money goes: a pathology report that arrives after the claim was built, so the pre-pathology finding is billed and the definitive diagnosis never reaches the claim — Account 22-9107's exact shape. A scanned outside report whose optical character recognition mangled the laterality. An operative report in a second system that the coder knew to open and the interface did not.

What the disciplined organization does: treats document completeness as a pre-coding check rather than a coding judgment. Define, per encounter type, which documents must be present before a chart is eligible for assisted coding; hold the ones that are not; and — the part that is usually missing — reconcile the engine's document inventory against the record's, periodically, on a sample. That is the same instrument as Chapter 37 §37.10's evidence test, pointed one step upstream: do not ask what the engine concluded, ask what it was looking at.


38.6 Natural language processing against a clinical note

Natural language processing (NLP) is the set of computational techniques for extracting structured meaning from unstructured human text. In this setting it means: turning a paragraph a physician typed at 6:42 p.m. into a list of clinical concepts a code can attach to.

The reason a whole section is spent on it is that the failures are systematic and they are predictable, which means a coder who knows where they occur can review an engine's output far faster than one who checks everything equally.

Finding a word is easy; deciding what the word does is the job

Every clinical concept an engine extracts has to be qualified along four axes before it can become a code. These four are where nearly all NLP coding errors live.

THE FOUR QUALIFIERS - and the sentences that break each one

  NEGATION      Is the concept asserted or denied?
                "no evidence of pneumonia" - "denies chest pain"
                "ruled out for pulmonary embolism"
                FAILURE: the concept is present in the text and
                         absent in the patient

  EXPERIENCER   Whose condition is it?
                "mother with colon cancer at 62"
                "father died of an MI"
                FAILURE: a family history becomes a diagnosis

  TEMPORALITY   When?
                "history of MI in 2019, no recurrence"
                "status post cholecystectomy 2014"
                FAILURE: a resolved condition becomes an active one -
                         and in ICD-10-CM these are DIFFERENT CODES in
                         different chapters, not variations of one

  CERTAINTY     How sure is the clinician?
                "probable pneumonia" - "rule out sepsis"
                "cannot exclude" - "concerning for"
                FAILURE: depends entirely on the SETTING. Inpatient,
                         "probable" at discharge is coded as established
                         (Ch. 9 §9.5). Outpatient, it is not. Same
                         four words, two opposite answers, and the
                         engine has to know which claim it is on

The certainty row is the one to remember for an exam and for a job, because it is the only one of the four where the correct handling flips depending on the setting rather than on the sentence.

Section detection, and the trap it sets

An engine identifies which part of a document a sentence sits in — history of present illness, review of systems, examination, assessment, plan, family history, procedure note, discharge instructions — and that judgment changes everything downstream. "Colon cancer" in a family history is not a diagnosis. "Chest pain" in a review of systems is not the reason for the visit. "Diabetes" in a problem list is not a condition addressed at this encounter.

Here is the trap, and it inverts what most people assume. An engine is more section-sensitive than a human, not less. A human reads the whole note. When a coder needs to know whether conservative therapy was tried, they read the history, the assessment, and the plan, and they find the sentence wherever it is. An engine that has been correctly configured to read medical necessity support from the assessment will do exactly that, correctly, and will not find a fact sitting in the history of present illness.

Which means Chapter 22 §22.6's finding has a second life here. The right fact in the wrong section is a documentation improvement opportunity when a human reads the chart. It is a systematic miss when software does. The remedy is the same — one sentence in the assessment — and the argument for it is now twice as strong.

The thing NLP structurally cannot do

An engine cannot flag the absence of a sentence nobody wrote. Natural language processing operates on text. Where there is no text, there is nothing to process. Chapter 37 §37.11's finding on Account 10-4471 — the note never states that the decision to inject was made during this visit — is not a hard problem for an engine. It is not a problem an engine is in a position to have.

Now the honest qualification, because a flat "software cannot find gaps" is wrong. Engines can and do flag a clinical indicator pattern without a corresponding diagnosis: oxygen orders and low saturations with no respiratory-failure statement; three low albumin results and a dietitian consult with no malnutrition diagnosis; a documented estimated blood loss and a transfusion with an assessment that reads only "blood loss." That is real, it works, and it is how modern CDI worklists are actually built. It is not the engine finding a missing sentence. It is the engine finding a present pattern the organization has told it to associate with a usually accompanying sentence.

And that capability carries §38.4's corruption in software. A tool configured to surface "clinical indicators for a major complication with no documented diagnosis" is a query-generation engine pointed in exactly one direction. Everything in §38.4 applies to it and applies harder, because the selection is now automatic, invisible, and scaled. If your indicator rules only fire in the direction that adds severity, your worklist is the finding. Build at least some rules that fire the other way — a documented diagnosis whose usual clinical indicators are absent from the record — and put the whole rule set on the assertion register with an owner who can say why each rule exists.

🔍 Check Your Understanding

  1. A discharge summary states: "Ruled out for pulmonary embolism. Mother with breast cancer. History of MI in 2019. Probable aspiration pneumonia, treated." Which qualifier is at issue for each of the four concepts, and which one of the four codes as an established diagnosis on this inpatient claim?
  2. An engine reports 96% accuracy on a facility's outpatient charts. Name two things you must know before that number means anything, and say which chapter of this book taught you to ask.
  3. Your CDI worklist is generated by twelve clinical-indicator rules. What is the one question to ask about the set as a whole, and what would a good answer look like?

Answers: (1) Certainty (the embolism — ruled out, not coded); experiencer (the mother's cancer — not this patient's); temporality (the 2019 infarction — an old infarction is a different code in a different place from an acute one); certainty again for the pneumonia — and on an inpatient claim "probable" at discharge codes as established under Chapter 9 §9.5, so that is the one. (2) The denominator and the definition of "accurate" — accurate against what standard, on which chart population, counting which codes. Chapter 29 §29.7's denominator problem, and Chapter 37 §37.3's rule that accuracy is reported with its denominator attached. (3) "Do any of these rules fire in the direction that removes or de-specifies a diagnosis?" A good answer names specific rules and shows the queries they produced.


38.7 Autonomous coding and where it currently works

Autonomous coding is the assignment of codes and the release of a claim without a human reviewing that particular chart. Not a suggestion accepted quickly. No human in that chart's loop at all.

This is where the field's marketing is loudest, so this section is deliberately flat. Autonomous coding is real, it is in production, it works well in a narrow set of domains, and it does not generalize from them. Both halves of that sentence are true and most discussions carry only one.

Where it genuinely works today

Radiology. A structured report, produced from a template, describing a study the ordering information already characterizes, with a small candidate code space — a chest radiograph, two views, is 71046 and the number of views is in the report. Pathology. Specimen-driven, with the code determined by the specimen type and the level of examination, and 88305 doing an enormous share of the work. Screening and routine laboratory. A defined service, a defined frequency, a diagnosis code that comes from the order. And parts of ophthalmology imaging, dermatology specimen work, and outpatient behavioral-health time-based reporting, for the same reasons.

The four properties, as a test you can apply

IS THIS DOMAIN AUTONOMOUSLY CODABLE? - four questions

  1  STRUCTURE       Is the source document produced from a template,
                     with the code-determining facts in named fields
                     rather than in prose?

  2  CODE SPACE      Is the realistic candidate set small - tens of
                     codes, not thousands?

  3  VARIABILITY     Do two clinicians documenting the same service
                     produce substantially the same document?

  4  FEEDBACK        Is there a fast, unambiguous signal when the
                     answer was wrong - a denial, an edit, a
                     reconciliation - that arrives in days, not
                     in an audit two years later?

  Four yeses: probably. Three: assisted, not autonomous.
  Two or fewer: the vendor is selling you 38.8.

Now apply the test to the work in this book and watch it fail.

Evaluation and management leveling fails property 1 and property 3 outright. The 2021 revision made the office-visit level turn on medical decision making — the number and complexity of problems addressed, the data reviewed and analyzed, and the risk of management — which is a judgment about a physician's reasoning as recorded in prose. Chapter 15 §15.13 leveled Account 10-4471 as a 99214 on two of three elements at moderate, and the argument took a page.

Surgical coding fails property 1. An operative report is a person describing, in prose, what they did, and Chapter 17 §17.3 taught reading it precisely because the code follows the documented objective, not the procedure's title.

Modifier 25 fails every property. It is not a pattern. It is an argument about separateness, built on this record from four elements, three of which have nothing to do with the knee.

And property 4 is the quiet one. A denial is a fast, unambiguous signal for a rejected code. There is no equivalent signal for a code that should have been reported and was not — Chapter 28 §28.8's underpayment does not deny, does not reject, appears on no exception report, and arrives as a payment. An autonomous system with only denial feedback learns to avoid the errors that get caught, which is not the same as learning to be right.

⚖️ Compliance Check

Nobody's software signs anything.

A claim submitted to a federal health program is a certification, and Chapter 5 §5.1 established whose it is: the provider's, and the organization's. Autonomy is a change in workflow, not a transfer of responsibility. There is no version of this in which "the engine coded it" is a defense, and no vendor contract can move a certification.

So the governance questions are the organization's, and they should be answered in writing before any scope goes live:

  1. Scope. Which document types, which code families, which payers — named, not "outpatient."
  2. Thresholds. At what confidence does a chart bypass review, who set that number, and on what evidence? This is a compliance decision wearing a configuration screen's clothes.
  3. Sampling. What share of autonomously coded claims is audited against the record, by whom, how often, and what is the standard? Chapter 37 §37.2's universe, sample, and standard, pointed at a population that no longer produces a natural human reader.
  4. Stop conditions. What measured result turns the scope off, who has the authority to do it, and has anyone ever done it in a drill?
  5. The register entry. Owner by name, evidence test dated, and the two events from Chapter 37 §37.10 that force an off-cycle review — a payer policy or edit change, and a system migration, upgrade, or vendor change — both of which apply to a coding engine with unusual force.
  6. The update calendar, which is item 6 because it is the one nobody assigns: ICD-10-CM changes every October 1, CPT every January 1, HCPCS Level II quarterly, and the NCCI edits quarterly. A model or rule set that is current on September 30 is stale on October 2, and the claims it releases in between are real.

Requirements, payer policies, and professional guidance in this area are moving. Verify current expectations with your compliance officer and the primary sources before relying on anything here.

This book will not tell you when

No date appears in this chapter. Not for autonomous coding generally, not for any specialty, not for the profession. Predictions in this field have a poor record in both directions — the confident ones from twenty years ago were wrong about the timeline and wrong about which parts would go first — and a textbook that names a year is teaching a reader to plan against a guess.

What you can watch instead are the four properties arriving in your own domain: documentation becoming structured, the code space narrowing, variability falling as templates standardize, and a faster feedback signal appearing. When three of the four move in your specialty within a couple of years, you are looking at the leading edge of a change, and you will see it before any announcement does.


38.8 Coder-in-the-loop, and what the human is actually for

Coder-in-the-loop describes a workflow in which software proposes and a credentialed human reviews, accepts, rejects, or corrects each suggestion before the claim goes out.

It is the dominant arrangement, it is the one most readers of this book will work in, and it has a specific failure mode that nobody advertises.

The failure mode: the loop that is not a loop

Automation bias is the well-documented human tendency to accept a system's recommendation more readily than the same conclusion reached alone — and to stop looking once a plausible answer is on the screen. In a coding workflow it produces rubber-stamping: a reviewer under a productivity standard, presented with a suggestion that is right most of the time, approves it.

The cruel part is that the engine's accuracy causes the problem. A tool that was wrong half the time would be checked. A tool that is right nine times in ten trains its reviewer, over months, that checking is wasted effort — and the tenth chart is the one that matters.

Three design decisions determine whether the loop is real. All three are usually made by someone who has never coded.

1. What the reviewer sees first. If the engine's answer appears before the note, it is an anchor, and every subsequent judgment is a judgment about whether to overturn it. Workflows that show the documentation first and the suggestions second produce different results. Ask which one you have.

2. Whether disagreement is cheap. If accepting a suggestion is one click and rejecting it requires a reason code, a free-text justification, and a supervisor's review, the measured agreement rate is a measurement of the interface, not of the engine. This is §38.4's lesson in a new medium, and the same rule applies: a number that can only move one way is not a measurement.

3. Whether the productivity standard was reset. Chapter 6 §6.9 introduced the productivity and quality standards a coder is held to. If a coder is now held to charts-per-hour computed on assisted throughput, with the quality standard unchanged and no allowance for the review work itself, the organization has budgeted for a rubber stamp and will get one. The loop is a staffing decision before it is a technology decision.

🔢 Code It

The engine's output on an inpatient record, and the coder's disposition. [Account 22-8891 — the constructed inpatient file]

```text ENGINE SUGGESTIONS - 4-day admission, discharged [constructed]

CODE CONF SOURCE TEXT THE ENGINE MATCHED CODER'S DISPOSITION ------- ---- ------------------------------ ------------------- J44.1 0.97 "acute exacerbation of COPD" ACCEPT - principal J96.01 0.93 "acute hypoxemic respiratory failure" ACCEPT - MCC, POA Y I50.32 0.88 "chronic diastolic CHF, stable" ACCEPT E11.22 0.81 "type 2 DM with CKD 3a" ACCEPT - linkage IS documented here N18.31 0.90 "CKD stage 3a" ACCEPT Z79.4 0.86 "insulin" ACCEPT - long-term use L89.153 0.79 "stage 3 sacral pressure ulcer" ACCEPT - but POA? J18.9 0.61 "cannot exclude superimposed pneumonia" REJECT ```

What the engine got right, and it is most of it. Six of the eight suggestions are correct and supported, including the two that decide the money: J44.1 as the condition occasioning the admission and J96.01 as the major complication or comorbidity that puts the stay in MS-DRG 190 rather than 191 — the \$1,867.44 Chapter 33 §33.10 published. Note also that the diabetes and kidney-disease linkage is documented on this record, which is why E11.22 is right here and E11.9 is right on Account 10-4471. Same coder, same conventions, two records, two answers.

The plausible wrong answer, named and rejected: accepting J18.9 at 0.61. "Cannot exclude superimposed pneumonia" is a certainty failure waiting to happen, and the reviewer who thinks "well, it's an inpatient claim and Chapter 9 §9.5 says probable codes as established" has applied the right rule to the wrong phrase. Probable, suspected, and likely at discharge are statements the clinician is making. Cannot exclude is a statement the clinician is declining to make. Reject it, and if the clinical picture warrants, that is an occasion for a §38.3 query — not for a coder's inference.

And the finding the engine cannot make: L89.153's POA indicator. The engine found the ulcer. Chapter 33 §33.7's question — was it present on admission? — is answered by the chronology of the record: the admission skin assessment documented intact skin and the ulcer appears on day 2, so POA = N, and under the hospital-acquired condition provision it cannot serve as a major complication or comorbidity. On this record that changes nothing, because J96.01 already holds the tier. The engine suggested a code. A person assigned a fact about time.

What the human is actually for

Not a pep talk. A list, each item demonstrated on a record you have already read.

One: the judgment that is an argument rather than a pattern. Modifier 25 on Account 10-4471 is supported by four elements, three of which have nothing to do with the knee (Chapter 14 §14.4). An engine can observe that an evaluation and management service and a minor procedure occurred on the same date. That observation is a co-occurrence, and appending a modifier to it is a configuration making an assertion nobody chose — which is Chapter 29's Case Study 2 exactly, in a new module.

Two: the rules that are not in the note. The payer's medical policy, the National Correct Coding Initiative edit and its modifier indicator, the contract, the timely filing calendar, the frequency limitation, the patient's benefit design. None of these is in the document the engine read, and several of them change quarterly. Chapter 22's Harborview file is the pure case: a claim that was never going to be covered, coded perfectly.

Three: the absence. The sentence nobody wrote. §38.6 explained why this is structural rather than temporary, and Chapter 37 §37.11's category-4 finding on Account 10-4471 is the book's own example — a finding that changed nothing on the claim and determined how Chapter 30's appeal had to be constructed.

Four: the accountability, and this one is not sentimental. A code is a legal attestation. The claim carries a certification, and a certification requires a certifier. When a reviewer asks in two years why this code was assigned, the answer has to be a path — main term, subterm, verify the Tabular, check the conventions, check the guidelines, check the edits — held by somebody who can explain it to a stranger. "The engine's confidence was 0.93" is not a path. It is not even a sentence about the patient.

And a fifth, which belongs to this chapter specifically: the query. An engine can raise a question. Only a person may ask a clinician a question about a patient, in a form governed by §38.3, with their name on it.

Making the loop measurable

Audit the accepts, not just the rejects. A quality program that samples the charts the coder changed is measuring the coder's disagreements. The exposure is in the agreements. Draw the sample from accepted suggestions, score against the record, and report the result by confidence band — because if accuracy in the highest confidence band is where you set your auto-accept threshold, you have just tested the threshold.

Score against the record, never against the engine. An "agreement with the engine" metric measures conformity. Chapter 37 §37.2's standard stack applies unchanged: the record, the code set, the guidelines, the edits, the coverage policy, as they stood on the date of service.

And put the whole arrangement on the register, with the owner, the evidence test, and a dated last-checked. An engine is the most consequential assertion-maker in a modern coding department and it is the one most likely to be treated as furniture.


38.9 Measuring an engine: precision, recall, and the cost asymmetry

You will be asked to evaluate a tool, or to sit in a room while somebody else does. Two numbers do most of the work, and almost nobody in the room will distinguish them.

Precision — of the codes the engine suggested, what share were supported by the record. A precision failure is a false positive: a code proposed that the documentation does not support.

Recall — of the codes the record supported, what share the engine found. A recall failure is a false negative: a code missed.

They trade against each other. Tune an engine to suggest only what it is sure of and precision rises while recall falls. Tune it to suggest everything plausible and recall rises while precision falls. A single blended number hides the trade, which is why a vendor quoting one number is quoting the one that flatters the product. Ask for both, and ask what population they were computed on.

🧮 Run the Numbers

An engine evaluation, and what the loop does to it. [constructed teaching figures — no product or benchmark is asserted]

A facility evaluates an engine on 500 inpatient records already coded by credentialed coders and re-reviewed to establish a reference standard.

```text Codes the record SUPPORTED (the reference standard) ....... 1,840 Codes the engine SUGGESTED ................................ 1,905

suggested AND supported   (true positives) ............... 1,712
suggested, NOT supported  (false positives) ..............   193
                                                           -----
                                                           1,905  OK

supported, NOT suggested  (false negatives) ..............   128
1,712 + 128 = 1,840  OK

PRECISION = 1,712 / 1,905 = 89.9% RECALL = 1,712 / 1,840 = 93.0% ```

Now put a coder in the loop and re-measure, because the engine's numbers are not the organization's numbers.

```text Of the 193 unsupported suggestions: rejected by the reviewer ................................. 171 ACCEPTED by the reviewer ................................ 22 ----- 193 OK

Of the 128 codes the engine missed: found independently by the reviewer ..................... 96 still missing at claim submission ....................... 32 ----- 128 OK

CODES ON THE FINAL CLAIMS supported ....... 1,712 + 96 = 1,808 unsupported ..... 22 ------------------ 1,830 codes

Post-review accuracy = 1,808 / 1,830 = 98.8% Post-review recall = 1,808 / 1,840 = 98.3% ```

What the arithmetic says. The loop turned 89.9% precision into 98.8% accuracy and 93.0% recall into 98.3%. That is a real and substantial improvement and it is the honest case for the assisted workflow.

And it left 22 unsupported codes on real claims. Twenty-two is the smallest number in the table and it is the only one with a False Claims Act attached to it. Thirty-two missed codes cost the facility money and distort its data. Twenty-two unsupported codes are twenty-two statements to a payer that the record does not support — and if they share a cause, Chapter 37 §37.6's extrapolation arithmetic is waiting, and consistency makes it worse rather than better.

The transferable habit: never accept a precision or recall figure without its denominator and its reference standard. Accurate against what, on which charts, counting which codes? Chapter 29 §29.7's denominator problem and Chapter 37 §37.3's rule that accuracy is reported with its denominator attached both apply verbatim.

The cost asymmetry, stated honestly

The two errors are not equally bad, and the way that fact is usually stated is wrong in a way that gets people into trouble. Both errors are real errors. Neither is safe.

A false negative — a missed code — costs money and truth. Understated severity, an understated case mix index, a distorted risk score (Chapter 36), quality data computed from claims that describe patients as healthier than they were, and an underpayment that never announces itself, never denies, and raises the net collection rate (Chapter 28 §28.8). Chapter 5 §5.8 settled the principle for this book: downcoding is not the conservative option. It is inaccurate, it underpays, and it is not a defense.

A false positive that survives review and reaches a claim is a different kind of object. It is a statement to a payer, and to a federal health program it is a certification (Chapter 5 §5.1). The False Claims Act supplies a well-developed theory for billing for what the record does not support. There is no corresponding theory for failing to bill for what it did.

So the two errors cost different currencies, and this is the sentence to carry: a missed code costs revenue and data accuracy; a fabricated code costs revenue, data accuracy, and legal exposure that compounds with volume and with consistency. Under the False Claims Act they are not symmetric at all.

Four design consequences follow, and they are the whole reason this section exists:

1. The review effort belongs disproportionately on the accept side. §38.8's sampling rule is not an arbitrary preference; it follows from the asymmetry.

2. The confidence threshold is a compliance decision. Raising it raises precision and lowers recall — fewer unsupported suggestions, more missed codes. That trade is made by a person with a name, on stated evidence, recorded on the register entry. It is not a default that arrived with the software.

3. And the asymmetry sets a trap: an organization that only checks accepts will let recall rot silently. Recall failures produce no denial, no edit, no exception report, and no complaint. The only instrument that finds them is a periodic human-coded reference sample of the kind the 🧮 above describes — which is expensive, which is why almost nobody runs one. Say so out loud when the budget is being set, because the alternative is discovering the number when somebody else computes it.

4. Neither error is fixed by coding low. An engine tuned toward silence is not a safe engine. It is an engine whose errors have been moved into the direction nobody measures.

Model drift

Model drift is the degradation of a model's performance over time as the world it was fitted to changes, with no change in the model itself.

Three drivers, in this domain specifically.

One: the code sets move on a calendar. ICD-10-CM changes every October 1, CPT every January 1, HCPCS Level II quarterly, and the NCCI edits quarterly. A model trained on last year's coded records does not know this year's codes and has learned last year's deleted ones. Code from the current year's book or encoder — never from a textbook, including this one.

Two: the documentation changes. A new template, a new electronic health record, a new service line, a new group of physicians with different habits. The input distribution shifts and nothing in the model notices. This is the most common driver and the easiest to anticipate, because somebody in the building always knows a template changed.

Three — and this is the one nobody monitors — the definition of "supported" changes without any text changing at all. A payer policy is revised, a coverage article is updated, a guideline is clarified, an edit pair is added. The notes read the same. The engine reads them the same. The answer is now different.

That third driver is Chapter 28's Case Study 1 with new clothes on: a configuration that was correct when written and decayed, with no event inside the organization to trigger a review. Which is precisely why Chapter 37 §37.10's two forcing events — a payer policy or edit change, and a system migration, upgrade, or vendor change — are on the register entry, and why the update calendar is item 6 of §38.7's governance list rather than an afterthought.

🎓 Exam Watch

What actually gets tested here, and what does not.

Precision and recall are not on the CPC, COC, or CCS exams. No credential in this field currently tests you on an F-measure. Learn them because you will be in the room when a tool is evaluated, not because a proctor will ask.

What IS tested, and heavily: the query rules. Expect a stem giving you a query's text and asking whether it is compliant, or giving you a documentation scenario and asking whether a query is appropriate. The distinctions candidates miss, in order:

  • Leading versus non-leading turns on whether the answer is supplied, the code or payment is named, or the options are stacked. A query can be polite, well-formatted, clinically sensible, and still leading.
  • "Query" versus "do not query." Do not query to obtain a higher-paying code; do not query when the answer is already in the record; do not query when the honest resolution is the less specific code (Chapter 4 §4.9).
  • Who owns the clinical statement. The correct answer is always the clinician. A stem in which a coder adds, moves, or infers a diagnosis is always wrong, no matter how obvious the clinical picture.
  • Concurrent versus retrospective review — know the definitions and that concurrent means during the encounter, not soon after it.
  • And CDI vocabulary generally: working DRG, query response, agreement rate, and the both-directions requirement.

The trap in the stem is usually a query that is compliant in every visible respect and contains one disqualifying clause — a payer's name, a dollar figure, a DRG, or an option set with no way to decline. Read the option list last and read it twice.


38.10 What to do with your career about all of this

This is the section you would skip to, so it will not waste your time with reassurance.

The honest map of exposure

The exposure is not uniform, and pretending it is uniform is what hurts people. The same four properties from §38.7 sort the work.

Most exposed — high volume, structured source documents, small code space, low variability, fast feedback, and no judgment about a rule outside the note:

  • Single-specialty production coding from structured reports (radiology, pathology, screening)
  • Charge entry from a completed encounter form
  • First-pass claim scrubbing against published edits
  • Routine denial triage on standard reason codes with standard fixes
  • Any task whose complete description fits in a sentence beginning "for every chart, if X then Y"

Least exposed — the work that is an argument, a conversation, or a decision somebody has to sign:

  • Evaluation and management leveling, surgical coding from operative reports, and modifier judgment
  • Anything requiring a payer's policy read against a specific record (Chapter 22)
  • Denials and appeals that turn on constructing an argument (Chapter 30)
  • Clinical documentation integrity and the query — a conversation with a clinician, with your name on it
  • Auditing, education, and building or evaluating the tools themselves
  • Any work where the deliverable is a defensible explanation rather than a code

One caution about that map, because it is the mistake people make with it: the exposed list is not a list of bad jobs. It is a list of jobs where the number of people needed changes. And the least-exposed list is not safe forever either — it is where the judgment currently is, and judgment work has moved before.

Six moves

1. Get the credential, and understand what it is for. Chapter 39 is the whole subject. The short version: the credential is what makes you legible for the review-and-defend roles rather than the production roles, and those are the ones on the second list.

2. Learn the adjacent thing you already touch. Every reader of this book is one step from a judgment role. A coder is one step from CDI, auditing, or education. A biller is one step from denial argument, payer policy, and contract reading. A front-desk person is one step from financial clearance and estimates. Pick the adjacent thing, learn it properly, and be the person who does both.

3. Learn to evaluate a tool, because almost nobody in the building can. After this chapter you can ask six questions most rooms have never heard: What are the precision and recall, on what population, against what reference standard? What is the confidence threshold and who set it? What share of accepted suggestions is audited against the record? What happens on October 1 and January 1? Who owns this on the assertion register and when did they last test it against a transmitted claim? Does anything in the worklist or the rule set fire in the direction that removes a diagnosis? The person who asks those questions is not competing with the tool. They are the reason it can be deployed.

4. Become the person who can explain the path. Chapter 40 calls it the reading habit. Read the Official Guidelines every year — they are free and they are republished annually. Read your payers' policy updates. Read one Federal Register rule a year in the area you work in. The ability to explain, two years later, to a stranger with a subpoena, exactly why a code was assigned is the thing this chapter's technology does not produce and cannot replace.

5. Write things down, and use the channel. Chapter 37 §37.1 counted the findings in this book that came from a person rather than a control, and the person column wins. The corollary from Chapter 19's Case Study 2 is the operational one: the absence of a place to say something is a control failure. If your organization has that channel, use it; if it does not, the observation that it does not is itself worth reporting.

6. Measure your own work. An unmeasured function is indefensible was written about departments and it applies to individuals with no modification. Keep your own accuracy record, your own denial and overturn record, your own list of findings you raised and what happened to them. When a conversation about staffing or outsourcing happens, the person who can produce twelve quarters of measured accuracy is in a different conversation from the person who can only say they work hard.

📞 On the Phone

The meeting where a computer-assisted coding implementation is announced.

What not to say: "So are we being replaced?" It is the honest question, it will be answered with a reassurance nobody can guarantee, and it puts you in the room as a problem rather than as the person who understands the thing being bought.

What to ask, in this order:

"What's the scope — which document types and which code families, exactly?" (Narrow scopes work. Broad ones are §38.8 with a different name.)

"Is anything being coded without a human in the chart, or is everything suggestion-and-review?" (Two completely different governance problems, and the answer decides everything after it.)

"What are the productivity and quality standards going to be after go-live, and were they recomputed for review work?" This is the most important question in the meeting, and it is the one a manager is least likely to have an answer to. It decides whether the loop is real.

"When we reject a suggestion, how many clicks is it?" (Ask it lightly. The answer tells you whether the agreement rate will mean anything.)

"Who's auditing the accepted suggestions, against the record, and how often?"

"And what happens on October 1?"

The failure mode: asking all six at once, in front of a vendor, as a challenge. Ask two in the meeting and the rest in writing afterward. The point is not to win the meeting. It is to be on the short list of people the organization has to consult before it changes anything — which is, in practice, the most durable position in a department.

What does not change

Four things in this chapter are not technology questions and will not be answered by one.

The certification on the claim is a person's and an organization's. No software signs.

The clinical statement belongs to the clinician. A coder may not write it, a CDI specialist may not write it, and an engine that writes it has put an assertion in the legal record over somebody else's signature.

The query rules do not relax because the query was drafted by software. Everything in §38.3 applies to a suggested query exactly as it applies to a typed one, and the name on it is still yours.

And a code remains a legal attestation. That has been this book's third theme since Chapter 5 and nothing in this chapter touches it. The tools change what is easy. They do not change what is true, and they do not change who answers for it.


🗂️ The Encounter — running the March 14 note through a coding engine

Thirty-eight chapters have coded, billed, denied, appealed, collected, and audited this file. Now hand it to the software.

The exercise. Feed the March 14 office note — printed in full at Chapter 4 §4.10 — to a computer-assisted coding engine configured for professional office coding, and read its output against what you know. (Nothing below changes the file. Account 10-4471's codes, claim, and money are exactly as canon has them.)

What this chapter contributes to the file: the engine's output, and the four findings it could not produce.

What the engine gets right, and it is most of it

ENGINE OUTPUT - Account 10-4471, DOS Mar 14, established patient  [constructed]

  SUGGESTED  CONF  SOURCE TEXT MATCHED               COMMENT
  ---------  ----  --------------------------------  --------------------
  99214      0.84  assessment; four problems; two    a LEVEL suggestion,
                   labs ordered; three medications   not a reading
  20610      0.96  "right knee intra-articular       correct: major joint
                   injection ... no imaging
                   guidance used"
  J1030      0.95  "methylprednisolone acetate       correct: 40 mg = 1 unit
                   40 mg"
  36415      0.93  "blood drawn in office by
                   venipuncture"
  M25.561    0.94  "right knee pain"                 correct for this DOS
  I10        0.92  "essential hypertension, at goal"
  E78.5      0.90  "hyperlipidemia"
  E11.9      0.91  "type 2 diabetes mellitus,
                   stable"                           right code. see below.

Seven of the eight are correct, and two of them deserve specific credit.

20610 rather than 20611 is a genuinely good piece of work, and it is good work the physician made possible. The engine's negation handling read "no imaging guidance used" — the documented negative §38.1 built its case on — and excluded the ultrasound-guided code on the record's own statement. This is exactly what an engine should do and exactly the kind of question a documented negative closes for every reader, human or otherwise.

J1030 at one unit is the other. The descriptor's dose is 40 mg, the note says 40 mg, the units match. Chapter 20's dosage arithmetic, performed correctly by pattern match, because that is what it is.

And E11.9 is where the chapter's argument lives. The engine returned the right code — Chapter 9 §9.7 established that E11.9 is correct for March 14 — but it did not do the reasoning that makes it right. It mapped a diabetes concept in an assessment to the unspecified default. It did not ask Section IV's question, which is whether the kidney disease on the reviewed problem list was addressed at this encounter. It happens that the answer is no and the code is therefore correct. The engine reached the right answer without performing the test, and a coder who cannot tell the difference cannot defend the code when somebody asks. That is the difference between an answer and a path, and Chapter 37 taught you why the path is the thing that survives an outside reading.

The four findings the engine cannot produce

1. Modifier 25 — because it is an argument, not a pattern. An engine can observe that an evaluation and management service and a minor procedure appear on the same date. Some are configured to append modifier 25 on that observation. That is a co-occurrence, and appending a modifier to it is a configuration making an assertion nobody chose — the mechanism this book has watched produce a dozen separate failures, and specifically Chapter 29's Case Study 2, where a template asserted a significant, separately identifiable service into two years of signed notes.

The real support is four elements, and Chapter 14 §14.4 froze them: three chronic conditions each separately assessed with a plan; prescription drug management across three medications; two laboratory tests ordered with stated clinical reasons; and a new problem with its own history, examination, and independent management decision. Elements one through three have nothing to do with the knee. That is the argument, it takes a paragraph to make, and no confidence score substitutes for it. Note the practical shape of this on a working day: confirming the four codes takes a coder a few minutes; the modifier 25 judgment takes longer than the four codes combined, and it is the only part a payer will ever question.

2. The diagnosis pointers. The claim points 36415 at B (E11.9), not at A, because the blood was drawn for the hemoglobin A1c — pointing the venipuncture at the knee would assert that a venipuncture treats knee pain (Chapter 25 §25.5). An engine that assigns pointers by proximity, or by first-listed default, points everything at A and produces a claim that is wrong in a way no edit catches and no denial reveals. This is a small, checkable, entirely automatable failure that most implementations get wrong because nobody specified it.

3. The diabetes and kidney-disease question — which is not a question about the text at all. Chapter 36 §36.11 owns it: the reviewed problem list carries chronic kidney disease, stage 3a, the assessment continues metformin without referencing renal function, and the right response was a query — one that asks whether renal function was evaluated or considered in managing the diabetes at this encounter, and that offers "not addressed" as a real answer. No engine can decide whether an encounter addressed a condition, because "addressed" is a property of the clinician's work, not of the words. And §36.11's limit governs the whole thing: a query cannot manufacture an encounter.

4. The sentence nobody wrote. The note never states that the decision to inject was made during this visit. Chapter 14 §14.4 found it, Chapter 37 §37.11 scored it as the file's single category-4 finding — supported, but fragile — and it determined how Chapter 30's appeal had to be built: as a constructed demonstration rather than as a quotation. The engine cannot flag it. There is no text to process. This finding required a person who knew what a modifier 25 defense needs and noticed that the record did not quite say it.

What the exercise settles, and what it does not

What it settles: on this record, an engine would have produced seven of eight codes correctly and saved real time doing it. That is not a small thing and this chapter will not pretend it is. An assisted workflow on this chart is faster, more consistent, and better documented than an unassisted one, and the source-text trail it leaves is better evidence than anything a manual process produces.

What it does not settle: every one of the four findings above is the kind that decides whether a claim survives — the modifier that gets audited, the pointer that misstates a clinical relationship, the query that should have been sent, and the sentence whose absence made an appeal harder to write. Three of the four are invisible to a text-reading system by construction, and the fourth is a configuration nobody specified. That is the honest answer to "why is the coder still there," and it is a list rather than a sentiment.

The open questions. Five of the file's six are closed: Q1 (modifier 25, Chapter 14), Q2 (the diabetes code, Chapter 9, with Chapter 36 answering whether it is sufficient), Q3 (no advance beneficiary notice, Chapter 22), Q5 (the knee, Chapter 22), Q6 (the \$185.00 charge, Chapter 23). Q4 — could the denial have been prevented? — remains open and belongs to Chapter 40, which is now two chapters away. Nothing in this chapter touches it, though everything in §38.1 is circling it.


Summary

Clinical documentation integrity is the discipline of making the record describe the patient accurately, completely, and in language the classification can read — corrected only by the clinician who owns the statement. It is not a revenue program, not a second coding department, and not an audit; the last distinction is timing, and timing is the whole method. A program that only ever moves in one direction is telling you about itself rather than about its patients.

Copy-forward has an operational answer and it is not a ban. Decide in writing which sections may carry and which may not — assessment, plan, examination findings for the treated problem, any time statement, any decision-to-proceed statement do not carry — display provenance, make the fields that decide codes structurally uncopyable (a required, blank-on-arrival discharge-time field being the instance Chapter 16 §16.4 asked for), measure note similarity on your own data, and then have the conversation the way Chapter 4 does: the note is underselling what you're actually doing.

A documented negative converts silence into evidence. "No aspirate obtained" and "no imaging guidance used" each close a coding question that would otherwise start a query, and they are the reason 20610 rather than 20611 survives an outside reading. Install the three or four that matter in your specialty as required fields — but never as default values, because a templated negative is a configuration making an assertion nobody chose and belongs on the assertion register the day it is built.

Concurrent review works because of when it happens. The physician's memory is current, the record is still being written, no amendment trail is created, and no bill exists yet. A working DRG is a review tool and never a bill; the reconciliation between it and the final coded DRG needs a written procedure and an adjudicator who is neither reviewer. Outside the hospital, pre-visit review surfaces chronic conditions before the clinician decides — after is a different activity with a different name.

A compliant query has seven elements: identification, the clinical evidence from this record, an open question, clinically reasonable options including "clinically undetermined" and a way to decline, a no-preference disclaimer, an instruction that the response land in the medical record, and retention. An option list can be leading with five options if every reasonable answer adds something. A retrospective query may ask a clinician to clarify what the record shows and may not ask them to supply what it never captured. Verbal queries are legitimate and get written down. And a query cannot manufacture an encounter.

Retention has two defensible models — inside the legal health record or as a separate business record — and three rules that hold under either: the response must land in the record itself (a code is never assigned from a query form), every query and response must be retrievable because the auditor's question is "show me the queries behind the shift," and retention must run at least as long as the audit look-back.

A leading query has five markers — it supplies the answer, names the code or classification, states or implies the money, offers no real alternative, or signals the preferred response — and repeated querying after a "no" is leading regardless of wording. The exposure is not a coding error; it is manufacturing documentation to support a code, a false-record theory, extrapolated more cleanly the more consistent it is. The claim is the organization's; the query has an individual's name on it.

Query rate and agreement rate are the two metrics that corrupt. A query rate target is met by sending queries, and the marginal query is the weakest one in the queue. Agreement rate corrupts more subtly and matters more: it is maximized by sending only the queries you expect to be answered your way, which is a leading program assembled entirely out of compliant queries. The selection is the leading act and no review of query text can detect it. Publish adds and removals as separate lines, publish "clinically undetermined" as its own line, watch both rates diagnostically, target neither, and never let compensation depend on the direction of a finding.

A computer-assisted coding engine ingests, segments, extracts, qualifies, maps, scores, and presents — and its most underrated output is the source text beside the code. It is an assertion-maker and it is the largest single entry most organizations will ever put on Chapter 37 §37.10's register: what it asserts, where it lands (a claim is correctable; a note over a physician's signature is not), an owner by name, and an evidence test run against a transmitted claim rather than a configuration screen.

Natural language processing turns text into concepts, and every concept must be qualified along four axes: negation, experiencer, temporality, and certainty — the last of which flips with the setting, since "probable" at discharge codes as established on an inpatient claim and does not outpatient. Section detection makes an engine more section-sensitive than a human, which gives Chapter 22 §22.6's right-fact-wrong-section finding a second and stronger life. And an engine cannot flag the absence of a sentence nobody wrote — though it can flag a clinical-indicator pattern with no corresponding diagnosis, which is how modern CDI worklists are built and which carries §38.4's corruption into software if every rule fires in one direction.

Autonomous coding is real in narrow domains and does not generalize. The test is four questions — structured source document, small code space, low variability, fast unambiguous feedback — and radiology, pathology, and screening pass it while E/M leveling, surgical coding, and modifier 25 fail it outright. Property four is the quiet one: denial feedback teaches a system to avoid the errors that get caught, which is not the same as being right. Nobody's software signs anything; scope, thresholds, sampling, stop conditions, the register entry, and the October 1 / January 1 / quarterly update calendar are the organization's to answer in writing. This book names no date, for any of it.

Coder-in-the-loop fails through automation bias, and the engine's accuracy is what causes it. Three design decisions decide whether the loop is real: what the reviewer sees first, whether disagreement is cheap, and whether the productivity standard was reset for review work. Audit the accepts, not the rejects, score against the record rather than against the engine, and report by confidence band.

What the human is for, as a list: the judgment that is an argument rather than a pattern; the rules that are not in the note; the absence; the accountability, because a certification requires a certifier and "confidence 0.93" is not a path; and the query, because only a person may ask a clinician a question about a patient.

Precision is what share of suggestions were supported; recall is what share of supported codes were found. They trade, and a single blended number hides the trade. On the worked evaluation the loop turned 89.9% precision into 98.8% accuracy and 93.0% recall into 98.3% — and left 22 unsupported codes on real claims, the smallest number in the table and the only one with a False Claims Act attached to it. Both errors are real; neither is safe; downcoding is not the conservative option. But a missed code costs revenue and truth, while a fabricated one costs revenue, truth, and legal exposure that compounds with volume and consistency — so review effort belongs on the accept side, the confidence threshold is a compliance decision with a person's name on it, and an organization that checks only accepts will let recall rot silently because nothing about a missed code ever announces itself.

Model drift has three drivers: the code sets move on a calendar, the documentation changes, and — the one nobody monitors — the definition of "supported" changes while every note reads the same. That third one is Chapter 28's Case Study 1 exactly: correct when written, decayed, with no internal event to signal it.

And on Account 10-4471 the engine gets seven of eight codes right, including a genuinely good 20610-not-20611 call that the physician's documented negative made possible — and misses the modifier 25 argument, the diagnosis pointer that keeps a venipuncture from treating knee pain, the query Chapter 36 owns, and the sentence nobody wrote. Three of the four are invisible to a text-reading system by construction. That is why the coder is still there, and it is a list rather than a sentiment.


Key Terms

Clinical documentation integrity (CDI) — the discipline of making the medical record describe the patient accurately, completely, and in language the classification systems can read; reviewed against the clinical evidence in the record and corrected only by the clinician who owns the statement. (Ch.38)

Concurrent review — review of a record while the encounter is still open, before the bill drops and while the clinician who wrote the note can still be asked; the timing is what distinguishes CDI from auditing. (Ch.38)

Working DRG — the Medicare Severity Diagnosis-Related Group a record would group to as it stands mid-stay, on incomplete documentation; a review and prioritization tool, never a bill. (Ch.38)

Compliant query — a written request to a clinician to clarify conflicting, ambiguous, incomplete, illegible, or clinically inconsistent documentation: presenting the record's own clinical evidence, asking an open question, offering clinically reasonable options including "clinically undetermined," naming no code and no payment consequence, and retained with its response. (Ch.38)

Leading query — a query that supplies the answer, names a code or classification, states or implies a financial consequence, offers no clinically reasonable alternative, or signals a preferred response; the exposure is manufacturing documentation to support a code rather than coding incorrectly. (Ch.38)

Query response — the clinician's answer; it counts only when it is documented in the medical record itself, since a code may never be assigned from an answer that exists only on the query form. (Ch.38)

Computer-assisted coding (CAC) — software that reads clinical documentation and proposes codes for a human to review, presenting each suggestion with the source text that produced it. (Ch.38)

Natural language processing (NLP) — the computational extraction of structured meaning from unstructured clinical text; in coding, the step that turns a physician's prose into clinical concepts a code can attach to. (Ch.38)

Negation detection — the natural language processing step that decides whether a concept found in the text is asserted or denied; the first of the four qualifiers, alongside experiencer, temporality, and certainty. (Ch.38)

Autonomous coding — assignment of codes and release of a claim without a human reviewing that particular chart; currently viable only where the source document is structured, the code space small, the variability low, and the feedback fast. (Ch.38)

Coder-in-the-loop — a workflow in which software proposes and a credentialed human reviews, accepts, rejects, or corrects each suggestion before submission. (Ch.38)

Automation bias — the tendency to accept a system's recommendation more readily than the same conclusion reached independently, and to stop looking once a plausible answer is displayed; the failure mode that turns a review loop into a rubber stamp. (Ch.38)

Precision — of the codes an engine suggested, the share supported by the record; a precision failure is a false positive. (Ch.38)

Recall — of the codes the record supported, the share the engine found; a recall failure is a false negative. (Ch.38)

False positive — a code proposed or reported that the documentation does not support; on a submitted claim it is a statement to the payer and, to a federal health program, a certification. (Ch.38)

False negative — a code the record supported that was never reported; it costs revenue and distorts severity, risk, and quality data, and it announces itself in no system. (Ch.38)

Model drift — degradation of a model's performance over time as the world it was fitted to changes, with no change in the model itself: revised code sets, changed documentation, and — the driver nobody monitors — a changed definition of what counts as supported. (Ch.38)


Spaced Review

  1. Write a compliant query on a constructed inpatient record documenting a hemoglobin of 7.8 g/dL, two units of packed red cells transfused, and an assessment reading only "blood loss." Then list the seven elements from §38.3 and check your draft against each. Finally, state which single change would make your query leading.

  2. A CDI program reports a 94% agreement rate, a query rate up eight points year over year, and a rising case mix index. Name what you would ask for before congratulating anyone, say what a defensible answer looks like, and explain how a program can reach 94% without ever sending a non-compliant query.

  3. (Chapter 37) An organization is about to let an engine code and release radiology claims without human review of individual charts. Using the assertion register's five fields, write the entry — and then name the two events that force an off-cycle review of it and why each one applies with unusual force to a coding engine.

  4. (Chapter 4) The March 14 note contains the clause "no imaging guidance used." Explain what that clause decides, what a coder faces without it, and why an absence of any statement about guidance is a different thing from a documented negative.

  5. (Chapter 36) The reviewed problem list on Account 10-4471 includes chronic kidney disease, stage 3a, and the assessment continues metformin without referencing renal function. State the query that should have been sent, name the answer that changes nothing, and explain — in one sentence — why no answer to any query could make the kidney disease reportable for March 14.


Next: Chapter 39. The credential. Two organizations, seven or eight credentials worth naming, what each exam actually tests, how to prepare week by week, the rules about annotating the code books you carry into the room, and how to choose between the paths honestly — including what §38.10's map of exposure implies about which credential to reach for first.