The Medical Necessity Sentence That Keeps the Claim Paid
When Maria watched a peer absorb a $14,200 recoupment letter after a high-frequency 90837 review, the failed notes were not empty. They were full of clinical content: symptoms described, interventions named, plans stated. What they lacked was one sentence, the sentence that ties symptoms to functional impairment to intervention to ongoing need for care in language a payer reviewer recognizes as medical necessity. Every billed note rests on that single load-bearing sentence, and most clinicians were never taught to write it deliberately. This lesson dissects the medical necessity sentence word by word, tests it against the InterQual and MCG criteria that Aetna, BCBS, Cigna, and UnitedHealthcare reviewers actually apply, runs it through a Medicare-style audit lens, and leaves you with a template bank keyed to your caseload, so that necessity documentation for a 90834 or 90837 becomes a sentence you can build, check, and defend.
What Medical Necessity Means to the Person Denying Your Claim
Medical necessity is not "the client benefits from therapy." Every payer contract defines it more narrowly, and behavioral health reviewers operationalize it through commercial criteria sets, most commonly InterQual and MCG (formerly Milliman Care Guidelines). The details vary by product, but the behavioral health outpatient logic converges on four findings the reviewer must be able to make from your note: an active, diagnosable condition; current symptoms of that condition; functional impairment caused by those symptoms; and a skilled intervention that is treating the condition at the level of care billed, with a reasonable expectation of improvement or, for maintenance cases, prevention of deterioration. Miss any link and the chain fails. A note can be clinically rich and still fail all four, which is exactly what happened to Maria's peer: "processed trauma material" names no symptom, no impairment, no specific skilled intervention, and no reason care must continue.
Understand the reviewer's working conditions and the sentence design follows. A concurrent-review clinician at a commercial payer is processing dozens of charts a day against a criteria screen. They are not reading your note as a colleague; they are scanning it for the four findings, often literally completing a checklist. The controlling analogy for this lesson: the medical necessity sentence is a load-bearing beam. The rest of the note is walls and finish work; you can decorate beautifully, but if the beam is missing, the structure fails inspection, and the inspector does not award points for the wallpaper. Your job is to put the beam where the inspector looks, in one or two sentences, usually in the Assessment section, written so the four findings are checkable without inference.
One more reviewer reality: the standard is per-episode and per-intensity, not per-diagnosis. A diagnosis of F33.1 justifies treatment in general; it does not by itself justify weekly 90837s this month. The reviewer is asking why this client needs this frequency and this session length now. That "now" is why the sentence must be rebuilt note by note from current data, and why a copy-pasted necessity sentence is a self-defeating artifact: identical sentences across twelve sessions tell the reviewer that nothing is changing, and a course of treatment in which nothing changes is, by the criteria, a course that either needs a different level of care or no longer needs this one.
The Sentence, Dissected Word by Word
Here is the worked sentence, built from Maria's F33.1 client in the shorthand lesson, and we will take it apart clause by clause: "Client's persistent depressive symptoms (PHQ-9 16, decreased from 19), including anhedonia and impaired sleep, continue to cause occupational impairment, evidenced by two missed workdays this week; weekly individual CBT targeting self-efficacy cognitions remains medically necessary to consolidate the partial treatment response and restore functioning to baseline."
"Client's persistent depressive symptoms" names current, active symptoms of the diagnosed condition: finding one and two in five words. "Persistent" matters because it asserts the condition is active now, not historical. "(PHQ-9 16, decreased from 19)" is the verifiable anchor, the detail no AI can supply and no reviewer can dismiss: an objective measure, with a delta showing trajectory. The delta does double duty: it proves the condition remains in the clinical range (16 is still moderate) and proves treatment is working (down three points), which together argue for continuation rather than either discharge or escalation. "Including anhedonia and impaired sleep" specifies symptoms in criteria language a reviewer can map to the diagnosis. "Continue to cause occupational impairment, evidenced by two missed workdays this week" is the impairment clause, and the phrase "evidenced by" is the hinge of the entire sentence: it converts an assertion into a checkable fact. Impairment without evidence is opinion; impairment with a countable fact (two missed workdays, a failed class, a child welfare referral, a canceled medical appointment the client could not face) is a finding.
"Weekly individual CBT targeting self-efficacy cognitions" names the skilled intervention, its frequency, and its specific target, which answers the reviewer's intensity question: why weekly, why this modality, aimed at what. "Remains medically necessary" uses the criteria's own term of art, which is not magic but is recognition vocabulary; reviewers scan for it. "To consolidate the partial treatment response and restore functioning to baseline" states the expectation of improvement with a defined endpoint, which closes the fourth finding and quietly answers the discharge question every reviewer carries: this clinician knows what done looks like. Forty-eight words, four findings, zero inference required. That is the beam.
Impairment without evidence is an opinion; impairment with a countable fact is a finding. The phrase "evidenced by" is the hinge that turns your clinical judgment into something a reviewer can check, and a checkable sentence is a payable one.
The Verifiable Detail Rule: What AI Cannot Supply
Run that sentence through an AI scribe's typical draft and you will see the failure pattern immediately. AI-drafted necessity language defaults to fluent abstraction: "Client continues to experience significant depressive symptoms impacting daily functioning; ongoing therapy is medically necessary to support symptom reduction." Every clause is plausible and every clause is unverifiable. No score, no delta, no countable impairment, no named modality target, no endpoint. A reviewer reads that sentence across twelve notes from twelve different clinicians every morning, because every AI tool generates a version of it, and reviewers have learned to discount it the way professors learned to discount "since the dawn of time" essay openers. Worse, an AI willing to invent specifics is more dangerous than one that stays vague: a hallucinated PHQ-9 score or a fabricated missed-workday count in a payer-facing document is a false claim, and patterned false claims are how a recoupment matter becomes a fraud matter.
So the rule, the same payer rule that governs this entire chapter: the verifiable details come from you, every time. For a 90837, the time in session, 53 minutes or more, comes from your calendar, because the reviewer's first question on a high-frequency 90837 audit is whether the sessions actually ran long enough to bill the code, and "AI wrote 55 minutes" is not an answer you can give under oath. The PHQ-9 or GAD-7 score and its delta come from the instrument you administered. The modality comes from what you actually did in the room, named at the technique level. The impairment evidence comes from what the client told you this week. Your shorthand template from the first lesson already captures all four; the necessity sentence is where they pay rent. The AI's legitimate role is assembly: given your four verifiable details, it can build the sentence in the canonical structure faster than you can type it, and that assembly-not-sourcing division is the entire game.
The prompt that does it: "Using ONLY the four details below, write a one-sentence medical necessity statement in this structure: current symptoms with measure and delta; functional impairment with 'evidenced by' and the countable fact; named intervention with frequency and specific target; 'remains medically necessary to' plus expected improvement with endpoint. Details: [score and delta] [impairment fact] [modality, frequency, target] [endpoint]." Four inputs, one beam, ten seconds, and nothing in it the model invented.
Testing the Sentence Against InterQual and MCG Logic
You will likely never read the licensed InterQual or MCG criteria text, because payers license them and do not hand them to network providers. But you can test your sentence against their published logic, because the structure is consistent: outpatient behavioral health continuation generally requires a covered diagnosis, active symptoms, functional impairment, treatment that is active and appropriate to severity, measurable response or a clinically justified plan adjustment if there is none, and absence of indicators for a different level of care. Convert that into a five-question self-audit and run your sentence through it. One: could the reviewer circle the active symptom? Two: could they circle the impairment evidence? Three: could they circle the named skilled intervention and its frequency? Four: could they circle the response data (the delta) or, if the client is not improving, the documented plan change that answers it? Five: does the sentence imply a level of care consistent with what you billed, neither "client is in crisis" on a routine weekly 90834 (which argues for escalation you did not provide) nor "client is stable and thriving" on a 90837 (which argues for discharge)?
Question four deserves its own paragraph because it is where honest clinicians get hurt. Clients plateau. A PHQ-9 that reads 16, 16, 15, 16 across two months is real clinical life, but a necessity sentence that keeps asserting "partial treatment response" against flat scores is now contradicted by its own evidence, and reviewers run exactly that longitudinal check. The criteria do not require constant improvement; they require that non-improvement be met with clinical reasoning: a documented plan adjustment (frequency change, modality change, adjunct referral, medication consult) or a defensible maintenance rationale (treatment is preventing deterioration, with the evidence being what happened during the last gap in care). The necessity sentence for a plateau reads differently: "Symptoms have plateaued (PHQ-9 stable at 16 across four sessions); clinician is introducing behavioral activation as the primary modality and adding a medication evaluation referral; continued weekly sessions remain medically necessary to implement the revised plan and prevent further occupational deterioration." That sentence survives because it shows the clinician saw the plateau and acted, which is the thing the criteria are actually screening for: active, responsive treatment rather than autopilot.
The Medicare-style lens adds one more test worth running even if you bill no Medicare: would this note demonstrate that the service was reasonable and necessary, that it required the skills of a licensed professional, and that the documentation supports the code billed? Medicare's documentation culture, time thresholds honored precisely, skilled-service language, longitudinal consistency, is the strictest common denominator, and a sentence that passes it passes nearly everything downstream.
Payer Dialects: Aetna, BCBS, Cigna, UnitedHealthcare
The four findings are universal; the emphasis shifts by payer, and knowing the dialects saves appeals. UnitedHealthcare (and its behavioral arm Optum) runs the most aggressive high-frequency and high-duration analytics in the commercial market; if you bill 90837 as your default code, expect the algorithm to flag you against peer norms, and expect the chart pull to focus on whether session length and clinical intensity justify the long code. Your defense is the verifiable pair: documented time in session from your records, and a necessity sentence whose severity language matches a 53-plus-minute clinical need (trauma processing that cannot be safely truncated, exposure work requiring full hierarchies, complexity such as F43.10 with co-occurring F43.12) rather than generic weekly support. Aetna's audit posture leans on its provider-manual documentation windows and on measurement; PHQ-9 and GAD-7 deltas embedded in the sentence align with where their reviews look. Cigna/Evernorth reviewers respond to functional language, work, school, relationships, ADLs, so the "evidenced by" clause should carry a functional noun. BCBS is thirty-plus independent plans, which means the dialect is local: your state plan's provider manual and its medical policy for outpatient psychotherapy are the source of truth, and citing your actual plan's language in a dispute beats any generic argument.
Two cross-payer constants. First, the sentence must agree with the rest of the chart: if your necessity sentence claims occupational impairment and your intake says the client is retired, the contradiction is the finding. AI assembly makes this failure easier, not harder, because the model will cheerfully build a beam from stale details if you feed it last month's facts; the four inputs must be this session's. Second, frequency claims are commitments: "weekly individual CBT remains medically necessary" obligates the chart to show weekly attendance or documented reasons for gaps, because a reviewer who sees "weekly" asserted and biweekly attendance delivered reads the sentence as boilerplate, and boilerplate is the gateway finding to a full chart audit.
Building the Template Bank Without Building Boilerplate
A template bank sounds like the boilerplate problem wearing a lab coat, so draw the line precisely: the structure repeats; the facts never do. Your bank holds sentence frames with mandatory variable slots, and a frame is only usable because its slots force this-session data into the load-bearing positions. Frame one, improving client: "[Symptoms] ([measure] [score], decreased from [prior score]) continue to cause [domain] impairment, evidenced by [countable fact this week]; [frequency] [modality] targeting [specific target] remains medically necessary to consolidate treatment response and [endpoint]." Frame two, plateau with plan change: "[Symptoms] have plateaued ([measure] stable at [score] across [n] sessions); clinician is [specific plan adjustment]; continued [frequency] sessions remain medically necessary to implement the revised plan and prevent [specific deterioration risk]." Frame three, maintenance: "[Symptoms] are partially remitted ([measure] [score]); during the [prior gap in care], client experienced [documented deterioration]; continued [frequency] sessions remain medically necessary to maintain functioning and prevent relapse, with step-down review at [date]." Frame four, acute exacerbation: "[Symptoms] have acutely worsened ([measure] [score], increased from [prior]) following [stressor], with impairment evidenced by [fact]; [frequency] [modality] remains medically necessary to restabilize, with [safety plan element documented separately] and level-of-care review if [threshold]."
Each frame ends differently because each answers a different reviewer suspicion: the improving client must show an endpoint, the plateau must show responsiveness, the maintenance case must show what happens without care, the exacerbation must show level-of-care awareness. Hand the frames to AI as the assembly instruction and your role collapses to supplying the slots, which is to say, to the clinical work you already did. The verification pass for a necessity sentence takes thirty seconds and asks exactly three things: is every number and fact in the sentence mine and from this session; do the four findings appear without inference; does the sentence agree with the rest of this note and the recent chart. Thirty seconds across a caseload is the cheapest recoupment insurance you will ever buy.
The Applied Problem: Your Medical Necessity Sentence Template Bank
Your artifact is a Medical Necessity Sentence Template Bank, one page, four frames, tuned to your caseload. Step one: copy the four frames from this lesson (improving, plateau-with-plan-change, maintenance, acute exacerbation) and rewrite each one in your own clinical voice without breaking the structure: symptoms with measure and delta, impairment with "evidenced by" and a countable fact, named intervention with frequency and target, necessity term plus expectation with endpoint. If your caseload includes a population with its own necessity logic (SUD with craving and use-day counts, ABA with skills-acquisition data, child work with school-functioning evidence), add a fifth frame for it, holding the same four-finding skeleton.
Step two: build the assembly prompt. Take the prompt from this lesson, embed your frames, and add the constraint block you already use everywhere: "Use ONLY the supplied details; if any slot is missing, output [CLINICIAN: SUPPLY] in that position rather than inventing a value." Test it deliberately by omitting the score and confirming the model placeholders instead of fabricating; an assembly prompt that invents a PHQ-9 under pressure is disqualified, because the one place you can never tolerate a hallucinated number is a payer-facing necessity claim.
Step three: run it against three real (de-identified) or fabricated cases, one improving, one plateaued, one maintenance, and put each output through the five-question InterQual/MCG-style self-audit and the three-item thirty-second verification. Then do the longitudinal check on your own chart: pull your last four notes for one client and read only the necessity sentences in sequence. If they are identical, you have found your boilerplate; if they tell a four-session story of measured response and responsive planning, you have found the beam. Done looks like: four to five frames in your voice, an assembly prompt that placeholders rather than invents, three test cases passing the five-question audit, and one real chart whose necessity sentences read as a trajectory. Staple it inside the front of the format decision tree; the bank is what fills the Assessment box no matter which format the tree chose.
Key Takeaways
- Medical necessity is a four-finding chain a reviewer must verify from your note: active diagnosed condition, current symptoms, functional impairment caused by those symptoms, and skilled intervention at the billed intensity with an expectation of improvement or prevention of deterioration. Miss one link and a clinically rich note still fails.
- The medical necessity sentence is the note's load-bearing beam, usually placed in the Assessment section, built so the four findings are checkable without inference. The hinge phrase is "evidenced by": impairment without a countable fact is opinion; with one, it is a finding.
- The verifiable details can only come from the clinician: the measure score and its delta (PHQ-9 16, down from 19), the time in session for a 90837 (53 minutes or more, from your calendar), the modality actually delivered with its specific target, and this week's impairment fact. AI's legitimate role is assembling your four details into the canonical structure, never sourcing them; a hallucinated number in a payer-facing claim turns recoupment risk into fraud risk.
- Test every sentence against the InterQual/MCG-style five-question self-audit: circleable symptom, circleable impairment evidence, named intervention with frequency, response data or documented plan change, and severity language consistent with the level of care billed. The Medicare-style lens (reasonable and necessary, skilled service, documentation supports the code) is the strictest common denominator.
- Plateaus are survivable; autopilot is not. Flat scores demand a sentence that shows the clinician saw the plateau and acted: a plan adjustment, an adjunct referral, or a maintenance rationale evidenced by what happened during the last gap in care. Identical necessity sentences across sessions are the boilerplate finding that triggers full chart audits.
- Payer dialects shift the emphasis: UnitedHealthcare/Optum runs high-frequency 90837 analytics and audits session time, Aetna leans on documentation windows and measurement deltas, Cigna reads for functional language, and BCBS necessity lives in your specific state plan's manual. The constants everywhere: the sentence must agree with the rest of the chart, and frequency claims are commitments the attendance record must honor.
- The template bank repeats structure, never facts: four frames (improving, plateau, maintenance, exacerbation), each with mandatory slots only this session's data can fill, assembled by a prompt that placeholders missing values instead of inventing them. The clinician supplies the slots, verifies in thirty seconds, and signs the note, because the signature, not the sentence, is what the payer ultimately audits.
Skill.re