The 90-Day Treatment Plan Review Cycle
Somewhere in your caseload right now is a treatment plan that expired without anyone noticing. The client kept coming, the sessions kept billing, and the document that justifies all of it went stale at day 91. When the payer's retrospective review pulls that chart, every session billed against the expired plan is a candidate for recoupment, and when the board pulls it after a complaint, the stale plan reads as treatment without a current clinical rationale. The 90-day treatment plan review is the most skippable-feeling document in behavioral health and one of the most expensive to skip. This lesson builds the review cycle as a system: an AI-drafted review that pulls the PHQ-9 and GAD-7 trajectory, session frequency, intervention summary, and outcome metrics from the chart; a clinician verification and decision pass that no software performs; and a signed document that is simultaneously payer-ready and board-defensible. By the end you will have the 90-Day Review Protocol, a repeatable procedure with a calendar trigger, a pull prompt, a decision framework, and a signature standard.
What the Review Actually Is: A Re-Justification, Not a Renewal
The controlling analogy for this lesson is the flight plan check. A pilot does not file a flight plan once and fly forever; at defined intervals the plan is checked against actual position, fuel burned, and weather, and the pilot either confirms the course, adjusts it, or diverts. The 90-day review is that check for treatment. It is not a renewal stamp on the original plan; it is a re-justification that answers four questions with evidence: Where did we say we were going (the original objectives, with their baselines and targets)? Where are we actually (current measure scores, session attendance, interventions delivered)? Does continued treatment at this frequency remain medically necessary (the trajectory argument)? And what changes (objectives met and retired, objectives revised, new problems added, frequency adjusted)?
Understanding what the document does for each of its three readers tells you what it must contain. The payer's concurrent or retrospective reviewer wants the medical-necessity arithmetic: diagnosis still supported, measurable progress or a documented clinical rationale for its absence, frequency justified by current severity. The board investigator, if it ever comes to that, wants evidence that treatment was directed rather than drifting: dated reviews showing the clinician repeatedly evaluated the course of care and made decisions. And the clinician, the reader everyone forgets, gets the one scheduled moment in a busy caseload to ask the question the weekly loop never forces: is this treatment actually working? A review cycle built only for the payer wastes its best clinical function.
The cadence itself is contract-driven and program-driven: 90 days is the common commercial and Medicaid convention, some programs run 30 or 180, and your payer contract or program rules control. What never varies is the structure of the obligation: a dated, signed re-evaluation at the defined interval, every interval, for every active client.
The Calendar Trigger: The Review That Schedules Itself
Reviews are missed for one boring reason: nothing fires at day 80. The first component of the protocol is therefore administrative, and it is the easiest AI touchpoint in the cycle. Your EHR or practice management system carries the plan date; a standing weekly task asks the AI (or a simple report) to list every client whose plan crosses day 75 within the next two weeks. The output is a queue, not a judgment: client initials, plan date, day count, review due date, and the next scheduled session inside the window. The verification act is a glance at the queue against your calendar; the decision of when in the window to conduct the review conversation with the client is yours. Day 75 is the right trigger because the review is not a document you generate about the client; it includes a conversation with the client, ideally a portion of a scheduled session, and you need a session inside the window to hold it in.
That conversation is worth naming as a clinical event, because the strongest reviews quote it. Fifteen minutes of a regular session: here is where we started, here is what the measures show, here is what I see, what do you notice, what should the next ninety days target? The client's own words about progress ("I rode the elevator twice last week; in January I could not enter the lobby") are evidence no instrument captures, and the review that contains them reads as collaborative care instead of paperwork. Schedule the measures to land just before the review session: the PHQ-9 or GAD-7 administered at the review-window session gives the trajectory its endpoint.
The Pull Prompt: Assembling the Evidence the Chart Already Holds
The review's raw material already exists if your session cycle has been running: field-entered measure scores, objective-numbered notes, named interventions, attendance records. The AI's job is assembly. Save this as your review pull prompt: "From the attached treatment plan, progress notes for the review period, and measure history, draft a treatment plan review with these sections: (1) Diagnosis as documented, unchanged unless I state otherwise. (2) Measure trajectory: every PHQ-9, GAD-7, or PCL-5 administration in the period, as a dated list with scores, and the net change from the plan's baseline; do not interpret the trajectory. (3) Session summary: number of sessions attended, missed, or cancelled, with CPT codes billed. (4) Intervention summary: the interventions named in the notes, with frequency of mention, quoting the notes; do not infer interventions not named. (5) Objective status table: each plan objective with its target, the relevant current measure score, and the supporting note citations; mark each objective [MET], [PROGRESSING], [NOT PROGRESSING], or [NOT ADDRESSED IN NOTES] strictly by comparing numbers to targets, with no clinical conclusions. (6) Open items: coordination tasks, referrals, or measures the notes show as incomplete. Use only the attached documents and cite the note date for every claim."
Read the architecture of that prompt. Sections one through six are evidence assembly: dates, scores, counts, quotes, citations. The two prohibitions ("do not interpret the trajectory," "no clinical conclusions") fence off the part of the review that is the clinical act. The status markers look like judgments but are defined arithmetically: a target of "PHQ-9 below 10" against a current score of 8 is [MET] by comparison, not by clinical reasoning. Even so, you verify every marker, because the comparison is only as good as the data pull, and a score attributed to the wrong date or instrument corrupts the whole trajectory. The verification pass on the pull: every score in section two checked against the EHR measure fields, the session count checked against the billing record, every quoted intervention spot-checked against its cited note, and every [NOT ADDRESSED IN NOTES] marker treated as a finding about your documentation, not necessarily about your treatment.
The AI assembles the evidence; the clinician renders the verdict. A review where the software wrote the clinical conclusions is not a review, it is a rubber stamp with a trajectory chart attached.
The Clinical Decision Pass: The Four Verdicts Only You Can Render
With the evidence assembled and verified, the clinical act begins, and it has four parts the AI never touches. First, the progress determination: what does this trajectory mean? A PHQ-9 that moved 16 to 9 is improvement; a PHQ-9 flat at 14 across ninety days is a clinical question with several honest answers (wrong modality, untreated comorbidity, external stressors, measurement insensitivity to the real work), and choosing among them is case formulation. Second, the necessity determination: does continued treatment at the current frequency remain justified? Improvement does not mean discharge, and non-improvement does not mean futility, but both require you to write the rationale in your own reasoning: "Symptom reduction has plateaued at moderate severity; I am revising the approach to incorporate behavioral activation targeting the documented withdrawal, and continued weekly frequency is indicated during the modality transition." Third, the plan revision: objectives met get formally retired with their evidence; objectives behind schedule get revised targets or revised methods; new problems that emerged in the period get added with new baselines. Fourth, the frequency and level-of-care call: step down, maintain, intensify, or refer out, stated explicitly with the reason.
The non-improvement scenario deserves special attention because it is where weak reviews die in audit. A reviewer who sees ninety days of flat scores and a review that says "continue current plan" has found their denial. The defensible version documents the clinical reasoning: what you considered, what you are changing, why continued care is necessary despite the plateau, and what marker at the next review would trigger a different decision. Paradoxically, the honest plateau review is stronger evidence of directed care than a row of cheerful improvement reviews, because it proves the reviews are real. The AI can format this paragraph after you compose its substance; it cannot supply the reasoning, and a generic "client would benefit from continued support" is the sentence reviewers are trained to disbelieve.
The Payer-Ready, Board-Defensible Document
Assemble the final document in a fixed shape so every review in your practice reads the same way: header (client identifiers, diagnosis, plan date, review period, review date); trajectory section (the dated measure list and net change, with the plan baseline restated); treatment summary (sessions, codes, interventions as delivered); objective-by-objective status with evidence; the clinician's progress and necessity determination in the clinician's own voice; the revised plan elements with new measurable targets and dates; the frequency decision with rationale; and signatures. The payer test for the finished document: could a reviewer who has never met the client reconstruct the medical-necessity argument from this document alone, with every factual claim traceable to a score, a date, or a cited note? The board test: does this document prove a licensed clinician evaluated the course of care and made decisions on this date? If either reader would have to take your word for anything important, the section that would answer them is missing.
Worked example, compressed. Client: F41.1, generalized anxiety disorder, plan dated 3/2. Trajectory: GAD-7 on 3/2: 17; 3/30: 15; 4/27: 12; 5/25: 10; net change minus 7 from baseline. Treatment summary: 12 sessions scheduled, 11 attended, one cancellation; 90834 x 11; interventions named in notes: cognitive restructuring (9 notes), worry exposure (5 notes), relaxation training (3 notes). Objective one (GAD-7 below 10 by week 16): current 10, [PROGRESSING], on pace. Objective two (return to weekly grocery shopping unaccompanied): client report in 5/18 note, "shopped alone twice," [MET]. Clinician determination: "Trajectory demonstrates consistent response to CBT with worry exposure. Objective two met and retired. Objective one revised target: GAD-7 at or below 7 by next review. Severity no longer supports weekly frequency; stepping down to biweekly sessions, with the client's agreement, and re-evaluating step-down tolerance at the next review." That document survives both readers, and notice the load-bearing material: numbers with dates, an honest step-down decision, and a clinician sentence no template wrote.
Signatures, Co-Signatures, and Keeping the Cadence Alive
The signature standard mirrors the rest of the program but with two review-specific points. The clinician signs and dates the review, and the date matters doubly here: it must fall inside the review window, because a review signed at day 120 documents that the plan ran stale for a month, which is the finding, not the fix. The client's participation gets documented, and where your program or payer requires it, the client signs the revised plan elements, which is the cleanest evidence of the collaborative review conversation. Pre-licensed clinicians route the review for supervisor co-signature inside the same window, and the review is exactly the document supervisors should be reading closely, because it is where a supervisee's case formulation is visible on one page; a supervisor who only ever co-signs progress notes is supervising the trees and never the forest.
Keeping the cadence alive across a full caseload is a systems problem, so give it a systems answer: the day-75 queue runs weekly without exception; every completed review immediately sets the next review date in the EHR; and a monthly self-audit asks one question, "do any active clients lack a current plan?", which should take five minutes and return zero. For Jordan's group practice, the same protocol scales with one addition: the compliance dashboard tracks review currency across all clinicians, because a single clinician's stale plans are a practice-level recoupment exposure under audit, and the malpractice carrier's questionnaire reads better when the answer to "how do you ensure treatment plans are current?" is a named protocol instead of a shrug.
The Applied Problem: The 90-Day Review Protocol
Your artifact is the 90-Day Review Protocol: a one-page procedure you run for every client, every interval. Build it with these five components. Component one, the trigger: the weekly day-75 queue (client, plan date, day count, due date, next session in window) and the rule that the review conversation books into a session inside the window. Component two, the measures: the relevant instrument (PHQ-9, GAD-7, or PCL-5 per the plan) administered at the review-window session so the trajectory has a current endpoint, verified against raw items before entry. Component three, the pull prompt, pasted in full from this lesson, with its two prohibitions and its citation requirement intact. Component four, the decision framework: the four verdicts (progress determination, necessity determination, plan revision, frequency call), each requiring a clinician-composed sentence, with the plateau rule in bold: flat trajectories get documented reasoning and a change or an explicit, marker-based rationale for staying the course. Component five, the close: signature and date inside the window, client participation documented, supervisor co-signature routed where required, next review date set in the EHR before the chart closes.
Now run the protocol against one real client this week, ideally your longest-running case, because long-running cases are where stale plans hide. Time it: the pull and verification should cost fifteen to twenty minutes, the review conversation fifteen minutes of a session, the decision pass and final document another fifteen. Under an hour for a document that used to either consume two hours or, to tell it true, not get done. Then check the protocol against the two reader tests: hand the finished review to a colleague and ask whether they could reconstruct the necessity argument cold, and whether they can point to the sentences only a clinician could have written. If the answer to the second is no, the AI wrote too much of your review.
"Done" looks like this: the protocol is saved and dated in your AI governance file, the day-75 queue is a standing weekly task, one real review has been produced and signed inside its window, the next review date is set, and your monthly "any stale plans?" audit returns zero. Ninety days from now, the system fires again without you having to remember it, which is the entire point.
Key Takeaways
- The 90-day review is a re-justification, not a renewal: it answers where the plan said you were going, where the measures say you are, whether continued treatment at this frequency remains medically necessary, and what changes. The cadence is contract-driven (90 days is the common convention; your payer or program rules control), and a stale plan exposes every session billed against it.
- The document serves three readers: the payer reviewer needs the medical-necessity arithmetic, the board investigator needs proof that care was directed rather than drifting, and the clinician gets the one scheduled moment to ask whether treatment is actually working. A review built only for the payer wastes its best clinical function.
- The system starts with a calendar trigger: a weekly day-75 queue listing every plan approaching expiry, so the review conversation books into a session inside the window and the measures land just before it to give the trajectory its endpoint.
- The pull prompt assembles evidence only: dated measure scores and net change from baseline, session and attendance counts with CPT codes, interventions quoted from notes, and an objective status table marked [MET], [PROGRESSING], [NOT PROGRESSING], or [NOT ADDRESSED IN NOTES] by arithmetic comparison, with citations for every claim and explicit prohibitions on interpreting the trajectory or drawing clinical conclusions.
- Four verdicts belong to the clinician alone: the progress determination, the necessity determination, the plan revision (retiring met objectives, revising stalled ones, adding new problems with baselines), and the frequency and level-of-care call. The plateau case is where weak reviews die: flat scores demand documented clinical reasoning and either a change or an explicit marker-based rationale, never "continue current plan."
- The finished document must pass two tests: a payer reviewer who never met the client can reconstruct the necessity argument with every claim traceable to a score, date, or cited note; and a board reader can see that a licensed clinician evaluated care and decided on this date. Signature and date must fall inside the window; supervisors co-sign reviews closely because the review is where a supervisee's case formulation is visible on one page.
- The artifact is the 90-Day Review Protocol: trigger queue, review-window measures, the pull prompt, the four-verdict decision framework with the plateau rule, and the close (signature in window, client participation documented, next review date set). Run faithfully, the review costs under an hour instead of two, and the monthly stale-plan audit should return zero.
Skill.re