Building the Roadmap: Prioritizing Use Cases by Clinical Impact and Risk
Jordan's readiness assessment is signed and dated, the regulatory gate is releasing, and now comes the question every vendor wants to answer for you: what do we actually roll out, and in what order? Get the sequence wrong and you burn your one chance at clinician trust on a use case that fails publicly, or worse, you lead with a client-facing tool that puts the practice on the wrong side of the Illinois WOPR Act line while the documentation problem that is actually bleeding money sits untouched. This lesson teaches the discipline of the behavioral health AI roadmap: plotting fifteen candidate use cases on a 2x2 of clinical impact against clinical risk, killing the high-risk/low-impact quadrant without sentiment, sequencing the survivors (documentation first, then measurement-based care, then admin and prior-auth, then client-facing), and tying 30/60/90/180-day milestones to three numbers a board or an owner can audit: claim denial rate, clinician retention, and MBC completion. You will leave with a sequenced roadmap for your own practice, built on your own baseline.
Why Sequence Beats Selection
Most practices approach AI as a selection problem: which tool? The roadmap reframes it as a sequencing problem: which use case earns the right to go first, and what does each later phase inherit from the one before it? The difference matters because trust, the scarcest resource in any clinical rollout, compounds or collapses based on the first deployment. A documentation scribe that visibly returns ninety minutes a night to clinicians builds the credibility the MBC rollout will spend six months later. A chatbot intake pilot that misroutes one client in crisis spends credibility the program never gets back, and in a 25-clinician practice the story travels to all three counties by Friday.
Think of the roadmap as triage. In an emergency department, you do not treat patients in order of arrival or in order of how interesting their cases are; you treat by acuity and survivability, with explicit criteria, and you re-triage as conditions change. The roadmap does the same thing to use cases. Acuity is clinical impact: how much measurable pain does this use case relieve, against the baseline metrics your readiness assessment froze (denial rate by payer, documentation time per note, the 72-hour backlog, turnover, MBC completion)? Survivability is the inverse of clinical risk: what happens when this use case fails, who is harmed, and which regulator, payer, or board gets involved? A use case with high acuity and high survivability goes first. A use case that would be fascinating but could kill someone does not get a bed.
The reframe also disciplines vendor conversations. A vendor sells a platform; a roadmap buys a use case. When Jordan evaluates a scribe vendor, the roadmap's first phase defines the job: reduce documentation time and raise note quality enough to move the Anthem denial pattern, measured at day 60 and day 90. Features outside that job are noise this quarter, whatever the demo shows.
The Impact Axis: What Actually Moves Your Three Numbers
Impact is not excitement. Impact is the measured distance a use case can move one of the three program metrics: claim denial rate, clinician retention, and MBC completion. Every candidate use case gets scored against those, using the baseline from your readiness assessment, and the scoring conversation is where vendor claims go to be tested. A scribe vendor claims time savings; your impact question is more specific: if clinicians stop writing thin notes at 9:54 PM from degraded memory, does the medical-necessity language improve enough that Anthem stops denying prior-auths where the PHQ-9 trajectory already justifies care? That is a testable claim with a dollar value, because every denied claim has one.
Retention deserves special weight in the impact scoring because its dollar value is the largest and least visible. Replacing a clinician costs $25,000 to $60,000 in onboarding plus six to nine months of suboptimal productivity, and the leading indicator of departure in a documentation-burdened practice is exactly the pattern your readiness assessment measured: unpaid evening hours, a growing unsigned backlog, after-9-PM signatures. A use case that removes five to ten unpaid hours a week from a clinician's life is a retention intervention wearing a documentation costume. When you score impact, ask of each use case: does this touch the reason clinicians leave?
MBC completion is the third axis number because it is both clinically meaningful and contractually valuable. Jordan's clinical director wants PHQ-9, GAD-7, and PCL-5 at every fourth session, billed as CPT 96127 where covered, and completed instruments feed outcome reports for payer and value-based-care conversations. A use case that lifts MBC completion (automated administration reminders, AI-drafted score-trend summaries the clinician verifies, flagging of overdue instruments) scores high on impact even though it looks administratively boring. Boring and measurable beats exciting and vague at every point on this map. One more impact rule from the playbook worth stating plainly: in CCBHC and HCBS environments, reporting-burden use cases (daily encounter documentation under the PPS methodology, quarterly narrative reporting, person-centered plan drafting) often have unexpectedly high ROI, because the burden is enormous, the format is structured, and the clinician verification step is clean. If you run one of these environments, score those use cases before you score anything glamorous.
The Risk Axis: What Happens When It Fails, and Who Comes Asking
Risk in this 2x2 is clinical and regulatory, not technical. The question is never "might the model make a mistake" (it will) but "when it does, what is the blast radius, and does the workflow contain it before harm?" A scribe that drafts a note a clinician reads and signs has a contained failure mode: the error dies at the clinician's verification pass, because the signature is a legal attestation and every word gets read before signing. A client-facing tool that interacts with a person in distress, unsupervised, has an uncontained failure mode, and it sits next to the lines the Illinois WOPR Act and Nevada AB 406 drew against AI providing therapy itself.
Score risk on three components. First, proximity to clinical judgment: anything that approaches the hard limits of this program scores maximum risk and is removed from consideration entirely, not merely deprioritized. AI never scores the CSSRS, never assigns a risk level, never makes the Tarasoff or duty-to-protect determination, never makes the mandated-report call. AI structures, transcribes, and formats after the clinician's determination; a use case that requires AI to make the determination is not a quadrant, it is a refusal. Second, containment: is there a mandatory human verification step between the AI output and any consequence, and is that step real (a clinician reading) or theatrical (a clinician clicking)? Third, audience: does the output go to a clinician (lowest risk), to a payer or court (medium, because errors become external representations), or to a client directly (highest, because no professional reads it first)?
Risk scoring is also where the practice's specific regulatory exposure re-enters. If any part of the practice is a Part 2 program, use cases touching SUD records inherit 42 CFR Part 2 constraints under the 2024 final rule and score higher risk with any vendor not architected for it. If the practice employs associates, any use case touching their clinical formulation work raises BBS supervision exposure unless the supervision agreements have caught up. The 2x2 is practice-specific; copying another practice's grid imports their risk profile and ignores yours.
Plotting the Fifteen Candidate Use Cases
Here is the candidate list a group practice like Jordan's should plot, and roughly where honest scoring lands each one. High impact, low risk (the go-first quadrant): (1) AI-drafted progress notes with mandatory clinician verification before signature; (2) AI-drafted treatment plan updates from clinician input; (3) prior-auth and appeal letter drafting with the clinician supplying the verifiable details AI cannot, the time-in-session minutes, the specific PHQ-9 delta, the modality named in session; (4) MBC administration support, reminders, and clinician-verified score-trend summaries (never AI scoring of the CSSRS); (5) for CCBHC/HCBS practices, daily encounter and quarterly narrative report drafting with clinical director sign-off.
High impact, higher risk (sequence later, with controls): (6) intake and biopsychosocial drafting from structured questionnaires, risk content always routed to the clinician untouched; (7) AI-assisted case-conceptualization support for licensed clinicians, which for associates requires supervisor knowledge and an updated supervision agreement first; (8) coordination-of-care letter drafting, external audience, so verification is stricter; (9) discharge summary drafting; (10) AI-supported supervision preparation, summarizing session themes for the supervisor's review.
Lower impact, low risk (fill-in work, fine whenever): (11) psychoeducation handout drafting at specified reading levels; (12) administrative email and scheduling templates; (13) translation support with bilingual-clinician verification. And the kill quadrant, high risk, low impact: (14) client-facing chatbots for between-session support, which sit closest to the WOPR/AB 406 line, carry uncontained failure modes, and move none of your three numbers; (15) AI-driven risk prediction or triage scoring, which violates the hard limits outright. The discipline of the 2x2 is that this quadrant dies in the meeting, in writing, with the reason recorded. High-risk/low-impact use cases do not get parked for later; they get a written no, because an undocumented maybe resurfaces every time a vendor pitches it.
The kill quadrant is the roadmap's most valuable output. What a practice writes down that it will not do, and why, is what protects it when a vendor, a board, or a plaintiff's attorney asks what it was thinking.
Documentation First: The Sequencing Logic
The survivors get sequenced in four phases, and the order is not arbitrary: documentation first, then measurement-based care, then admin and prior-auth at scale, then (only if it ever earns its place) anything client-facing. Documentation goes first for four compounding reasons. It is the largest pain, the five to ten unpaid weekly hours your readiness assessment measured. It has the most contained risk, because the clinician verification pass before signature is already the law of the practice. It produces the fastest visible win, which buys the trust every later phase spends. And it is the substrate for everything downstream: better notes carry the medical-necessity language that prior-auth letters cite, the time-in-session detail that survives an Optum 90837 audit, and the documented score trends that MBC reporting aggregates. A practice that starts anywhere else is building the second floor first.
Measurement-based care comes second because it rides on documentation's rails. Once clinicians trust the scribe workflow and notes reliably capture instrument scores, the MBC cadence (PHQ-9, GAD-7, PCL-5 on schedule, 96127 billed where covered) becomes an extension of an existing habit rather than a new burden. Admin and prior-auth scale third: the practice now has months of richer documentation to draw on, so AI-drafted prior-auths and appeals cite real PHQ-9 trajectories instead of adjectives, which is precisely the fix for Jordan's Anthem denial pattern. Client-facing applications come last or never, and the bar is explicit: a contained failure mode, a clear position on the right side of the WOPR/AB 406 line, a consent architecture, and an impact case stated in the three numbers. Most group practices will correctly conclude that nothing client-facing clears that bar in the first year, and the roadmap should say so rather than leave the question open.
Sequencing also has a people layer. Phase one belongs to volunteers and champions, including Jordan's twelve self-adopters now operating under BAAs and policy. Phase two recruits the persuadable middle on the strength of phase one's visible results. The principled skeptics are invited to audit, not to adopt: give them the error log and the kill-quadrant memo, and let the program's discipline do the persuading.
The 30/60/90/180-Day Milestones, Tied to Numbers
A roadmap without dated milestones is a mood board. Tie each phase gate to the three metrics, measured against the frozen baseline. Day 30: documentation pilot live with five volunteer clinicians; consent addenda in client files; AI-use log running (which tool, which session, when, edit distance); measure time-per-note against baseline and target a visible reduction; zero unverified notes signed, audited by spot-check. The day-30 gate is operational: does the workflow hold, do clients consent, do clinicians keep verifying when tired?
Day 60: expand to the willing majority; first denial-rate reading by payer, expecting movement on documentation-sensitive denials like Anthem's medical-necessity pattern (denial data lags, so day 60 is a first signal, not a verdict); unsigned-past-72-hours backlog should be visibly shrinking; clinician satisfaction pulse taken, because retention is a slow metric and satisfaction is its fast proxy. Day 90: documentation phase judged against its targets and formally closed or extended; MBC phase opens with the instrument cadence and 96127 billing where covered; MBC completion baseline established for the caseloads in scope; the kill-quadrant memo reviewed and re-affirmed, because ninety days is exactly when the first vendor will pitch the chatbot again.
Day 180: the full first-year roadmap judged on the three numbers. Denial rate by payer against baseline, with the prior-auth drafting use case now live and appeals citing PHQ-9 trajectories. Retention read directly: trailing turnover, plus the satisfaction pulse trend, priced against the $25,000 to $60,000 per-clinician replacement cost and six to nine months of suboptimal productivity that every retained clinician avoids. MBC completion against its day-90 baseline. The day-180 review also re-runs the readiness assessment from the previous lesson, because the practice the roadmap was built for no longer exists; the assessment, the 2x2, and the milestones form a loop, not a line. Whatever the numbers say goes in writing to the owner, the clinical director, and the compliance officer, in the same document set the malpractice carrier's next questionnaire will be answered from.
Governing the Roadmap: Re-Triage, Escalation, and the Paper Trail
A roadmap is a governance instrument, which means it needs an owner, a cadence, and a record. The owner is the named program lead from the readiness assessment, not a committee. The cadence is monthly: metrics reviewed, milestone status updated, any new use-case proposal scored on the same 2x2 before it gets a meeting. The record is a dated roadmap document with version history, because the sequence will change and the reasons it changed are exactly what an auditor, a carrier, or a board inquiry will want to see. Re-triage is expected: if the day-60 denial reading shows the Anthem pattern unmoved, the response is diagnostic, not cosmetic. Pull ten denied claims, read the notes against the denial reasons, and determine whether the problem is the drafting, the verification, or a payer policy no note quality can fix; the roadmap adjusts on that finding.
Escalation rules belong in the roadmap itself. Any AI output error that reaches a signed note triggers the incident process and a verification-step review. Any use case that drifts toward the hard limits, an MBC tool that starts offering risk scores, a scribe feature that proposes a CSSRS interpretation, gets shut off and documented, because vendors ship features faster than practices ship policy, and the roadmap's quadrant assignments must be re-checked at every vendor product update. And the paper trail discipline applies to the kill quadrant most of all: when the board member's nephew forwards a pitch for a client-facing wellness chatbot, the answer is not a new debate, it is the dated memo from the quadrant meeting, with its reasons, and an invitation to bring new evidence to the next monthly review.
Finally, the roadmap stays honest about what it is not. It is not a promise that AI fixes the practice; it is a sequence of testable bets, each priced against three numbers, each killable on its data. The budget lesson that follows prices those bets: per-clinician licensing against enterprise against EHR-native, held against the unpaid hours, the denials, and the retention dollars this roadmap is built to move.
The Applied Problem: The Sequenced AI Roadmap with the 2x2 and Named Milestones
Your deliverable is the Sequenced AI Roadmap: a three-part document for your own practice. Part one is the 2x2: all fifteen candidate use cases (adjust for your environment; CCBHC and HCBS practices add their reporting-burden cases) plotted on impact against risk, with a one-line evidence note per placement tying impact to one of your three baseline numbers and risk to a named failure mode and audience. Part two is the kill-quadrant memo: every high-risk/low-impact use case listed with the written reason it dies, signed by the program owner and the clinical director. Part three is the four-phase sequence (documentation, MBC, admin/prior-auth, client-facing-if-ever) with 30/60/90/180-day milestones, each milestone stating the metric, the baseline value, the target, and who measures it.
Build it in this order. First, copy your baseline metrics from the readiness assessment: denial rate by payer with stated reasons, documentation time per note, the 72-hour backlog, trailing turnover, MBC completion if measured. Second, convene the same four raters (owner, clinical director, compliance officer, billing manager) plus one champion and one principled skeptic, and score the fifteen use cases independently before reconciling, exactly as you scored readiness; divergences are findings. Third, draft with help if you want it, using an administrative prompt with no PHI: "Here are fifteen AI use cases for a behavioral health group practice, each with an impact score tied to denial rate, retention, or MBC completion, and a risk score with failure mode and audience. Produce a 2x2 placement table, a kill-quadrant memo with reasons, and a four-phase rollout sequence with 30/60/90/180-day milestones, each milestone naming its metric, baseline, target, and owner."
Then run the verification pass yourself, because this artifact gets audited like a note: check that no surviving use case violates the hard limits (no AI scoring of the CSSRS, no risk levels, no Tarasoff or mandated-report determinations, no client-facing therapy-like interaction near the WOPR/AB 406 line); check that every milestone has a number, a date, and a named owner; check that the kill memo gives reasons an outside reader could evaluate, not adjectives. Done looks like this: one document, dated and versioned, with a 2x2 a board member can read in two minutes, a kill memo a carrier underwriter would respect, and milestones a billing manager can measure without interpretation. File it next to the readiness assessment; the budget model in the next lesson prices its first two phases.
Key Takeaways
- AI rollout is a sequencing problem before it is a selection problem. The roadmap works like ED triage: impact is acuity measured against your three numbers, risk is the inverse of survivability, and the first deployment either compounds or collapses the trust every later phase will spend.
- Impact means measured movement on claim denial rate, clinician retention, and MBC completion, scored against the frozen baseline from your readiness assessment. Retention carries the largest hidden dollar value: replacing a clinician costs $25,000 to $60,000 in onboarding plus six to nine months of suboptimal productivity.
- Risk is clinical and regulatory, scored on proximity to clinical judgment, containment by a real human verification step, and audience. Any use case requiring AI to score the CSSRS, assign risk levels, or make Tarasoff or mandated-report determinations is a refusal, not a quadrant placement.
- Of the fifteen candidates, high-impact/low-risk goes first (verified note drafting, treatment plan updates, prior-auth drafting with clinician-supplied verifiable details, MBC support, CCBHC/HCBS reporting). High-risk/low-impact dies in writing: the kill-quadrant memo with signed reasons is the roadmap's most defensible artifact.
- Sequence documentation first, then measurement-based care, then admin and prior-auth, then client-facing if ever. Documentation is the largest pain, the most contained risk, the fastest visible win, and the substrate every later phase rides on; most practices will correctly conclude nothing client-facing clears the bar in year one.
- Milestones are dated and numeric: day 30 proves the workflow and consent hold, day 60 reads the first denial signal and the backlog, day 90 closes documentation and opens MBC with 96127 billing where covered, day 180 judges the program on all three numbers and re-runs the readiness assessment.
- The roadmap is a governance instrument: a named owner, a monthly re-triage cadence, version history, escalation rules for any drift toward the hard limits, and a paper trail that answers the vendor pitch, the carrier questionnaire, and the board inquiry with dated reasoning instead of recollection.
Skill.re