โ†
AI for Mental & Behavioral Health Clinicians
Visionary ยท M12 ยท lesson 12 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Scaling What Works, Killing What Doesn't
๐Ÿ“–
now learning

Scaling What Works, Killing What Doesn't

15 min

The most expensive AI tool in behavioral health is not the one that fails its pilot. It is the one that passes a twelve-clinician pilot, gets rolled out to the full workforce on a victory lap, and then quietly stops working at scale: adoption sags, hallucination incidents climb with the long tail of session types nobody piloted, the consent process frays in the satellite offices, and two years later the organization is paying enterprise licensing for a tool a third of its clinicians have abandoned, with nobody empowered to say the word "kill." This lesson builds the machinery that prevents both failure modes: a stage-gate model that scales a successful pilot in controlled expansions with evidence required at every gate, and off-ramp criteria, written before scaling begins, that let the organization stop a tool that does not survive a real caseload. By the end you will have built the Stage-Gate Scorecard with Off-Ramp Criteria, the document that turns "should we expand this?" from a politics question into an evidence question.

The Locks on the Canal

Carry this analogy through the lesson: scaling a clinical AI tool is moving a ship through canal locks, not opening a floodgate. A floodgate releases everything at once and cannot be closed against the current; that is the enterprise rollout announced the week after a good pilot. A canal moves the ship in stages: the vessel enters a lock, the gate closes behind it, the water level is checked and adjusted, and only when the level is verified does the next gate open. At every lock the ship can be held, and at every lock the ship can be turned back. The gates are not obstacles to the journey; they are what makes the journey survivable, because the canal was built by people who knew that water at different levels destroys ships that move between them too fast.

The water-level difference is real in your organization. A pilot population and a full caseload are at different levels in every way that matters: the pilot had volunteers, the workforce has skeptics and strugglers; the pilot excluded high-risk sessions, Part 2 records, and group work, the caseload includes all of them; the pilot had a weekly huddle watching thresholds, scale has whatever monitoring you deliberately build; the pilot had twelve clinicians whose incidents could be hand-investigated, scale has incident volumes that need a system. A tool that performed beautifully in the pilot lock can be wrecked, or do wrecking, when the gate opens onto water it has never floated in. The stage-gate model exists to change the levels gradually and to check, with instruments rather than optimism, before every gate.

Picture where this lands. Jordan's practice finished a defensible twelve-week scribe pilot, the kind built in the previous lesson: composite primary outcome met, two tier-two thresholds breached and remediated, incident register clean in the back half. The vendor's account executive proposes enterprise deployment in thirty days with a discount that expires Friday. The clinical director wants it yesterday. Two supervisors are skeptical. Without a stage-gate model, this becomes a contest of enthusiasm and rank. With one, it becomes a sequence everyone agreed to before the pilot ended.

The Four Gates: Pilot, Beachhead, Expansion, Standard of Work

A workable stage-gate model for a behavioral health organization has four stages, each with an entry gate that demands evidence and an exit gate that demands more.

Gate one, pilot to beachhead. The pilot's prespecified decision criteria produced a scale recommendation; the gate verifies it. Evidence required: the primary composite met its margin, every tier-one event has a verified remediation, tier-two threshold history is documented, and the consent decline rate is within the prespecified band. The beachhead itself is deliberately modest: roughly double the pilot population, drawn this time to include the skeptics and the average, not just volunteers, still excluding the populations the pilot excluded. The beachhead's question is selection bias: does the tool work for clinicians who did not choose it?

Gate two, beachhead to expansion. Evidence required: the composite holds in the unselected population, adoption fidelity (fraction of eligible sessions actually using the tool) stays above a prespecified floor, the incident rate per hundred sessions does not exceed the pilot rate by more than a stated margin, and the monitoring system that replaced the weekly huddle is demonstrably catching what the huddle caught. Expansion brings in the harder territory deliberately and one increment at a time: additional offices, then additional session types, and only with separate sub-gates the populations the pilot excluded. Part 2 caseloads enter only after the data architecture is verified against the 2024 final rule's consent and redisclosure requirements; group and couples work enters only behind a group-built consent design; sessions involving elevated risk remain governed by the standing rule that the tool formats after the clinician's determination, never before it, and AI never scores the CSSRS, never assigns a risk level, never makes the duty-to-protect or mandated-report call.

Gate three, expansion to standard of work. The tool stops being a project and becomes infrastructure: written into onboarding, the documentation policy, the consent packet, the supervision agreements, and the malpractice-renewal narrative. Evidence required: stable or improving composite across all expanded populations, denial rates on tool-assisted documentation at or below the organization's baseline, an incident process that runs without the pilot team, and a budget line that survives contact with the CFO at full licensing cost.

Gate four is permanent: the standing re-verification gate. Standard of work is not tenure. Annually, and after any major vendor model change, the tool re-passes a slim version of gate three: sampled note-quality audit, incident-rate review, denial-rate check, consent-process spot check. A tool that was defensible in 2026 is not automatically defensible in 2028; the re-verification gate is how the organization notices.

Off-Ramps, Written Before the On-Ramp

Here is the discipline that separates a stage-gate model from a slide about one: the off-ramp criteria, the conditions under which the tool is stopped, are written before scaling begins, with the same numeric specificity as the pilot's stopping rules, and they apply at every stage including standard of work. Off-ramps written in advance are evidence-based; off-ramps improvised during a crisis are political, slow, and usually too late.

Three categories of off-ramp, mirrored from the pilot but tuned for scale. Immediate suspension: any tool-generated risk determination reaching a chart; any PHI breach outside the BAA chain; a vendor change that breaks the BAA or adds an unapproved subprocessor; loss of the regulatory basis for operation (a new statute in your state, a board guidance change). Suspension means capture stops organization-wide pending investigation, and the protocol names who can order it without a meeting. Threshold-triggered kill review: the composite metric below its floor for two consecutive measurement periods; incident rate per hundred sessions exceeding the stated ceiling; adoption fidelity below the floor for a quarter (a tool clinicians have abandoned is already dead and is now only a cost and a liability); denial rates on tool-assisted notes rising above baseline; consent declines or revocations trending past the band. Strategic kill: the vendor is acquired and the data terms change; the renewal price destroys the ROI math; a superior alternative passes its own pilot. Strategic kills are decisions, not emergencies, but they belong in the written criteria so that someone owns watching for them.

Every off-ramp needs three more things written down. The decision-maker: a named role, not a committee that must be assembled. The kill execution plan: how capture is turned off, how clients are notified that the consent they gave no longer applies to new sessions, how data return and deletion certification are extracted from the vendor under the contract's exit clause, and how documentation workflow reverts without a note backlog catastrophe (the reversion plan is rehearsed, like the pilot's fire drill, because forty clinicians losing their scribe on a Tuesday is itself a clinical operations event). And the no-sunk-cost clause, in writing: prior spending on licenses, integration, and training is explicitly excluded from kill deliberations, because the money is gone either way and the only question is whether future clients and clinicians are served by the tool's continued operation.

An organization that cannot name, in writing, the conditions under which it would kill a tool has not adopted the tool; the tool has adopted the organization.

The Scorecard: Instruments, Not Vibes

Each gate decision runs on a one-page scorecard, completed by the same kind of small group that ran the frontier scan: a clinical leader, a compliance lead, a working clinician, plus, from gate two onward, the billing manager, because denial data is gate evidence. The scorecard's sections mirror what this chapter has built. Clinical performance: composite metric current value versus floor, note-quality audit sample results, incident rate per hundred sessions versus ceiling, with the risk-workflow line item always present (count of tool-generated risk language events, target zero). Adoption reality: fidelity percentage, abandonment count, time-to-proficiency for new users. Consent integrity: decline rate, revocation count, spot-check results on whether the clinician-voiced consent is actually being delivered (this line becomes the bridge to the next lesson, where a platform tries to deliver it for you). Financial: per-clinician cost at current scale, denial-rate delta, the ROI calculation rerun with real numbers rather than the vendor's. Regulatory: any statute or board-guidance change since the last gate, BAA and subprocessor status, Part 2 architecture status if applicable.

Every line has three columns: the number, the prespecified threshold, and a green/amber/red status that is computed, not negotiated. The gate rule is mechanical: all green, the gate opens; any red, the gate holds and the off-ramp question is formally asked; amber items get a remediation owner and a date, and two consecutive ambers on the same line convert to red. The mechanical rule is the point. The moment status colors become negotiable, the scorecard is decoration and the loudest voice in the room is the real governance.

The scorecard is also where the kill decision gets its dignity. A kill executed from a red scorecard line, against prespecified criteria, with a named decision-maker and a rehearsed reversion plan, is not a failure story; it is the governance system working exactly as designed. The organizations that cannot kill tools are the ones that never defined what killing would look like, so every kill proposal arrives as an accusation against whoever championed the tool. Prespecified criteria depersonalize the decision: the tool did not survive a real caseload, the scorecard says so, and the people who ran the process did their jobs well.

What Scale Breaks That Pilots Cannot See

Know the standard failure patterns so the scorecard watches for them. The long tail of sessions: a pilot on individual outpatient therapy never met the bilingual session, the telehealth session with a client in a car, the session where a parent walks in mid-capture, the intake that becomes a crisis evaluation in minute forty. The incident mix at scale is dominated by session types the pilot never sampled, which is why the incident register stays open forever and why expansion adds session types one increment at a time.

Consent drift: in the pilot, twelve carefully briefed clinicians delivered the consent script with fidelity. At scale, consent becomes a line in the intake packet, the script gets paraphrased, new hires learn it from whoever trained them, and one office quietly treats the EHR checkbox as the consent. The scorecard's spot-check line exists because consent quality decays silently and is the single likeliest place where a scaled tool generates a board complaint.

Supervision gaps: at scale, pre-licensed associates use the tool under supervisors who were never in the pilot and may not know what reviewing AI-assisted documentation requires. Gate two evidence should include updated supervision agreements; the supervisor signs the supervision log, and an associate's unreviewed AI-drafted note is the supervisor's exposure.

Monitoring decay: the pilot's weekly huddle dies at scale by design, and what replaces it is usually nothing. The expansion gate requires the replacement to exist and to be demonstrably catching events: an incident intake channel every clinician knows, automated threshold dashboards where feasible, and a quarterly review with the same three-question discipline the huddle had. Vendor drift: models update, subprocessors change, terms get revised at renewal. The standing re-verification gate plus a contract clause requiring notice of material changes is the counter; the scorecard's regulatory section is where vendor drift surfaces.

Killing Well: The Operational Craft

Because nobody teaches it, walk through what a good kill actually looks like. The trigger: two consecutive measurement periods with the composite below floor, plus adoption fidelity at 41 percent against a 60 percent floor. The named decision-maker, the clinical director, convenes the kill review within the protocol's clock, hears the remediation history (two attempts, documented, neither moved the numbers), and issues the kill decision in writing, citing the scorecard lines. No villain is named, because the criteria were prespecified and the process was followed.

Execution follows the rehearsed plan. Capture is disabled on a stated date with two weeks' notice to clinicians and a reversion workflow that includes documentation support during the transition (template libraries, smart phrases, temporary note-time allowances), because the kill's biggest clinical risk is a documentation backlog among clinicians who structured their day around the tool. Clients are informed at their next session, in the clinician's own voice, that the tool is no longer in use, and the chart documents it; consent given for a tool that no longer operates should not silently linger in the record as if it does. The vendor exit clause is invoked: data return in a usable format, deletion certification for the rest, subprocessor confirmation, all on the contract's timeline with the compliance lead tracking each deliverable. The final act is the autopsy memo: two pages, what the tool was supposed to do, what the evidence showed at each gate, why it died, and what the next evaluation should screen for differently. The memo goes in the governance binder next to the Frontier Scan Decision Sheets, because an organization's kill history is its institutional immune memory.

And note what a good kill buys you culturally: the next pilot gets honest data, because clinicians have seen that reporting problems leads to remediation or a clean exit rather than blame; the next vendor negotiates differently, because your organization demonstrably walks; and the board trusts the next scale recommendation more, because it comes from a body with a proven willingness to say no to itself.

The Applied Problem: Build the Stage-Gate Scorecard with Off-Ramp Criteria

Your artifact is the Stage-Gate Scorecard with Off-Ramp Criteria: the one-page-per-gate instrument plus the written off-ramps, built for the tool that just exited your pilot. Build it in four steps.

Step one: define the four stages and their gates for your context. For each gate, write the evidence list: composite metric with its floor, incident rate with its ceiling, adoption fidelity with its floor, consent decline band, denial-rate condition, and the gate-specific items (skeptic inclusion at beachhead; monitoring-system demonstration, supervision-agreement updates, and Part 2 architecture verification before any excluded population enters at expansion; policy integration and full-cost budget survival at standard of work; the annual slim re-verification thereafter).

Step two: write the off-ramp criteria in three categories with actual numbers. Immediate suspension triggers (tool-generated risk determination in a chart, PHI breach outside the BAA chain, BAA-breaking vendor change, regulatory-basis loss). Threshold kill-review triggers (composite below floor two consecutive periods; incident rate above ceiling; fidelity below floor for a quarter; denial delta above baseline; consent metrics past the band). Strategic kill watch items (acquisition and data-term changes, renewal pricing versus ROI, superior alternative with passed pilot). For each, name the decision-maker role and the clock.

Step three: build the scorecard template. Five sections (clinical performance, adoption reality, consent integrity, financial, regulatory), every line with number, threshold, and computed green/amber/red, the risk-workflow zero-tolerance line always present, and the mechanical gate rule printed at the bottom: all green opens, any red holds and asks the off-ramp question, two consecutive ambers convert to red. Add the no-sunk-cost clause verbatim and the kill execution checklist: capture-off date and notice, clinician documentation support plan, client notification in the clinician's voice with chart documentation, vendor data return and deletion certification, autopsy memo to the governance binder.

Step four: the verification pass. Run last quarter's real numbers (or the pilot's final numbers) through the scorecard and confirm every status computes without judgment calls; any line that needs discussion to color is rewritten until it does not. Then rehearse the reversion plan on paper: walk one Tuesday-morning kill through capture-off, client notification, and the first week of documentation support, and fix whatever the walkthrough exposes. Done looks like this: four gates with evidence lists, off-ramps in three categories with numbers, owners, and clocks, a scorecard whose colors compute mechanically, a rehearsed kill checklist, and signatures from the clinical director, compliance lead, and billing manager filed in the governance binder before any expansion begins.

Key Takeaways

  • Scaling is canal locks, not a floodgate: four stages (pilot, beachhead, expansion, standard of work) with an evidence gate at each, plus a permanent annual re-verification gate, because a tool defensible in 2026 is not automatically defensible in 2028.
  • The beachhead exists to answer selection bias: roughly double the pilot population, deliberately including skeptics and average adopters, still excluding what the pilot excluded. Expansion adds offices, session types, and excluded populations one increment at a time, with Part 2 caseloads entering only behind verified data architecture and group work only behind a group-built consent design.
  • Off-ramp criteria are written before scaling begins, with numbers, in three categories: immediate suspension (tool-generated risk determinations, PHI breaches, BAA-breaking vendor changes, regulatory-basis loss), threshold-triggered kill review (composite below floor, incident ceiling, fidelity floor, denial delta, consent band), and strategic kill (acquisitions, pricing, superior alternatives). Each has a named decision-maker and a clock.
  • The scorecard computes, it does not negotiate: every line has a number, a prespecified threshold, and a green/amber/red status; all green opens the gate, any red holds it, two consecutive ambers convert to red. The risk-workflow line (tool-generated risk language, target zero) is always present, because AI never scores the CSSRS, never assigns risk levels, never makes protective determinations at any scale.
  • Scale breaks what pilots cannot see: the long tail of session types, consent drift toward a checkbox, supervision gaps around associates whose supervisors never piloted, monitoring decay after the huddle dies, and vendor drift in models and subprocessors. The scorecard's lines exist to watch for exactly these.
  • A good kill is governance working, not failure: prespecified criteria depersonalize the decision, the no-sunk-cost clause keeps spent money out of it, the rehearsed reversion plan protects clinicians from a documentation backlog, clients are told in the clinician's voice, the vendor exit clause extracts data return and deletion certification, and a two-page autopsy memo becomes institutional immune memory.
  • Killing well pays forward: honest incident reporting in the next pilot, stronger vendor negotiating position, and board trust in the next scale recommendation, because the governance body has proven it can say no to itself.