AI for Mental & Behavioral Health Clinicians
Strategic · M12 · lesson 12 of 23 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Designing a Clinical Pilot for an AI Scribe
📖
now learning

Designing a Clinical Pilot for an AI Scribe

15 min

The vendor passed your 12-question RFI. The BAA covers your tier, the subprocessor list is clean, the SOC 2 Type II is current, and the demo, of course, was lovely. This is the moment most practices sign, and it is exactly the moment Jordan almost did, until the compliance officer asked one question nobody in the room could answer: "How will we know, thirty days from now, whether this thing actually worked?" Not felt good. Worked. This lesson is the answer: a 30-day, 5-clinician pilot with paired metrics (time-per-note, clinician satisfaction NPS, documentation error rate, client opt-out rate, claim denial rate), each measured at baseline before the scribe arrives and again during the pilot, with go/no-go thresholds named in writing before the first session is recorded. You will finish with the AI Scribe Pilot Protocol, a document that turns the purchase decision into something you would recognize from your own clinical training: a measurement-based decision made on data, not vibes.

Treat the Pilot Like Measurement-Based Care, Because That Is What It Is

Hold one analogy for this entire lesson: the pilot is measurement-based care for your practice, and the AI scribe is the intervention. You already know this protocol clinically. You would never start a client on a new treatment, skip the baseline PHQ-9, run twelve sessions, and then decide whether it worked by asking the client "so, vibes?" You administer the measure before treatment, you re-administer on a schedule, you define in advance what response and remission look like, and you change course when the data says the intervention is not working, even if everyone in the room likes each other. Every discipline you have about outcome measurement transfers directly to vendor evaluation, and every shortcut you would never tolerate in clinical care (no baseline, no fixed measure, a decision criterion invented after the results are in) is precisely the shortcut practices take when piloting software.

The reason vibes-based pilots fail is not that feelings are worthless; it is that feelings are systematically biased in a known direction. A pilot has novelty energy. The vendor's customer success team is attentive in a way it will never be again. The clinicians who volunteered are the enthusiasts. The notes feel faster because attention is on them. Every one of these forces inflates the perceived benefit during exactly the window in which you are deciding. The paired-metric design exists to put a measurement floor under that inflation: numbers collected the same way before and during, compared directly, with the decision rule fixed before anyone could be tempted to move it. In clinical language, you are pre-registering the trial.

One scoping note before the design work: a pilot presumes the diligence is done. The BAA is signed for the pilot period, the consent addendum is in place for participating clients, and the data path passed your RFI gates. A pilot is never a substitute for diligence, because no 30-day trial reveals whether a subprocessor signs a BAA at your tier. The pilot tests the one thing the paperwork cannot: whether the product, in your practice, with your clinicians and your payers, produces the outcomes the vendor's case studies promised.

The 30-Day, 5-Clinician Frame: Why These Numbers

Thirty days is long enough to outlast the novelty arc and short enough to keep the practice's attention. Week one of any tool is unrepresentative in both directions: setup friction makes it look worse than it is, and novelty enthusiasm makes it feel better than it is. By weeks three and four, the clinicians have settled into the workflow they would actually live with, the early bugs have either been fixed or revealed the support model, and at least one full billing cycle has begun moving claims that carry pilot-period documentation. Shorter pilots measure the honeymoon. Much longer pilots without a decision point become de facto adoptions, the software equivalent of a treatment that drifts on without a review because nobody scheduled one.

Five clinicians is the smallest number that gives you variance instead of an anecdote. Choose them deliberately, the way you would compose any sample you intend to learn from: not five enthusiasts. The composition that produces honest data is roughly two volunteers who want the tool, two neutral clinicians chosen for representative caseloads, and one constructive skeptic, the senior clinician whose objections are clinical rather than reflexive. The skeptic matters twice over: their data is the stress test, and their eventual verdict is the most credible artifact your rollout communication will ever have. If the skeptic's time-per-note dropped and their error audit came back clean, the rest of the practice will believe the pilot. If only the enthusiasts improved, you have learned that the tool serves enthusiasts, which is worth knowing before you buy twenty-five licenses.

Caseload composition matters as much as clinician composition. A pilot that only touches uncomplicated adult anxiety clients is the vendor demo wearing your letterhead. The participating caseloads should include the sessions your practice actually documents: high-frequency 90837s, couples sessions, clients with risk content, and, if your practice holds them, the records where 42 CFR Part 2 applies, handled under the protections your RFI already verified. Clients must consent through your AI addendum before their sessions touch the scribe, and the opt-out rate you are about to measure only means something if clients were actually given the choice.

The Five Paired Metrics: What You Measure and How

Paired means every metric is collected twice with the same method: a baseline window before the scribe (use the two weeks before launch, or pull the prior month from your EHR where the data already exists) and the pilot window itself. The comparison is each clinician against their own baseline, which controls for the fact that your five pilots have different caseloads, different documentation habits, and different starting speeds. Here are the five, with the collection method that makes each one real.

Metric one: time-per-note. The headline claim of every scribe is time, so measure it rigorously: each pilot clinician logs minutes from session end to signed note for every note in both windows, using a shared log sheet rather than memory. Self-timing is imperfect, but it is imperfect the same way in both windows, which is what pairing is for. Capture the distribution, not just the average; a tool that saves twenty minutes on routine notes but adds thirty to complex ones has a shape your average will hide. Metric two: clinician satisfaction NPS. Ask the standard question (how likely are you to recommend this tool to a colleague, 0 to 10) at baseline about the current documentation workflow, then at day 10, day 20, and day 30 about the scribe workflow. The trajectory matters more than any single reading: satisfaction that climbs as novelty fades is signal; satisfaction that peaks at day 10 and slides is the honeymoon ending on schedule.

Metric three: documentation error rate, the metric vendors never volunteer. Define an error audit before launch: a supervisor or designated reviewer samples a fixed number of notes per clinician per window (ten is workable) and counts defined defects: factual inaccuracies against the clinician's recollection or recording, omissions of clinically material content, insertions of content that did not occur (the hallucinated "client denied suicidal ideation" is the canonical catastrophic example), and template language that fails the payer-detail standard, such as a 90837 note missing time-in-session minutes or the named modality. Baseline notes get the same audit, because human 9:54 PM notes have error rates too, and the honest question is whether the scribe's errors are fewer, different, and more dangerous or less. Metric four: client opt-out rate. Track the percentage of clients asked who decline the scribe, and have clinicians note the stated reason. Opt-out is your client-acceptability thermometer, and a rising rate or a cluster of trauma clients declining tells you something the other four metrics cannot. Metric five: claim denial rate. Pull your denial rate for pilot clinicians' claims at baseline from the prior period and track pilot-period claims as adjudication returns. Thirty days will not give you complete adjudication, which is fine: the protocol names denial rate as a trailing metric with a 60-day look-back checkpoint, and the go/no-go at day 30 weighs the four leading metrics with denial data marked preliminary.

You would never judge a treatment without a baseline measure and a pre-defined response criterion. Give the vendor's product exactly the scrutiny you give your own interventions: paired data, locked thresholds, and a decision the day the data arrives.

Pre-Naming the Go/No-Go Thresholds: The Pre-Registration Step

Here is the step that separates a pilot from a prolonged demo: before day one, the practice writes down the numbers that will decide, and signs them. This is the pre-registration discipline from clinical research, and it exists because human beings, including conscientious practice owners, will rationalize any result they are allowed to interpret after the fact. A 12 percent time saving will be read as "promising" by whoever liked the demo and "marginal" by whoever did not, unless the protocol already said what counts.

Write three bands per metric: Go, Conditional, and No-Go. A workable starting set, which you should adjust to your practice's economics before adopting: time-per-note must improve by at least 30 percent on average with no clinician worse than baseline (Go), 15 to 30 percent triggers Conditional, under 15 percent is No-Go, because a marginal saving will not survive the loss of novelty attention. Clinician NPS at day 30 must average 7 or higher with no participant below 5 for Go; any participant at 3 or below triggers a structured exit interview before any decision. Error rate is the asymmetric one: the audited error rate must be at or below baseline, and a single uncorrected fabricated clinical statement in a signed note, or any error touching risk content, is an automatic No-Go regardless of every other metric, because no time saving prices in a fabricated risk assertion. Client opt-out above 20 percent triggers Conditional with a consent-language review; a pattern of opt-outs concentrated in trauma or adolescent caseloads triggers clinical review regardless of the rate. Denial rate must show no increase at the 60-day look-back, with any new denial citing documentation quality investigated individually.

Notice what the asymmetry encodes: time and satisfaction are negotiable in degree; fabrication touching clinical content is not negotiable at all. That is the clinical spine showing through the procurement process, and it should. The thresholds go into the protocol document with the date and the signatures of the owner, the clinical director, and the compliance officer, for the same reason the RFI's gate rule went above the scoring table: so that a future meeting, under a vendor discount deadline, cannot quietly soften what the practice decided when it was thinking clearly.

Running the 30 Days: Cadence, Roles, and the Things That Go Wrong

A pilot needs an owner, and it should not be the most enthusiastic person in the building. Name a pilot coordinator (often the compliance officer or practice manager) who collects the logs weekly, runs the NPS pulses, schedules the error audits, and owns the day-30 meeting. The cadence is simple: a 30-minute launch huddle to train the five on logging (not just on the tool; the vendor trains the tool, you train the measurement); weekly 15-minute check-ins where the coordinator collects data and surfaces friction; the day-10 and day-20 NPS pulses; error audits at mid-point and end; and the day-30 decision meeting with the protocol document open on the table.

Plan for the predictable failure modes, because they are predictable. Logging fatigue: clinicians stop recording time-per-note around day 8; the coordinator's weekly collection and a friction-free log (a two-column sheet or a form that takes ten seconds) are the countermeasures, and a clinician with fewer than 60 percent of notes logged is excluded from the time metric rather than guessed at. Vendor intervention: the customer success team, sensing a structured pilot, offers extra training, custom templates, and weekly calls; accept what any customer would get at your tier and decline the rest, because you are piloting the product you would buy, not the white-glove version that disappears after signature. Cross-contamination: non-pilot clinicians start asking for access mid-pilot; hold the line, since an uncontrolled rollout destroys the comparison and, more practically, those clinicians have no baseline. Mid-pilot product updates: note the date of any significant update in the log, because a tool that changes materially in week three has effectively reset its own trial, and the day-30 meeting should know it.

And the failure mode that matters most: a serious incident. If the error audit, or any clinician, surfaces a fabricated risk statement, a Part 2 handling failure, or a consent breach mid-pilot, the protocol's incident clause activates: the affected note is corrected immediately under your documentation-amendment procedure, the incident is reported through the vendor's clinical escalation path (the one your RFI question ten verified exists), and the pilot pauses for a 48-hour review rather than coasting to day 30. A pilot is a controlled exposure, and controlled means you can stop it.

The Day-30 Meeting: Deciding on Data, Not Vibes

The decision meeting has a fixed agenda because unstructured decision meetings revert to vibes within minutes. First, the coordinator presents the paired numbers, metric by metric, against the pre-named bands, with no commentary: time-per-note distributions per clinician, NPS trajectory, error audit results with every defect itemized, opt-out rate and reasons, and the preliminary denial picture flagged as trailing. Second, the room classifies each metric into its band, Go, Conditional, or No-Go, by reading the protocol, not by debating it. Third, the decision rule executes: all five in Go bands (with denial pending) is a Go, proceeding to phased rollout with the 60-day denial checkpoint as a condition subsequent. Any automatic No-Go trigger, fabrication touching clinical content above all, ends the engagement regardless of the other numbers. A mixed picture lands in Conditional, which is not a soft yes; it is a named list of deficiencies, a written remediation request to the vendor, and a defined re-test window of two to four weeks on the specific failing metric, after which the bands apply again. One Conditional cycle, not an endless one; a vendor that cannot clear the bar on the second pass has answered the question.

Document the decision the way you would document a clinical decision: the numbers, the bands, the outcome, the signatures, and, if the answer is no, the specific data behind it. The no-decision memo has real value. It protects the practice when the same vendor returns next year with a new version (you have a baseline of what failed), it gives your malpractice carrier and any future auditor evidence of governed adoption, and it tells your clinicians, who watched five colleagues spend a month on this, that the practice's measurement culture is real. The strongest message a leadership team can send about measurement-based care is to visibly subject its own purchasing decisions to it.

One senior-supervisor reframe to close: a no-go pilot is not a failed pilot. The pilot succeeded if it produced a defensible decision, in either direction, for the cost of thirty days and five clinicians' patience. The failed pilot is the one that ends in a shrug and a signature, where the practice spent the month and bought the vibes anyway. The protocol you are about to write exists to make that ending impossible.

The Applied Problem: Write the AI Scribe Pilot Protocol

Your artifact is the AI Scribe Pilot Protocol: a signed, dated document that any practice could hand to a new compliance officer and run without you in the room. Build it in four steps.

Step one, the frame page: the product and tier under pilot (carried over from the RFI), the 30-day window with dates, the five named clinicians with their selection rationale (two volunteers, two representative, one constructive skeptic), the caseload-inclusion statement (routine, complex, risk-content, and Part 2 sessions as applicable, all under signed consent addenda), and the pilot coordinator by name. Step two, the measurement section: for each of the five paired metrics, write its definition, its baseline source (two-week log or prior-month EHR pull), its collection method and cadence (per-note time logs collected weekly; NPS at baseline, day 10, 20, 30; ten-note error audits at mid-point and end with the defect taxonomy listed, including the fabricated-content category; opt-out tracking with stated reasons; denial rate with the 60-day trailing checkpoint).

Step three, the decision section, which is the heart: a table of the five metrics with Go, Conditional, and No-Go bands filled in with your practice's numbers, the automatic No-Go triggers written in bold prose (any uncorrected fabricated clinical statement in a signed note; any error touching risk content; any Part 2 or consent breach), the Conditional procedure (named deficiencies, written remediation request, one re-test window of two to four weeks), the incident clause (immediate note correction, vendor clinical escalation, 48-hour pause), and the day-30 meeting agenda in its fixed order. Then the signature block: owner, clinical director, compliance officer, dated before launch.

Step four, the verification pass. Read the protocol as the skeptic clinician would: is any metric collectible only by heroic effort (fix the log), is any band vague enough to argue about later (replace adjectives with numbers), and could the day-30 meeting reach a decision from this document alone if every author were on vacation? Then run the pre-registration test: hand the unsigned protocol to someone who has not seen the vendor demo and ask them to state, from the document, exactly what result ends the engagement. If they can say it in one sentence ("a fabricated clinical statement, or less than 15 percent time saving"), the protocol is done. Sign it, date it, and only then schedule the launch huddle.

Key Takeaways

  • The pilot is measurement-based care applied to your own purchasing: baseline before the intervention, the same measures during, response criteria defined in advance, and a decision made on the data. Every shortcut you would never tolerate clinically (no baseline, post-hoc criteria) is the shortcut that makes software pilots worthless.
  • The 30-day, 5-clinician frame is deliberate: thirty days outlasts the novelty arc that inflates week-one impressions, and five clinicians, composed as two volunteers, two representative caseloads, and one constructive skeptic, produce variance instead of an anecdote. The skeptic's clean data is the most credible rollout asset you will ever have.
  • The five paired metrics are time-per-note (per-note logs, distribution not just average), clinician satisfaction NPS (baseline, day 10, 20, 30, trajectory over any single reading), documentation error rate (ten-note audits against a defect taxonomy including fabricated content), client opt-out rate (with stated reasons), and claim denial rate (a trailing metric with a 60-day look-back checkpoint).
  • Paired means each clinician is compared against their own baseline with the same collection method in both windows, which controls for caseload and habit differences and makes self-timing imperfections cancel out.
  • Go/no-go thresholds are pre-named, banded (Go, Conditional, No-Go), and signed before launch, with asymmetric rules: time and satisfaction are negotiable in degree, but a fabricated clinical statement in a signed note or any error touching risk content is an automatic No-Go no other metric can offset.
  • Run the pilot with a named coordinator, weekly collection, mid-point and final audits, and countermeasures for the predictable failures: logging fatigue, vendor white-glove inflation, cross-contamination from non-pilot clinicians, and mid-pilot product updates that reset the trial. A serious incident (fabricated risk content, Part 2 or consent breach) pauses the pilot under the incident clause; controlled exposure means you can stop it.
  • The day-30 meeting reads the protocol instead of debating it: numbers, bands, decision. Conditional is one remediation-and-re-test cycle, not an endless one, and a no-go pilot that produces a documented, defensible decision is a successful pilot. The failure is the shrug and the signature.