Running Defensible Clinical AI Pilots
Somewhere in your organization's recent history, there is a "pilot" that was actually a demo: a vendor configured the tool, picked the friendliest clinicians, ran it for six weeks with no control group, no predefined outcomes, and no stopping rules, and then a slide deck announced that documentation time fell 40 percent and clinician satisfaction was high. That deck became the evidence for a three-year contract, and nobody can now say what the tool actually does to a real caseload, because nothing about the exercise was designed to find out. This lesson teaches you to run a clinical AI pilot with IRB-equivalent rigor: a written hypothesis, a comparison condition, primary and secondary outcomes named before launch, predefined stopping rules that protect clients, and a protocol clean enough to publish. By the end you will have built the Pilot Protocol with Stopping Rules, the artifact that separates an organization that tests AI from one that merely buys it slowly.
The Demo That Wore a Lab Coat
Start with the story, because every operator has lived a version of it. Jordan's group practice in Sacramento agreed to "pilot" an ambient scribe with twelve volunteer clinicians. The vendor ran the onboarding, chose the configuration, and checked in weekly. Six weeks later the vendor presented results: average note time down from 18 minutes to 9, satisfaction 4.6 out of 5, zero reported incidents. Jordan's board approved the enterprise contract that month. Eight months later the picture looked different: the volunteers had been the practice's most tech-comfortable clinicians, note time for the full workforce settled around 14 minutes, two clinicians had quietly stopped using the tool after hallucinated content appeared in drafts, no one had measured note quality or audit-readiness at all, and the one metric the payers cared about, documentation that survives a 90837 medical-necessity review, had never been defined as an outcome. The pilot had not failed. There had never been a pilot. There had been a demo wearing a lab coat.
Carry this lesson's controlling analogy from here: a defensible pilot is a clinical trial, not a test drive. On a test drive, the dealer picks the route, the weather is good, and the question being answered is "do you like it?" In a trial, the protocol is written before the first participant enrolls, the outcomes are named in advance, the comparison condition exists precisely because human enthusiasm is a confound, and there are rules, written down, for stopping the whole thing if someone is being harmed. Nobody would approve a new clinical intervention for their caseload on test-drive evidence. An AI tool that touches session audio, clinical documentation, or risk-adjacent workflows is a clinical intervention in everything but billing code, and it deserves trial-grade discipline.
The phrase to internalize is IRB-equivalent rigor. Most operational pilots are not human-subjects research requiring actual IRB review, and this lesson is not a substitute for legal advice on when review is required (if you intend to publish generalizable findings, or your setting is academic or grant-funded, consult your IRB before launch). IRB-equivalent means you voluntarily impose the discipline an IRB would demand: a protocol, defined risks, consent handling, data protections, predefined endpoints, and a stopping mechanism, because the people in your pilot, clients and clinicians both, deserve those protections whether or not a committee requires them.
The Hypothesis Comes First, in Writing, Before the Vendor Call
A trial begins with a falsifiable hypothesis, and so does a defensible pilot. Not "we want to try the tool," but a sentence with a direction and a number: "Among clinicians using the ambient scribe for outpatient individual therapy, average documentation time per session will decrease by at least 30 percent relative to the comparison group, without a decline in note-quality audit scores." Write it before the vendor configures anything, because the vendor's instinct is to define success as adoption, and adoption is not a clinical outcome.
A workable hypothesis forces three definitions. First, the population: which clinicians, which session types, which client populations are in scope, and, critically, which are excluded. Most pilots should exclude the highest-risk contexts at the start: sessions involving active suicidality follow the organization's existing risk protocols without the new variable, 42 CFR Part 2 records stay out until the data architecture is verified, couples and group sessions wait for consent designs built for them. Exclusion is not timidity; it is the same logic that keeps a phase-one drug trial out of the ICU. Second, the intervention: the exact tool, the exact configuration, frozen for the pilot's duration. A vendor who pushes a model update mid-pilot has changed the intervention, and the protocol should say what happens when they do (log it, and either restart the measurement window or treat it as a protocol deviation). Third, the comparison: against what?
The comparison condition is where most pilots quietly die. A before-and-after measurement on the same volunteers confounds the tool with enthusiasm, novelty, and selection: the clinicians who volunteer are the ones predisposed to succeed with it. The defensible designs, in ascending order of effort: a concurrent comparison group of similar clinicians who continue usual documentation practice while the pilot group uses the tool; a stepped-wedge rollout where clinicians cross from usual practice to the tool on a schedule, so each cohort serves as comparison for the others; or, simplest and still far better than nothing, a matched historical baseline using each pilot clinician's own prior-quarter metrics, with the explicit caveat in the protocol that this design cannot separate the tool from secular trends. Pick one, name it in the protocol, and never let the vendor talk you down to "we will survey the users."
Outcomes: Primary, Secondary, and the Ones That Protect People
A trial names its primary outcome before enrollment because outcomes chosen afterward are chosen to flatter. Your pilot does the same. The primary outcome is the single measure on which the scale-or-kill decision will turn, and for most documentation-AI pilots it should be a composite the payers and the board both respect: documentation time per session AND note quality, because time savings purchased with thin notes is a recoupment subsidy. Note quality needs an instrument, not an impression: a structured audit rubric scoring whether each note contains the verifiable detail AI cannot supply, the elements you have carried through this program, such as time-in-session minutes supporting a 90837, the specific PHQ-9 delta, the modality actually named in session, the risk documentation reflecting the clinician's own determination. Two trained reviewers score a random sample of notes from both arms, blinded to arm where feasible, at baseline and at close.
Secondary outcomes capture what else matters: clinician burnout (a brief validated measure at start and end), same-day note closure rate, claim denial rate on piloted sessions, client consent decline rate and any client complaints, and adoption fidelity (what fraction of eligible sessions actually used the tool, because a tool used in 30 percent of sessions has told you something the satisfaction survey will not). Define each measure's source and collection cadence in the protocol. If a measure cannot be collected without heroics, drop it now rather than pretend later.
Then come the safety outcomes, and this is where a clinical AI pilot differs most from an IT rollout. Predefine the events you will count as harms and near-misses: hallucinated clinical content found in a draft (an invented score, diagnosis, or event), a consent failure (capture without documented consent, or capture continuing after revocation), a confidentiality event (PHI routed outside the BAA chain), a risk-workflow intrusion (any instance of the tool generating risk language, scores, or determinations rather than formatting the clinician's own), and any client-reported distress attributable to the tool's presence. Every event gets logged in a pilot incident register with date, description, severity, and disposition. The register is not bureaucracy; it is the data that feeds the stopping rules.
A pilot without predefined stopping rules is not an experiment; it is exposure with a start date and no brakes.
Stopping Rules: The Brakes You Install Before You Drive
Stopping rules are the heart of IRB-equivalent rigor, and they must be written, numeric where possible, and assigned to a named decision-maker before launch. Three tiers work for most behavioral health pilots.
Tier one, immediate halt, no meeting required: any single event in which the tool generated a risk determination, score, or safety-relevant clinical conclusion that reached a chart or influenced a clinical decision; any confidentiality breach involving PHI outside the BAA chain; any capture of a session without valid consent. One event stops the pilot the same day, pending investigation. The protocol names who can pull this cord (any pilot clinician, the pilot lead, the compliance lead) and states explicitly that pulling it carries no penalty, because a stopping rule nobody dares invoke is decoration.
Tier two, threshold-triggered review within five business days: hallucinated clinical content found in more than a defined fraction of audited drafts (set the number in advance, for example 5 percent of sampled notes); client consent declines exceeding an expected band, suggesting the consent process is misfiring; note-quality audit scores falling below baseline in the pilot arm; more than a set number of clinicians abandoning the tool. Crossing a threshold convenes the pilot governance group, which must either remediate with a documented change or stop.
Tier three, scheduled interim look: at the midpoint, a planned review of all outcomes and the incident register, with the explicit authority to stop early for futility (the primary outcome is clearly not moving and the remaining weeks cannot change that) as well as for harm. Futility stopping is the discipline most organizations lack: they run failing pilots to the end because the calendar says so, burning clinician goodwill on a tool the data already condemned.
Write each rule with its number, its trigger source (the incident register, the audit sample, the consent log), its decision-maker, and its clock. Then rehearse the tier-one halt once before launch, the way you rehearse a fire drill, so the first real pull of the cord is not also the first test of whether anyone knows what happens next.
Consent, Data, and the People Inside the Pilot
Clients in a pilot are not research subjects in the regulatory sense, but they are people whose sessions are being processed by a system your organization has not yet decided to trust, and the consent process must say so plainly. The pilot consent addendum, in the clinician's own voice at session one, covers: what is captured and by what tool, that the tool is being evaluated (not established practice), what happens to audio and drafts, the BAA protections in place, that declining changes nothing about the care, and how to revoke mid-episode. Track declines as data, not as friction: a high decline rate is a finding about the tool's acceptability, and the protocol should say in advance what decline rate triggers a tier-two review.
Clinicians inside the pilot need their own protections, stated in writing: participation is voluntary; pilot metrics will not be used for individual performance evaluation; the pre-signature review obligation is unchanged (the clinician signs the note, and the signature is a legal attestation, not a formatting step); and any clinician can invoke the tier-one halt without penalty. Supervisors of pre-licensed associates in the pilot countersign this understanding, because the supervisee's AI-assisted note lands on the supervisor's license, and a pilot that quietly conscripts associates without their supervisors' documented agreement has created exactly the exposure this program has warned about since Level 1.
Data handling gets its own protocol section: where audio lives and for how long, who at the vendor can access it, the subprocessor list as of launch date (attached as an exhibit, because subprocessor drift mid-pilot is a protocol deviation), confirmation that pilot data is excluded from model training or the contractual basis for any exception, and the deletion certification you will demand at pilot end for non-retained data. If any piloted caseload touches SUD records, 42 CFR Part 2 architecture is verified before launch or those records are excluded; the 2024 final rule's consent and redisclosure requirements do not pause for innovation.
Running the Weeks: Fidelity, Logging, and the Vendor at Arm's Length
Execution is mostly the discipline of writing things down. The pilot lead keeps a running log: enrollment counts, sessions captured versus eligible, protocol deviations (a clinician using the tool on an excluded session type, a vendor-pushed update, a missed audit window), and every incident-register entry. The log is boring on purpose. Its value appears at week nine, when someone asks why adoption dipped in week four, and the log shows the vendor shipped an interface change then.
Keep the vendor at arm's length from the measurement. The vendor onboards, trains, and supports; the vendor does not collect your outcomes, does not run your satisfaction survey, does not draft your interim report. Vendor-collected evidence is the demo-as-evidence trap in its purest form: the entity with the strongest interest in a positive result controls the instrument. Share the protocol with the vendor before launch (a serious vendor will respect it; a vendor who pushes back on stopping rules or independent measurement has told you something important), but the data pipeline runs through your team.
Hold the weekly fifteen-minute pilot huddle: pilot lead, one clinician representative, compliance lead. Three questions only: any incident-register entries this week, any threshold approaching, any deviation to log. This cadence is what makes the stopping rules real rather than theoretical, because thresholds are only protective if someone is actually watching the numbers against them.
The Publication-Ready Standard, Even If You Never Publish
Hold the protocol to a publication-ready standard: written so that a peer organization could replicate the pilot from the document alone. This is not vanity. The publication-ready test is a forcing function for every weakness pilots hide: vague populations, undefined outcomes, missing comparison conditions, stopping rules that live in someone's head. If the methods section would embarrass you in front of a journal reviewer, it should embarrass you in front of your board, because the board is making a bigger bet on it than any journal would.
The protocol document, in order: background and hypothesis; population with inclusion and exclusion criteria; intervention and configuration (frozen, versioned); comparison condition and its acknowledged limits; primary outcome with its instrument; secondary and safety outcomes with sources and cadence; consent procedures for clients and clinicians; data handling, BAA chain, and subprocessor exhibit; stopping rules in three tiers with numbers, owners, and clocks; analysis plan (who computes what, when, and the pre-stated decision criteria for scale, extend, or kill); timeline and roles. Ten to fifteen pages. Signed by the pilot lead, the clinical director, and the compliance lead before the first session is captured.
The analysis plan deserves one more sentence of emphasis: state the decision criteria before the data exists. "We will recommend scaling if the primary composite improves by the prespecified margin with no unresolved tier-one events and tier-two thresholds unbreached; we will recommend killing if the primary outcome fails or any tier-one event lacks a verified remediation; anything between goes to a defined extension with new criteria, not an indefinite limbo." The next lesson builds the stage-gate machinery that consumes this output; your job here is to produce an output worth consuming.
The Applied Problem: Build the Pilot Protocol with Stopping Rules
Your artifact is the Pilot Protocol with Stopping Rules, built for a real candidate tool, ideally one that survived your Frontier Scan Decision Sheet from the previous lesson. Work through it in four passes.
Pass one: the science skeleton. Write the falsifiable hypothesis with a direction and a number. Define the population with explicit exclusions (active-risk sessions, Part 2 records pending architecture verification, couples and group pending consent design). Choose and name the comparison condition, and write one honest sentence about its limits. Name the primary outcome as a composite (time AND audited note quality) and build or adopt the audit rubric, including the verifiable-detail items: session minutes supporting the billed code, instrument deltas, modality named, risk documentation reflecting the clinician's own determination.
Pass two: the safety architecture. Draft the incident-register template (date, event type, severity, description, disposition). Write the three tiers of stopping rules with actual numbers: tier one's same-day halt triggers, tier two's thresholds (for example, hallucinated content in more than 5 percent of audited drafts; consent declines above your prespecified band), tier three's midpoint interim look with futility authority. Assign a named decision-maker and a clock to each tier, and write the no-penalty clause for anyone who pulls the tier-one cord.
Pass three: the people and the data. Draft the client pilot-consent language in a clinician's voice, the clinician participation memo (voluntary, no performance use, signature obligations unchanged, halt authority), the supervisor countersignature line for any associate participants, and the data-handling section with the subprocessor exhibit and the end-of-pilot deletion certification requirement.
Pass four: the verification pass. Read the whole protocol against the publication-ready test: could a peer organization replicate this from the document alone? Then run the adversarial read: hand it to your most skeptical clinician and ask them to find the demo hiding inside it, any place where the vendor controls measurement, where an outcome is undefined, where a stopping rule has no number or no owner. Done looks like this: a ten-to-fifteen page document, every outcome named before launch, every stopping rule numeric with a named owner, three signatures at the bottom, and a scheduled fire-drill rehearsal of the tier-one halt on the calendar before the first session is captured.
Key Takeaways
- The demo-as-evidence trap is the default failure mode of clinical AI adoption: vendor-run, volunteer-staffed, no comparison condition, no predefined outcomes, no stopping rules, and a slide deck mistaken for data. A defensible pilot is a clinical trial, not a test drive.
- IRB-equivalent rigor means voluntarily imposing trial discipline (protocol, defined risks, consent handling, predefined endpoints, stopping mechanism) even when no IRB review is legally required; consult your IRB if you intend to publish or operate in an academic or grant-funded setting.
- The hypothesis comes first, in writing, with a direction and a number, before the vendor configures anything. It forces definition of population (with explicit exclusions for active-risk sessions, Part 2 records, and group or couples work), a frozen intervention, and a named comparison condition.
- The primary outcome should be a composite of documentation time AND audited note quality, scored on a rubric that checks the verifiable detail AI cannot supply: session minutes for the billed code, instrument deltas, the modality named in session, and risk documentation reflecting the clinician's own determination.
- Stopping rules come in three tiers: same-day halt for any tool-generated risk determination, consent failure, or PHI breach; threshold-triggered five-day review for hallucination rates, consent declines, or quality drops; and a midpoint interim look with the authority to stop for futility, not just harm. Every rule has a number, an owner, a clock, and a no-penalty clause.
- Clients get pilot-specific consent in the clinician's own voice with a mid-episode revocation path; clinicians get written protections (voluntary participation, no performance use, unchanged signature obligations); supervisors countersign for any pre-licensed associate, because the associate's AI-assisted note lands on the supervisor's license.
- Keep the vendor at arm's length from measurement and hold the protocol to a publication-ready standard: if the methods would embarrass you before a journal reviewer, they should embarrass you before the board, which is betting more on them than any journal would.
Skill.re