Running Pharmacy AI Pilots to Evidence
A regional specialty pharmacy ran what its leadership proudly called a successful AI pilot, and the slide deck was glowing: prior-authorization turnaround down forty percent, technicians thrilled, the vendor delighted. Then the chief pharmacy officer asked three questions that turned the celebration into a quiet, uncomfortable pause. How many of the AI-assembled justifications did a pharmacist actually verify before submission, and did anyone count the criteria the AI got wrong? What was the turnaround before the pilot, measured the same way, and how do you know the improvement is the AI and not the two extra technicians you added the same month? And when the pilot ends, what exactly are we now confident is true, written down, that a board or an accreditor could read? The room had no answers, because the pilot had measured enthusiasm, not evidence. It had proven that people liked the tool, which is not the same as proving the tool was fast, safe, and ready to scale. This lesson is about the difference, because at the enterprise level a pilot that produces a good feeling and no evidence is worse than no pilot at all: it spends real money and real patient exposure to manufacture false confidence, and false confidence is exactly what scales a hidden safety problem across an entire network.
The Difference Between a Pilot and a Demo
The first thing a transformer has to fix is a confusion that runs through most pharmacy AI efforts: the difference between a demo and a pilot. A demo answers the question "can this tool do the thing at all," and it is answered in a vendor's controlled environment with cherry-picked examples; it is theater, useful only for deciding whether something is worth a real test. A pilot answers a far harder question: "does this tool produce real value, safely, in our actual operation, with our actual staff, on our actual patients, measured against what we did before." The two are constantly conflated, and the conflation is dangerous, because a successful demo creates pressure to skip the pilot and scale straight to the network, which is precisely how an unverified tool reaches patients at scale. A transformer's job is to insist, every time, that an impressive demo earns a disciplined pilot and nothing more, and that scaling is earned by pilot evidence, not by demo enthusiasm.
The reason this matters so much in pharmacy specifically is the patient-safety asymmetry that governs the whole program. In a low-stakes domain, a sloppy pilot that overstates a tool's value costs some wasted effort when the tool underdelivers at scale. In pharmacy, a sloppy pilot that fails to measure the tool's error rate, the criteria it fabricated, the doses it misread, the interactions it missed, can scale a hidden safety defect across every site, where it reaches patients before anyone has the evidence to catch it. The discipline of a real pilot is therefore not bureaucratic caution; it is the safety control that stands between an enthusiastic demo and a network-wide patient-safety event. A transformer who treats the pilot as a formality to rush through has misunderstood its entire purpose, which is to find the tool's failure modes deliberately, in a contained setting, before they can find a patient at scale.
A demo proves a tool can work in a vendor's hands. A pilot proves it works safely in yours, measured against what you did before. Scaling is earned by pilot evidence, never by demo enthusiasm.
Designing a Pilot That Produces Evidence
A pilot produces evidence only if it is designed to before it starts, and the design has a small number of non-negotiable parts. It needs a baseline: the same metric, measured the same way, before the AI is introduced, because an improvement you cannot compare to a starting point is a number, not evidence. The specialty pharmacy's forty-percent improvement was meaningless precisely because no one had measured the before, the same way, on the same population. It needs a defined scope: a specific workflow, at a specific site or sites, for a specific time, so that what the pilot proves is bounded and clear rather than a vague impression. It needs success criteria written down in advance, the specific thresholds on the specific metrics that would make the tool worth scaling, set before the data arrives so that the pilot tests a hypothesis rather than rationalizing whatever happened. And it needs a measurement plan that captures both sides of the ledger, the value and the risk, because a pilot that measures only the time saved and not the errors introduced is measuring the easy half and ignoring the half that can hurt a patient.
That last point is where pharmacy pilots most often fail, so it deserves emphasis. It is natural to design a pilot around the exciting metric, the turnaround that dropped, the hours that were freed, and to forget that in a clinical setting the load-bearing measurement is the error rate. A prior-authorization pilot must measure not only how much faster the assembly became but how often the AI fabricated a criterion, misstated a clinical fact, or matched the request to the wrong payer rule, because those are the failures that turn a fast PA into a denial, a delay, or a compliance problem. A verification-support pilot must measure not only how many signals the tool surfaced but how many were wrong, and how many a pharmacist would have caught anyway. The whole point of a safe pilot is to make the failure modes visible and counted, so that the decision to scale is made with the risk quantified rather than assumed away. A transformer designs the measurement plan so that the tool cannot pass the pilot by being fast while being quietly unsafe.
The Safety Gate Built Into the Pilot
Because a pilot in pharmacy runs on real patients, it cannot be a free-running experiment; it has to carry its own safety gate, a structural guarantee that nothing the AI produces reaches a patient without a competent human verifying it first. This is the cardinal rule, AI supports the pharmacist's judgment and never replaces it, operationalized as a pilot design constraint. During the pilot, every AI-assembled justification is verified by a pharmacist before submission, every AI-surfaced clinical signal is treated as a prompt to think and not a verdict to accept, and every counseling draft is fact-checked before a patient hears it. The safety gate is what allows the pilot to look for failure modes without those failures reaching patients, because the verification step catches them and, critically, records them as data. The errors the pharmacists catch during the pilot are not an embarrassment to hide; they are the most valuable output the pilot produces, because they are the measured error rate that tells you whether the tool is safe to scale and where it needs guardrails before it is.
There is a discipline here that separates a real safety-gated pilot from a theatrical one. In a theatrical pilot, the verification step exists on paper but is performed as a rubber stamp, because everyone wants the pilot to succeed and a careful check slows the impressive number down. That is the worst of all outcomes: it produces a glowing result and an uncounted error rate, manufacturing exactly the false confidence that scales a hidden defect. A transformer protects against this by making verification a measured, accountable step rather than a formality, by counting and reviewing the errors caught, and by treating a pilot in which the pharmacists found and logged many AI errors as more trustworthy, not less, because it proves the safety gate was real. The uncomfortable truth is that a pilot reporting zero caught errors is usually not a pilot with a flawless tool; it is a pilot with a rubber-stamp verification step, and a transformer reads a suspiciously clean error log as a red flag rather than a triumph.
Reading the Evidence Honestly
When the pilot ends, the evidence has to be read honestly, which is harder than it sounds because everyone involved has an incentive for the pilot to have succeeded. The vendor wants the contract, the team that championed the tool wants to be right, and leadership wants the win it can report. Against all of that pull, the transformer's job is to read the data for what it actually shows, including the parts that are inconvenient. The honest reading asks several hard questions. Did the value materialize against the baseline, measured the same way, or did it shrink once you controlled for the confounds like the extra staff added the same month? Was the error rate acceptable, and acceptable specifically for a clinical setting where the asymmetry means a single dangerous error is not offset by a thousand minutes saved? Did the safety gate hold under real pressure, or did verification degrade into rubber-stamping as the team rushed to hit the impressive number? And were the failures the kind you can guardrail against, or the kind that reveal the tool is fundamentally unsuited to the task?
Honest reading also means being willing to reach an inconvenient conclusion, and a transformer establishes before the pilot that "do not scale" is a legitimate, respected outcome rather than a failure. A pilot that produces clear evidence the tool is not ready, or not safe, or not worth the cost, has succeeded completely, because it spent a small, contained amount of money and patient exposure to prevent a large, network-wide mistake. The framing that a pilot must end in a deployment to have been worthwhile is exactly backward and exactly dangerous, because it pressures the team to scale a tool the evidence does not support. The most valuable pilots an enterprise runs are sometimes the ones that kill a tool everyone was excited about, on the strength of evidence the demo had hidden, and a transformer who has built a culture where that outcome is celebrated rather than punished has built the single most important condition for safe AI adoption: a pilot whose conclusions can be trusted because the team was free to reach them.
Evidence That Satisfies Safety, Not Hype
The standard a pilot's evidence must meet is set by patient safety and, increasingly, by external accreditation, not by what makes a good internal announcement, and the gap between those two standards is where enterprises get into trouble. Hype-grade evidence is a single headline number, the turnaround dropped, presented without a baseline, an error rate, or a controlled comparison, and it is enough to impress a meeting and nothing more. Safety-grade evidence is the full ledger: the value measured against a real baseline, the error rate measured and judged against the clinical stakes, the safety gate documented as having held, and the confounds acknowledged and controlled for as far as possible. The difference is not academic. Hype-grade evidence scales tools that should not be scaled, on the strength of a number that did not mean what the room thought it meant. Safety-grade evidence scales only what is genuinely proven, which is the only kind of scaling the patient-safety asymmetry permits.
There is a further reason a transformer holds pilots to the safety-grade standard, and it connects this lesson to the accreditation the whole program is built toward. URAC's Health Care AI Accreditation, with its user track for organizations that deploy AI, runs on evidence: a reviewer asks not whether you believe a tool is safe but whether you can show the disciplined process by which you established that it is. A pilot run to a safety-grade standard, with its baseline, its measured error rate, its documented safety gate, and its honest scale or no-scale decision, is not only the right way to protect patients; it is precisely the artifact an accreditor wants to see, the documented evidence that the organization tests its AI rather than merely adopting it. So the discipline of a rigorous pilot does double duty: it protects the patient and it produces the accreditation-grade record. A transformer who runs pilots this way is not adding a compliance burden on top of the real work; the rigorous pilot is the real work, and the accreditation evidence is its natural byproduct. The next lesson takes a tool that has earned its scaling decision through evidence and confronts the harder problem of scaling it across many sites without letting the verification and governance that made the pilot safe quietly erode along the way.
The Pilot as the Enterprise Learning Engine
Beyond any single tool, the disciplined pilot is how an enterprise learns, and a transformer treats the pilot process itself as a capability worth building rather than a one-off event run differently every time. An organization that runs its pilots to a consistent standard, baseline, scope, written success criteria, a real safety gate, a measured error rate, and an honest reading, accumulates something more valuable than any individual tool: a repeatable method for converting AI claims into trustworthy evidence. That method is what lets the enterprise evaluate the next tool, and the one after, with steadily improving judgment, because each pilot teaches the organization not only about the tool but about how to test tools. The first few pilots are clumsy and the measurement plans miss things; by the tenth, the enterprise has a pilot playbook that catches the failure modes the early ones missed, and that playbook is a competitive asset, because it lets the organization adopt good tools faster and reject bad ones earlier than competitors who run every pilot from scratch on a glowing demo.
This is the deeper reason a transformer invests in pilot discipline rather than treating it as overhead. The goal is not merely to evaluate one tool correctly; it is to build an organization that cannot be fooled by a good demo, that reflexively asks for the baseline, the error rate, and the controlled comparison, and that treats "do not scale" as a respected outcome. An enterprise with that reflex is structurally protected against the most common way pharmacy AI goes wrong, which is not a malicious vendor but an enthusiastic team scaling an unproven tool on the strength of a feeling. The pilot, run with discipline, is the institutional habit that turns enthusiasm into evidence before evidence is needed to protect a patient, and building that habit across the enterprise is the transformer's real deliverable in this part of the work. A single well-run pilot protects the patients of one workflow; a culture of well-run pilots protects every patient the enterprise will ever touch with a tool it has not yet seen.
Key Takeaways
- At the enterprise level a pilot that produces a good feeling and no evidence is worse than no pilot, because it spends real money and patient exposure to manufacture false confidence, and false confidence is what scales a hidden safety problem across a network.
- A demo answers "can this tool work at all" in a vendor's controlled environment; a pilot answers "does it produce real value, safely, in our operation, measured against what we did before." Scaling is earned by pilot evidence, never by demo enthusiasm.
- A pilot produces evidence only if designed to in advance: a baseline measured the same way, a defined scope, success criteria written before the data arrives, and a measurement plan that captures both the value and the risk.
- In pharmacy the load-bearing measurement is the error rate: how often the AI fabricated a criterion, misstated a fact, or matched the wrong rule, because a pilot that measures only time saved is measuring the easy half and ignoring the half that can hurt a patient.
- The pilot carries its own safety gate, the cardinal rule operationalized: nothing AI produces reaches a patient without a competent human verifying it first, and the errors caught are the pilot's most valuable output, not an embarrassment to hide.
- A pilot reporting zero caught errors usually has a rubber-stamp verification step, not a flawless tool; a transformer reads a suspiciously clean error log as a red flag and treats a pilot that caught and logged many errors as more trustworthy, not less.
- Read the evidence honestly against the incentive to succeed: confirm the value held against a controlled baseline, the error rate was acceptable for a clinical setting, and the safety gate held; establish before the pilot that "do not scale" is a respected, successful outcome.
- Safety-grade evidence, baseline, measured error rate, documented safety gate, controlled confounds, is both the only kind the patient-safety asymmetry permits scaling on and exactly the documented record URAC's user-track accreditation wants to see, so the rigorous pilot does double duty.
Skill.re