Redesigning Clinical Workflows Around AI
A large primary-care group bought an ambient scribe and, in a single sentence at the kickoff meeting, doomed the whole project: "We will bolt it onto the existing visit and get twenty percent of our documentation time back." Eight months later the time savings were real, the physicians were less exhausted, and the coding audit was a disaster. Notes had been signed faster, but they had also been read less. The AI had drafted, the clinician had clicked, and the verification step that used to happen naturally when a tired doctor typed the note by hand had quietly evaporated. The tool did exactly what it promised. The workflow around it is what failed, because nobody redesigned it. This lesson is about the discipline that group skipped: putting humans on judgment and AI on throughput without hollowing out the clinical accountability that keeps patients safe.
Bolting On Versus Redesigning
There are two ways to introduce AI into a clinical workflow, and they look almost identical on the purchase order and completely different on the floor. The first is bolting on: you take an existing process and drop the AI into one step, changing nothing else, and harvest the efficiency. The second is redesigning: you rethink the whole flow around the fact that a new kind of worker, tireless, fast, confident, and occasionally confidently wrong, has joined the team, and you decide deliberately where the human judgment now lives and where the verification gate now sits. Bolting on is faster to deploy and far more common. It is also how you hollow out accountability without noticing, because the efficiency arrives immediately and the erosion of the human check arrives quietly, over months, in the audit and the malpractice file.
The reason bolting on is so seductive is that the old workflow already contained safety checks that were invisible because they were incidental. When a physician wrote a note by hand, the act of writing was also an act of thinking: reconstructing the visit, noticing the gap, catching the wrong laterality because their own hand was forming the word. That verification was free, so nobody ever named it as a step. When you replace the writing with an AI draft and a signature, you remove the labor, which is the point, but you also silently remove the thinking that rode along with the labor. If you have not deliberately rebuilt that verification as an explicit, defended step, you have not made the workflow faster; you have made it faster and blinder. The redesign discipline is precisely the work of finding every safety check that used to be incidental and deciding, on purpose, how it survives the efficiency gain.
Humans on Judgment, AI on Throughput
The organizing principle of a good redesign is a clean division of labor: AI carries throughput, humans carry judgment. Throughput is the high-volume, pattern-heavy, tireless work that machines do well and humans do worse as they fatigue: transcribing a conversation into a draft, pulling a first-pass summary from years of notes, flagging the images that probably need a closer look, drafting a routine reply for review. Judgment is the irreducibly human work of deciding what is true about this patient, what it means, and what to do, the part that carries the license, the liability, and the moral weight. A workflow that assigns throughput to AI and judgment to humans is playing to the strengths of each. A workflow that lets AI drift into judgment, or that buries the human's judgment under so much throughput that they cannot exercise it, is dangerous however impressive its metrics.
The trap is that throughput and judgment are not cleanly separable in every task, and the boundary is exactly where safety is won or lost. An AI that summarizes a chart is doing throughput, but the moment it decides which abnormal value is worth surfacing and which to drop, it has crossed into judgment, and if the human downstream treats the summary as complete, the machine has effectively made a clinical decision no one verified. The redesign question is therefore not just "what does the AI do and what does the human do" but "at exactly which point does the machine's throughput hand off to the human's judgment, and is that handoff a real gate or a rubber stamp?" Draw that line in the wrong place, or leave it blurry, and you get a workflow where accountability has technically stayed human but functionally migrated to the model.
A useful way to find that boundary is to ask, for each thing the AI produces, whether a wrong version of it would change what happens to the patient. If the answer is no, it is throughput and the human can lean on it: a first-draft transcript, a formatting pass, a list of prior visits to skim. If the answer is yes, the AI has reached into judgment and the output cannot be allowed to become action or record without a human verifying it against the source. A chart summary is the classic mixed case: assembling the timeline is throughput, but selecting which of forty results matters is judgment, because a dropped potassium or a hidden trend changes management. The redesign has to split that single "summary" task at the seam, letting the AI assemble but requiring the human to confirm that nothing decision-changing was omitted, rather than treating the whole summary as one trustworthy block. Getting this seam right is most of the work, because it is invisible on the demo, where the summary looks complete and correct, and only shows its teeth on the case where the one value that mattered is the one the AI decided to leave out.
Keep the Verification Gate
The single most important artifact of a redesigned workflow is the verification gate: the explicit point where a competent human reviews the AI output against the source before it becomes action or record. The primary-care group's failure was a missing gate. In their old workflow the gate was implicit and free, riding on the act of typing; in the new one, nobody rebuilt it, so the signature became a reflex rather than an attestation. Redesigning around AI means making the gate explicit, reachable, and hard to skip, so that the efficiency gain does not come at the cost of the check that makes the output safe and the record defensible.
A real gate has properties that a rubber stamp does not. It presents the AI output next to the source it can be checked against, rather than asking the human to trust it in isolation. It lands at a moment when the human actually has the cognitive room to check, not at the end of a twelve-hour shift when they will click anything to go home. It is proportionate to risk, deep for a discharge decision or a dose, light for a scheduling suggestion, because a gate that treats every output as equally dangerous trains people to treat all of them as equally trivial. And it produces a record: the attestation, the note that says what was verified and why the clinician agreed or overrode. A gate you can skip without a trace is not a gate. A gate that produces a reconstructable decision is the load-bearing structure of the whole redesign.
These properties translate into concrete design decisions you can inspect, not vague intentions. Consider four levers. First, defaults: if the AI output arrives pre-accepted and the human must actively opt out to reject it, the gate is already lost, because the path of least resistance is to do nothing and the tired clinician will take it; a safe gate makes acceptance an act, not the absence of one. Second, friction placement: the small amount of friction the design can afford should sit on the high-risk elements, the medication change, the abnormal value, the laterality, and nowhere else, so the clinician's attention is spent where harm lives. Third, source proximity: the check is only real if the thing to verify against is on the same screen at the same moment; a gate that requires the clinician to leave the note, open another system, and hunt for the source is a gate that will be skipped under load. Fourth, the trace: the attestation should capture not just that a signature occurred but what was confirmed and any override reason, because a signature with no content is indistinguishable from a reflex and proves nothing to a later auditor. When you evaluate a proposed gate, walk these four levers explicitly. A design that fails any one of them, a pre-accepted default, friction smeared evenly across trivial and critical outputs alike, a source two clicks away, or an attestation that records only a click, is a rubber stamp wearing the vocabulary of a gate.
The Gate Must Fit the Real Shift, Not the Demo
Here is where many well-intentioned redesigns still fail: the gate is designed for the conditions of a vendor demo, a rested clinician, a single unhurried patient, a clean case, and then deployed into the conditions of an actual shift, a fatigued clinician, a packed panel, and a dozen competing demands. A gate that is realistic at 9 a.m. on the demo laptop is a fiction at hour eleven of a short-staffed floor, and the workflow inherits the difference as risk. The redesign discipline therefore includes a stress test that most projects skip: walk the gate through the worst plausible conditions your clinicians actually work in, not the best. Ask what the gate becomes when the person at it is exhausted and behind, because that is the person who will be standing there most often, and the gate has to survive them, not an idealized version of them. A verification step that only holds when everyone is fresh is not a safety control; it is a hope.
This is also why proportionality is not a nicety but a survival requirement for the gate itself. A clinician can sustain deep verification on the handful of genuinely high-risk outputs per shift, the doses, the discharge decisions, the abnormal values that change management. A clinician cannot sustain deep verification on every one of two hundred low-stakes AI touches, and a workflow that demands it will get neither, because the exhausted human will start skipping uniformly. Concentrate the human's finite verification budget where the harm lives, and spend it lightly where it does not, and the gate becomes something a real person on a real shift can actually hold. Design the gate for the tired clinician and it protects everyone; design it for the demo and it protects no one when it matters.
Efficiency that removes the labor but not the thinking is the good kind. Efficiency that removes both, and leaves a signature where a judgment used to be, is a liability wearing the costume of progress.
Avoid Deskilling
There is a slower, subtler failure than a missing gate, and it does not show up for years: deskilling. When AI does a task well for long enough, the humans who used to do it lose the ability to do it, or to check it. A radiology workflow that leans on AI triage can, over time, produce radiologists who are excellent at confirming the AI and worse at the unaided read the AI never flagged. A generation of clinicians trained with the AI drafting every note may never build the documentation fluency that let their predecessors catch the AI's errors by instinct. The efficiency is real, but it is quietly spending down a capability, and the day the AI is wrong, or unavailable, the human who was supposed to be the backstop has been hollowed out along with the workflow.
Deskilling is an organizational-design problem, not an individual failing, which means it has organizational-design answers. A workflow can be built to preserve the underlying skill deliberately: rotating clinicians through unaided reads, keeping the human's independent assessment upstream of the AI's suggestion rather than downstream of it so the human forms a view before seeing the machine's, and treating the human's ability to catch the AI as a competency to be maintained rather than assumed. The executive who redesigns a workflow purely for maximum throughput, with no thought to what skill the workflow is silently retiring, is optimizing for this quarter's efficiency and mortgaging the safety of the year the tool fails. The verification gate only works if the human at the gate is still able to verify, and that ability is not free; it is a thing the workflow either preserves or erodes.
The upstream-versus-downstream point deserves a concrete illustration, because it is the single most effective anti-deskilling move and it is almost always designed away for the sake of speed. Picture two radiology workflows. In the first, the AI pre-reads the study and presents its findings, and the radiologist reads the AI's markup and confirms or corrects it. This is downstream: the human sees the machine's answer before forming their own, and over months their eye stops doing the independent search because the AI has already done it, so they become excellent at agreeing with the AI and blind to what the AI never flagged. In the second workflow, the radiologist records their own impression first, and only then is the AI's read revealed, and any disagreement is surfaced for a deliberate second look. This is upstream: the human forms an independent view every single time, the skill is exercised rather than retired, and the AI becomes a genuine second opinion that catches the human's misses instead of a first opinion that the human rubber-stamps. The two workflows use the identical AI and produce nearly identical throughput. Only one of them keeps the radiologist able to work when the AI is wrong or absent. That ordering, human first then machine, is a design choice the throughput-maximizer will delete because it feels redundant, and it is precisely the redundancy that keeps the backstop alive.
Deskilling also has to be monitored, not just designed against once, because it is a slow drift that a launch-day design cannot prevent by itself. A workflow can track leading indicators: the rate at which clinicians override the AI (a rate that falls steadily toward zero is a warning that the human check is decaying into agreement), the time spent at the verification gate (collapsing to near-instant means the gate has become a reflex), and periodic unaided competency checks that confirm the humans can still do the task the AI usually does. None of this is exotic instrumentation, but it requires someone to own the signal and treat a decaying override rate as the safety problem it is rather than the efficiency win it superficially resembles. A falling override rate looks like the tool getting better and the humans trusting it more. It can just as easily be the humans getting worse and the check quietly dying, and only the monitoring, plus the occasional unaided read, can tell you which.
A Worked Example: Two Redesigns of the Same Visit
Take the ambient-scribe primary-care visit and design it two ways. In the bolt-on version, the AI listens, drafts the note, and drops it into the queue; the physician's workflow is unchanged except that the blank note is now a full draft, and the physician signs to close the encounter. The metric dashboard is glorious: documentation time down, throughput up, burnout scores improving. And the failure is invisible until the coding audit: notes with confabulated exam findings the physician never performed, dropped pertinent negatives, the occasional wrong laterality, all signed, all now the legal record, because the signature had become a reflex and the verification gate was never built.
Now the redesigned version. The AI still listens and drafts, carrying the throughput. But the workflow explicitly inserts a verification gate before signing: the physician is shown the draft alongside the specific high-risk elements to confirm, the exam findings, the medication changes, the plan, presented against what is actually in the record, so checking is fast and structured rather than a wall of text to trust or ignore. The gate is proportionate: the physical exam and the medication list get real scrutiny, the pleasantries do not. The attestation captures what was verified. And the human's judgment is protected by keeping the visit itself, the history and the exam, a human act the AI documents rather than performs. The efficiency is nearly the same, the documentation time is still down, but the check survived the redesign. The difference between the two versions is not the tool. It is whether someone did the work of designing the gate back in.
Walk the redesigned visit through the four levers to see why it holds. The default is that the exam findings and medication changes are unconfirmed until the physician actively confirms them, so signing without checking is not the easy path but a deliberate override that leaves a mark. The friction is concentrated: the tool highlights the three or four elements that can hurt a patient and lets the rest pass, so the physician's scarce attention lands on the medication reconciliation and the abnormal result, not on the greeting. Source proximity is built in: each high-risk element is shown next to what the record actually contains, so confirming laterality or a dose is a glance, not an expedition into another system. And the attestation records what was confirmed and any override, so the note carries evidence that a human was meaningfully in the loop. This is the difference between a redesign and a slogan: you can point at each lever and show how it survives a real, tired clinician on a full panel. A gate that cannot survive that walkthrough is decoration.
Watch the two audits diverge. The bolt-on group defends signed notes full of things that did not happen, with no record of any verification, a malpractice and fraud exposure with the clinician's name on every one. The redesigned group produces, for each note, an attestation of what was checked, a defensible record that the human was meaningfully in the loop. Same efficiency, same tool, opposite risk posture, and the entire difference is the redesign the first group skipped in the name of getting twenty percent back faster.
The Executive Discipline of Redesign
For the leaders who own workflow, the lesson is a warning against buying efficiency without buying the redesign that makes it safe. When a vendor promises time savings, the executive question is not only "is the savings real?" but "which incidental safety check does this efficiency remove, and how will we rebuild it on purpose?" A deployment plan that budgets for the tool and the training but not for redesigning the workflow around the verification gate is a plan to harvest the efficiency and eat the risk. The redesign is not overhead on the AI project; it is the part of the project that makes the AI safe to have.
The discipline is teachable and repeatable. For any AI you are about to introduce, map the current workflow and find the safety checks that are incidental and free, name the throughput the AI will carry and the judgment the human must keep, place the verification gate where the human can actually reach it and make it proportionate to risk, decide what skill the workflow must preserve so the human at the gate stays able to verify, and require an attestation that makes the decision reconstructable. Do that, and AI makes your clinicians faster and your record more defensible at once. Skip it, and you have bolted a fast, confident, occasionally wrong new worker onto a process that no longer checks its work, which is not a redesigned workflow. It is an accident that has not been audited yet.
There is one more executive habit worth naming, because it is where the redesign either lives or dies after go-live: treat the workflow as something you keep watching, not something you ship and forget. A gate that was proportionate and reachable on launch day can decay as volumes rise, as staffing thins, as clinicians learn to route around it, and as the model itself changes underneath the workflow. The Clinical AI Lead who owns adoption and the AI Safety Officer who owns monitoring should be reading the same signal from two angles: is the gate still being used as designed, or has it quietly eroded back into a rubber stamp under the pressure of the real shift? The redesign is not a one-time act of architecture; it is a standing commitment to keep the human check alive as the conditions around it change. A workflow that was safe at launch and was never looked at again is simply a bolt-on with a better origin story.
None of this is anti-efficiency, and it is worth ending on that, because the reflexive worry is that all this gating slows the very speed the AI was bought to deliver. It does not. A well-placed, proportionate gate costs seconds on the outputs that matter and nothing on the outputs that do not, while the efficiency of the AI carrying the throughput remains almost entirely intact. What the redesign changes is not how fast the clinician moves but whether the record they leave behind is true and defensible. You get the twenty percent, and you keep the check. That is the whole promise of doing this well: speed and safety are not a trade if you redesign the workflow instead of bolting the tool on and hoping.
Key Takeaways
- There are two ways to add AI to a clinical workflow: bolting it onto one step and harvesting the efficiency, or redesigning the whole flow around it. Bolting on is faster and far more common, and it is how accountability gets hollowed out without anyone noticing.
- Old workflows contained safety checks that were incidental and free, like the thinking that rode along with writing a note by hand. Replacing the labor with an AI draft removes the labor and silently removes the thinking unless you deliberately rebuild it.
- The organizing principle of a good redesign is a clean division of labor: AI carries throughput (high-volume, pattern-heavy, tireless work), humans carry judgment (deciding what is true, what it means, and what to do, which carries the license and the liability).
- Safety is won or lost at the handoff between throughput and judgment; the redesign question is exactly where the machine's output hands off to the human's judgment, and whether that handoff is a real gate or a rubber stamp.
- The verification gate is the load-bearing artifact: an explicit, reachable, proportionate-to-risk review of the AI output against its source, landing when the human has room to check, and producing an attestation. A gate you can skip without a trace is not a gate.
- Deskilling is the slow failure: when AI does a task well for long enough, humans lose the ability to do it or check it, so the backstop is hollow the day the AI is wrong. It is an organizational-design problem with organizational-design answers, like keeping the human's independent read upstream of the AI's and monitoring the override rate for the slow decay of the check.
- The worked example shows two redesigns of the same ambient-scribe visit with nearly identical efficiency and opposite risk: one produces signed notes full of unverified fabrications, the other produces an attestation of what was checked. The difference is the gate, not the tool.
- The executive discipline: when buying efficiency, ask which incidental safety check it removes and how you will rebuild it. A plan that budgets for the tool but not the workflow redesign is a plan to harvest the efficiency and eat the risk, and a gate must be monitored after launch, since a falling override rate can signal a decaying check rather than a better tool.
Skill.re