AI in the Emergency Department Workflow
It is 2 a.m., the waiting room holds twenty-nine people, and the board is a wall of red. A triage nurse is on her fourth hour without sitting down. An ambient scribe is running in a room down the hall, quietly turning a chest-pain interview into a note. On the radiology worklist, an imaging triage tool has just re-ordered the queue, floating one head CT to the top with a small colored flag beside it. Every one of these tools is trying to give the team back the one thing the emergency department never has enough of: time. And that is exactly why the emergency department is the single most dangerous place in the hospital to use AI, because time pressure is the fuel that automation bias runs on, and the ED is built out of nothing but time pressure. This lesson is about using AI here, in the highest-acuity, highest-volume, highest-stakes environment there is, without letting the speed the tools give you quietly override the safety the patient is owed.
Why the Emergency Department Is the Worst Case for Automation Bias
You already know from earlier in this program what automation bias is: the human tendency to accept an authoritative machine output and skip the verification you would otherwise perform, and the fact that it gets worse, not better, as the tool becomes more reliable. Now take that failure mode and drop it into the one environment engineered to maximize every condition that intensifies it. The research is consistent that automation bias is amplified by time pressure, cognitive load, high volume, interruption, and fatigue. The emergency department is not a place where those conditions occasionally appear. It is a place where they are the permanent operating state, present on every shift, for every clinician, all night long.
Consider what a single ED clinician is holding at once. Multiple patients at different stages of workup. A waiting room that grows while they work. Constant interruption, the phone, the overhead page, the paramedic radio, the family at the desk. Undifferentiated complaints where the diagnosis is genuinely unknown and the cost of missing the dangerous one is catastrophic. Decisions made on incomplete information under a clock. This is the cognitive environment in which an AI tool now offers a confident, well-formatted, plausible answer, and offers it precisely at the moment the human has the least spare attention to scrutinize it. Everywhere else in the hospital, automation bias is a serious risk. In the emergency department, it is the house always winning, because the house is time pressure and the ED never stops applying it.
There is a cruel symmetry here worth naming plainly. The exact reason a busy ED clinician reaches for AI, to save time under crushing load, is the exact reason they are least able to check what it gives them. The tool is leaned on hardest at the moment the human guarding it is weakest. That is not a marginal concern to manage. It is the central fact of AI in the emergency department, and every safe workflow in this lesson is built to survive it.
It helps to be concrete about the physiology of the mistake. Automation bias is not laziness and it is not incompetence. It is what a normal, competent brain does when it is running near capacity and something offers to reduce the load. A confident, well-formatted AI output is a cognitive off-ramp, and a brain that has been sprinting for nine hours will take the off-ramp. The clinician who signs a confabulated note at 3 a.m. is very often the same clinician who, rested and unhurried in a morning conference, would have caught the error instantly and been appalled by it. The error does not reveal a bad clinician. It reveals a predictable interaction between a plausible machine output and a depleted human under load. That is precisely why the fix cannot be "be a better clinician" or "care more." The clinicians this happens to already care and are already good. The fix has to live somewhere the erosion cannot reach it, which is why this whole lesson keeps returning to fixed structure rather than in-the-moment resolve.
Time pressure is the permanent weather of the emergency department, and it is exactly when the human check is weakest. That is not an occasional storm to wait out. It is the climate you practice in every shift.
The Three Places AI Shows Up in the ED
AI enters the emergency department through three main doors, and each one carries its own version of the same danger. Knowing the specific failure mode of each is the difference between using the tool and being used by it.
Ambient Documentation: The Scribe That Confabulates
The first door is the ambient scribe, the tool that listens to the encounter and drafts the note. In a specialty drowning in documentation, this is a genuine gift, and it is already mainstream. But the ED is the hardest possible environment for an ambient scribe to get right. Encounters are interrupted and fragmented. Multiple people talk at once. The room is loud. History is taken in pieces between other tasks. Into that mess, a generative model will do what generative models do: it will produce a smooth, complete, professional note, and where the audio was ambiguous or absent, it will fill the gap with something plausible rather than leave it blank. That is the failure mode you must hold in your mind: a confabulated exam finding, a documented normal that was never examined, a stated pertinent negative that was never asked, a history detail invented to make the narrative flow. The note reads perfectly. That is what makes it dangerous. Under a full waiting room, a perfect-reading note is precisely the one a tired clinician signs without cross-checking against what actually happened in the room.
Triage and ESI Decision Support
The second door is triage support, tools that suggest an acuity level or an Emergency Severity Index score, or that flag a patient as higher or lower risk. Triage is the highest-leverage decision in the entire department, because it determines who is seen first and who waits, and a wrong answer at triage propagates through everything downstream. The danger here is the downgrade: an AI suggestion that a patient is lower acuity than they are, accepted under volume pressure because the queue is enormous and the suggestion offers relief. A triage nurse who is calibrated treats an ESI suggestion as one input weighed against the patient in front of them, the vital signs, the gestalt, the story that does not fit. A triage nurse under automation bias lets the number decide, and the sick patient who did not look sick waits in the lobby.
Imaging Triage: A Flag Is Not a Read
The third door is imaging triage, and it is the one most widely misunderstood, so it deserves the most care. FDA-authorized imaging AI is heavily concentrated in radiology, which is where roughly three-quarters of recent device authorizations cluster, and a large class of these tools does something specific and narrow: it scans studies as they arrive and re-orders the worklist, pushing a study it thinks may contain a critical finding, an intracranial hemorrhage, a pulmonary embolism, a large-vessel occlusion, toward the top of the radiologist's queue so it gets read sooner. That is genuinely valuable. Minutes matter in a bleed or a stroke, and moving the right study up the list can save brain and save life. It is worth pausing on why the number of these tools has grown so fast: FDA clearance authorizes a defined intended use on the data the device was tested against, and it is not a promise that the tool performs the same way on your population, your scanner protocols, or your acuity mix. A cleared tool is a starting point for local validation, not a finished guarantee, and that distinction matters more in the ED than almost anywhere, because the ED sees the widest and least predictable case mix in the building.
But you must understand exactly what the tool does and, more importantly, what it does not do, because the two most dangerous misreadings of an imaging triage tool are both seductive and both wrong. First: a flag is not a read. The tool raising a flag is a prioritization signal, a suggestion that a human should look at this one sooner. It is not a diagnosis, not an interpretation, and not a substitute for the radiologist's read. The study still has to be read by a qualified human, and the flag does not tell you the finding is real. Second, and this is the one that kills people: the absence of a flag is not a normal study. These tools are tuned for a particular finding, and they have a sensitivity below one hundred percent. A study with no flag has not been cleared. It has simply not been prioritized by the algorithm, which is a completely different thing. A clinician who starts treating unflagged studies as reassuring, as probably fine, has quietly converted a worklist-ordering tool into a diagnostic one it was never built to be, and has handed the model a veto over their clinical suspicion that it was never authorized to hold. If the patient's story and exam say stroke, the absence of an AI flag does not lower that concern by one degree.
An imaging triage tool reorders the worklist. It does not read the film. A flag is not a diagnosis, and the absence of a flag is not a normal study. It is an unread study that the algorithm did not move.
A Worked Example: The Same Shift, Two Clinicians
Watch the difference between a clinician the ED erodes and a clinician who came in already defended. Same department, same brutal night, same tools. The gap between them is not talent or dedication. It is structure.
Before: The Suggestion That Sails Into the Chart
A clinician is running a full waiting room at the worst hour of the night. A middle-aged patient came in with chest discomfort and left before the workup was finished. The ambient scribe produces a clean, complete note, and among its tidy lines is a documented finding: lungs clear to auscultation bilaterally, no lower-extremity edema. The clinician never listened to that patient's lungs and never examined the legs; the encounter was cut short and chaotic, interrupted twice by a trauma activation next door. But the note reads like every other note, authoritative and unremarkable, and there are eleven other charts waiting. The clinician signs. The confabulated exam is now in the legal record, attesting to an examination that did not happen, and if this patient returns in extremis, or if a plaintiff's expert reads that chart in two years, it says the clinician examined and found normal what they never touched at all. Worse, the confabulated normal actively hides the very findings that might have changed the disposition. A documented "no lower-extremity edema" quietly argues against the unilateral leg swelling nobody looked for, and a clean lung exam softens the picture just enough that the next reader relaxes. The AI did not simply add a false line. It tilted the whole chart toward reassurance, and it did so in the fields a busy reader trusts most.
Consider where this same AI genuinely helped, because the point is not that the scribe is the enemy. On four other patients that shift, the ambient scribe let the clinician stay at the bedside instead of turning to a keyboard, captured a medication list the clinician would have had to reconstruct from memory, and produced a first draft that was eighty percent correct and saved real minutes on each chart. That is the whole reason the tool is in the room. The danger is not that AI is useless in the ED; it is precisely useful enough, often enough, that trusting it becomes a habit, and the habit does not distinguish the eighty percent it gets right from the confabulated twenty percent that can end a career or a life. The tool earns trust on the easy cases and then spends that trust on the hard one. That is the pattern to hold in your mind: the failure rides in on the success.
In a parallel version of the same night, a triage AI downgrades a patient with a vague complaint to a lower acuity, the nurse under thirty-in-the-lobby pressure accepts it, and the downgrade sails into the record and out to the lobby with it. Notice the shared anatomy of both failures: an authoritative AI output, a clinician with no spare attention, and no fixed checkpoint between the suggestion and the chart. Speed overrode safety, and nobody decided to let it. It just happened, gently, the way automation bias always happens.
After: The Clinician Who Came in Defended
Now the informed clinician, same night, same load, same tools. The difference is that their safety does not depend on being sharp at 2 a.m., because they know they will not be. They run a fixed habit that fires by rule, not by mood. Every ambient note gets a specific check before signing: the exam findings and the pertinent negatives are confirmed against what actually happened in the room, because those are exactly the fields a scribe confabulates. When the note claims a clear chest exam they did not perform, the rule catches it, and they correct or delete it before attestation. Every triage suggestion is treated as an input, not a verdict; the number is weighed against the vitals and the gestalt, and a downgrade that does not match the patient is overridden with a one-line reason. And the imaging flag is used for what it is: a study with a flag gets looked at sooner, and a study without a flag changes nothing about the clinical suspicion, so the stroke workup proceeds on the exam, not on the algorithm's silence. When they agree or disagree with an AI suggestion on anything that matters, they leave a short note saying why, because a sentence of clinical reasoning is what turns an AI-assisted decision into a defensible one and proves a human was in the loop. Same erosion of attention across the shift. Different outcome, because the check was never riding on the attention in the first place.
This is the whole lesson in one contrast. You cannot out-vigilance the emergency department; the department will win. What survives the night is not willpower but structure: a small number of fixed verification habits, tied to the specific things AI gets wrong, that fire whether or not you are sharp. The defended clinician is not more careful in the moment. They are more careful in advance.
The Check Belongs to the Team, Not One Hero
The emergency department is not a solo performance, and its AI safety is not one clinician's job. The record that AI now touches passes through many hands, and each role holds a distinct part of the check. Making that ownership explicit is how a department keeps the human in the loop when any single human is overwhelmed.
The triage nurse is the first and highest-leverage line, owning the acuity decision and the calibration of any triage or ESI suggestion against the living patient. The ED physician or advanced practice clinician owns the diagnostic and disposition decisions and the attestation of the ambient note, which means they own catching the confabulated finding before it becomes part of the signed record. The ED technician and nursing staff hold ground truth about what was actually done, the vitals actually taken, the exam actually performed, and are often the ones positioned to notice when a note claims something that did not happen. The radiologist owns the actual read of the image, and it is worth stating clearly that the imaging AI serves the radiologist's prioritization; it does not replace the read, and the flag or its absence never overrides the radiologist's interpretation or the ED clinician's clinical suspicion. When these roles understand that AI has entered their shared workflow and that each of them owns a specific checkpoint, the department gains something no individual can provide: redundancy. A confabulation the exhausted physician might sign, the tech might question. A downgrade the nurse might accept, the physician might catch at the bedside. The team is the loop, and a team that talks about where AI touches its work is a team that catches what a lone tired clinician would miss.
Why ED Verification Gates Are Built Differently Than the Clinic's
Much of what this program teaches about verifying AI output was framed around the clinic: the primary care visit, the specialist consult, the scheduled encounter where a physician has a few minutes after the patient leaves to read the ambient note carefully and correct it before signing. That model is real and it works, but it quietly assumes something the emergency department cannot provide: a moment of calm at the end. In the clinic, verification is a step you take when the pressure has briefly dropped. In the ED, the pressure never drops, so a verification gate designed for the clinic will simply never fire. The department will roll straight over it. This is why ED verification has to be engineered differently, and understanding that difference is what separates a policy that looks good on paper from one that survives a Saturday night.
The first difference is timing. A clinic gate can afford to be a careful, comprehensive review because there is space for it. An ED gate has to be fast, narrow, and targeted, because a slow gate is a gate that gets skipped. Instead of "review the whole note," the ED version is "confirm the exam findings and the pertinent negatives, the two fields the scribe confabulates, and move on." A good ED gate checks the few things most likely to be wrong and most likely to hurt, and deliberately does not try to check everything, because a check that demands more time than the shift can spare is a check that will be abandoned the first busy night and never resumed.
The second difference is who holds the gate. In the clinic, verification is largely one physician's task at the end of one encounter. In the ED, no single overwhelmed human can reliably be the gate, so the check is distributed across roles and moments, as the previous section described. The clinic can lean on an individual; the ED must lean on a system, because in the ED the individual is predictably depleted at exactly the hour the check matters most.
The third difference is the cost of a miss. A missed confabulation in a clinic note for a stable follow-up patient is a serious documentation and liability problem. A missed downgrade at ED triage, or a stroke workup deferred because an imaging tool stayed silent, can be a dead or disabled patient within the hour. The undifferentiated, high-acuity, time-critical nature of ED presentations means the tail risk of an unverified AI output is fatter and faster than in almost any clinic. That is why the ED cannot simply import the clinic's gates and hope. It has to build gates that assume the worst conditions, fire by rule rather than by available attention, and target the specific failures that turn an AI error into a patient harm before anyone has a chance to catch it downstream.
A verification gate designed for a calm moment at the end of a clinic visit will never fire in an emergency department, because the calm moment never comes. ED gates must be fast, narrow, distributed, and automatic, or they do not exist.
The Iron Rule, at the Speed of the ED
Everything in this lesson reduces to the rule that anchors the entire program, applied to the one place that tests it hardest: AI assists, the clinician decides, the record proves it. In the emergency department, each clause carries a specific weight. AI assists means the scribe drafts, the triage tool suggests, and the imaging tool reorders the queue, and not one of those is a decision. The clinician decides means the acuity, the diagnosis, the disposition, and the read remain human judgments that the tools inform but never make, no matter how confident the output or how full the lobby. The record proves it means the note reflects what actually happened, the confabulations are gone, the overrides are documented with a reason, and the chart would defend the clinician rather than indict them if a surveyor, an auditor, or a plaintiff's expert read it in two years.
The temptation the ED manufactures, relentlessly, all night, is to let the first clause swallow the other two, to let assist quietly become decide because deciding takes time you do not have. Resisting that is not about working harder inside a shift that is already breaking you. It is about building the small, fixed, boring checkpoints in advance so that speed and safety stop being a trade-off you have to make freshly at 2 a.m. under pressure. The fastest safe ED is not the one that trusts AI most. It is the one whose verification habits are so well-built that they cost almost nothing and fire every time, so the human check survives the very conditions designed to destroy it.
Key Takeaways
- The emergency department is the single highest automation-bias-risk environment in the hospital, because time pressure, cognitive load, volume, interruption, and fatigue are its permanent operating state, and those are exactly the conditions that intensify over-trust of AI.
- The cruel symmetry: the reason a busy ED clinician reaches for AI to save time is the same reason they are least able to verify what it gives them. The tool is leaned on hardest when the human is weakest.
- Ambient scribes in the noisy, interrupted ED will confabulate: a documented exam that never happened, a pertinent negative never asked, a history detail invented to smooth the narrative. The note reads perfectly, which is what makes it dangerous to sign unchecked.
- Triage and ESI decision support carry the downgrade risk: an acuity suggestion accepted under volume pressure that sends a sick patient to the lobby. Treat the number as one input weighed against vitals and gestalt, never a verdict.
- Imaging triage tools reorder the worklist; they do not read the film. A flag is not a diagnosis and still requires a human read. Critically, the absence of a flag is not a normal study; it is an unread study the algorithm did not prioritize, and it must never lower a clinical suspicion the exam supports.
- You cannot out-vigilance the ED. The defenses that survive are fixed verification habits tied to what AI specifically gets wrong (confirm exam findings and pertinent negatives before signing, weigh triage suggestions against the patient, document overrides with a reason) that fire by rule, not by in-the-moment attention.
- AI safety in the ED is a team property. The triage nurse, ED physician, technician, and radiologist each own a specific checkpoint, and making that ownership explicit gives the department redundancy no single overwhelmed clinician can provide.
- The iron rule holds at the speed of the ED: AI assists, the clinician decides, the record proves it. The fastest safe department is not the one that trusts AI most, but the one whose verification habits are so well-built that they cost almost nothing and fire every time.
Skill.re