AI for Healthcare & Clinical Practice
Capable · M23 · lesson 23 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
What the Scribe Gets Wrong and How to Catch It
📖
now learning

What the Scribe Gets Wrong and How to Catch It

15 min

A cardiologist reviews an ambient note and finds a beautifully documented cardiovascular exam: regular rate and rhythm, no murmurs, rubs, or gallops, no jugular venous distension, no peripheral edema. It is a model exam. There is only one problem. The patient came in for a rash, the cardiologist listened to the heart for about four seconds out of courtesy, and never examined the neck veins or the legs at all. The note did not lie exactly. It did something more insidious: it wrote down the exam a cardiologist usually performs, not the exam this cardiologist actually did. This lesson is a field guide to that specific kind of error and its cousins, the recurring, recognizable ways an ambient scribe gets a note wrong, and the concrete habits that catch each one.

Why a Catalog of Failures Is a Safety Tool

You catch errors faster when you know their shape in advance. A radiologist who has seen a thousand examples of a particular artifact recognizes the thousand-and-first instantly; a clinician who knows that ambient scribes tend to fabricate complete exams will spot a fabricated exam in a note the way a proofreader spots a familiar typo. The failures in this lesson are not random. They cluster into a small number of recognizable types, each rooted in the pipeline you already understand: capture drops, transcription mishears, and generation invents or smooths. Learning the catalog turns verification from a vague sense that you should be careful into a targeted search for specific, known problems. That is the difference between hoping you will notice something wrong and actively hunting the errors you know this tool tends to make. Hoping fails on a busy shift; hunting a known list does not.

This matters more every month, because ambient documentation is no longer a pilot at the edge of the organization. By the end of 2025 ambient scribes reached roughly 30 percent market penetration; in 2026 UCSF reported that about 70 percent of its physicians were using an AI scribe. Kaiser Permanente logged 7,260 physicians across 2.5 million encounters over fourteen months. Northwell rolled out system-wide across tens of thousands of clinicians and dozens of hospitals. The technology clearly helps: one 2025 multi-system study found physician burnout fell from 51.9 percent to 38.8 percent after just thirty days on an ambient scribe, with documentation the single largest driver of that burnout in the first place. But read those numbers the way this program teaches every number, as a figure to verify rather than to repeat blindly. A tool that drafts most of your notes and lightens a real burden is exactly the tool whose output you are most tempted to sign without reading, and that temptation is where the catalog earns its keep. The more notes the scribe writes, the more disciplined the human reading them has to be.

Keep one principle in mind as we go: every error type below is silent. The note does not underline its fabrications or footnote its omissions. It presents every line, true or invented, in the same confident, professional register. Your defense is not the note flagging its own mistakes; it is you knowing the mistakes exist and looking for them where they live. And there is a second reason to look, beyond patient safety: when you sign that note, it becomes the legal record. Your attestation says these words describe what you did. A confabulated exam is not the scribe's statement; once signed, it is yours, and "the model wrote it" is never a defense to a board, a plaintiff, a family, or a Joint Commission surveyor.

The Confabulated Exam Finding

The most characteristic ambient scribe error is the confabulated physical exam: a normal, complete, plausible exam finding for a system that was never actually examined. The cardiologist's note above is the archetype. The mechanism is pure generation. A clinical note of a given type usually contains a full exam, so the generative model, doing exactly what it was built to do, produces a full exam, filling in the expected normal findings for systems the conversation never mentioned. The result is dangerous precisely because it is normal and plausible. A bizarre finding would catch your eye. A textbook-normal cardiovascular exam attracts no attention at all, which is exactly why it sails into the signed record.

The catch strategy is a single disciplined question asked of every documented exam: did I actually perform this. Not "is this finding plausible," not "is this what I would expect," but "did my hands and my stethoscope actually do this maneuver on this patient today." If the note documents a neurological exam and you tested only gait, the rest is confabulated and must be deleted or corrected to what you did. This feels almost paranoid the first few times, because the findings are usually normal and usually right. Do it anyway, because the day the confabulated exam is wrong, when the note says "no peripheral edema" on a patient whose legs you never looked at and who is quietly in heart failure, is the day it matters, and you cannot know in advance which day that is.

Consider the moment this error actually gets signed, because it is rarely a moment of carelessness. It is the end of a long clinic session. You are twelve notes behind, the waiting room is full, and the ambient note for the rash patient reads cleanly, professionally, completely. The exam section says everything you would want an exam to say. Your eye slides over the cardiovascular block because it is normal, and normal is what you expected, and you have forty seconds per note if you want to leave before dark. This is automation bias in its purest clinical form: the tendency to accept an authoritative, fluent output under time pressure without applying the scrutiny you would apply to your own writing. Automation bias predicts precisely this failure. The more plausible and polished the output, the less you check it, which is exactly backward from what safety requires. The reason to name it is that once you can feel it happening, that little internal nudge of "it looks fine, just sign it," you can treat the nudge itself as the signal to slow down on the exam block for two more seconds and ask the one question that catches the error.

Wrong Laterality

Left becomes right. Upper becomes lower. The lesion on the right forearm is documented on the left. Laterality errors are a distinct and dangerous category because a single flipped word can misdirect an entire downstream chain: imaging ordered on the wrong side, a referral that sends a specialist looking at the wrong limb, in the worst case a wrong-site procedure. Laterality can flip at transcription, when the ASR mishears "left" as a similar sound, or at generation, when the model reconstructs a sentence and reassigns the side. Either way, the word that ends up in the note is small, confident, and possibly wrong.

The catch strategy is to treat every stated side as a checkpoint. Wherever the note names left or right, upper or lower, proximal or distal, your eye stops and confirms it against what you actually found. This is a two-second habit with an enormous safety return, because laterality is both easy to flip and catastrophic to get wrong. A useful reinforcement is to cross-check the side against the orders and referrals generated from the visit: if the note, the imaging request, and the referral all name the same side, you have real corroboration; if they disagree, you have caught something before it reached the patient.

Picture the checkpoint working in a busy emergency department. A patient arrives with a deep laceration to the left shin from a fall off a curb. The triage nurse, reviewing the ambient note before the physician orders imaging to rule out a foreign body, notices the note reads "laceration to the right lower leg." She stops at the word "right," because she was trained to stop at every stated side, and she remembers cleaning and packing the left. That two-second catch matters enormously here, because the very next click was going to be an X-ray order, and the order pulls its laterality straight from the note. Had the "right" stood, the imaging request would have named the right leg, the tech would have positioned the wrong limb, and the patient with a real wound on the left would have received a clean film of an uninjured side. The nurse flags it, the physician confirms left, the note and the order are corrected together, and the wrong-site cascade never starts. Notice that no brilliance was required and no deep clinical judgment. It required one person who treats every "left" and "right" in an AI note as a checkpoint rather than a fact.

The Misattributed Speaker, Hallucinated Specifics, and Stale Carry-Forward

Three more error types round out the catalog, and they share a family resemblance: each is a plausible-sounding piece of the note that is not anchored to what actually happened in the room. Learn to distrust the parts that sound like generic clinical history, because that is where these three live.

The Misattributed Speaker

Real encounters are crowded. A patient, a spouse, an adult child, sometimes an interpreter, all talk, and the scribe has to figure out who said what. It often gets this wrong. The spouse's report that "he's been confused since Tuesday" gets attributed to the patient as if he described his own confusion. The clinician's own musing out loud, "this could be a small stroke," gets recorded as the patient's stated concern. A symptom the daughter denied on her father's behalf gets logged as the patient denying it himself. Speaker misattribution corrupts the note in a subtle but consequential way, because in medicine who reported something changes what it means: a collateral history from a spouse and a patient's own account are different kinds of evidence, weighted differently.

The catch strategy is to read the subjective section asking not only what was said but who said it. When the note attributes a statement to the patient, check it against your memory of the room: did the patient actually say that, or did a family member, or was it your own thinking rendered as the patient's words. This matters most for history that drives a diagnosis. A patient who spontaneously reports his own confusion and a patient whose wife reports confusion he does not perceive are telling you two different clinical stories, and a note that collapses them into one has lost real information. An interpreter's clarifying question logged as the patient's statement is the same error in a different costume, and the fix is the same: reassign the words to the person who actually spoke them, or remove them.

Hallucinated Specifics and Stale Carry-Forward

The next error type is the hallucinated specific: an invented detail with the ring of authenticity, a duration ("for the past three weeks") the patient never gave, a number, a family history the visit never covered. These are pure generation, plausible details supplied to make the note read complete, and they are hard to catch because they are individually reasonable. The catch is skepticism toward specifics you do not remember: if the note contains a precise detail you cannot place, treat it as suspect until confirmed, rather than assuming you simply forgot. A hallucinated family history is especially insidious, because it reads as legitimate risk information and can quietly change a screening or risk-stratification decision downstream.

The last error type is stale carry-forward: content that migrates from a previous note or a template into today's note, describing a state that is no longer true. A resolved problem still listed as active, an old medication still in the plan, last visit's assessment reappearing as if it were today's. This is the ancient sin of copy-forward wearing a new coat, and AI can accelerate it by fluently reconstructing a note that looks current but rests on stale material. The catch is to read the note as a description of today specifically: is this the problem list, the medication list, the assessment for this encounter, or a ghost of a prior one. Anything that describes a past state as present is a carry-forward error waiting to mislead the next reader, who may be a covering physician with no memory of the room to correct against.

The scribe does not lie in obvious ways. It writes the exam a clinician usually performs, attributes the room's words to the likeliest speaker, and supplies the negatives a note like this usually contains. Every one of those defaults can be wrong, and none of them announce it.

Dropped, Added, or Flipped Pertinent Negatives

You met this failure in the last lesson; here is its full anatomy, because it takes three distinct forms. First, the added negative: the note documents "denies chest pain" for a question no one asked, manufacturing a data point out of the expected shape of a review of systems. Second, the dropped negative: a pertinent negative that was reported simply vanishes, so a patient who explicitly denied fever leaves no record of having been asked. Third, and most dangerous, the flipped negative: "reports chest pain" becomes "denies chest pain," or the reverse, inverting the actual clinical picture. A flipped negative is worse than a missing one, because a missing negative merely leaves a gap, while a flipped negative actively asserts the opposite of the truth and looks completely normal doing it.

The catch strategy is to treat the pertinent negatives that matter to your clinical reasoning as load-bearing and verify them directly. You do not need to audit every negative in a boilerplate review of systems, but any negative you actually relied on to rule something in or out deserves a direct check against what happened. If your assessment rests on the patient denying chest pain, confirm the note reflects a real denial, actually elicited, correctly recorded. The negatives that support your decisions are the ones an inverted word can quietly destroy.

Watch this one play out on a hospital floor. A hospitalist admits an older man for dizziness. During the interview the patient clearly reports intermittent chest pressure with exertion over the past week. The ambient note, drafting the review of systems in its usual tidy form, records "denies chest pain." The hospitalist, mid-admission and juggling three other patients, is about to write an assessment that leans toward a benign vestibular cause and defers a cardiac workup to the morning. But she relies on that chest-pain line, so by her own rule she verifies it directly, and she stops cold: this is a flipped negative, and it is load-bearing. The note asserts the exact opposite of what the patient told her, and it does so in the calm, standard language of a normal review of systems, which is why it very nearly justified doing less. She corrects it to "reports exertional chest pressure," escalates the workup, and the flipped word never gets a chance to steer the admission wrong. A dropped negative here would have merely left a gap she might have noticed; the flip did something more dangerous, it handed her false reassurance in the register of routine.

A Common Root and a Fast Routine

Step back and notice that four of these six error types share a single underlying cause, which is worth naming because it makes them easier to anticipate. The confabulated exam, the added negative, the hallucinated specific, and much of speaker misattribution all come from the same generative instinct: the model fills in the expected shape of a clinical note. A note of this type usually has a complete exam, so it writes one. A review of systems usually contains a set of negatives, so it supplies them. A history usually has durations and details, so it invents plausible ones. A statement in the room usually comes from the patient, so it attributes it to the patient. In every case the model is not maliciously fabricating; it is doing exactly what it was trained to do, producing the most likely text for a note like this, and "most likely" is a statement about notes in general, not about what happened in your room today.

This is a genuinely useful mental shortcut. Whenever you are reading a scribe note and you reach a part that is exactly what such a note usually contains, a full exam, a tidy set of negatives, a clean duration, a confident attribution, raise your guard, because that is precisely where the model's "fill the expected shape" instinct operates most strongly. The parts of the note that look most standard and most complete are, paradoxically, the parts most likely to have been supplied by the model rather than observed in the encounter. A messy, incomplete, specific detail is more likely to be real than a smooth, complete, generic one. Learning to feel that inversion, to be more suspicious of the polished parts than the rough ones, is one of the most valuable instincts a scribe user can develop. It is the same instinct a seasoned editor has when a paragraph reads too smoothly to be a first draft, or a radiologist has when a shadow falls exactly where the scanner's artifact usually falls: the very familiarity of the pattern is the reason to look twice.

A Thirty-Second Routine

You can compress the whole catalog into a fast routine that runs in well under a minute on most notes. One: for every documented exam or finding, ask "did I do this." Two: at every left, right, or anatomic side, stop and confirm it. Three: in the history, ask "who actually said this" for anything that drives the assessment. Four: for any pertinent negative your plan relies on, confirm it was elicited and recorded correctly. Five: for any precise specific you cannot place, treat it as suspect. Six: ask whether the problem list, medications, and assessment describe today or a prior visit. Six quick questions, each aimed at one known error type. Run them until they are automatic, and the catalog stops being a list you remember and becomes a reflex you perform. The routine is deliberately short because it has to survive a real clinic day; a verification habit that takes five minutes per note will be abandoned by the third patient, but one that takes thirty seconds and targets exactly the six things that go wrong can run on every note, on the busiest shift, without ever becoming the thing you skip to get home.

A Worked Example: One Note, Four Errors

Bring it together on a single encounter. An older woman is seen for a fall with a laceration to her left shin. Her son, in the room, mentions she has seemed more forgetful lately. The clinician cleans and closes the wound, checks her gait, and orders a head CT to be safe. The ambient note comes back polished, and it contains four classic errors, one of each major type. A confabulated exam: a full normal neurological exam including cranial nerves and sensation, when only gait was tested. A wrong laterality: the laceration documented on the right shin. A misattributed speaker: "patient reports increasing forgetfulness," when it was the son who raised it and the patient herself was unaware. And a hallucinated specific: "forgetfulness for approximately six months," a duration no one ever stated.

Now watch a clinician who knows the catalog. She reads the exam and asks "did I do this," catching and trimming the confabulated neuro exam to what she performed. She stops at "right shin," confirms it was the left, and fixes it, then glances at the CT order to confirm the laceration side is consistent. She reads "patient reports forgetfulness," recognizes it was collateral from the son, and corrects the attribution, which materially changes the picture: a patient unaware of her own cognitive change is a different and more concerning story than one complaining of it. And she flags "approximately six months," a specific she cannot place, deletes it, and notes only what was actually said, that the son reported recent forgetfulness of uncertain duration. Four errors, four known types, four targeted catches. None of this required brilliance. It required a clinician who knew what to look for, looking for it.

That is the entire promise of this field guide: the errors are predictable, and predictable errors are catchable. The scribe will keep producing confabulated exams, flipped sides, misattributed histories, and hallucinated specifics for as long as the technology works the way it does, which is to say indefinitely. What changes, once you have internalized the catalog, is that these stop being surprises that occasionally slip through and become a short list of known suspects you check for on every note, in the same reflexive way you confirm a patient's identity before a procedure. The tool does not get safer. You get better at reading it. And every one of those catches, remember, protects two things at once: the patient in front of you, and the signed record that will carry your name and your attestation into the next visit, the next audit, and, if it ever comes to it, the next courtroom. Verify, do not repeat blindly; assist, decide, and let the record prove that you did.

Key Takeaways

  • Ambient scribe errors are not random; they cluster into a small number of recognizable types, each rooted in the pipeline: capture drops, transcription mishears, and generation invents or smooths. Knowing the catalog turns verification into a targeted search for known problems.
  • Ambient documentation is now mainstream (roughly 30 percent penetration by end of 2025, about 70 percent of UCSF physicians in 2026, burnout falling from 51.9 to 38.8 percent in 30 days). Read those numbers as figures to verify, not repeat, and remember that the tool doing most of your notes is the tool you are most tempted to sign unread.
  • Every error type is silent: the note presents invented and true lines in the same confident register and never flags its own mistakes. Once you sign, it is your attestation and the legal record; "the model wrote it" is never a defense.
  • The confabulated exam finding is the signature scribe error: a complete, normal, plausible exam for a system never examined. Catch it by asking of every documented exam, "did I actually perform this," and beware automation bias, which predicts you will skip exactly the polished, plausible block that most needs a look.
  • Wrong laterality is a distinct, dangerous category: a single flipped left or right can misdirect imaging, a referral, or a procedure. Treat every stated side as a checkpoint and cross-check it against orders and referrals before the wrong-side cascade starts.
  • Speaker misattribution corrupts the history, because who reported something changes what it means; hallucinated specifics invent plausible details you cannot place; stale carry-forward describes a past state as present. Read the subjective section asking who said each thing, distrust specifics you do not remember, and read the note as an account of today specifically.
  • Pertinent negatives fail three ways: added, dropped, and flipped. The flipped negative is worst because it asserts the opposite of the truth and looks normal, and it can hand you false reassurance in the register of routine. Verify the negatives your reasoning relies on directly.
  • The skill is not brilliance but pattern recognition, compressed into a thirty-second, six-question routine: a clinician who knows the catalog and looks for each type will catch a confabulated exam, a flipped side, a misattributed history, and a hallucinated specific in the seconds before signing.