Managing the Skeptics and the Over-Trusters
On the same medicine floor, two physicians received the same well-validated sepsis alert about the same kind of deteriorating patient. The first, a committed skeptic, had long since decided the model cried wolf and dismissed it without a glance; her patient's decline was caught late, the old-fashioned way, with a worse outcome than necessary. The second, an enthusiastic adopter, saw the alert, trusted it completely, and started the bundle without pausing to confirm the clinical picture, which happened this time to be a mimic that the model had misread; he treated a patient who did not need it and missed what was actually wrong. One refused a good tool. One deferred to it blindly. Both were wrong, in opposite directions, and a leader who only worries about one of them has misunderstood the adoption problem entirely.
Two Opposite Failure Modes, One Adoption Problem
Most change-management thinking treats adoption as a single dial, running from resistance to enthusiasm, with the leader's job being to turn the dial up. That model is dangerously incomplete, because it treats enthusiasm as the goal and resistance as the only enemy. In clinical AI there are two failure modes, and they sit at opposite ends of the dial. The skeptic refuses or ignores a useful, well-verified tool, leaving its value, and often the clinician's own sanity and time, on the table. The over-truster defers to the tool blindly, accepting its outputs without the verification that keeps a model error from becoming a patient harm, which is automation bias by another name. Turning the adoption dial up moves people from the first failure toward the second. The goal is not high on the dial; it is the right place on the dial, which is calibrated use: trusting the tool enough to capture its value, distrusting it enough to keep checking it.
Seeing adoption as two-tailed rather than one-tailed changes what a leader optimizes for. If you measure success by usage, you will celebrate the over-truster and worry only about the skeptic, and you will systematically push your workforce toward the more dangerous of the two errors. If instead you measure success by calibrated, verifying use, you recognize that both the clinician who refuses the tool and the clinician who defers to it without checking have failed the same underlying task, which is to weigh the tool's output for what it is actually worth on this patient. Neither blanket refusal nor blind trust is acceptable, because neither is thinking. The entire discipline of managing adoption is moving both tails toward the calibrated middle.
It helps to lay the two-tailed picture out as a map, because leaders who can name where a clinician or unit actually sits can choose an intervention instead of pushing indiscriminately. The four positions below are the terrain.
| Position | What it looks like | Direction to move |
|---|---|---|
| Legitimate skeptic | Distrusts a tool that was never locally validated or whose limits were hidden | Do not move them; meet the demand with evidence, honesty, and protected time |
| Managing skeptic | Refuses or ignores a tool that is genuinely well-validated, out of habit or alarm fatigue | Toward center with evidence in their terms and by fixing the noise behind the fatigue |
| Calibrated user | Uses the tool and still verifies high-stakes outputs against the patient | Hold here; defend the verification structure |
| Over-truster | Accepts outputs as answers, signs drafts unread, follows scores as verdicts | Toward center with structural verification the comfort will not supply |
Read the map and the management strategy falls out of it. If you believe there is one enemy, resistance, you spend your energy pushing, and every unit of push moves the whole distribution toward the enthusiasm end, which means toward over-trust. You will get exactly what you optimized for: a workforce that uses the tool a great deal and checks it very little, and you will not notice the problem, because it will look on paper like a resounding success. If instead you believe there are two enemies at opposite ends, you stop pushing indiscriminately and start doing something more surgical: you identify which position a given clinician or unit is actually in, and you apply the intervention that moves that specific tail toward center. The refuser and the blind adopter are not on a spectrum where one is closer to the goal than the other; they are equidistant from calibration in opposite directions, and pretending otherwise is how well-run rollouts quietly become unsafe.
Adoption is not a single dial from resistance to enthusiasm. It has two failure modes at opposite ends: the skeptic who refuses a good tool and the over-truster who defers to it blindly. The goal is not high on the dial. It is calibrated in the middle.
Understanding the Skeptic (and Why Some Skepticism Is Right)
The first move with a skeptic is to resist the temptation to treat all skepticism as an obstacle to be overcome, because a great deal of clinical skepticism about AI is correct and valuable. A clinician who distrusts a model that was never locally validated, whose failure modes were never disclosed, that adds work without removing more, is not being difficult; they are applying exactly the professional judgment the whole program is built to cultivate. That kind of skepticism is a feature, and a leader who tries to steamroll it with mandates is destroying the very calibration they should be building. The earlier lesson on earning trust is the correct response to this skeptic: give them the evidence, the local validation, the honesty about limits, and the protected time, and the skepticism dissolves because its legitimate demands were met.
The skeptic who actually needs managing is a different animal: the one who refuses or ignores a tool that has been well-validated, whose value is demonstrated, out of habit, blanket distrust, alarm fatigue, or a prior bad experience generalized too far. This skeptic is leaving real value on the table, and sometimes leaving a patient less safe, as the physician who ignored the good sepsis alert did. The management here is not a mandate but evidence delivered in the skeptic's own terms: the local data showing the tool works here, the cases where it caught something, the peers they respect who use it well. Alarm fatigue in particular has to be addressed at the source, by fixing the flood of false positives that trained the dismissal, not by ordering the clinician to stop being fatigued. You cannot argue someone out of alarm fatigue; you have to earn back the signal by cleaning up the noise. The reflex to reach for a mandate is strongest with this skeptic and is almost always wrong, because a coerced skeptic does not become a calibrated user; they become a resentful one who complies mechanically, which is how you accidentally manufacture an over-truster.
The distinction between the two skeptics is not academic; it determines whether your intervention helps or backfires. Treat the legitimate skeptic like a managing skeptic, and you steamroll a correct judgment and lose a person who was actually doing the program's work. Treat the managing skeptic like a legitimate one, and you keep supplying evidence that has already been supplied while the real barrier, alarm fatigue or a generalized bad memory, goes unaddressed. So the diagnostic question comes first: is this clinician distrusting a tool that genuinely has not earned trust, or refusing one that has? The answer decides everything that follows, and getting it wrong is one of the most common ways a leader wastes months.
Understanding the Over-Truster (the Quieter, Deadlier Risk)
The over-truster is the more dangerous of the two, and the harder to see, because their failure looks like success on every dashboard a leader is likely to watch. High usage, fast adoption, enthusiastic feedback, all the metrics that make a rollout look triumphant, are exactly the metrics an over-truster produces. The over-truster accepts the AI's output as an answer rather than an input, signs the draft without reading it against the source, follows the risk score as a verdict, and stops performing the verification that renders the model's inevitable errors inert. This is automation bias in its native habitat, and it is intensified by precisely the conditions a busy clinical environment guarantees: time pressure, high volume, fatigue, and a tool that has been right often enough to earn a trust the rare failure then betrays.
Managing the over-truster is harder than managing the skeptic for a structural reason: the over-truster feels safe. The skeptic knows they are in tension with the tool; the over-truster feels the comfortable ease of a job made lighter, and that comfort is the danger. So the interventions are different. You cannot rely on the over-truster to self-correct through discomfort, because they feel none. Instead you build the verification into the workflow as a structural requirement rather than a personal choice, exactly as the automation-bias lesson taught: fixed verification steps on high-stakes outputs, forcing functions, an attestation that requires the clinician to affirm they checked. You teach the specific metacognitive skill of treating the comfortable feeling that the tool is always right as the alarm rather than the reassurance. And you keep it culturally legitimate, even expected, to question and verify the AI, so that the path of least resistance is verification rather than deference. The over-truster will not build these guards for themselves, because from the inside nothing feels wrong, so the leader and the system must build them.
A particularly cruel feature of the over-truster problem is that the best tools produce the worst version of it. A clunky, mediocre tool never earns enough trust to make anyone defer blindly; the over-truster is created by the excellent tool, the one that is right so often that checking it comes to feel like a pointless ritual. This is the same paradox that runs through the automation-bias foundation of this program: reliability is what dismantles vigilance, so the tools you have most successfully adopted are precisely the ones your over-trusters have most stopped watching. For a leader, this means the over-trust risk is highest exactly where the rollout looks most triumphant, on the mature, beloved, high-adoption tool that everyone has come to rely on. The instinct to relax oversight on a tool that has performed well for a year is understandable and exactly backward. The mature tool needs its verification structure defended, not dismantled, because it is carrying the most trust and therefore the most over-trust.
The Standard of Care Cuts Both Ways
There is a legal spine underneath this that turns calibration from a management preference into a clinical duty, and it is worth making explicit because it is the reason both tails are genuinely dangerous rather than merely one of them being inconvenient. The evolving standard of care now runs in two directions at once. A clinician can be found liable for following a wrong AI recommendation that a reasonable clinician would have questioned, and a clinician can also be found liable for ignoring an accurate AI recommendation that a reasonable clinician would have heeded. The skeptic and the over-truster each expose the clinician and the institution to opposite ends of the same liability, and a leader who understands this stops thinking of the skeptic as merely under-adopting and the over-truster as merely over-adopting. Both are standing on legal risk.
The table makes the symmetry concrete, and it also shows why documentation is the through-line that protects against both.
| Failure mode | Liability exposure | What defends against it |
|---|---|---|
| Over-truster follows a wrong AI recommendation | Liable for acting on an output a reasonable clinician would have questioned | A verification step and a short note showing the clinician confirmed the picture before acting |
| Skeptic ignores an accurate AI recommendation | Liable for disregarding an output a reasonable clinician would have heeded | A short note showing the clinician considered the alert and gave a clinical reason for disagreeing |
Notice what the right-hand column shares. In both directions, the thing that converts a defensible clinician into an indefensible one is the absence of a recorded, reasoned engagement with the AI output. The over-truster who signs blindly cannot show they weighed the output; the skeptic who dismisses reflexively cannot show they considered it. Calibrated use is not only safer for the patient, it is the posture that generates the record, a brief note explaining why the clinician agreed or disagreed with the AI, that materially strengthens the defense in either direction. This is why the two-tailed problem is not a soft cultural preference. It is the operational form of a legal reality, and it is why a leader who tolerates either tail is tolerating exposure.
This legal spine also changes how a leader should talk about the program to clinicians who fear that using AI at all increases their liability. The honest answer is that both using it uncritically and refusing it categorically carry exposure, and the only posture that reduces liability in both directions is calibrated, documented use. A clinician who confirms the clinical picture before acting on an alert and writes a one-line reason has done the thing that defends the decision if the alert was wrong; a clinician who considers an alert, disagrees for a stated clinical reason, and records that has done the thing that defends the decision if the alert was right. In neither case did the AI make the clinician safer by itself; the documented human judgment did. Framing calibration this way, as the liability-reducing posture rather than an extra burden, is often what finally moves a wary clinician, because it connects the abstract idea of calibration to the concrete thing they most want to protect, which is their license and their patient.
It is worth being explicit that this is not a call to write a paragraph on every AI output; that would be its own kind of failure, drowning the record and the clinician in defensive documentation. The note that matters is short and specific, and it is reserved for the high-stakes moments where the clinician either overrode the tool or relied on it in a consequential decision. A leader who asks for calibrated documentation everywhere will get resentment and boilerplate; a leader who targets it at the decisions that actually carry weight will get a record that means something. The skill being cultivated is judgment about when the engagement needs to be visible, which is itself part of the calibrated professional stance the whole lesson is trying to build.
Moving Both Toward Calibrated Use
The unifying insight is that both tails are moved toward the calibrated middle by the same three forces, applied differently: evidence, design, and culture. Evidence moves the skeptic by showing the tool genuinely works here, and moves the over-truster by showing, concretely, the errors the tool makes and the verification that catches them, so that trust becomes specific rather than blanket. Design moves the skeptic by removing the friction and false alarms that fed the refusal, and moves the over-truster by building in the verification steps that their comfort will not supply. Culture moves both by making the calibrated stance, using the tool and checking it, the normal, respected, expected behavior, so that neither reflexive refusal nor blind deference is the social path of least resistance.
The table below shows how each force is aimed differently at each tail, which is the practical core of the whole lesson: the forces are shared, but the application is opposite.
| Force | Applied to the skeptic | Applied to the over-truster |
|---|---|---|
| Evidence | Local data and respected peers showing the tool works here | Concrete examples of the tool's errors and the checks that catch them |
| Design | Remove false alarms and friction that fed the refusal | Build fixed verification steps and attestations the comfort will not supply |
| Culture | Make attending to a good alert the respected norm | Make questioning and verifying the AI the respected, expected norm |
Notice that a mandate appears nowhere in that table, and its absence is deliberate. A mandate is the tool that most tempts a leader facing a skeptic and most endangers a leader facing an over-truster. It coerces the skeptic into resentful mechanical compliance, which is not calibration but a manufactured over-trust, and it signals to the whole workforce that the point is usage rather than good judgment, which is precisely the wrong lesson. The leader's job is not to force the dial to a number but to cultivate, in a whole workforce, the professional judgment to place each AI output correctly: neither refused out of hand nor swallowed whole, but weighed. That judgment is a culture as much as a skill, and it is built with evidence and design, not decree.
There is a leadership honesty required here that is easy to dodge. Cultivating calibrated judgment in a whole workforce is slower, harder, and less legible than mandating usage and watching a number climb. It does not produce a clean adoption chart to show the board next quarter. It produces something better and quieter: a workforce that uses good tools well and bad tools warily, that catches the errors the tools inevitably make, and that will tell you when a tool starts failing in the field. A leader who wants the fast chart will reach for the mandate and manufacture over-trust; a leader who wants a safe program will do the patient work of evidence, design, and culture and accept that calibration cannot be decreed into existence. The two-tailed problem does not have a one-lever solution, and the leaders who look for one are the ones who end up, months later, wondering why their beautifully-adopted tool produced a string of quiet safety events that no dashboard warned them about.
A Worked Example: Two Interventions on the Same Unit
Return to the medicine floor and the two physicians. The skeptic who ignored the good sepsis alert does not need a disciplinary note about alert compliance; that would harden her refusal into resentment. She needs to be shown, by a respected colleague and with local data, that this particular alert, after the false-positive problem was cleaned up, now fires with enough precision to be worth her attention, and the specific cases where it caught a deterioration she would have seen late. Her skepticism was partly earned by an earlier era of noisy alerts, so the intervention respects that history and answers it with evidence rather than overriding it with authority. Move her by earning back the signal, and she becomes a calibrated user who attends to the alert and still applies her own judgment, which is exactly what you want.
The over-truster who ran the bundle on the mimic needs the opposite intervention, and crucially he needs one that does not depend on him noticing he has a problem, because he does not feel that he does. The unit builds a required verification step: acting on the sepsis alert prompts a brief, structured confirmation of the clinical picture, an attestation that the clinician considered the alternatives, before the bundle proceeds. The teaching accompanying it names the trap directly: the alert is an input to your judgment, not a replacement for it, and the moment it feels most obviously right under time pressure is the moment to confirm rather than defer. The structural step catches the error even on the day his vigilance is gone, which is the whole point, because you cannot solve a failure of attention with more attention. Same alert, same unit, two clinicians at opposite failure modes, two different interventions, one shared destination: calibrated use, where the tool is trusted enough to help and checked enough to be safe.
The final thing to notice about this worked example is that the two interventions are not symmetric in what they rely on. The skeptic's intervention works through her judgment: it gives her better information and trusts her to reweigh the alert accordingly, which is appropriate, because her problem was never a failure of care or attention but a rational response to old evidence. The over-truster's intervention deliberately does not rely on his judgment in the moment, because his problem is precisely that his in-the-moment judgment feels fine while being unguarded; so the fix is structural, a step that fires whether or not he is paying attention. This asymmetry is the practical heart of the lesson. You move the skeptic by respecting and informing their judgment, and you protect against the over-truster by not depending on theirs when the stakes are high. A leader who confuses the two, who tries to fix the over-truster with a persuasive talk or the skeptic with a forcing function, will fail both. The right intervention matches the shape of the failure, and the two failures have opposite shapes.
Key Takeaways
- Adoption has two opposite failure modes, not one: the skeptic who refuses a useful, well-verified tool and leaves value and safety on the table, and the over-truster who defers blindly, which is automation bias; the goal is calibrated use in the middle, not high usage.
- Measuring success by usage rewards the over-truster and pushes the workforce toward the more dangerous error; measure calibrated, verifying use instead, and diagnose which of four positions a clinician sits in before intervening.
- Distinguish the legitimate skeptic (distrusting an unvalidated tool, whose demands you meet with evidence and honesty) from the managing skeptic (refusing a proven tool), because treating one like the other backfires.
- Fix the managing skeptic with evidence in their own terms and by cleaning up the false-positive flood behind alarm fatigue, never by decree, because a coerced skeptic becomes a resentful, manufactured over-truster.
- The over-truster is the quieter, deadlier risk because their failure looks like success on every dashboard, and the best, most-adopted tools produce the worst over-trust, so defend the mature tool's verification structure rather than relaxing it.
- Manage the over-truster structurally: fixed verification steps, forcing functions, and attestations on high-stakes outputs, plus the metacognitive habit of treating the comfortable feeling that the tool is always right as the alarm.
- The standard of care cuts both ways: a clinician can be liable for following a wrong AI recommendation and for ignoring an accurate one, and in both directions a short, reasoned note engaging the AI output is what defends the decision, so calibrated use is a legal posture, not only a cultural one.
- Both tails move toward the middle through evidence, design, and culture applied in opposite directions; a mandate belongs nowhere in the toolkit, because it manufactures over-trust and teaches that usage, not judgment, is the point.
Skill.re