AI for Mental & Behavioral Health Clinicians
Proficient · M5 · lesson 5 of 30 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Chain-of-Thought for Differential Diagnosis
📖
now learning

Chain-of-Thought for Differential Diagnosis

15 min

A 34-year-old new client arrives at intake with a history that refuses to sit still: two years of recurrent depressive episodes, a college transcript full of missed deadlines, three or four stretches she calls "my productive weeks" when she slept five hours and reorganized her apartment, and a prior provider's chart listing three diagnoses across three years. Bipolar II (F31.81), MDD recurrent (F33.x), and adult ADHD (F90.0) all have a claim on this presentation, and the cost of guessing wrong is not academic: an antidepressant alone can destabilize an unrecognized bipolar II, and a misdiagnosed label follows a client through every payer review for a decade. This lesson teaches chain of thought prompting for clinicians: how to make a large language model show its diagnostic reasoning step by step, criterion by criterion, so you can audit that reasoning against your own and find the question you forgot to ask. You will finish with a reusable Differential-Reasoning Prompt with a built-in verification pass, and one rule carved above everything: the model reasons out loud to sharpen your thinking, and the diagnosis is yours. AI never diagnoses.

What Chain-of-Thought Prompting Actually Is, and Why Clinicians Should Care

Chain-of-thought prompting is the technique of instructing a model to produce its intermediate reasoning steps before its conclusion, instead of jumping straight to an answer. Ask a model "which diagnosis fits this presentation?" and you get a verdict: confident, fluent, unauditable. Ask the same model to "walk through each candidate diagnosis criterion by criterion, state what evidence supports or contradicts each criterion, identify what is missing, and only then summarize where the evidence points," and you get something categorically different: a visible reasoning trail you can check line by line. The verdict is a black box. The chain-of-thought answer is a glass box, and for clinical work the difference is everything, because a clinician cannot ethically use a conclusion whose reasoning cannot be inspected.

Why does forcing the steps change the quality of the output and not just its format? Because a language model generates text sequentially, and each step it writes becomes context that conditions the next. When the model must lay out criterion A, then evidence for and against it, then criterion B, its eventual summary is constrained by the trail it just built, the same way your own case formulation improves when you write it out rather than carry it as a hunch. The structure disciplines the generation. This is also why chain-of-thought output fails more visibly: an error sits in a specific numbered step where you can see it, name it, and correct for it, rather than hiding inside a confident final sentence.

Here is the controlling analogy, and it carries through every section: chain-of-thought prompting turns the model into a case-presenting supervisee, not a consulting diagnostician. Think of the sharpest associate you ever supervised, the one who presents at the 8 AM consultation group by laying out every criterion, every piece of supporting and contradicting evidence, and every open question, then waits for you to weigh in. You would never let that associate sign the diagnosis. But you would absolutely let the presentation expose the question you had not thought to ask. That is the entire value proposition, and the entire boundary, of chain of thought prompting for clinicians: a tireless, criterion-complete case presentation that you, the licensed clinician, then judge.

Why Bipolar II, MDD Recurrent, and Adult ADHD Form the Hard Triangle

This differential is one of the most consequential and most commonly fumbled triangles in adult outpatient work. All three conditions present, on the surface, with overlapping complaints: low mood, concentration problems, sleep disruption, irritability, functional impairment, a chart history of "depression" diagnosed somewhere along the way. The distinguishing data hides where a rushed intake does not reach: the texture of the "good weeks" (hypomanic episodes meeting duration and symptom criteria for F31.81, or just relief between depressive episodes of F33.x?), the developmental timeline (adult ADHD, F90.0, requires symptoms present before age 12, so the differential turns partly on report cards and childhood collateral), and the episodicity question (ADHD is trait-like and persistent; mood disorders are episodic with inter-episode baselines).

Each misdiagnosis pathway has its price. Call bipolar II "recurrent MDD" and the antidepressant that follows carries a known destabilization risk, while the hypomanic stretches keep getting coded as recovery instead of pathology. Call adult ADHD "depression" and the client gets years of treatment aimed at mood while the executive-function deficits actually driving the work failures go unaddressed. Call a mood disorder "ADHD" and a stimulant lands on an unstable mood architecture. And the conditions co-occur, which breaks either-or thinking: a client can carry both F90.0 and F33.x, and the question becomes not "which one" but "what is the full diagnostic picture and what is primary for treatment planning."

This complexity is exactly why the differential is a strong use case for chain-of-thought prompting and a terrible one for verdict prompting. A verdict ("this sounds like bipolar II") compresses a dozen unresolved evidentiary questions into false confidence. A criterion-by-criterion walk forces every one of them to the surface: Did the productive weeks last at least four days? Was there a distinct change observable by others? Were the attention problems present in third grade or did they arrive with the first depressive episode? You may know the answers. The point of the exercise is discovering which ones you do not.

The Differential-Reasoning Prompt: Actual Text, Line by Line

Here is the prompt, in full, written for a de-identified case summary you prepare first (no names, no dates of birth, no identifying detail; a BAA-covered tool if any client-derived content is involved; a fully fabricated vignette for skills practice). Read it slowly, because every line is doing structural work:

"You are assisting a licensed mental health clinician with a structured differential-diagnosis reasoning exercise. You do not diagnose; the clinician retains all diagnostic authority. Using the de-identified case summary below, walk through a structured differential between bipolar II disorder (F31.81), major depressive disorder, recurrent (F33.x), and adult ADHD (F90.0). For each candidate: (1) list the DSM-5-TR criteria relevant to the differential; (2) for each criterion, state what in the summary supports it, what contradicts it, and what is unknown; (3) name the missing information and the intake questions that would resolve it; (4) address episodicity versus trait persistence, the age-of-onset requirement for ADHD, and the duration and observability requirements for hypomania; (5) consider comorbidity: state whether the evidence could support more than one diagnosis simultaneously. Show all reasoning before any summary. Do not state a diagnostic conclusion. End with: (a) a ranked list of open questions for the clinician, and (b) a list of any place where your reasoning relied on an assumption not present in the case summary."

Notice what the prompt never asks for: a diagnosis. It asks for criteria, evidence mapping, missing-information flags, the three discriminating axes (episodicity, onset age, hypomania duration and observability), the comorbidity check, and, critically, a self-audit of assumptions. That last instruction matters more than it looks: language models pattern-complete, and when a summary says "productive weeks," a model may silently assume decreased need for sleep or goal-directed activity the summary never stated. Forcing the model to list its own assumptions converts the most dangerous failure mode, invisible fabrication, into a visible line item you can strike out. The opening sentence, the role and authority statement, is not decoration either: it sets the frame that everything downstream is reasoning support, and it is the sentence you would want in the transcript if this exchange ever appeared in a board inquiry about your AI use.

The model's reasoning exists to sharpen the clinician's thinking, never to replace it. AI never diagnoses. The differential is yours, the conclusion is yours, and the signature on the diagnosis is yours.

Reading the Output: A Worked Before-and-After

To see what the structure buys you, compare two outputs on the same fabricated vignette (the 34-year-old above). First, the verdict prompt: "Given this presentation, what is the most likely diagnosis?" A typical answer: "This presentation is most consistent with bipolar II disorder. The periods of high productivity with reduced sleep suggest hypomanic episodes, and the recurrent depressive episodes complete the picture. Adult ADHD is less likely given the prominence of mood symptoms." Three sentences, total confidence, and at least three buried failures: it assumed the productive weeks met the four-day duration criterion (never stated), it never asked whether anyone else observed a change in functioning (required), and it dismissed ADHD without touching the age-of-onset question, the single most discriminating data point available.

Now an excerpt from the chain-of-thought output on the same vignette: "F31.81, Criterion: hypomanic episode requires a distinct period of elevated, expansive, or irritable mood and increased energy lasting at least 4 consecutive days. Supports: 'productive weeks' with reduced sleep and increased activity. Contradicts: nothing directly. Unknown: duration (the summary says 'weeks' but does not confirm 4 or more consecutive days of mood change), whether the change was observable by others, whether mood was elevated or merely energy was high, and whether any period was substance-induced. F90.0, Criterion: several symptoms present prior to age 12. Supports: 'college transcript full of missed deadlines' suggests early adult impairment. Unknown: childhood history entirely absent from summary; report cards, parent report, or childhood collateral would be decisive. Assumption audit: I assumed 'reorganized her apartment' reflected goal-directed activity increase; the summary does not state whether this was one day or sustained." The second output is not smarter. It is checkable, and it hands you a ranked intake agenda: confirm hypomania duration and observability, obtain childhood collateral, screen for substance-induced episodes.

This is the supervisee analogy doing its work. The verdict output is the associate who announces "it's bipolar II" in the hallway. The chain-of-thought output is the associate who presents properly and ends with "here is what I still don't know." You would correct the first associate's process before debating the conclusion. The prompt builds the second associate by construction.

The Verification Pass: Auditing the Reasoning Before You Use It

Chain-of-thought output is auditable, which means you must actually audit it. The verification pass has four moves, and they take less time than they sound like because the structure makes errors local. Move one: criterion fidelity. Open the DSM-5-TR (the actual text, not your memory and not the model's) and spot-check the criteria the model listed. Models paraphrase criteria fluently and occasionally wrongly: a duration threshold shifted, a symptom-count requirement softened, a specifier omitted. Any criterion you intend to lean on gets checked against the text, because you are the only party in the exchange holding the authoritative source.

Move two: evidence fidelity. For every "supports" and "contradicts" line, confirm the cited fact actually appears in your case summary. The assumption audit helps, but do not rely on the model's self-report alone; scan for evidence lines that quietly upgraded "client mentioned poor sleep" into "client reported decreased need for sleep," which are different clinical facts pointing at different diagnoses (insomnia with fatigue suggests depression; decreased need for sleep without fatigue is a hypomania flag). Move three: omission hunting. Ask what the model left out entirely. Common omissions in this triangle: thyroid and other medical rule-outs, substance-induced mood episodes, trauma history that mimics all three conditions, and the borderline personality presentation that overlaps with bipolar II's affective instability. If the differential is missing a candidate your clinical experience insists on, that is your judgment outperforming the tool, which is the correct hierarchy.

Move four: the second-pass adversarial prompt. Feed the model's own output back with: "Now argue the opposite. For each candidate diagnosis you found support for, construct the strongest case against it using only facts from the case summary, and identify which single piece of missing information would most change the picture." This is the clinical equivalent of asking your supervisee to steelman the diagnosis they like least, and it reliably surfaces the confirmation drift that accumulates in any single reasoning pass, human or machine. When the adversarial pass and the original pass converge on the same open questions, you have a trustworthy intake agenda. When they diverge, the divergence is the most valuable finding on the page.

Where the Line Is: Reasoning Support Versus Diagnosis, Stated Without Softening

Now the boundary, stated the way a supervisor states it once and means it permanently. The model does not diagnose. Not "the model diagnoses and you confirm." Not "the model suggests and you usually agree." The diagnostic act, selecting and recording F31.81 versus F33.1 versus F90.0 in the chart, attaching it to a treatment plan, transmitting it to a payer, is a licensed clinical judgment that belongs to you alone, the same way the CSSRS score, the duty-to-protect determination, and the mandated-report call belong to you alone. The chain-of-thought exercise is reasoning support: it organizes criteria, maps evidence, exposes gaps. The moment its conclusion starts functioning as the conclusion, you have outsourced the one act your license exists to authorize, and no disclaimer in a prompt protects you from that.

There are practical corollaries. First, what goes in the chart is your formulation in your words: reasoning you own and can defend in a peer review or a board inquiry, never a pasted model transcript; your practice's AI-use policy and log govern how AI assistance is recorded. Second, the model never meets the client. Everything it reasons over is your summary, so every weakness of your summary propagates: if you did not capture the affective texture of the "productive weeks," the model cannot reason about it, and its confident criterion-mapping will be confidently incomplete. The clinician is load-bearing at both ends. Third, the privacy frame applies fully: de-identify before anything client-derived touches a model, use BAA-covered tools for anything that is or could be PHI, and prefer fabricated vignettes when the goal is sharpening skills rather than working a live case.

Finally, name the seduction so you can resist it. Chain-of-thought output is persuasive precisely because it shows its work; a wrong conclusion wrapped in ten plausible steps is more dangerous than a naked wrong verdict, because the steps borrow your trust. The discipline is to grade the steps, not admire them. A supervisee who presents beautifully and reasons wrongly still reasons wrongly, and a supervisor charmed by the presentation has stopped supervising.

When to Run the Exercise, and When Not To

Chain-of-thought differential work earns its time in specific situations. The complex new intake where the history spans multiple prior diagnoses and you want a criterion-complete map before session two. The case you are preparing for consultation group, where the gap list becomes your "open questions" slide. The stalled-treatment review where you suspect the working diagnosis, inherited from a prior provider, was never properly differentiated. The supervision context, where a supervisor assigns the exercise to an associate and reviews both the model's reasoning and the associate's audit of it, teaching differential thinking and AI verification in one move. And deliberate practice: running fabricated vignettes through the prompt and grading the output against the DSM-5-TR is a skills gym for your diagnostic reasoning, with zero privacy exposure.

Where it does not belong: anywhere speed pressure would convert reasoning support into verdict adoption. Do not run it ten minutes before the session and skim the summary line. Do not run it on a crisis presentation; active suicidal ideation, abuse disclosure, and acute risk are human-only territory across this entire program. Do not use it to settle a disagreement with a colleague by appeal to the model's output; the model is not an authority, it is a presentation, and presentations get argued with. And do not let it substitute for the instruments and collateral that actually discriminate this triangle: validated screeners administered and interpreted by you, childhood records, family interviews, longitudinal mood charting. The model organizes evidence; it does not generate it.

A scheduling note from the trenches: the exercise works best between sessions one and two of a complex intake, when you have a rich first-session picture, an unsettled differential, and a second appointment in which to ask what the gap list surfaces. Used there, the ranked open questions become your session-two agenda, which is the most concrete payoff this technique offers: not an answer, but a better next hour of questions.

The Applied Problem: Build Your Reusable Differential-Reasoning Prompt with Verification Pass

Your artifact is a saved, reusable Differential-Reasoning Prompt with Verification Pass: a two-part prompt block stored with your working prompts, ready for any complex differential, with this lesson's triangle as its test case. Build it in four steps.

Step one, write the case-summary template that feeds the prompt: a de-identification checklist header (no names, no dates of birth, no locations, no unique identifying details, BAA-covered tool if client-derived) and a structured skeleton: presenting complaints, episode history with durations as precisely as known, sleep pattern in clinical terms (insomnia with fatigue versus decreased need for sleep), functional timeline back to childhood, substance history, medical rule-outs already done, prior diagnoses with sources, collateral available. The skeleton matters because the model reasons only over what you feed it; a disciplined summary template is half the technique.

Step two, install the main prompt, generalized: replace the three named candidates with bracketed fields ([Candidate 1 with ICD-10 code], [Candidate 2], [Candidate 3]) while keeping every structural instruction intact: criteria listing, supports-contradicts-unknown mapping, missing-information ranking, the discriminating-axes instruction, the comorbidity check, the no-conclusion rule, the assumption self-audit. Step three, append the verification pass as a standing checklist: (1) spot-check every load-bearing criterion against the DSM-5-TR text; (2) trace every supports and contradicts line back to the summary, hunting for silent upgrades; (3) hunt omissions, including medical, substance, trauma, and personality-presentation rule-outs; (4) run the adversarial second-pass prompt ("argue the opposite using only facts from the summary") and compare gap lists.

Step four, test the artifact on the fabricated 34-year-old from this lesson. Run the prompt, run the verification pass, and grade yourself on one criterion: did the exercise produce at least three intake questions you had not already planned to ask? If yes, it works. Done looks like a two-part prompt block with a summary template, a generalized reasoning prompt, a four-move verification checklist, and a top line you never delete: the model presents, the clinician diagnoses, and the diagnosis in the chart is yours alone.

Key Takeaways

  • Chain-of-thought prompting instructs the model to show its intermediate reasoning, criterion by criterion, before any summary. It converts an unauditable verdict into a glass-box presentation whose errors are local, visible, and correctable: the only form of model reasoning a clinician can ethically work with.
  • The bipolar II (F31.81) versus MDD recurrent (F33.x) versus adult ADHD (F90.0) triangle is hard because the surface symptoms overlap, the discriminating data hides in episodicity versus trait persistence, the before-age-12 onset requirement for ADHD, and the four-day duration and observability requirements for hypomania, and the conditions can co-occur.
  • The Differential-Reasoning Prompt demands criteria listing, supports-contradicts-unknown mapping per criterion, a ranked missing-information list, the comorbidity check, a ban on diagnostic conclusions, and an assumption self-audit that converts silent fabrication into visible, strikeable line items.
  • The verification pass has four moves: check criteria against the actual DSM-5-TR text, trace every evidence line back to the case summary to catch silent upgrades like "poor sleep" becoming "decreased need for sleep," hunt omitted rule-outs (medical, substance, trauma, personality presentations), and run an adversarial second pass that argues the opposite.
  • The controlling frame: the model is a case-presenting supervisee, never a consulting diagnostician. Its presentation can expose the question you forgot to ask; it can never sign the diagnosis. The diagnostic act belongs to the licensed clinician alone. AI never diagnoses.
  • Chain-of-thought output is dangerous in proportion to its persuasiveness: ten plausible steps wrapped around a wrong conclusion borrow your trust. Grade the steps, never admire them, and remember the model never met the client, so every weakness in your summary propagates.
  • Run the exercise between sessions one and two of a complex intake, in consultation prep, in stalled-treatment reviews, and as deliberate practice on fabricated vignettes. Never under speed pressure, never on crisis presentations, never as a substitute for validated instruments, collateral history, and your own observation.