Forensic Evaluation Write-Ups, with Hard Limits
A licensed psychologist finishes a child custody evaluation: three home visits, collateral interviews with two teachers and a pediatrician, an MMPI-3 on each parent, forty hours of work, and a fifty-page report due to the court in nine days. Thirty of those pages are recitation: who was interviewed and when, which instruments were administered and what they are, the procedural history of the case, the demographic and developmental background. Ten pages are the part the judge actually reads: the findings, the opinions, the recommendations. The temptation is obvious, and so is the catastrophe: an evaluator whose opinions were drafted by a language model has handed opposing counsel the cross-examination of the decade. This lesson draws the line that keeps forensic work defensible: AI may structure the boilerplate, methodology, instruments administered, demographic and historical recitation, and AI never drafts opinions, conclusions, or psycho-legal recommendations. We anchor the line in the AAPL Practice Guidelines and the APA Specialty Guidelines for Forensic Psychology, and in the Greenberg and Shuman role distinction that explains why the same clinician cannot be both therapist and forensic evaluator for the same person. By the end you will build a Forensic-Report Skeleton with every AI-permitted section explicitly marked, an artifact you can defend under oath.
Two Different Jobs: The Greenberg and Shuman Role Distinction
Before a single prompt is written, get the roles straight, because every AI rule in this lesson flows from them. Greenberg and Shuman's classic analysis of the therapeutic and forensic roles is the field's foundational text on why these are different jobs that cannot be held by the same clinician for the same client. The treating therapist works for the patient: the relationship is built on alliance and advocacy, the data is the patient's subjective report accepted in the service of healing, the standard is clinical utility, and confidentiality is the room's load-bearing wall. The forensic evaluator works for the court or the retaining party: the relationship is explicitly non-therapeutic, the data is verified across collateral sources precisely because self-report is treated with skepticism, the standard is legal relevance and admissibility, and the evaluee is warned at the outset that nothing said is confidential and that the evaluator is not their clinician.
Mixing the roles corrupts both. A therapist who opines forensically on their own client trades the alliance for an opinion the court should not trust, because therapy never gathered the adversarially tested data a forensic opinion requires. An evaluator who slides into a helping stance compromises the neutrality the court retained them for. This is why the rule is stated flatly: the same clinician cannot serve both roles for the same client. The APA Specialty Guidelines for Forensic Psychology and the AAPL Practice Guidelines both institutionalize this boundary, and the next lesson in this chapter, the subpoenaed treating therapist, lives entirely on the therapeutic side of it. Today we are on the forensic side, where you were retained as an evaluator, custody evaluation, fitness-for-duty, independent medical examination (IME), and the product is a report headed for an adversarial proceeding.
Hold this lesson's controlling analogy: a forensic report is a load-bearing bridge, and the evaluator is the structural engineer who stamps it. Contractors can pour the standard footings, the approach roads, the guardrails, the parts built to known specifications. But the span calculations, the judgment about whether this bridge holds this load over this river, carry the engineer's stamp and the engineer's liability, and no engineer outsources the stamp. In this lesson, AI is the contractor on the standard sections. The opinions are the span. Nobody but the evaluator touches the span.
Why the Forensic Report Is the Most Adversarially Read Document You Will Ever Write
A progress note might never be read by anyone but you and an auditor. A forensic report is read, line by line, by people paid to destroy it. Opposing counsel will depose you on your methodology, your data sources, the basis for every opinion, and, increasingly, on your process: who drafted this report, what tools were used, what was generated versus authored. Discovery requests now routinely probe drafting process, and "did artificial intelligence draft any portion of your opinions, Doctor?" is a question already being asked in depositions. If the answer is yes for the opinion sections, the cross-examination writes itself: the expert's opinion is the entire product the court is paying for, and an opinion assembled by a text predictor is an opinion the expert cannot fully trace, defend, or own. The witness who answers "the methodology and history sections were formatted with AI assistance from my verified notes; every opinion and recommendation was authored solely by me" survives. The witness who hesitates does not.
The professional anchors are explicit. The AAPL Practice Guidelines for forensic psychiatric evaluation and the APA Specialty Guidelines for Forensic Psychology both place the integrity of the evaluator's reasoning at the center of forensic practice: opinions must rest on the evaluator's own examination of sufficient data, the reasoning chain from data to opinion must be transparent, and the evaluator must be able to account for the basis of every conclusion. A language model breaks that chain in a way no disclosure can repair, because the model's contribution to an opinion is not traceable even by the person who prompted it. The boilerplate sections, by contrast, are exactly the kind of standardized recitation the guidelines treat as procedural scaffolding: accurate, complete, and formatted to convention, with no judgment embedded.
There is also a confidentiality problem that arrives before the drafting problem. Forensic files contain the most litigation-sensitive material in clinical practice: psychological test data, collateral statements, criminal and child-protective records. Any AI tool touching a forensic file needs the same contractual scrutiny as a clinical tool, a BAA where health information is in play, plus the practical question of whether evaluation data may flow into vendor analytics or model improvement at all. Many forensic evaluators conclude the only acceptable configuration is a tool with zero retention and no training use, or no tool at all for certain matters. Decide this per-matter, in writing, before intake.
The Anatomy of the Report: Sorting Sections into Contractor Work and Span Work
Walk the standard report skeleton, custody, fitness-for-duty, or IME, and sort each section. Referral question and procedural history: contractor work. This section recites who retained you, the legal question as posed, the order or stipulation under which you act, and the case's procedural posture. It is factual recitation from documents; AI can structure it from the evaluator's verified source list. Notification and informed consent: contractor work with a caveat. The section documenting that the evaluee was told the evaluation is non-confidential, non-therapeutic, and court-bound is standardized language, but the fact that the warning occurred is the evaluator's attestation; AI formats the paragraph, the evaluator confirms the event.
Sources of information: contractor work, and one of AI's most useful jobs. A forensic report lists every interview with date and duration, every document reviewed, every collateral contact, every instrument administered. AI excels at converting the evaluator's log into a complete, consistently formatted source list, and at cross-checking that every source cited in the body appears in the list and vice versa, an error class that opposing counsel hunts for sport. Instruments administered: contractor work for the descriptions. Boilerplate descriptions of the MMPI-3, the PAI, a structured interview protocol, what the instrument is, what it measures, its general use, are standardized text the evaluator maintains in a verified library. AI assembles and formats them. The scores, the validity-scale interpretation, and what the results mean for this evaluee: span work, untouched.
Demographic, developmental, and psychosocial history: contractor work at the recitation level. The history section organizes verified facts, education, employment, relationships, medical and psychiatric history, legal history, from records and interviews into chronological narrative. AI structures the timeline from the evaluator's extracted facts, and only from them: a standing instruction prohibits the model from inferring, characterizing, or filling gaps, because a hallucinated date or an invented inconsistency in a forensic history is a credibility wound that bleeds across the whole report. Behavioral observations, mental status findings, test interpretation, the clinical formulation, the answers to the psycho-legal questions, and the recommendations: span. All of it. AI does not draft opinions, conclusions, or psycho-legal recommendations, full stop, and it does not draft the "discussion" tissue that connects data to opinion, because that tissue is the opinion.
The opinion section of a forensic report is the one product the court actually retained you for; an opinion a language model helped compose is an opinion you cannot fully trace, defend, or own under oath.
The Hard Limits, Stated Plainly
Write these into your forensic practice policy verbatim, because under deposition you will want to testify that they existed in writing before this case. One: AI drafts forensic boilerplate only, methodology descriptions, instrument descriptions, source lists, procedural history, demographic and historical recitation from evaluator-verified facts. Two: AI never drafts opinions, conclusions, psycho-legal recommendations, test interpretations, behavioral observations, or the reasoning connecting data to opinion. Three: AI never scores or interprets any psychological instrument; scoring follows the test publisher's authorized procedures, and interpretation is the evaluator's. Four: AI is never used to draft testimony, deposition answers, or talking points; the witness's voice must be the witness's, a rule we expand in the next lesson. Five: every AI-assisted section is generated from evaluator-supplied, verified content, reviewed word by word, and the evaluator can identify, on the record, which sections received formatting assistance and attest that no opinion content did.
Notice what the limits protect. They protect the evaluee, whose liberty, employment, or parental rights ride on an opinion that must be a human expert's. They protect the court, whose gatekeeping over expert testimony assumes the expert's reasoning is the expert's. And they protect you: the evaluator who can answer the AI deposition question cleanly, with a written policy, a marked skeleton, and a verification log, has converted a cross-examination trap into a demonstration of rigor. The evaluator who cannot is one exhibit away from a Daubert-flavored challenge to everything they have ever signed.
One more limit, easy to miss: AI never sees what it does not need. The contractor pours footings from the engineer's drawings, not from the engineer's entire file cabinet. Prompts for boilerplate sections carry the minimum content required, the verified fact list for the history section, the instrument names for the descriptions, never the raw test data, never the collateral interviews wholesale, never the protected records of third parties who appear in the file. Forensic files are full of other people's information, the other parent, the complaining employee, the children, and none of those people consented to a vendor's servers.
The Worked Workflow: A Fitness-for-Duty Report, Section by Section
Make it concrete with a fitness-for-duty evaluation: an employer-retained evaluation of a law enforcement officer following a critical incident. The evaluator has completed the work: review of personnel and incident records, two clinical interviews, collateral interviews with a supervisor, MMPI-3 administered and scored through the publisher's system, notes verified. Drafting begins with the skeleton, and the skeleton begins with marks: each section header carries a bracket tag, [AI-PERMITTED: format from verified inputs] or [EVALUATOR ONLY], assigned before any drafting, so the boundary is an artifact of the file, not a memory.
For the permitted sections, the prompt pattern is constant: "Format the following evaluator-verified content into the [section name] section of a fitness-for-duty evaluation report. Use formal forensic report register. Include only the content provided. Do not infer, characterize, summarize beyond the given facts, or add transitional commentary that evaluates the evaluee. Flag any gap rather than filling it." The evaluator feeds the referral letter facts for the referral section, the dated source log for sources of information, the instrument list for the standardized descriptions, and the extracted, verified history facts for the background section. Each draft gets a word-by-word read against the source material, the same discipline as every signature in this program, because the evaluator's name on the report attests every sentence, formatted or not.
Then the contractor goes home. The evaluator writes, personally and from scratch: behavioral observations from the interviews, the mental status findings, the test interpretation built from the publisher-scored profile, the formulation, the opinion on the psycho-legal question (fitness for duty, with or without conditions), and the recommendations. No model in the loop, no "improve my phrasing" pass on opinion text, because a rephrased opinion is a co-authored opinion. The last AI task is the one place the model returns: a verification audit of the assembled report. "List every factual assertion in sections 1 through 5 and the source list entry it should correspond to; flag any assertion without a source, any source without a citation, and any inconsistency in dates or names." The model checks the contractor's work. It still never inspects the span: the audit prompt covers the recitation sections only, and the opinion sections go to the evaluator's own final read and, where practice standards suggest, a peer-consultation review.
Answering the Deposition Question Before It Is Asked
Build the answer into the file from day one. Three artifacts make the AI question a non-event. First, the written forensic AI policy, the five hard limits above, dated before the case began. Second, the marked skeleton: the report template with section-by-section [AI-PERMITTED] and [EVALUATOR ONLY] tags, showing the boundary was structural. Third, the drafting log: a one-page record noting, for each AI-assisted section, the date, the tool, the input class (verified source log, instrument library, extracted history facts), and the evaluator's verification read. With those three documents, the deposition exchange runs: "Did AI draft any portion of your opinions?" "No. My written policy prohibits it, my report template marks the opinion sections evaluator-only, and my drafting log shows AI assistance was limited to formatting the source list, instrument descriptions, and history recitation from facts I verified. Every opinion, interpretation, and recommendation in this report I authored personally."
That answer does more than survive; it lands as rigor. The expert who can produce a documented quality-control process around drafting looks more careful than the expert who typed everything into Word with no process at all. This is the recurring pattern of the whole program: the clinicians and evaluators who fare best with AI are not the ones who avoid it entirely or adopt it wholesale, but the ones whose boundaries are written, structural, and provable. The forensic context just raises the stakes, because here the document's entire purpose is to withstand attack.
A closing word on role purity. If you are the treating therapist and someone asks you for a custody recommendation, a parenting-capacity opinion, or a fitness conclusion about your own client, the answer is not a better prompt; the answer is Greenberg and Shuman. The treating clinician does not perform the forensic role for their own client, does not opine on parenting capacity, and refers the forensic question to an independent evaluator. The next lesson teaches exactly how to hold that line when the subpoena arrives with your name on it.
The Applied Problem: The Forensic-Report Skeleton with AI-Permitted Sections Marked
Your artifact is the Forensic-Report Skeleton: a reusable report template for your evaluation type (custody, fitness-for-duty, or IME) in which every section carries an explicit tag, [AI-PERMITTED: format from verified inputs only] or [EVALUATOR ONLY: no AI contact], plus the standing prompt block for the permitted sections and the three-line drafting-log format. This artifact is simultaneously a workflow tool and deposition exhibit A.
Step one, list the sections of your standard report in order: referral question and procedural history; notification and informed consent; sources of information; instruments administered (descriptions); demographic, developmental, and psychosocial history; behavioral observations and mental status; test results and interpretation; clinical formulation; opinions on the psycho-legal questions; recommendations. Tag each one yourself, on paper, applying the rule: recitation of verified facts and standardized descriptions is contractor work; anything containing observation, interpretation, judgment, or the connective reasoning between data and conclusion is span work. If a section mixes both (instruments administered often does), split it into a tagged pair. Step two, have AI format the skeleton: "Format this evaluator-tagged section list into a forensic report template. Preserve every tag exactly as written. Under each AI-permitted section, insert this standing prompt: 'Format the following evaluator-verified content into this section. Include only the content provided. Do not infer, characterize, or fill gaps; flag gaps instead.' Under each evaluator-only section, insert the line: 'Drafted personally by the evaluator. No AI contact, including rephrasing.' Do not alter any tag assignment."
Step three, verify and complete. Confirm no tag drifted in formatting. Confirm the opinion, interpretation, observation, and formulation sections all read [EVALUATOR ONLY]. Append the drafting-log template (section, date, tool, input class, verification read completed) and a copy of your five written hard limits as the skeleton's cover page. Done looks like this: a template you can open for the next retention, hand to a trainee with the boundaries already built in, and produce in discovery as proof that the line between contractor and engineer existed before this case, in writing, with your signature under it.
Key Takeaways
- Therapeutic and forensic roles are different jobs that cannot be held by the same clinician for the same client, per Greenberg and Shuman: the therapist works from alliance and accepted self-report; the evaluator works from skepticism, collateral verification, and explicit non-confidentiality. Mixing them corrupts both.
- AI drafts forensic boilerplate only: methodology, instrument descriptions, source lists, procedural history, and demographic and historical recitation built strictly from evaluator-verified facts. AI never drafts opinions, conclusions, psycho-legal recommendations, test interpretations, behavioral observations, or the reasoning that connects data to opinion.
- The AAPL Practice Guidelines and the APA Specialty Guidelines for Forensic Psychology center the transparency of the evaluator's reasoning: every opinion must rest on the evaluator's own examination of sufficient data, with a traceable chain the evaluator can defend; a language model in the opinion chain breaks traceability in a way no disclosure repairs.
- The deposition question "did AI draft any portion of your opinions?" is already being asked. The survivable answer is built in advance from three artifacts: a written forensic AI policy, a section-tagged report skeleton, and a drafting log showing what was formatted versus authored.
- Minimum necessary applies to prompts: AI receives only the verified fact list a section needs, never raw test data, wholesale collateral interviews, or third parties' protected records. Forensic files are full of other people's information, and none of them consented to a vendor's servers.
- AI never scores or interprets psychological instruments, and AI never drafts testimony, deposition answers, or talking points; the witness's voice must be the witness's, and scoring follows the test publisher's authorized procedures.
- The treating clinician never opines on parenting capacity or any forensic question about their own client; the forensic question goes to an independent evaluator. The engineer stamps the span personally; the contractor pours only the standard footings, to the engineer's drawings, under the engineer's inspection.
Skill.re