Recognizing Bad AI Output Before It Goes Out
The two previous lessons taught you to get good output. This one teaches you to catch the bad output that will slip through anyway, because no matter how well you prompt, the model will sometimes produce something confidently wrong, and the moment that matters is the moment before you hit send. A professional who can spot bad AI output in the ten seconds before it leaves their hands is worth far more than one who writes perfect prompts and trusts everything that comes back. This lesson builds that radar. It catalogs the specific tells of bad AEC output, teaches you to read a draft the way an editor reads for a particular kind of error, and ends by having you red-team real AI-drafted submittal transmittals against your own log until catching the defects becomes reflex.
The Skill That Separates Safe Users From Burned Ones
Here is the uncomfortable truth that organizes this lesson: good prompting reduces bad output but never eliminates it, so the decisive skill is not prompting, it is detection. The model can produce a flawless-looking RFI with one fabricated sheet number buried in it, and no prompt guarantees that never happens. What protects you is the habit of reading the output critically before it goes out, with your eye tuned to the specific ways AEC output goes wrong, so the fabrication that survived the prompt does not survive your review. The professionals who get burned are not the ones who write bad prompts; they are the ones who write good prompts and then trust the polished result without the final read.
This reframes where you spend your attention. The prompt is the front end and the review is the back end, and the back end is where safety actually lives, because the prompt can only bias the model toward good output while the review is where you catch what got through. The good news is that this detection is learnable and fast, because AEC output fails in a small number of recognizable ways, and once you know the tells you can scan a draft for them in seconds, the same way an experienced plan reviewer flags the usual problems without reading every line at equal depth. The radar is a small set of patterns held in mind while you read, and this lesson installs them.
Good prompting reduces bad output but never eliminates it. The decisive skill is detection, the critical read in the ten seconds before you hit send, with your eye tuned to the specific ways AEC output goes wrong.
The Tells of Bad AEC Output
Bad AEC output has a recognizable vocabulary of failure, and naming the tells turns a vague unease into specific things to check. The first and most dangerous is confident wrongness: the output states something false in exactly the same authoritative tone it uses for true statements, because the model's confidence carries no signal about its accuracy. There is no waver, no hedge, no tell in the tone itself, which is precisely why you cannot rely on how sure the output sounds and must check the substance. Confidence is the disguise every other tell hides behind.
The specific fabrications cluster in references. A plausible-sounding spec section that is formatted perfectly and does not exist or does not govern. An invented submittal number that fits your numbering scheme and corresponds to nothing in your log. A made-up product cut sheet or model number that sounds like a real product and is not the specified one or is not real at all. A fabricated dimension or quantity stated with false precision. These are the references that drive action, the things a sub orders from or builds to, which is exactly why their fabrication is expensive. The pattern across all of them is the same: the model fills a reference-shaped hole with something that has the right shape and may have the wrong or no content, and the shape is convincing while the content is unverified.
There are also structural tells worth knowing. Over-confident completeness, where the output presents itself as a finished, comprehensive answer when it actually skipped something the source required, because the model would rather produce a complete-looking answer than flag a gap. Smoothed-over contradictions, where the output reconciles two things that actually conflict by quietly ignoring the conflict, which is dangerous because the conflict was the whole point. And generic drift, where the output slides toward boilerplate that could apply to any project rather than addressing the specifics of yours, a sign the model ran out of real information and filled with filler. Each of these is a recognizable shape, and recognizing the shape is what lets you catch it fast.
How to Read a Draft for Defects, Fast
Knowing the tells is half the skill; the other half is a reading method that surfaces them efficiently, because reading every word at equal depth is too slow for a busy professional and not even the most effective approach. The method is to read with a specific question in mind for each pass rather than reading passively. The first and most important pass is the reference pass: ignore the prose entirely and hunt only for the reference claims, every spec section, sheet number, submittal number, product, dimension, and check each against the source or mark it for verification. This is fast because references are a small fraction of any document and they cluster in predictable places, and it catches the most expensive defects directly.
The second pass is the specificity check: ask whether this draft actually addresses your project or could have been written for any project, because generic drift signals the model filled with boilerplate where it lacked real information. The third is the contradiction check: look for places where the draft asserts something that conflicts with what you know to be true about the situation, because a smoothed-over contradiction is the model papering over exactly the problem you needed surfaced. Reading in these targeted passes, references, specificity, contradictions, surfaces the high-value defects quickly and lets you spend your limited review time where the risk actually is, which is the editor's skill applied to AI output. You are not proofreading; you are hunting specific, known failure shapes.
Calibrating Trust by Consequence, Not by Polish
A subtle but critical part of detection is calibrating how hard you look to what the output will actually do, not to how good it looks. Polish is not a signal of accuracy, as the few-shot lesson showed, so you cannot let a draft that reads beautifully buy itself a lighter review. Instead, you scale your scrutiny to consequence: a draft headed for an internal note gets a light pass, a draft headed for an RFI on the record gets the full reference pass, and a draft headed toward anything that touches the cardinal rule's protected categories, a stamp, a schedule, a pay app, a safety plan, gets the most rigorous read you can give it, regardless of how clean it looks.
This consequence-based calibration is what makes detection sustainable, because you cannot read everything at maximum depth and should not try. The skill is to instantly assess where a given output is headed and apply the matching level of scrutiny, heavy where the cost of a surviving defect is high and light where it is low, so your finite review attention goes where it protects the most. The failure mode to avoid is the inverse, letting polish substitute for consequence, where a beautiful draft headed for a pay app gets a casual read because it looked done. The draft's destination, not its appearance, sets the review depth, and holding that discipline is what keeps the detection skill both rigorous where it matters and fast everywhere else.
Why Detection Is Hard, and How to Beat the Two Traps
If detection were easy, nobody would get burned, so it is worth naming why it is hard and how to counter it. There are two psychological traps that defeat even careful professionals. The first is fluency bias: well-written text feels true. Your brain evolved to treat articulate, confident communication as a signal of competence, because for most of history only competent people could produce it, and the model exploits that wiring by producing articulate, confident text regardless of accuracy. The countermeasure is to consciously decouple your judgment of accuracy from your reaction to the writing, by checking facts at the source rather than asking yourself whether the draft "sounds right," because sounding right is exactly what a fabrication does best.
The second trap is completion pressure: when a draft looks done and you are busy, the pull to just send it is enormous, and that pull is strongest at exactly the moments, deadline, fatigue, end of day, when your detection is weakest. The countermeasure here is structural rather than willpower-based, because willpower loses to deadline pressure reliably. You make the final read a fixed step that the document cannot skip, the same gate-not-mood principle from the cardinal rule, so that the reference pass happens not because you remembered to be careful but because it is simply the last station every AI-drafted document passes through before it leaves. The professional who turns detection into a non-skippable habit beats both traps at once, because the habit fires even when fluency bias and completion pressure are both pushing hardest, which is precisely when an undetected defect would otherwise get through.
Using the Model to Help Check Itself, Carefully
There is a useful and slightly paradoxical move worth knowing: while the model cannot be trusted to verify its own facts, it can be used to help surface candidates for your review, as long as you remain the one who confirms. Asking the model to "list every spec section, sheet number, and product reference you used in this draft" produces a clean checklist of exactly the items you need to verify, which is faster than hunting them yourself, and the model is reliable at this extraction task even though it is not reliable about whether those references are real. You are using it as a fast index of its own claims, then verifying each against the source yourself.
The boundary on this is strict and worth stating clearly: the model can list what it asserted, but it cannot confirm whether what it asserted is true, because asking it "are these correct?" just produces more confident prediction, often a doubling-down on its own fabrication. So the safe pattern is to use the model to enumerate its references, which it does well, and to use yourself and the source documents to verify them, which only you can do. This turns the reference pass from a hunt into a checklist-driven verification, speeding up detection without ceding any of the actual verification to the tool. It is a small efficiency that preserves the hard boundary: the model helps you find what to check, and you do the checking, every time, against something other than the model.
The Applied Problem: Red-Team Three Submittal Transmittals
Here is the exercise that builds the radar into reflex. Take three AI-drafted submittal transmittals, the kind a model produces when asked to draft transmittals for items going to the design team, and red-team each one against your project's actual submittal log. Red-teaming means you read not to approve but to attack, actively hunting for every defect as if your job were to find the thing that would embarrass you if it went out.
Run the passes on each transmittal. Reference pass: does every submittal number match your actual log, or did the model invent one that fits the scheme but corresponds to nothing? Does every cited spec section govern the item, or is it a plausible fabrication? Is every product the specified one, or a made-up cut sheet? Specificity check: does the transmittal address these actual items, or did it drift into generic submittal boilerplate? Contradiction check: does anything in the transmittal conflict with the spec or the submittal register? For every defect you find, flag it with the citation that disproves it, the log entry that shows the real submittal number, the spec section that governs, exactly as you flagged fabrications in the Level 1 audits. Tally the defects per transmittal, because the count is the proof of why this read is non-negotiable.
The deliverable is the three transmittals marked up with every defect flagged and a tally, and the lasting product is a tuned radar, the internalized set of tells and the habit of the targeted passes, so that detection becomes the automatic last step before any AI output leaves your hands. This is the skill that makes everything else in this level safe to use, because the hands-on lessons ahead all produce AI drafts that someone has to catch before they go out, and that someone is you, reading for the known failure shapes in the ten seconds that protect your name, your project, and your firm. Prompting gets you a good draft; detection is what makes it safe to send. Master both and you have the complete loop: a front end that produces strong drafts and a back end that catches what slips through, which together let you use AI fast and safely on the volume of documents that fill a real week, without ever being the professional whose name went on the fabrication nobody caught. The prompt and the few-shot examples raise the odds that the draft is good; the detection read guarantees that a bad one does not leave your hands, and it is the guarantee, not the odds, that lets you actually trust the speed and put your name on what the model helped you write.
Key Takeaways
- Good prompting reduces bad output but never eliminates it, so the decisive skill is detection: the critical read in the ten seconds before you hit send. The professionals who get burned write good prompts and then trust the polished result without the final read.
- The most dangerous tell is confident wrongness: the model states false things in the same authoritative tone as true ones, so you cannot rely on how sure it sounds. Confidence is the disguise every other tell hides behind.
- The expensive fabrications cluster in references: plausible-sounding spec sections, invented submittal numbers, made-up product cut sheets, and fabricated dimensions, all reference-shaped holes filled with convincing shape and unverified content.
- Structural tells include over-confident completeness (a finished-looking answer that skipped something), smoothed-over contradictions (papering over the conflict that was the point), and generic drift (boilerplate where real information ran out).
- Read in targeted passes, not passively: the reference pass hunts only the citations against the source, the specificity check asks whether it addresses your actual project, and the contradiction check looks for papered-over conflicts. You are hunting known failure shapes, not proofreading.
- Calibrate scrutiny by consequence, not polish. Scale the read to where the output is headed (light for an internal note, full for an RFI on the record, maximum for anything touching a stamp, schedule, pay app, or safety plan), because polish is not a signal of accuracy.
- The artifact: red-team three AI-drafted submittal transmittals against your real log, flag every defect with the disproving citation, and tally them, building the tuned radar that makes detection the automatic last step before any AI output leaves your hands.
Skill.re