System Prompts for Disclosure Contexts
A disclosure lead opens a fresh chat with the model on a Monday and asks it to draft a paragraph on the company's Scope 3 transport emissions. It produces a clean, confident sentence with a tidy figure and a plausible emission factor. On Tuesday a colleague opens a different chat and asks the same model the same kind of question, and gets a different factor, a different boundary assumption, and a number that does not reconcile with Monday's. Nobody told the model which framework it was working under, where the organizational boundary sat, or that it must cite a source or refuse. So it guessed, twice, differently. The fix is not better luck on each prompt. It is a system prompt: a standing instruction that loads before every task and locks the model into assurable behavior by default.
What a System Prompt Actually Is
Most reporting professionals meet AI through the chat box, where every request starts from nothing. You type a question, the model answers, and the next question starts fresh. A system prompt changes that. It is a block of instructions that sits above the conversation and applies to every task in it, before you type a single word. Think of it less as a question and more as a job description handed to a contractor on day one: here is the framework you report under, here is the boundary you respect, here is the one rule you never break. The model reads it first, every time, and it shapes everything that follows.
The reason this matters for disclosure is the reason the whole program matters: every figure you publish is now an audited figure, and the model, left to its defaults, behaves in exactly the way that fails an audit. Its default is to be helpful and fluent, to fill a gap with a plausible number, to answer rather than refuse, to pick a reasonable-sounding emission factor rather than admit it does not have one. None of those defaults are malicious. They are simply not the defaults a disclosure context needs. A disclosure context needs a model that refuses when it lacks a source, that holds one framework steady across a hundred tasks, that never silently moves the boundary. The system prompt is where you install those defaults so you do not have to re-type them, and re-remember them, on every single request.
Without one, the discipline lives only in your head and your typing fingers. You know you should tell the model to cite or refuse, so you do, on the prompts you remember to do it on. You know it should work under ESRS, so you mention it, when you think of it. The gaps are not the prompts where you remembered. They are the Tuesday prompt where you were rushed, or the colleague's prompt who never learned the discipline, or the third prompt in a long session where the instruction from the first one has faded. A system prompt closes those gaps by making the discipline structural rather than remembered. It is the difference between a control that depends on a person being careful every time and a control that holds whether or not they are.
The Three Things It Must Lock
A disclosure system prompt can do many things, but three are load-bearing, in the sense that if any one is missing the output drifts toward the failure modes that end careers. Lock the framework, lock the boundary, and lock the cite-or-refuse rule, and you have converted a general-purpose model into one that defaults to assurable behavior.
Lock the Framework
The model needs to know which rulebook it is operating under, because the rulebook changes the answer. A disclosure drafted under ESRS (the European Sustainability Reporting Standards, the datapoint set behind CSRD) is structured around double materiality and a specific datapoint architecture. A disclosure under ISSB (IFRS S1 and S2, the global baseline now adopted or planned across more than thirty jurisdictions) is built on financial materiality and the investor lens. A CBAM declaration of embedded emissions for imported steel speaks the language of actual values versus default values and the authorised declarant. These are not interchangeable. A model that does not know which one it is in will blend them, borrowing an ESRS concept into an ISSB answer or a generic carbon-footprint framing into a CBAM line, and the blend is wrong in a way that an assurer or a customs authority will catch. Naming the framework in the system prompt holds the model in one rulebook so it stops improvising across three.
Lock the Boundary
The boundary is the perimeter of what the company is counting: which legal entities, which facilities, which operations are inside the inventory and which are outside. The GHG Protocol gives two main approaches, the organizational boundary (drawn by control or by equity share, deciding which entities you consolidate) and the operational boundary (which sorts emissions into Scope 1, 2, and 3). An undocumented or shifting boundary is one of the classic assurance findings, because a number is only meaningful against a stated perimeter. If the model does not know your boundary, it will assume one, and its assumption will not match yours, and the figures it drafts will quietly count something you exclude or exclude something you count. Stating the boundary in the system prompt means every figure the model touches is interpreted against the right perimeter, and the model can flag when a request seems to cross it rather than silently absorbing the discrepancy.
Lock Cite or Refuse
This is the rule the whole program rests on, installed as a default. Cite the source or refuse means the model may state a figure or a factor only when it can name where the figure came from, and when it cannot, it must say so and decline rather than invent. This single instruction is the direct countermeasure to the hallucinated emission factor, the fabricated activity figure, and the invented target, because each of those is the model answering when it should have refused. The default model wants to help, and helping, to a model, means producing an answer. Cite-or-refuse redefines helping: in a disclosure context, the most helpful thing the model can do when it lacks a source is to tell you it lacks one. You would far rather see ten honest refusals you can go resolve than one fluent fabrication you do not catch.
A system prompt does not make the model smarter. It makes the model's defaults match the assurer's expectations, so that being lazy with a single prompt no longer means being unsafe.
A Worked Example: The Actual System Prompt
Here is a system prompt a disclosure team could load before drafting and lookup work. It is written in plain language, not code, because the model reads plain language. Read it once, then we will take it clause by clause and show why each line is doing real work.
You are a disclosure assistant supporting the sustainability reporting team of a large undertaking in scope of CSRD. You assist with drafting and lookups; you do not decide what is disclosed. A human reviewer signs off every output.
FRAMEWORK: Unless I state otherwise in a specific task, all work is under ESRS (the datapoint set behind CSRD). If a task is for ISSB (IFRS S1 and S2) or for a CBAM embedded-emissions declaration, I will say so, and you must apply that framework and not blend it with others. If you are unsure which framework applies, ask before proceeding.
BOUNDARY: Our organizational boundary is the operational-control approach. Our reporting entities are the parent and its majority-controlled subsidiaries; joint ventures we do not control are outside the boundary. The operational boundary follows the GHG Protocol scopes. Interpret every figure against this boundary. If a request appears to count something outside the boundary or to exclude something inside it, flag it and ask rather than proceeding.
SOURCING RULE (most important): State a figure, an emission factor, or a quantitative claim only if you can name its source. Cite the source inline: for a factor, the named database, table, version, and year; for activity data, the file, page, and line. If you do not have a source for a number, do not produce the number. Say what is missing and refuse to fill the gap. Never invent, estimate silently, or guess a factor, an activity figure, or a target. The phrase "the AI estimated it" is not a source.
ESTIMATES: If I explicitly ask for an estimate, label it clearly as an estimate, state the method (for example spend-based or average-data), and note its uncertainty. Never present an estimate as measured data.
NARRATIVE: When drafting narrative, do not soften, omit, or invent. Do not add a target, commitment, or achievement the source material does not support. If asked to describe a negative impact, describe it as the evidence states it.
OUTPUT: When you state a number, show the source beside it. When you refuse for lack of a source, say exactly what evidence would let you proceed.
Why Each Clause Is Load-Bearing
The role line ("you assist, you do not decide; a human signs off") sets the accountability frame from the first sentence. It tells the model it is an assistant inside a controlled process, not the author of record. This is not decoration. It primes the model away from the confident, final-sounding outputs that tempt a reviewer to wave them through, and it encodes the cardinal rule that accountability stays human. Remove it and the model writes as if it were the discloser, which is exactly the posture you do not want a reviewer reading.
The framework clause stops the cross-framework blending described above. The crucial detail is the default plus the override: ESRS unless told otherwise, and an explicit instruction to ask if unsure. That structure means the common case is handled automatically while the exceptions are surfaced rather than guessed. Remove it and the model silently picks a framing per task, and your hundred outputs are no longer consistent with each other, which is itself an assurance problem because inconsistency reads as a lack of control.
The boundary clause does two jobs. It states the perimeter, and it instructs the model to flag boundary crossings rather than absorb them. That second job is the subtle one. A model that merely knows the boundary still tends to quietly fit a request into it; a model told to flag mismatches turns the boundary into an active check. Remove this clause and every figure the model handles floats free of a perimeter, and an assurer's first question, against what boundary, has no answer in the workflow.
The sourcing rule is the heart of the prompt and the reason it is marked as most important. Notice it does three things at once: it permits a number only with a source, it specifies what a source looks like for a factor versus for activity data, and it forbids the specific fabrication moves by name. The specificity matters. "Be accurate" is advice the model cannot act on; "name the database, table, version, and year, or do not state the factor" is an instruction it can follow and you can check. The closing line, that "the AI estimated it" is not a source, closes the loophole where the model treats its own guess as provenance.
The estimates clause handles the legitimate case the sourcing rule would otherwise over-block. Estimation is allowed in disclosure; fabrication is not, and the line between them is labeling. This clause lets the model estimate when you explicitly ask, but forces the estimate to wear a label, a method, and an uncertainty, so a transparent estimate never gets laundered into measured data. Without it, you either lose estimation entirely or you reopen the door to silent estimation, and you want neither.
The narrative clause targets the drafting-specific failure modes: the softened negative impact, the omitted bad news, the invented target or commitment. These are the greenwashing-by-fluency risks, and a general model drifts toward them because softening reads as polish. The clause forbids the drift in plain terms. Remove it and the model writes a nicer story than the evidence supports, which is the precise thing an assurer is trained to catch and a regulator is empowered to act on.
The output clause makes the discipline visible in the result rather than buried in the process. By requiring the source beside every number and a specific statement of what is missing on every refusal, it turns each output into something a reviewer can check at a glance and something that drops cleanly into the evidence file. A refusal that says "I cannot state this factor; provide the 2026 grid factor table and I will" is far more useful than a blank, because it tells the human exactly what to go get.
How to Build and Maintain Yours
The worked prompt is a template, not a finished artifact, because the load-bearing facts in it are your facts. The framework is whichever you report under, the boundary is your actual consolidation approach and your actual entity list, and the sourcing detail should name the factor databases your team actually uses. Building your version is mostly the work of filling those specifics in correctly, and that work is itself valuable because it forces the team to state, in one place, the framework, boundary, and sourcing rule it is supposed to be applying anyway.
A few maintenance disciplines keep it trustworthy over time. Treat it as a controlled document. The system prompt encodes your basis of preparation in miniature, so it should be versioned, dated, and owned, not pasted from memory and quietly edited. When the boundary changes because the group restructured, or the framework default changes because a new obligation came into scope, the prompt changes too, and the change is recorded. Keep it in the evidence trail. If an assurer asks how the team kept AI output inside the framework and the boundary, the dated system prompt is part of the answer: here is the standing instruction every task ran under. Do not let it become a substitute for verification. A good system prompt sharply reduces the rate of bad output, but it does not eliminate it, because a model can still misread a source or slip a fabrication past its own instruction. The prompt lowers the base rate; the human checks catch the remainder. The two are layers, not alternatives.
One more caution worth stating plainly. A system prompt is an instruction, not a guarantee. The model is generally good at following it and occasionally fails to, especially under long conversations or unusual requests where the standing instruction competes with the immediate ask. This is exactly why the role line insists a human signs off every output and why the rest of this chapter builds structured output and verification checklists on top of the system prompt. The prompt makes the safe behavior the default. The human and the checklist make it the verified outcome. You install the default so the verification has less to catch, not so it has nothing to catch.
Before and After: The Same Question, Two Worlds
Return to the Monday and Tuesday problem. Without a system prompt, the disclosure lead asks for a Scope 3 transport paragraph, and the model, helpful and unconstrained, produces a fluent sentence with a confident factor it cannot source, under whatever framing it inferred, against a boundary it assumed. The colleague's version differs because the assumptions differed. Neither output declares what it cannot support, so both look finished, and the unsourced factor is now one careless review away from a public, assured disclosure.
With the system prompt loaded, the same request runs under ESRS by default, against the stated operational-control boundary, with cite-or-refuse in force. The model drafts the paragraph but, reaching the transport factor it has no source for, stops and says: "I cannot state a transport emission factor without a named source. Provide the factor database, table, and year, and I will complete the sentence and cite it." The output is less satisfying in the moment and far safer in the file. Monday and Tuesday now agree, because both ran under the same standing instruction, and the gap that would have failed assurance has been surfaced as a question instead of buried as a number. That is the entire value of the system prompt: it moves the failure from the published disclosure, where it is a crisis, to the draft, where it is just a thing to go resolve.
Key Takeaways
- A system prompt is a standing instruction that loads before every task, turning disclosure discipline from something you remember to type into something that holds whether or not you remember.
- The default behavior of a general model, being helpful, fluent, and answer-seeking, is exactly the behavior that fails assurance; the system prompt resets the defaults to match the assurer's expectations.
- Three things are load-bearing and must be locked: the framework (ESRS, ISSB, or CBAM, never blended), the boundary (organizational and operational, stated and defended), and cite-or-refuse (a number only with a source, otherwise a refusal).
- The cite-or-refuse rule is the direct countermeasure to the hallucinated factor, the fabricated activity figure, and the invented target, because each of those is the model answering when it should have refused.
- An estimates clause keeps legitimate estimation alive by forcing every estimate to wear a label, a method, and an uncertainty, so a transparent estimate is never laundered into measured data.
- A narrative clause blocks the softened impact, the omitted bad news, and the invented commitment, the greenwashing-by-fluency risks a general model drifts toward.
- The system prompt is a controlled document: versioned, dated, owned, kept in the evidence trail, and updated when the boundary or framework changes, so it can answer an assurer's question about how AI output stayed in scope.
- The prompt is an instruction, not a guarantee; it lowers the base rate of bad output so the human sign-off and the verification checklist have less to catch, but it never replaces them.
Skill.re