Reading an AI Materiality Suggestion Skeptically
An AI tool finishes its run and presents a verdict: "Based on the inputs, we recommend the following six topics as material: climate change, water, workforce health and safety, business conduct, biodiversity, and circular economy." It is confident. It is clean. It even names the ESRS standards each maps to. A tired analyst on a deadline feels a wave of relief and reaches to accept it. A skilled one feels something closer to suspicion, and starts asking questions, because they know that an AI materiality suggestion is the start of the analysis, not the end of it. This lesson is the set of questions that turn that suspicion into a method, the five interrogations that separate a useful suggestion from a misleading one.
Why Skepticism Is the Skill, Not Speed
By this point in the program you can make AI cluster inputs and draft rationales fast. The scarce skill is no longer producing an output. It is reading one well. A materiality suggestion arrives looking finished, and looking finished is precisely the danger, because the polish has no relationship to whether the conclusion is right. The model has pattern-matched text and produced the kind of answer that materiality outputs usually contain. Whether that answer survives an external assurer is a separate question the output cannot answer about itself.
The stakes make the skepticism non-optional. Under the Corporate Sustainability Reporting Directive (CSRD), which survived the 2025 to 2026 Omnibus simplification and still binds the largest undertakings (more than 1,000 employees and more than EUR 450M turnover), your double materiality assessment is part of the disclosure, and an external assurer reads it. Roughly 73% of large global companies now obtain external assurance on at least some sustainability disclosures, up from 51% in 2019. Double materiality is the ESRS rule that a topic is material if it is significant from either impact materiality (how your company affects people and the environment) or financial materiality (how a sustainability matter affects your company's finances), and you must assess both. The suggestion the tool just handed you will, if you accept it, become an audited conclusion. So you read it the way the assurer will: not asking "does this look reasonable" but "can this be defended."
One framing keeps you honest. Treat the AI materiality output as input to your judgment, never as the answer. The model is a fast research assistant who has read the whole pile and written up an opinion. You would never let a junior assistant's first-pass opinion become the company's audited disclosure without checking it. The fact that the assistant is a machine, and writes with machine confidence, makes the checking more necessary, not less.
Notice the trap in that confidence, because it is a specific cognitive trap and naming it helps you avoid it. A human junior analyst who is unsure will usually signal it: they hedge, they ask, they flag the topic they are not certain about. The model does not. It renders the topic it is confident about and the topic it invented in exactly the same calm, fluent, professional register. There is no tremor in the prose where the evidence runs out. This means you cannot use the texture of the output as a guide to where the risk is; the surface is uniformly smooth over solid ground and over nothing. The only way to find the holes is to probe the ground yourself, topic by topic, which is what the five questions force you to do. An analyst who reads for tone will be reassured by a confident invention. An analyst who reads for evidence will catch it.
The Five Questions That Separate Useful From Misleading
Here is the interrogation. For any AI materiality suggestion, ask these five questions in order, and do not accept the suggestion until each has a real answer.
One: What Evidence Is This Built On?
Ask the suggestion to show its inputs. For each topic it calls material, which stakeholder comments, survey responses, risk-register entries, and documents fed the conclusion. A defensible suggestion can be traced back to specific evidence; an indefensible one cannot, and the inability to trace is itself the finding. Watch for the topic that appears with no inputs underneath it, the invented theme: a clean, plausible material topic the model produced because it belongs in this kind of list, not because any stakeholder raised it. If you open "biodiversity" and find no comment, no register entry, nothing actually raised it, the suggestion is not surfacing a topic from your evidence; it is decorating the list from its training data.
Two: What Threshold Was Applied?
A topic is material because it crosses a line. So ask: what line. The materiality threshold is the documented standard, set in advance and usually built from severity and likelihood for impacts and from magnitude and likelihood for financial effects, that separates material from not material. An AI tool does not know your threshold unless you gave it to the model, and even then it cannot exercise the judgment your threshold requires. So when a suggestion declares a topic material, ask against what threshold, and whether that threshold is yours and documented. If the answer is that the tool applied some implicit internal sense of importance, you have a suggestion with no governance behind it, and "the model thought it was important" is not a threshold an assurer accepts.
Three: What Was Dropped, and Why?
The most dangerous part of a materiality suggestion is not what it includes. It is what it silently leaves out. A topic raised by a small but important group, an indigenous community near a mine, a workforce in one high-risk site, can be the most material impact in the assessment and the least frequent comment in the pile, and a frequency-leaning model will quietly drop it. So ask the suggestion the inverse question: what topics appeared in the inputs that you did not call material, and why. A suggestion that cannot account for its exclusions is hiding its most likely error. An undocumented exclusion is one of the most common assurance findings precisely because dropping a topic quietly looks exactly like concealing one.
This question is harder than the others for a simple reason: it asks you to see what is not there. The included topics sit in front of you, named and rationalised; the dropped one is an absence, and absences do not announce themselves. The only reliable way to find them is to work from the inputs forward rather than from the suggestion backward. Take the full list of topics that appeared anywhere in your raw inputs, compare it against the suggestion's material list, and force a documented reason for every topic on the first list that is missing from the second. The reason might be entirely legitimate: assessed and judged below threshold, with severity and likelihood recorded. But it has to exist and be written down, because "we considered it and excluded it for these reasons" is a defensible position and "it simply was not on the list" is a finding. The discipline turns the invisible exclusion into a visible, reasoned decision, which is the difference between an assessment that survives the engagement and one that does not.
Four: Does This Match the IROs You Already Know?
You are not coming to this blind. Your company has a risk register, prior-year reports, known impacts, and a sense of its own business that predates the AI run. The ESRS frame these as IROs: impacts, risks, and opportunities. So hold the suggestion against what you already know. If your business runs three plants in water-stressed basins and the suggestion does not surface water, that gap is a loud signal that something in the inputs or the clustering went wrong. If the suggestion surfaces a topic with no plausible connection to your operations or value chain, that is a signal of invention. The known IROs are your reality check: a suggestion that contradicts what an informed insider knows about the business is wrong somewhere, and you find out where before you accept it.
Five: Would an Assurer Accept the Basis?
The final question collapses the other four into the only test that ultimately matters. For each topic, could you hand an external assurer a basis they would accept: the evidence, the threshold, the documented exclusions, the named human decision, all reconstructable without you in the room. If the honest answer for a topic is "no, I could not defend this," then the suggestion is not yet a conclusion for that topic, no matter how confident the tool was. This question is the one that keeps the whole exercise pointed at reality, because the assurer is the reality. A suggestion that cannot survive this question is a suggestion you rework, not one you file.
The phrase "reconstructable without you in the room" is the sharpest part of this question, and it is worth taking literally. The test is not whether you, personally, could explain a topic if the assurer asked. You can explain almost anything in the moment, drawing on context you carry in your head. The test is whether a competent successor, handed only your file and none of your memory, could rebuild the conclusion: see what evidence fed it, what threshold applied, why other topics were excluded, who decided. If the answer depends on something only you know, the basis is incomplete, because the basis is supposed to live in the file, not in you. This matters in practice because people leave, reorganise, and forget, and an assurance file that only works while its author is available is a file with a single point of failure. The discipline that survives is the one where the reasoning is written down at the moment it is made, not promised for later.
An AI materiality suggestion is input to judgment, never the answer. The questions are simple: what evidence, what threshold, what was dropped, does it match the IROs you know, would an assurer accept the basis. If any answer is missing, the suggestion is not done.
Before and After: Reading a Suggestion the Right Way
Watch the five questions turn a confident output into a corrected, defensible set. The setting is the same mid-cap food manufacturer, and the tool has just returned its six-topic recommendation.
Before (the suggestion, accepted at face value): The tool recommends climate change, water, workforce health and safety, business conduct, biodiversity, and circular economy, each mapped to an ESRS standard, each with a one-line rationale. A rushed analyst accepts all six, drops them on the matrix, and moves on. The list looks complete and credible. It would file without incident, right up until the assurer starts asking questions.
After (the same suggestion, interrogated): You run the five questions. On evidence, five of the six trace cleanly to inputs, but biodiversity has nothing underneath it; no stakeholder, no register entry, no document raised it. It is an invented theme, and you remove it, documenting that it was assessed and found to have no basis in the inputs. On threshold, you find the tool applied no documented line, so you take each surviving topic to your own severity-and-likelihood threshold and confirm five clear it, recording the reasoning. On what was dropped, you ask for the topics the tool excluded and find that human rights in the tier-two supply chain was raised by two NGOs but not surfaced, because few voices raised it; severity, not frequency, governs, and you add it as a seventh material topic. On known IROs, you check against the business and confirm the three water-stressed plants are reflected and no surfaced topic is alien to your operations. On the assurer's basis, you confirm each of the now-six material topics has traceable evidence, an applied threshold, a documented exclusion (biodiversity), and a workshop decision. The confident six became a defensible six: one invention removed, one buried impact restored, every topic now carrying a basis.
The AI still earned its place. It read the whole pile and surfaced five of the right topics fast. What it could not do was notice that biodiversity had no evidence, apply your threshold, surface the buried supply-chain impact, or reconcile the list against what an insider knows. That was the five questions, run by you. The tool produced a suggestion. The skeptical reading produced the conclusion.
Building the Habit
The five questions are not a one-time audit you run on the final output. They are how you read every AI materiality artifact, every cluster, every draft, every recommendation, from the first one. Make them a written checklist the whole team uses, so the skeptical reading does not depend on which analyst happened to be tired that day. Keep the answers, because the answers are the basis: the evidence you traced, the threshold you applied, the exclusions you documented, the reconciliation against known IROs, and the assurer-ready judgment are exactly the file the engagement will ask for. Skepticism, recorded, is not just caution. It is the audit trail.
There is an order to the questions that makes the habit efficient rather than exhausting. Start with evidence, because a topic that fails the evidence question, an invented theme with nothing underneath it, can be removed immediately, and you do not want to spend threshold and exclusion analysis on a topic that should not be on the list at all. Then apply your threshold to what survives. Then run the drop question across the full input set, which is the one most analysts skip because it requires looking at what is not on the list, the hardest thing to see. Then reconcile against your known IROs, which is fast because you already carry that knowledge. Finally, for each surviving topic, ask the assurer's question and confirm the file would stand without you. Run in that order, the five questions take a fraction of the time the manual all-by-hand assessment would have, because the AI did the reading and you are doing the judging, which is the division of labour the whole chapter has been building toward.
A closing caution about momentum. The greatest danger with a confident suggestion is not that you will fail to catch an error if you look; it is that the suggestion's polish, combined with deadline pressure, will quietly talk you out of looking. The list is clean, the standards are mapped, the rationales are written, the meeting is in an hour, and accepting it feels not just easier but reasonable. Resist that specific feeling, because it is the exact feeling that precedes most AI-related disclosure failures: not a dramatic mistake, but a quiet decision to trust an output because it looked trustworthy. The five questions are your defence against your own understandable urge to be done. Run them even when, especially when, the suggestion looks like it does not need them.
And carry the right disposition into the room. The analyst who accepts AI suggestions arrives at assurance defending an output they did not really read. The analyst who interrogates them arrives with a thinner list, perhaps, but one where every topic survived five hard questions and every answer sits in the file. To the assurer, the second analyst looks faster and more in control, because the discipline that makes a suggestion defensible is the same discipline that makes any assessment assurable. The tool gave you a head start. The skeptical reading is what you were hired for.
Key Takeaways
- An AI materiality output is input to your judgment, never the answer; it looks finished, but polish has no relationship to whether the conclusion can be defended to an assurer.
- Double materiality means a topic is material if significant from either impact materiality (your effect on people and planet) or financial materiality (the effect on your finances), and a suggestion must be read against both.
- Ask what evidence the suggestion is built on; a topic with no inputs underneath it is an invented theme the model decorated the list with, not a topic your stakeholders raised.
- Ask what threshold was applied; the materiality threshold is your documented line from severity and likelihood, and "the model thought it was important" is not a threshold an assurer accepts.
- Ask what was dropped and why; the silent exclusion of a rare but severe impact is the most likely error and a common assurance finding, because dropping a topic quietly looks like concealing it.
- Ask whether the suggestion matches the IROs you already know; your risk register and business knowledge are the reality check that catches both missing topics and invented ones.
- Ask whether an assurer would accept the basis: traceable evidence, an applied threshold, documented exclusions, a named decision, all reconstructable without you in the room.
- Make the five questions a written checklist and keep the answers, because skepticism recorded is the audit trail; the tool gives the head start, the skeptical reading is the conclusion.
Skill.re