AI-Assisted Supplier-Questionnaire Drafting and Triage
A carbon accountant at a mid-cap manufacturer opens her inbox in late January and counts the damage. She sent one generic supplier questionnaire to four hundred suppliers in October. Ninety came back. Of those ninety, a third answered a different question than the one she asked, a third attached a corporate sustainability brochure instead of a number, and the rest gave her figures with no unit, no period, and no way to tell whether the supplier measured them or guessed. Scope 3 is roughly 75% of her footprint, and 79% of reporters say supplier-data availability is their top barrier. She is now that statistic. The questionnaire was supposed to be the instrument that collected assurable evidence from her value chain. Instead it produced a pile of correspondence she cannot put in an inventory. This lesson is about using AI to fix both ends of that pipe: drafting a questionnaire that is actually an evidence-collection instrument, and triaging the responses so the good data surfaces and the holes become visible instead of disappearing into an average.
The Questionnaire Is an Instrument, Not a Mailing
The single most expensive mistake in value-chain data collection is treating the supplier questionnaire as a mailing rather than as an evidence-collection instrument. A mailing goes out, gets a response rate, and is judged a success if enough people reply. An instrument is judged by whether the data it returns can survive an assurance engagement: whether each datapoint arrives with a clear value, a unit, a reporting period, an indication of whether it is primary (measured or supplier-reported activity data) or secondary (estimated, modelled, or industry-average), and a thread back to a source the supplier can produce if the assurer asks. Most questionnaires fail not at the response stage but at the design stage, because they were written to be answered, not to be assured.
Primary data is activity data the supplier actually measured or has on record: metered electricity for the facility that made your goods, fuel consumed by the trucks that shipped them, mass of material in the units you bought. Secondary data is anything estimated, modelled, allocated, or drawn from an industry average. The reason this distinction matters more than almost anything else in the supplier pipeline is that an assurer treats the two completely differently, and a defensible inventory has to know, datapoint by datapoint, which it is holding. A questionnaire that does not force the supplier to tell you which kind of number they are giving you has thrown away the most important piece of metadata before the response even arrives. You care because at the end of the chain, when you build the inventory, a primary figure and a secondary estimate cannot look identical in the file, and you can only honour that distinction if you captured it at the source.
So the design goal is not "ask the supplier about emissions." It is "extract, per relevant Scope 3 category, a value, a unit, a period, a primary-or-secondary flag, and a pointer to evidence, in a form a machine can ingest and an assurer can trace." Everything AI does well here serves that goal, and everything it does dangerously fails it.
Where AI Genuinely Helps: The Drafting End
The first job is drafting, and here AI earns its place quickly, because the bottleneck has never been writing one questionnaire. It has been writing the right questionnaire for each kind of supplier and not having the hours to do it. A logistics provider, a steel mill, a contract manufacturer, and a software vendor sit in different Scope 3 categories, hold different data, and should be asked different questions. A single generic form asks the steel mill about office electricity and asks the software vendor about smelting, and both ignore most of it. The realistic alternative, hand-tailoring four hundred questionnaires, never happens, so everyone sends the generic one and gets generic noise back.
Tailoring by Category and Supplier Type
AI collapses that tradeoff. Given a supplier's category, sector, and the Scope 3 category they fall into, a model can draft a tailored questionnaire that asks the logistics provider for tonne-kilometres and fuel type, the manufacturer for the mass and material of the components you bought, and the software vendor for data-centre energy and your allocation of it. The model is good at this because it is a language and templating task: take a known structure, vary it sensibly by context, and produce clean prose. You give it the category map (which suppliers fall into Category 1 purchased goods, Category 4 upstream transport, and so on), and it drafts a fit-for-purpose instrument for each cluster in minutes rather than weeks.
The discipline that makes this assurable rather than just fast is that the draft must build in the evidence structure from the first line. Every quantitative question should request the value, the unit, and the reporting period explicitly, never leaving the supplier to guess the format. Every quantitative question should ask the supplier to state whether the figure is measured or estimated and to name the basis. And the questionnaire should ask, for material figures, what evidence the supplier can provide on request, an invoice, a meter reading, a methodology note, so that the thread back to source exists before you ever need to pull it. A questionnaire drafted this way is doing the assurer's preparation for you, at the moment of collection, when it is cheapest.
Plain Language and a Reason to Answer
AI also helps with the unglamorous half of response rate: clarity and tone. A supplier sustainability contact, often a single overstretched person, abandons a form that is confusing, jargon-heavy, or visibly indifferent to their effort. The model can render each question in plain language, explain in one line why the figure is needed and what "primary data" means in terms the supplier understands, and keep the instrument short enough to finish. None of this changes the data you are asking for; it changes whether you get it back in a usable state. A clearer instrument is not a softer instrument. It is a more precise one.
A questionnaire is not judged by its response rate. It is judged by whether the data it returns can be put in an inventory and defended to an assurer. Design the instrument to collect evidence, or you have built a very efficient way to collect noise.
Where AI Genuinely Helps: The Triage End
The second job is triage, and this is where the inbox-full-of-correspondence problem gets solved. When responses come back, they come back heterogeneous: clean spreadsheets, scanned PDFs, free-text emails, attachments that answer a question you did not ask. A human reading four hundred of these in sequence is slow, inconsistent by the afternoon, and prone to quietly accepting a half-answer because the deadline is close. AI triage is a sorting and first-pass-reading task, which is squarely in the zone where the technology is reliable, provided you ask it to sort and flag rather than to decide.
Useful triage sorts every incoming response into a small number of buckets. A complete and clean bucket: the supplier gave a value, a unit, a period, and a primary-or-secondary indication, and the figure is plausible. A partial bucket: some fields are present, others missing, for example a value with no unit or a number with no statement of whether it was measured. An off-target bucket: the supplier answered a different question, sent a brochure, or attached unusable material. A non-response bucket: nothing came back at all. And a flag-for-review bucket: something is present but suspicious, a figure an order of magnitude off what the supplier reported last year, a unit that does not match the question, a number that implies an implausible intensity. The point of the buckets is not to let AI grade your suppliers. It is to route the analyst's scarce attention to where it matters, the partials worth chasing and the flags worth investigating, instead of spending it equally on four hundred items most of which are either fine or empty.
Triage Extracts and Flags, It Does Not Decide
The line that keeps triage assurable is that the model sorts and surfaces; the analyst decides and the file records the decision. AI can read an email and propose that it contains a Category 1 figure of a certain value in a certain unit for a certain period, flagged as the supplier's own measured data. What it must never do is silently promote that proposal into the inventory. Every figure it surfaces is a candidate that a human confirms against the actual response, and the act of confirming, or rejecting, is what the assurance file later shows. We will build that confirmation step out fully in the next two lessons on parsing and accountability; here the essential idea is that triage is a way of seeing your inbox clearly, not a way of letting the AI fill your inventory while you look away.
The Trap: Losing Track of What Is Primary
The specific way this whole pipeline fails, the failure this lesson exists to prevent, is losing track of what is primary versus what is estimated. It happens because both ends of the pipe blur the distinction if you let them. At the drafting end, a questionnaire that does not ask the supplier to state whether a figure is measured or estimated returns numbers with no tier, and a number with no tier defaults, in a hurried analyst's hands, to being treated as if it were real. At the triage end, an AI that summarises a messy response into a clean figure can launder an estimate into something that looks measured, because the summary is tidy and the original caveat got dropped. By the time the figure reaches the inventory, nobody can say whether the supplier metered it or guessed it, and the inventory has quietly inflated its share of primary data, which is exactly the thing an assurer probes.
The defence runs through the whole pipeline and starts at design. Force the tier at the question (ask "is this measured or estimated, and on what basis"). Carry the tier through triage (the bucket and the candidate datapoint both record it). Preserve it into parsing (the next lesson). Never let a tidy summary erase a supplier's own statement that a number was estimated. The data hierarchy that assurers care about, primary above secondary, is only as good as your weakest handoff, and AI introduces a new handoff, the summarisation step, that is very good at producing clean text and very willing to drop an inconvenient caveat. Treat every AI summary of a supplier response as a draft that may have smoothed away the one word, "estimated," that you most needed to keep.
A Worked Example: One Supplier Programme, Two Ways
A company is collecting Category 1 (purchased goods and services) data from four hundred suppliers for the reporting year. Watch the same AI tools produce a noise machine and then an instrument.
Before (the mailing, what actually happened the first year): The analyst asks AI to "write a supplier emissions questionnaire" and sends the single generic result to all four hundred suppliers. It asks broadly about "your carbon emissions" with no unit specified, no period stated, and no question about whether figures are measured or estimated. Responses dribble in. The analyst, overwhelmed, opens them one by one. A supplier writes "our emissions are about 1,200 tonnes" with no unit confirmation, no period, and no indication of basis; the analyst, tired, records 1,200 tonnes CO2e as if it were the supplier's measured Category 1 figure for the reporting year. Dozens of responses are brochures and get a mental shrug. Non-responders are forgotten rather than logged. The result is a partial dataset of unknown tier, an unknown number of silent gaps, and no record of which figures were measured. When the assurer asks "how much of your supplier data is primary, and show me the basis," the file cannot answer.
After (the instrument, the second year): The analyst gives AI the category map and has it draft tailored questionnaires by supplier cluster. The Category 1 manufacturers receive a form that asks, per purchased product, for the value, the unit (explicitly tonnes CO2e or the activity quantity and unit), the reporting period (the stated reporting year), whether the figure is measured or estimated and on what basis, and what evidence can be provided on request. The questions are in plain language with a one-line reason each. When responses return, AI triages them into complete, partial, off-target, non-response, and flag-for-review buckets, surfacing for each a candidate datapoint with its tier preserved, never writing anything to the inventory. The "about 1,200 tonnes" reply now lands in the partial bucket, flagged for missing period and missing basis, and routes to the analyst to chase rather than into the inventory as a phantom primary figure. The brochures land in off-target and are logged as such. The non-responders are counted and named, becoming a visible gap the team will handle openly. At year end the analyst can tell the assurer exactly what share is primary, point to the basis on each material figure, and name the gaps. Same AI, same four hundred suppliers. One produced correspondence; the other produced evidence.
The difference was not effort and not a better model. It was treating the questionnaire as an instrument designed to collect a value, a unit, a period, a tier, and a thread to evidence, and treating triage as a way to route attention rather than a way to fill the inventory unsupervised.
Working Rules for AI-Assisted Supplier Collection
A handful of rules keep the supplier pipeline on the assurable side from the first questionnaire to the last triaged reply. Design the questionnaire as an evidence-collection instrument, not a mailing: every quantitative question must request a value, a unit, a reporting period, a measured-or-estimated tier with its basis, and a pointer to evidence on request. Use AI to tailor by Scope 3 category and supplier type so each supplier is asked only for data they hold, which lifts both relevance and response quality. Keep the instrument in plain language with a one-line reason per question, because a clearer form is a more precise instrument, not a softer one. Use AI triage to sort responses into complete, partial, off-target, non-response, and flag-for-review buckets so your attention goes to the partials and flags, not equally to everything. Hold the line that triage sorts and surfaces while the human confirms and the file records the decision; never let a triage summary write to the inventory. Force the primary-or-secondary tier at the question and carry it through every handoff, and treat any AI summary of a response as a draft that may have quietly dropped the word "estimated." Log non-responses as named gaps rather than forgetting them, because a counted hole is a managed risk and a forgotten one is a silent misstatement waiting for the assurer to find it. Follow these and AI turns the 79% supplier-data bottleneck into a faster pipeline that is also a cleaner one. Ignore them and it turns into a faster way to collect noise of unknown tier.
Key Takeaways
- The supplier questionnaire is an evidence-collection instrument, not a mailing; judge it by whether the data it returns can be put in an inventory and defended to an assurer, not by its response rate.
- Every quantitative question must capture five things at the source: a value, a unit, a reporting period, a primary-or-secondary tier with its basis, and a pointer to evidence the supplier can produce on request.
- Primary data is supplier-measured activity data; secondary data is estimated, modelled, or industry-average; an assurer treats them differently, so the tier must be captured at the question and never blurred.
- AI genuinely helps at the drafting end by tailoring questionnaires per Scope 3 category and supplier type, asking each supplier only for data they hold, which the hand-tailored alternative never delivers because nobody has the hours.
- AI genuinely helps at the triage end by sorting heterogeneous responses into complete, partial, off-target, non-response, and flag-for-review buckets, routing the analyst's scarce attention to the partials and flags that matter.
- The line that keeps triage assurable is that AI sorts and surfaces candidate datapoints while the human confirms and the file records the decision; triage must never write to the inventory unsupervised.
- The central failure mode is losing track of what is primary versus estimated, because a tidy AI summary will launder an estimate into something that looks measured by dropping the supplier's own caveat.
- Log non-responses as named, counted gaps rather than forgetting them; a visible hole is a managed risk an assurer respects, while a forgotten one is a silent misstatement waiting to be found.
Skill.re