ESRS Datapoint Tagging and Digital Reporting
The narrative is perfect. The disclosure lead has spent weeks getting the climate-transition section to read exactly right, every figure verified, every claim sourced. Then the digital-reporting specialist applies the machine-readable tags, and one value, gross Scope 1 emissions, gets mapped to the location-based element instead of the gross direct element. The prose is flawless. The number is correct. And the disclosure is now wrong, because the only version of that figure a regulator's software will ever read is the one carried by the tag. This lesson is about treating the tag as what it actually is: a disclosure in its own right, with the same audit weight as the words.
The Tag Is the Disclosure, Not Decoration
Sustainability reporting has gone digital. Under the ESRS regime the statement is not only written for humans, it is marked up so machines can read it, and increasingly it is the machine-readable layer that downstream users, regulators, investors, data aggregators, actually consume. Digital tagging is the act of attaching a machine-readable label to each reported value or block of text, identifying which defined element of the reporting taxonomy it represents. Why you care: once the report is filed digitally, the tagged value is the canonical version of your disclosure for any automated reader, and if the tag is wrong, those readers receive the wrong disclosure no matter how clean your prose is.
The mental shift this demands is large. Reporting professionals are trained to pour their care into the narrative, the words an assurer and a board will read line by line. Tagging gets treated as a clerical afterthought, a formatting step handed to whoever runs the software. That instinct is exactly backwards for the digital layer. The tag is not a description of the disclosure; for a machine, the tag is the disclosure. A flawless number under the wrong tag is a wrong disclosure, full stop.
An ESRS datapoint, recall, is a specific, defined unit of disclosure the standard requires: a named quantitative or narrative element such as gross Scope 1 emissions in tonnes CO2e. Why you care here: each of those thousands of datapoints maps to a precise element in the digital taxonomy, and tagging is the act of binding your value to the right element. Get the binding wrong and you have not made a typo, you have misstated a datapoint.
It helps to see why this layer exists at all. The point of machine-readable disclosure is comparability and consumption at scale. A regulator supervising hundreds of large undertakings cannot read every statement by hand; an investor screening a portfolio cannot parse the prose of two hundred companies to find each one's market-based Scope 2. Tagging is what lets software pull your gross Scope 1, your net Scope 1, your transition-plan disclosure, and line them up against everyone else's, instantly and at population scale. That capability is exactly why a tagging error is so consequential: the moment your number enters the comparable, machine-read universe, it is being read by systems that take the tag at face value and have no way to notice that the surrounding paragraph said something different.
So the reporting professional inherits a second surface of correctness. There is the surface they know well, the prose an assurer reads line by line, and there is the surface they often ignore, the tagged layer a machine reads element by element. Both are filed. Both are the disclosure. A statement can be perfect on the first surface and defective on the second, and in the digital regime the second surface is the one that travels furthest and fastest.
The Digital Taxonomy, in Plain Terms
Behind the tags sits a digital taxonomy: a structured, machine-readable dictionary that defines every reportable element, what it means, its unit, its data type, and how it relates to the others. When you tag your gross Scope 1 figure, you are pointing it at one specific entry in that dictionary, the entry that means exactly "gross Scope 1 GHG emissions, in tonnes CO2e, for the reporting period," and not at the neighbouring entries that mean location-based emissions, or net emissions, or a prior-period restatement.
Think of the taxonomy as a vast set of labelled boxes. Each disclosure value has exactly one correct box. The boxes next to it look similar and are often only one qualifier apart: gross versus net, market-based versus location-based, current versus restated, consolidated versus a single entity. A human reader glides over those distinctions because the surrounding prose makes the meaning obvious. A machine has no prose; it has only the box you put the number in. Put the right number in the wrong box and every automated reader downstream inherits your error silently, with no sentence nearby to correct it.
A mistagged datapoint is not a formatting slip. It is a misstatement that no human will see and every machine will believe.
This is why the digital layer carries genuine audit weight. As assurance extends across sustainability disclosure, the tagging is part of what can be examined, because it is part of what is filed. An assurer or a regulator can compare the human-readable statement to the tagged data and find the place where they diverge. When they diverge, the question is not academic: which one is the disclosure? For an automated user, the answer is always the tag.
The Anatomy of a Tag
It is worth slowing down on what a single tag actually carries, because the detail is where the errors hide. A tag binds a value to a taxonomy element, but it also typically fixes the context around that value: the reporting entity it belongs to, the period it covers, the unit it is expressed in, and the sign and scale of the figure. Each of those is a place to go wrong independently of the element itself. Tag the right element but the wrong period, and a current figure is filed as a prior-year value. Tag the right element but the wrong entity context, and a subsidiary's number is attributed to the group. Tag the right element but mishandle the unit or scale, and 121,400 tonnes becomes 121.4 or 121,400,000 in a downstream system that trusts the declared scale.
None of these are visible by looking at the number. They are visible only by reading the tag's full context against the meaning of the value, which is precisely what verification has to do. A practitioner who treats a tag as just "which box" and ignores the period, entity, unit, and scale context is checking a fraction of where the error lives. The discipline is to verify the whole tag, not just the element label.
AI Assists the Tagging, the Human Verifies
Tagging at the scale of a full ESRS statement is brutal manual work. A large undertaking may face thousands of datapoints, each needing the correct element from a taxonomy with many near-identical neighbours. This is precisely the kind of high-volume, pattern-heavy matching that AI does quickly, and it is a legitimate use: let the model propose the taxonomy element for each value, and let a human verify.
The division of labour matters as much here as anywhere in the program. The model is good at the first pass: reading "gross Scope 1 emissions, 121,400 tCO2e" and proposing the gross-direct element, mapping hundreds of values in the time a person would map a handful. What the model cannot be trusted to do is settle the close calls, and the close calls are exactly where mistaggings live. Was this figure market-based or location-based? Is this a current value or the restated prior year? Does this number belong to the consolidated group or one subsidiary? The model will often guess plausibly and sometimes guess wrong, and a plausible wrong tag is the dangerous one, because it does not look like an error.
So the human verifies each tag the way they verify each number: against the definition. The verifier opens the proposed element's taxonomy definition and confirms it means exactly what this value means, qualifier by qualifier. Where the model proposed an element and the human changed it, that override is logged. The override log is the proof that a person actually examined the tagging rather than accepting the machine's first pass, and it is precisely what an assurer wants to see when they ask how the digital layer was controlled.
Why Not Just Trust the Model
It is tempting, at thousands of datapoints, to let AI auto-apply tags and move on. The reason not to is the same reason you do not let AI publish an unverified number: the cost of being wrong is asymmetric and the error is invisible at a glance. An auto-tagged statement can be 99% right and still carry a handful of mistagged material datapoints that no human ever looked at, each of which is a disclosure error sitting in the filed, machine-readable record. Auto-tagging also destroys the override log, so even where the machine happened to be right, you cannot demonstrate that a human controlled the process. Speed is real, but it is speed on the proposal, not on the verification. The verification is the part that makes the tag assurable.
Prioritising Verification at Scale
Verifying every tag against its full definition is the ideal, and on a manageable statement it is achievable. On a statement with thousands of datapoints under deadline, the practical move is to verify by risk rather than uniformly, and to document that you did. Two factors drive priority. The first is materiality: the datapoints that move an investor's or regulator's read of the company, the headline emissions figures, the targets, the financial-effect metrics, get full verification because a mistag there is a material misstatement. The second is qualifier-sensitivity: any value whose correct element turns on a fine distinction, market-based versus location-based, gross versus net, current versus restated, consolidated versus single entity, is exactly where the model is most likely to err, so those get focused human attention regardless of size.
What you do not do is treat the obvious, low-risk mappings as if they carried the same error probability as the close calls. A clearly unique narrative datapoint with one plausible element is a different risk than a Scope 2 figure sitting next to its location-based twin. Allocating verification effort toward materiality and qualifier-sensitivity, and recording the basis for the allocation, is itself a control an assurer can evaluate. It shows the human effort went where the disclosure risk was, not spread thin and blind across everything.
A Worked Example: One Tag, Two Outcomes
Take the datapoint: gross Scope 2 GHG emissions, market-based, FY2025, 64,200 tCO2e. The narrative is verified and correct. Now it has to be tagged.
The outcome you do not want. The team runs the tagging tool and lets it auto-apply across the statement to hit the filing deadline. For this datapoint the model maps 64,200 to the location-based Scope 2 element, a neighbour in the taxonomy that differs by a single qualifier, because the prose around it mentioned both methods and the model latched onto the wrong one. Nobody verifies; the override log is empty because no overrides were made. The statement is filed. An investor's data platform ingests the tagged file and records the company's market-based Scope 2 as a location-based figure, comparing it against peers on the wrong basis. Months later an assurer reconciling the human and digital layers finds the divergence and raises it. The number was always right. The disclosure, for every machine that read it, was wrong, and there is no log to show a human ever checked.
The outcome you want. The same tool proposes the location-based element. A verifier, working through the tags, opens the proposed element's definition, sees it means location-based, and checks it against the value, which the verified narrative and the basis-of-preparation both confirm is market-based. The mismatch is caught. The verifier overrides the tag to the gross market-based Scope 2 element and the override is logged with the reason and the date. When the assurer later asks how the digital layer was controlled, the team shows the override log: here is where the model proposed an element, here is where a human checked the definition, caught the mismatch, and corrected it. The tagged figure now matches the narrative, and both say market-based. The disclosure is consistent across every reader, human or machine.
Same value, same taxonomy, same model proposal. The difference was a human who verified the tag against its definition and a log that proves it happened.
Fitting Tagging Into the Disclosure Workflow
Tagging is not a standalone chore that happens once the report is otherwise done. It belongs inside the datapoint-to-disclosure path as a distinct, scheduled stage with its own verification, sitting after a number has been verified to source and before the datapoint is signed off. The sequence matters. You verify the number first, because there is no point tagging a figure that is itself wrong. You confirm the tag second, binding the verified value to its correct element with the full context checked. Then the accountable owner signs off a datapoint that is now both correct and correctly tagged. Skipping the middle stage, or bolting tagging on at the very end as a formatting pass run by whoever has the software open, is how a clean number ends up in the wrong box at the last minute with nobody scheduled to catch it.
This also changes how a reporting close is staffed and timed. If tagging is treated as a true verification stage, then the time and the people for it are planned, the override log is captured as part of each datapoint's evidence packet, and the digital layer arrives at sign-off already controlled. If it is treated as a formatting afterthought, it gets compressed into the final hours, verification gets skipped under deadline pressure, and the mistaggings that the close calls invite slip through. The discipline that protects the narrative, slow, source-tied verification with an accountable owner, is the same discipline the digital layer needs. The tag is a disclosure, so it earns a disclosure's process.
One more practical point: the tagging evidence connects directly to the rest of the assurance file. The override log sits alongside the numeric verification packet for each datapoint, so when an assurer examines a figure they can see, in one place, that the number was verified to source and that the tag was verified against its definition, with the human decisions on both recorded. A datapoint whose evidence packet covers both surfaces, the value and the tag, is fully reconstructable. One that covers only the number leaves the digital layer unexplained, and the digital layer is the version every machine reads.
What Good Looks Like to an Assurer
Step back and picture the engagement from the assurer's side. They have the human-readable statement and the tagged file, and a limited budget of attention. What reassures them is not a claim that the tagging is correct, it is evidence that a controlled process produced it. So they look for two things. First, do the human and digital layers reconcile on the datapoints they sample, the narrative and the tag saying the same thing, market-based where the prose says market-based, current where the prose says current. Second, when a tag was chosen, can you show that a human checked it, which is what the override log provides. A team that can put both in front of the assurer, a reconciling sample and a populated override log, has demonstrated control of the digital layer. A team that can show neither is asking the assurer to take the tagging on faith, and faith is not what an assurance engagement runs on.
This is why the worst outcome is not a mistag you caught, it is a mistag nobody looked for. A caught mistag, overridden and logged, is the system working: the model proposed, the human checked, the error was corrected, and the log proves the loop closed. An uncaught mistag in an auto-applied statement is the system absent: no check, no log, a filed misstatement that surfaces only when an external party reconciles the layers and asks the question you never asked yourself. The presence of overrides in your log is not a sign of a sloppy process; it is a sign of a working one. An override log that is suspiciously empty across thousands of datapoints is the thing that should worry you, because it almost certainly means the verification did not happen, not that the model was flawless.
The reporting professional's instinct to lavish care on the words is exactly right; the error is stopping at the words. The digital layer is the second half of the same disclosure, read by the systems that compare you to everyone else, and it deserves the same care, the same verification against an authoritative definition, and the same accountable record. Treat the tag as a disclosure, give it a disclosure's process, and the machine-readable statement becomes as defensible as the narrative you worked so hard to get right.
Key Takeaways
- Under the ESRS digital regime, the machine-readable tagged value is the canonical version of your disclosure for any automated reader. If the tag is wrong, those readers receive the wrong disclosure no matter how clean the prose.
- Digital tagging binds each value to one defined element in the digital taxonomy, a structured dictionary of every reportable element, its meaning, unit, and data type. Each datapoint has exactly one correct element.
- A mistagged datapoint is a disclosure error in its own right, not a formatting slip, even when the underlying number is perfect. It is a misstatement no human will see and every machine will believe.
- The dangerous mistaggings live in the close calls one qualifier apart: gross versus net, market-based versus location-based, current versus restated, consolidated versus single entity.
- The tag carries the same audit weight as the narrative because it is part of what is filed and what an assurer or regulator can examine and reconcile against the human-readable statement.
- AI is a legitimate accelerator for the high-volume first pass: let the model propose the taxonomy element for each value across thousands of datapoints. The model is good at the obvious cases and unreliable on the close calls.
- The human verifies each tag against its taxonomy definition, qualifier by qualifier, exactly as they verify each number against its source, and logs every override.
- Do not let AI auto-apply tags unverified. Auto-tagging hides invisible material errors and destroys the override log that proves a human controlled the digital layer.
Skill.re