A GreenDIA and Klimatkollen technical report separates extraction accuracy from source-report quality and proposes 30 errors and issues that human annotators and Climate NLP systems can assess.
30 Data Quality Problems in Corporate GHG Emissions Reporting
30 Data Quality Problems in Corporate GHG Emissions Reporting
About the source
The technical report “Typology of Data Quality Problems in the Corporate Reporting of GHG Emissions” was written by Andreas Dimmelmeier (LMU Munich), Felicitas Sommer (Technical University of Munich), and Alexandra Palmquist (Klimatkollen). It is work by the Green Data, Indicators, Algorithms (GreenDIA) consortium and Klimatkollen, funded by the Bavarian Research Institute for Digital Transformation (bidt) and Klimatkollen. It is not our research.
Read the full technical report and access the typology, annotation framework, and examples on OSF.
The problem: correct extraction can still produce unreliable data
AI systems are increasingly used to turn corporate sustainability reports into structured emissions databases. Much evaluation therefore asks whether a model extracted the number printed in a report correctly.
That is necessary, but it is not enough. The disclosed number may itself be incomplete, inconsistently calculated, poorly labelled, or presented without the context needed to interpret it. A pipeline can reproduce a source value accurately while still delivering a misleading result.
The report argues that extraction accuracy and quality of the source disclosure must be assessed together. It also warns about circular evaluation when an extraction system is benchmarked against commercial datasets that may themselves have been produced by opaque automated pipelines.
A typology of 30 errors and issues
The authors organise 30 data quality problems into 15 errors and 15 issues.
Errors include factual inconsistencies and departures from GHG Protocol guidance, such as:
- reporting Scope 1 and Scope 2 only, without Scope 3;
- combining scopes instead of disclosing them separately;
- calculation errors or different values for the same metric in different parts of a report;
- reporting only emissions intensity rather than absolute emissions;
- showing emissions only in charts without numerical values.
Issues cover disclosures that may not be factually wrong but make values difficult to assess, such as:
- unclear organisational or operational boundaries;
- missing reasons for omitted Scope 3 categories;
- no indication of whether values are estimated or calculated;
- no explanation of the consolidation approach;
- technical notes placed far away from the emissions values they qualify.
The framework then classifies these problems by annotation difficulty, GHG Protocol principle, broad cause, and likely effect on extraction or measurement accuracy. Its finer categories cover boundary specification, aggregation, calculation, format, method descriptions, missing data, and emissions categorisation.
Two practical uses
The report proposes two complementary applications:
- Report-level assessment: an expert annotator answers a checklist derived from the 30 problems and records evidence and comments.
- Problem-level annotation: text passages and report screenshots are labelled with a specific quality problem for human training or evaluation of text and multimodal AI systems.
The accompanying repository contains the full typology, spreadsheet framework, and annotated text and image examples. The authors describe these resources as work in progress and invite researchers and practitioners to test and extend them.
What this adds to the sustainability report AI benchmark
The typology is directly relevant to our AI Benchmark for Sustainability Report Analysis and its open research collaboration.
A useful benchmark should distinguish at least three questions:
- Did the system find and reproduce the disclosed value correctly?
- Did it recognise evidence that the source disclosure may be unreliable or incomplete?
- Do expert annotators agree on the quality problem and its supporting evidence?
Keeping these questions separate avoids penalising a model for faithfully extracting a flawed disclosure, or rewarding it merely because its output matches another database with unknown provenance. The report’s checklist could inform a dedicated source-quality track, while its annotated passages and screenshots could support error-detection and multimodal evaluation tasks.
What the typology means for XBRL and machine-readable reporting
The report also sharpens the debate around mandatory XBRL tagging for ESRS sustainability reporting.
Machine-readable tagging can address part of the problem: it can make values, units, periods, and categories easier to locate and process consistently. This reduces dependence on repeated PDF extraction and can remove some format ambiguity.
But XBRL does not automatically make the underlying disclosure complete or correct. A tagged value can still reflect an incomplete boundary, an unexplained methodology, an inconsistent restatement, or a calculation error. Structured reporting and source-quality assessment therefore solve different layers of the same problem:
- XBRL provides structure and traceability at the source.
- Quality controls test whether the tagged disclosure is coherent, complete, and credible.
- AI can help analyse narrative context and flag problems that tags alone cannot resolve.
This supports a combined approach rather than an “XBRL or AI” choice. See also the MSCI Institute comparison of XBRL and AI extraction.
Limits and next steps
The typology is an initial framework, not a finished measure of reporting reliability. The report notes that counting identified problems is only a rough quality indicator: different problems can have very different effects, and some require substantial domain expertise to detect.
The more valuable next step is empirical validation: test the categories across sectors and reporting regimes, measure expert agreement, evaluate whether models can identify each problem, and study how strongly individual problems affect extracted values and estimates of underreporting.
Access the report and materials
- Dimmelmeier, A., Sommer, F., & Palmquist, A. Typology of Data Quality Problems in the Corporate Reporting of GHG Emissions. GreenDIA and Klimatkollen.
- Supplementary typology, annotation framework, and examples.