Mandatory XBRL Tagging for ESRS vs AI Claims
Why machine-readable sustainability data matters
Industry associations and data analytics providers have questioned whether mandatory XBRL tagging is necessary for digital sustainability reporting in the EU. Their argument is that AI systems can extract the same information from documents such as PDFs, avoiding additional tagging requirements for companies.
This European Call for Evidence gathers research to test that claim in the context of the European Sustainability Reporting Standards (ESRS) and Corporate Sustainability Reporting Directive (CSRD). Research on AI-based information extraction, including work on sustainability reports, continues to document gaps in reliability, comparability, and cost. EU policymakers need evidence on those trade-offs before treating AI extraction as a substitute for structured disclosures.
From an engineering perspective, receiving structured, machine-readable data is preferable to rebuilding that structure through PDF parsing and large language model pipelines. Mandatory ESRS XBRL digital tagging can make sustainability information directly available to civil society, academia, public-interest organisations, and other stakeholders without requiring each user to operate a separate extraction system.
The questions we are investigating
We are inviting researchers, engineers, sustainability specialists, and data users to contribute evidence on:
- How well current AI tools analyse and extract information from sustainability reports
- What it costs to build, operate, and validate these systems
- Which technical and methodological limitations remain unresolved
- How AI-based extraction compares with parsing machine-readable financial and non-financial disclosures
- What the consequences of each approach are for access, reproducibility, and independent scrutiny
The aim is not to reject useful AI applications. It is to distinguish tasks where AI adds value from a basic reporting infrastructure question: whether companies should publish structured data at the source.
Public statement and research collaboration
The first activity is a public statement representing the perspectives of engineers and researchers on the advantages of machine-readable sustainability data. Evidence gathered through this call will also inform collaborative research papers on information extraction from company reports.
Read the working brief and Call for Evidence or submit evidence and register your interest.
Connection to our sustainability report benchmark
The AI Benchmark for Sustainability Report Analysis is an activity that can provide evidence for this work. It examines model performance across reports, industries, criteria, and languages, as well as human agreement and pipeline robustness.
Those evaluations help answer a central policy question: where does AI extraction work reliably, and where does it remain a costly or incomplete substitute for data that could be published in a structured form? The XBRL initiative gives the benchmark a direct regulatory application, while the benchmark provides an empirical basis for the public statement.
Evidence we are building on
Recent peer-reviewed and industry work already constrains the claim that AI PDF extraction can replace mandatory tagging:
- Forster et al., Nature Communications (2026) — open RAG–LLM extraction across 600 European firms: useful scale and strong average validation, but Supplementary Fig. S2 Spearman correlations vs Refinitiv span ~0.17–0.90 (median ~0.78; GHG-reduction % weak at n = 40), and Table S5 sMAE ranges from ~0 on several emissions fields to >0.8 on metrics such as female top-management share — plus expensive expert gold labels and ambiguous missingness.
- ChatReport and Climate Finance Bench — published extraction / QA accuracies remain well below what would be needed to treat unstructured PDFs as a full substitute for tagged source data.
- MSCI Institute on XBRL vs AI — once set up, XBRL processing of digital sustainability filings was up to 10× faster and cheaper than AI PDF extraction in their India comparison.
Connection to EU Better Regulation
This project also connects with our EU Better Regulation collaboration. Decisions about digital sustainability reporting should be based on evidence about implementation costs, data quality, accessibility, and policy outcomes.
The Better Regulation work argues for outcome-first lawmaking, balanced participation, transparent use of evidence, and careful evaluation of proposed simplification. Applying those principles here means comparing the full consequences of structured reporting and downstream AI extraction rather than assuming that a new technology makes reporting infrastructure unnecessary.
Opinion and further reading
Our opinion piece Is XBRL Tagging for Sustainability Reports “Archaic Nonsense” — or a Chance for Digital Reporting and AI? lays out the short public framing — and invites NLP researchers to discuss it at NLP4Climate on 12 August.
Contribute evidence
We welcome empirical studies, benchmark results, technical assessments, cost analyses, and concise practitioner statements. Contributions should help policymakers understand the capabilities and limitations of current systems without overstating either the promise of AI or the burden of machine-readable reporting.
Fill in the contribution form to join the initiative.
Sustainability Report AI Benchmark
Benchmark AI extraction quality, costs, limitations, and robustness to establish the evidence base for comparing AI pipelines with machine-readable disclosures.
Call for Evidence
Gather research and practitioner evidence on AI extraction quality, costs, limitations, and comparison with machine-readable disclosures.
Public Statement
Draft a joint statement representing the perspectives of researchers and engineers on the value of mandatory machine-readable sustainability data.
Collaborative Research
Develop research with participating institutions on information extraction from company reports.
Related Research and Policy Work
View all posts »Is XBRL Tagging for Sustainability Reports "Archaic Nonsense" — or a Chance for Digital Reporting and AI?
Is XBRL tagging for sustainability reports "archaic nonsense" — or a chance for digital reporting and AI? If you are an NLP researcher and want to discuss this with us, join NLP4Climate on 12 August.
Assessing Corporate Sustainability with LLMs: Nature Communications Evidence from Europe
Nature Communications (July 2026): an open RAG pipeline over PDFs of 600 European firms yields millions of ESRS-aligned ESG observations — with strong agreement on validated subsets, and clear limits on what PDF extraction can replace.
When Old Tech Beats New Tech: MSCI Institute on XBRL for Sustainability Reporting
MSCI Institute finds that once set up, XBRL-based sustainability disclosures can be processed up to 10× faster than AI PDF extraction — and argues digital tagging and AI work best together, not as substitutes.
AI Benchmark for Sustainability Report Analysis
Building open-source tools to benchmark AI models in sustainability report analysis, with focus on detecting greenwashing and improving transparency.
Benchmarking the Benchmarks: Critical Findings on Climate-Related NLP Dataset Quality
A groundbreaking reproducibility study on 29 climate-related NLP datasets reveals that most tasks rely on surface-level keyword patterns rather than deep understanding, with 96% of datasets containing annotation issues that compromise evaluation reliability.
CHATREPORT: Democratizing Sustainability Disclosure Analysis through LLM-based Tools
Horizon Europe Consortium Partnership
We partner in Horizon Europe consortia across several roles: combining a business perspective for sustainable long-term project outcomes with the concrete technical work of software architecture, product research, open source strategy, development, AI/ML training, and helping put the consortium and its funding together in the first place.
NLP4Climate: The Research Community for Natural Language Processing in the Climate Domain
NLP4Climate is a growing international research community connecting people who use natural language processing to understand, analyse, and act on climate change. Monthly calls, self-organised focus groups, open talks, and a Slack community.
The Adaptation Exchange: Building a Climate-Resilient Economy Together
As a founding partner of the Adaptation Exchange, Climate+Tech brings scientific rigour and solution discovery to a cross-sector platform that translates climate risk awareness into coordinated resilience investment—connecting corporates, insurers, finance, and solution providers.