Seven climate risk vendors assessed the same dummy portfolio and the same flood-prone stretch of road outside Paris — and disagreed almost completely. A detailed Investor Leadership Network study, alongside a 13-vendor GARP benchmark, traces that disagreement to specific, nameable methodological choices, not just irreducible uncertainty.
Why Climate Risk Vendors Disagree — and What That Means for Data Quality
Why Climate Risk Vendors Disagree — and What That Means for Data Quality
About the Two Studies
Both sources here are industry practitioner studies rather than peer-reviewed academic research: an Investor Leadership Network (ILN) report and a GARP Risk Institute benchmark. Both are methodologically transparent and produced by credible institutions, and both are written from an investor-side perspective on evaluating climate risk vendors.
Introduction
Buy a physical climate risk assessment for the same building, road, or company from two different vendors, and you may get two very different answers — not slightly different, but different enough to change a lending, insurance, or adaptation decision. It would be easy to read that as evidence that physical climate risk simply can’t be modeled reliably. The two studies covered here argue, in detail, that this isn’t quite right: a meaningful share of the disagreement is not irreducible uncertainty about a genuinely unknowable future, but the product of specific, traceable, and nameable methodological choices made at each step of a long modeling chain — from how an asset’s location is resolved, to which climate models are used, to how physical exposure gets translated into a financial number.
That distinction matters for anyone buying, regulating, or building on top of climate risk data. Treated as noise, disagreement gets averaged away and the signal is lost. Treated as an unquestionable black box, buyers pick a vendor on brand or feature count and never learn whether the number they’re paying for is fit for their specific decision. The ILN’s “Why Vendors Disagree,” authored by climate risk specialist Nik C. Steinberg and published in June 2026, and a separate GARP Risk Institute benchmark of 13 vendors, both push toward a third option: treat climate risk data quality as something that can be interrogated and compared.
The ILN Study: Testing Seven Vendors Against a Shared Portfolio
The ILN’s Climate Change Advisory Committee — representing 13 global institutional investors managing over USD $10 trillion collectively — commissioned this study after members repeatedly reported difficulty evaluating vendors from demo calls and marketing material alone. Between June 2025 and April 2026, seven vendors were given an identical dummy investment portfolio: three real assets (an NYC apartment building, a Singapore warehouse, and a Milan hotel), five listed equities, and two linear assets (including a stretch of toll road near Paris), plus a 24-question methodology survey spanning eight areas of the modeling chain, each scored for rigor, completeness, and transparency.
No Consensus on the Top Hazard
Asked to rank the most potentially damaging hazards facing each of the three real estate assets, no two of the seven vendors agreed on the top two hazards for any of the three properties. The only partial exception: six of seven vendors agreed that the Singapore warehouse faces significant extreme-heat exposure. Even there, the disagreement in the margins is telling — one vendor flagged meaningful wildfire risk at that same warehouse, despite the site sitting in an almost entirely industrial area with essentially no burnable vegetation nearby, suggesting a wildfire model driven purely by fire-weather variables with no land-cover check. A different vendor flagged tropical cyclone risk for Singapore, a location that sits outside the primary tropical cyclone tracks affecting the western North Pacific and is rarely directly hit by cyclones at all.
Flood Risk Near Paris: Agreement Indistinguishable From a Coin Flip
The study’s starkest finding concerns eleven sites along a stretch of the A4 toll road following the Seine and Marne rivers outside Paris. Vendors were asked whether each site would be exposed to a 1-in-100-year flood event by mid-century under a high-emissions scenario. All vendors correctly agreed on the one unambiguous hotspot — a site sitting 28 meters from a river confluence, on a roadway just 5 meters above the riverbank, immediately upstream of a flood-control structure where water is likely to pool. That’s the entire consensus.
Across the remaining ten sites, vendor agreement was statistically indistinguishable from random chance. The study reports a Fleiss’ Kappa of 0.008 — a formal measure of inter-rater agreement beyond what chance alone would produce, where 0 represents chance-level agreement and 1 represents perfect agreement. A score of 0.008 means that, outside the one obvious case, these vendors are essentially no more likely to agree with each other about flood exposure than if their answers had been assigned at random.
Counting Assets Is Already Hard
The study also asked vendors how many physical sites they could identify for five publicly listed companies of varying size. Reported asset counts differed substantially across vendors for every company tested — before any hazard modeling even begins. The study attributes this to a mix of internal data-processing capacity, how much third-party location data a vendor can afford to license, and methodological choices about whether to include subsidiary-owned sites or apply stricter criteria for linking a facility to its parent company. A vendor working from a materially incomplete asset inventory cannot produce a reliable company-level risk picture, no matter how sophisticated its downstream hazard modeling is.
Why the Numbers Diverge: Inside the Eight Methodological Areas
The ILN study’s real contribution is tracing why this happens, area by area:
- Model skill and agreement. Global and regional climate models operate natively at roughly 50–250 km resolution. Most vendors draw on established CMIP6 multi-model ensembles (commonly via NASA’s downscaled NEX-GDDP product), but none of the seven vendors tested model skill across all variables, regions, and time steps, and none formally measure inter-model agreement to communicate uncertainty — despite this being standard practice in the wider climate modeling community. One vendor took a notably different approach: selecting the single model with the greatest severity increase per hazard (the driest model for drought, the wettest for flood), which the study flags as a way of representing a bounding worst case rather than a central estimate.
- Spatial resolution. The study offers concrete guidance: building-scale hazards need roughly 30×30 meter resolution, neighborhood-scale effects roughly 1×1 km, and broader landscape hazards can be assessed above that. Six of seven vendors do apply additional high-resolution layers (flood defenses, land cover) on top of downscaled climate data. But the study is explicit that finer resolution does not automatically mean more accurate — some leading digital elevation models used to sharpen flood maps carry vertical uncertainty of up to a full meter, and over-precise outputs built on an uncertain elevation layer can look more authoritative than they are.
- Exposure and materiality. At least one vendor in the study placed the coordinates of a land-based asset in the ocean — evidence, the authors note, of essentially no location validation at all. Only a minority of vendors apply multi-step verification to catch coordinate drift, duplicate records, or misplacements in the messy commercial location datasets the whole industry ultimately relies on.
- Vulnerability indicators. Approaches range from broad building-type archetypes (three to five categories, borrowing heavily from FEMA’s HAZUS damage functions — which FEMA itself warns should not be applied to individual assets, only portfolios) to genuinely granular, component-level damage curves covering hundreds of thousands of asset-type-and-hazard combinations. More granularity isn’t a free win either: the study notes it can raise the risk of over-fitting without necessarily improving predictive skill.
- Financial metrics. Nearly all vendors produce some form of Average or Expected Annual Loss figure, but the study calls out a specific credibility problem: some vendors report these figures to the fourth decimal place, a level of numerical precision the underlying uncertainty simply cannot support. No vendor in the study models indirect financial effects such as supply-chain disruption or utility outages tied to a hazard event.
- Physical risk reduction. This is where the study finds vendors weakest overall. Risk-reduction recommendations are almost entirely engineering-focused (flood barriers, wind-rated construction, defensible space for wildfire), with essentially no vendor addressing governance, insurance strategy, or community- and municipal-level adaptation options that may be more cost-effective in many contexts.
What GARP’s Separate 13-Vendor Benchmark Adds
The ILN study builds directly on a companion effort: the GARP Risk Institute’s benchmarking study of 13 physical risk vendors, published for the UK’s Climate Financial Risk Forum in October 2025 — the larger and earlier of the two studies. Where the ILN study probes methodology qualitatively across eight domains with 7 vendors, GARP’s benchmark quantifies dispersion directly across 13 vendors and 100 shared properties, plus a dedicated geocoding test.
GARP’s geocoding test surfaced a striking concrete failure mode: for one property, described only by a well-known store name in Boston, a single vendor’s submitted coordinates landed 1,507 kilometers away — matched instead to a similarly named road in Atlanta. That kind of error happens upstream of any climate model, and no amount of sophistication in the hazard layer compensates for attaching the analysis to the wrong building. GARP also found flood damage-ratio correlations between vendor pairs ranging from just 0.2 to 0.9 for the same 100 properties, and inconsistent hazard metrics across vendors (some report cyclone wind speed as a 1-minute sustained average, others 10-minute, others as 3-second gusts, and at least one instead reports the probability of reaching a given Saffir–Simpson category) — differences that make direct vendor-to-vendor comparison difficult even before asking which number is more accurate. GARP’s full study also sets out a practical vendor-selection checklist.
What Buyers Should Actually Do
Both studies converge on a similar practical message, and it’s less “pick the best vendor” than “learn to ask the right questions”:
- Match the tool to the specific use case. The ILN study notes that most of the methodologies it examined are suitable for portfolio-level screening and regulatory disclosure, but not for site-specific capital allocation decisions, which require deeper, bottom-up analysis and independent third-party review.
- Ask how asset locations are validated, not just how hazards are modeled — a coordinate error can invalidate an otherwise sophisticated analysis before it starts.
- Ask about model skill testing and inter-model agreement, not just which climate scenarios are offered — a vendor using dozens of models without any skill assessment isn’t necessarily better positioned than one using a smaller, carefully validated set.
- Treat extreme numerical precision (four decimal places on a loss estimate) as a yellow flag, not a sign of rigor — GARP and ILN both note this level of false precision is incompatible with the uncertainty inherent at that stage of the modeling chain.
- Prefer vendors who are explicit about what they don’t model — indirect and cascading impacts, compounding hazards, non-engineering adaptation options — over vendors who imply comprehensive coverage without saying so directly.
A Proposed Way Forward: A CMIP-Style Vendor Ensemble
The ILN study’s conclusion doesn’t stop at “buyers should do more diligence.” It proposes a specific structural fix worth taking seriously: a voluntary vendor ensemble, modeled loosely on the Coupled Model Intercomparison Project (CMIP) that climate science itself uses to compare global climate models under shared protocols. Rather than treating vendor disagreement purely as a private due-diligence problem for each buyer to solve alone, a shared, standards-anchored ensemble would let investors see where vendor methods and findings converge, understand why they diverge where they do, and identify where deeper independent inquiry is genuinely warranted — while letting vendors demonstrate the distinct strengths of their approaches without disclosing proprietary methodology. It’s a proposal, not something that exists yet, but it’s a meaningfully different frame from “investors should just do more homework”: it treats disagreement as an industry-level transparency problem, not only an individual buyer’s risk to manage.
Conclusion
Both studies land on the same underlying point from different angles: physical climate risk vendors disagreeing is not, by itself, proof that the field is broken. What it demonstrates is that asset-level physical climate risk assessment runs through a long chain of real methodological choices — model selection and skill testing, downscaling and spatial resolution, location validation, vulnerability granularity, and hazard-to-loss translation — and that current market practice too often compresses all of that into a single confident-sounding output. The ILN’s Fleiss’ Kappa of 0.008 on flood exposure near Paris, and GARP’s 1,507-kilometer geocoding miss, are not edge cases dredged up to embarrass anyone; they’re concrete illustrations of exactly how and where that compression breaks down. For buyers, the honest response isn’t “which vendor is right,” but “which vendor’s methods, assumptions, and limitations are documented well enough for me to judge whether the output is right for my decision” — and, at an industry level, the ILN’s CMIP-style ensemble proposal is a serious attempt to make that judgment easier for everyone at once, rather than leaving each buyer to re-derive it alone.
Access the Source Material
- Steinberg, N. C. (2026). Why Vendors Disagree: A Practical Guide to Evaluating Physical Climate Risk Data Vendors. Investor Leadership Network. Full PDF · investorleadershipnetwork.org
- Paisley, J., & Nelson, M. (2025). Comparing Climate Risk Vendors: A User’s Guide to Physical Risk Assessments. GARP Risk Institute, for the Climate Financial Risk Forum. garp.org
Related Work
- A Risk Professional’s Guide to Physical Risk Assessments: The GARP 13-Vendor Benchmark
- Climate Risk Companies Don’t Always Agree: CarbonPlan’s Vendor Comparison
- The 2023 Climate Risk Landscape: UNEP FI’s Multi-Vendor Technical Supplement
- Climate Risk Data Quality Is a Market Problem, Not Just a Methods Problem
- Climate Risk Assessment Services
- AI Climate Risk Assessment Tools
- Climate Risk Data Quality — topic hub
- Decision Making in Deep Uncertainty