INDUSTRY REPORTS

GARP tested 13 physical climate risk vendors against the same 100 properties worldwide. The correlation between vendors' flood damage-ratio estimates ranged from 0.2 to 0.9, one vendor geocoded a Boston property 1,507 km away in Atlanta, and dispersion behaves in opposite directions depending on which statistic you use to measure it.

A Risk Professional's Guide to Physical Risk Assessments: The GARP 13-Vendor Benchmark

· 9 min read

A Risk Professional’s Guide to Physical Risk Assessments: The GARP 13-Vendor Benchmark

About the Source

This is an industry practitioner report, not peer-reviewed academic research. It was produced by Jo Paisley and Maxine Nelson of the GARP Risk Institute for the Financial Resilience Working Group of the UK’s Climate Financial Risk Forum (CFRF) — a body convened and facilitated by the Prudential Regulation Authority (PRA) and Financial Conduct Authority (FCA), though the report explicitly states it does not represent regulatory guidance.

It is the earlier and larger of two closely related benchmarks: published October 2025, testing 13 vendors against 100 properties, versus the Investor Leadership Network’s “Why Vendors Disagree” study of 7 vendors against a smaller dummy portfolio in June 2026, which explicitly builds on this one. The two are independent efforts with different designs, so they read as complementary rather than duplicate evidence.

Introduction

If two physical climate risk vendors are given the exact same set of buildings, the same climate scenario, and the same time horizon, how different should their risk estimates reasonably be? The GARP Risk Institute set out to answer that empirically rather than rhetorically. Working with 13 vendors — anonymized throughout the study, though named collectively in the acknowledgments as Climate X, Fathom, First Street, ICE, JBA Risk Management, Jupiter Intelligence, Moody’s, MSCI, Planetrics (a McKinsey & Company solution), Riskthinking.AI, S&P Global, Twinn by Haskoning, and XDI — the study built a shared portfolio of 100 real-world-style properties and a shared geocoding test, then measured exactly how far apart the resulting numbers landed.

The headline finding is not subtle: vendor estimates disperse substantially, for both the hazard itself (say, flood depth) and the resulting financial damage. But the more useful contribution, as with the ILN study, is where in the modeling chain that dispersion originates — starting well before any climate model is even invoked.

Location Errors Happen Before the Climate Modeling Starts

GARP’s asset location survey gave 13 vendors 20 properties worldwide with deliberately varying levels of address completeness — from just an asset name and city, up to a full postal address with coordinates — and asked each vendor to independently geocode them. For each property, GARP calculated the median of all vendors’ submitted coordinates, then measured how far each individual vendor’s answer was from that median.

The most extreme result: for a property named after a well-known store in Boston, one vendor’s submitted coordinates were 1,507 kilometers away — matched instead to a similarly named road in Atlanta. That’s not a climate modeling disagreement; it’s a data-matching failure that invalidates every subsequent hazard and financial estimate for that property, however sophisticated the vendor’s downstream modeling is.

More broadly, the study found that better address information generally does reduce geocoding dispersion — but not to zero. Even with the most complete addresses provided (asset type, name, city, county, postcode, and country), meaningful dispersion remained across vendors. And a separate, accidental case reinforces the same point: a data-entry error on GARP’s own side left the negative sign off one property’s longitude, placing it in the North Sea instead of Scotland. Vendors handled this inconsistently — some flagged it, some returned null results, some reported hazard data for the North Sea location as given, and some used the other address fields to correct it and place it back on land. There’s no universally “right” way to handle malformed input, but the range of vendor responses to the same error is itself informative about how much manual judgment sits inside what looks like an automated pipeline.

Damage Estimates for Identical Properties: Correlations as Low as 0.2

For the core quantification exercise, vendors assessed defended and undefended flooding, coastal flooding, tropical cyclones, windstorms, heat, and wildfire across the 100 properties, all under a single shared scenario (RCP 8.5) to remove scenario choice as a source of variation.

Even with the scenario held constant, the results diverged sharply. For combined (pluvial and fluvial) flooding at a 1-in-200-year severity in 2030, the pairwise correlation between different vendors’ damage-ratio estimates for the same 100 properties ranged from 0.2 to 0.9 — meaning some vendor pairs were reasonably aligned on which properties would be hit hardest, while others were nearly uncorrelated. For heat, the dispersion was sometimes extreme at the level of a single property: for one property in Hong Kong, estimates of the number of annual days above 35°C ranged from 0 to 192 days. The zero, on inspection, didn’t mean “no days exceeded the threshold” — it meant that particular vendor’s model classified the property as not exposed to heat hazard at all, a categorically different kind of disagreement than a difference in degree.

Wildfire proved to be the hardest hazard to compare at all. GARP had hoped to standardize on two possible metrics (a fire weather index, or an annual probability/count measure), but vendors’ actual outputs varied so much — fire weather indices, counts of days over various thresholds, wildfire probabilities, and flame length — that the study concluded it couldn’t draw meaningful cross-vendor benchmarking conclusions for wildfire at all.

A Counterintuitive Result About How Disagreement Evolves Over Time

One of the study’s more interesting findings concerns how vendor disagreement changes as you look further into the future or further into the tail of the risk distribution — and it depends entirely on which statistic you use to measure “disagreement.”

Measured by plain standard deviation, dispersion among vendors’ flood-depth estimates increases the further out you project (from 2025 to 2100) and the more severe the return period you examine (from a 1-in-20-year event to a 1-in-1,000-year event). That’s the intuitive result: further out, or further into extremes, should mean more uncertainty.

But measured by the coefficient of variation (standard deviation divided by the mean) — a better metric when comparing series with very different average magnitudes — the result reverses: dispersion decreases as you project further into the future or into more severe return periods. The explanation GARP offers is that the mean of vendors’ estimates is rising faster than the standard deviation around it, which is consistent with all 13 vendors independently agreeing that physical risk is rising over time due to climate change, even while disagreeing considerably on the exact magnitude at any given point. It’s a useful reminder that “how much do vendors disagree” doesn’t have a single answer — it depends on whether you’re asking about absolute or relative disagreement, and the two can point in opposite directions.

The study also checked whether vendors that assumed more aggressive global warming (their estimates of degrees of warming under the shared RCP 8.5 scenario ranged from a 0.7°C spread by 2025 to a 3.2–5.7°C spread by 2100) were also the ones producing the highest physical risk estimates. They largely weren’t — differences in warming assumptions didn’t meaningfully predict a vendor’s ranking on flood, windstorm, or cyclone risk, which points back to modeling and methodology choices, not just differing climate assumptions, as the dominant driver of disagreement.

A Practical Checklist, Not Just a Diagnosis

GARP’s report is explicitly written to be operational for financial institutions selecting or validating a vendor relationship, structured around seven areas: use case, peril and geographic coverage, asset data quality, scenario needs, granularity and resolution, methodology and outputs, and estimates of uncertainty. A few of the sharper, more specific questions it recommends:

  • Does the vendor treat your asset as a point, a fixed buffer around a point, or the actual building/site footprint — and does that match how you need it modeled?
  • What baseline period does the vendor use for “degrees of warming” under a shared scenario — differences here alone can shift results even when the named scenario is identical?
  • Does the vendor provide any measure of uncertainty at all? (In this study, 4 of the 13 vendors provided none.)
  • What averaging or “centering” window does the vendor apply when reporting a hazard “in 2050” — a 10-year centered average and a 20-year centered average are not the same number?

Conclusion

GARP’s study doesn’t try to declare a “best” vendor, and it explicitly resists the idea that one exists independent of use case. What it does instead is give financial institutions — and anyone else relying on this kind of data — an evidence base for something that’s easy to suspect but hard to substantiate without a controlled test: that physical climate risk numbers from different vendors, even for the identical asset under the identical scenario, can disagree enough to change a decision. A geocoding error of 1,507 kilometers, a damage-ratio correlation of 0.2 between two vendors assessing the same 100 buildings, and a metric (wildfire) too inconsistently reported across vendors to benchmark at all are concrete, reproducible findings, not general skepticism about the field. Read alongside the ILN’s later, complementary study, the picture that emerges is consistent: current market practice compresses a long chain of real methodological choices into a single confident number, and the gap between vendors is often traceable to specific, nameable decisions — which means it’s also addressable, with better buyer questions and, potentially, more standardized industry practice over time.

Access the Source Material

  • Paisley, J., & Nelson, M. (2025). A Risk Professional’s Guide to Physical Risk Assessments: A GARP Benchmarking Study of 13 Vendors. GARP Risk Institute, for the Climate Financial Risk Forum (published via the UK Financial Conduct Authority). Full PDF
  • Paisley, J., & Nelson, M. (2025, October 23). Comparing Climate Risk Vendors: A User’s Guide to Physical Risk Assessments. GARP Risk Institute (summary article of the same study).