INDUSTRY REPORTS

UNEP FI, backed by 44 banks and investors, ran an identical dummy portfolio past 15 climate risk vendors. The same portfolio scored 76.9, 63, and 53.4 out of 100 for physical risk from three different vendors — and physical Value-at-Risk estimates ranged from 0.6% to over 21%.

The 2023 Climate Risk Landscape: UNEP FI's Multi-Vendor Technical Supplement

· 7 min read

The 2023 Climate Risk Landscape: UNEP FI’s Multi-Vendor Technical Supplement

About the Source

This is a report by the UN Environment Programme Finance Initiative (UNEP FI), produced by its 2022 Climate Risk and TCFD Programme working group with pro bono participation from 15 climate risk vendors and support from 44 banks and investors. It’s not peer-reviewed academic research, and UNEP FI is explicit that it is not an endorsement of any vendor: “the report does not endorse or judge any specific methodology or results and does not make any claims on the superiority of one methodology or approach over others.” It sits alongside the ILN, GARP, and CarbonPlan comparisons as a coalition-produced, methodologically transparent look at the same market.

Introduction

Where the ILN and GARP studies test vendors against portfolios of individually specified real or dummy properties, UNEP FI took a different, complementary approach: it built one large, realistic fictional investment portfolio collaboratively with the banks and investors participating in its working group, then invited climate risk vendors to run their own tools against it, pro bono, and report back their results using whatever metrics and methodology they normally use. The result is less of a controlled experiment than the ILN and GARP studies (vendors weren’t constrained to a single shared metric or scenario), but it captures something the more tightly controlled studies can’t: what buyers actually see when they ask several vendors, in the vendors’ own native outputs, “what’s my risk?”

Building a Shared, Realistic Test Portfolio

Participating banks voted on which sectors, regions, and asset classes to include, settling on agriculture, real estate, energy, oil and gas, and transportation as the top five sectors of interest. The resulting dummy portfolio held 358 securities across 44 countries (skewed toward North America, reflecting participants’ actual holdings), spanning corporate loans, equity, mortgage loans, real estate, and municipal and sovereign bonds. Notably, 15 of the corporate loan holdings were deliberately made unlisted companies with minimal public data, specifically to test how vendors handle imperfect, incomplete inputs — a real-world condition that cleaner, purpose-built test portfolios (like ILN’s or GARP’s) don’t always capture. Fifteen vendors and tools participated in the 2022 working group, spanning both dedicated physical/transition risk specialists (XDI, Munich RE, Moody’s/RMS) and larger financial data providers (MSCI, S&P Global, BlackRock’s Aladdin Climate).

The Same Portfolio, Three Very Different Physical Risk Scores

For a headline comparison, three vendors — S&P Global Sustainable1, Moody’s, and ISS ESG — each returned a physical risk score for the identical portfolio on a shared 1–100 scale (100 being highest risk):

VendorPhysical risk scoreHazards scoredScenarioCoverage
S&P Global Sustainable176.98 hazardsSSP3-RCP7.0, out to 2050272/358 holdings
ISS ESG63.06 hazardsRCP4.5, out to 2050302/358 holdings
Moody’s53.46 hazardsRCP8.5, out to 2030–2040260/358 holdings

That’s a spread of more than 23 points on a 100-point scale for the same 358-security portfolio — comparable to the difference between a “moderate” and a “high” risk classification, depending on where a vendor draws its thresholds. Part of this is genuinely not comparable on its face: S&P Global’s score is normalized against global upper and lower hazard thresholds rather than a specific baseline year, Moody’s measures against “Moody’s universe” of assessed companies, and ISS ESG’s number is a sector-relative score, meaning 63 indicates lower risk than sector peers, not an absolute risk level. That’s precisely the report’s point: buyers who see “76.9,” “63,” and “53.4” without understanding these different reference frames could easily — and wrongly — read them as three measurements converging on roughly the same conclusion, or as meaningfully different assessments of the same risk, when they’re not measuring quite the same thing at all.

Physical Value-at-Risk: A 35x Spread

The dispersion is starker still for physical Value-at-Risk (PVaR), a metric intended to represent the percentage of portfolio value at risk from physical climate hazards. Three vendors reported PVaR for the same portfolio weighting:

VendorPVaRScenarioConfidence level
ISS ESG0.6%RCP4.5Not applicable
CLIMAFIN2.47%–2.78%RCP4.5 / RCP8.5, 2030–208095th percentile
MSCI9.98%–21.36%MSCI Average / Aggressive scenario50th / 95th percentile

ISS ESG’s 0.6% and MSCI’s 21.36% both purport to answer the same underlying question — how much of this portfolio’s value is at risk from physical climate hazards — and differ by more than a factor of 35. Some of this gap is attributable to genuinely different definitions (ISS ESG frames its number as a change in company valuation from physical risk costs, not a financial Value-at-Risk in the traditional sense) and different confidence levels and scenario severities. But the report is candid that a portion of the divergence isn’t explainable by definitional differences alone, and that a buyer comparing vendor proposals side by side, without understanding these mechanics, has no way to know which number — if either — is closer to their institution’s actual exposure.

A Cited Academic Cross-Check: Six Tools, 11 Sectors, No Consensus

The report cites an external academic comparison (Hain et al.) of six commercial physical risk scoring tools — Trucost, Carbon4 Finance, Southpole, Truvalue Labs, and two academically derived measures — applied to 408 US corporations and aggregated to sector rankings for 11 broad sectors under RCP8.5 by 2050. The standard deviation of sector rank across the six tools ranged from 1.9 (Health Care, Consumer Staples) to a striking 4.0 (Utilities) — meaning that for a sector like Utilities, different physical risk tools couldn’t even agree on whether it belonged near the top or bottom of an 11-sector risk ranking. This is an independent line of evidence, using different vendors and a different portfolio entirely, pointing at the same underlying problem the UNEP FI portfolio exercise surfaces directly.

Six Things Financial Institutions Actually Want (and Aren’t Fully Getting)

Rather than ranking vendors, UNEP FI closes with six capability gaps voiced by the participating banks themselves:

  1. Tailoring to risk profile — the ability to input an institution’s own risk thresholds and vulnerability assumptions, rather than accepting a vendor’s defaults.
  2. Balancing maximum vs. mean scores — most tools default to worst-case framing; banks want the option to see average risk too, to avoid distorted prioritization.
  3. Sector and industry comparability — better heatmapping and hotspot tools to compare exposure across a diversified portfolio, not just within single-asset assessments.
  4. Incorporating adaptive capacity — most tools’ adaptation assumptions are backward-looking (based on existing flood defenses, for instance) rather than forward-looking about planned resilience investment.
  5. Secondary risk analysis — cascading effects (the report cites Australia’s 2019–2020 bushfires causing a measurable ~5% drop in GDP, per Moody’s) are rarely modeled, even though they can dwarf direct asset damage.
  6. Data reliability and transparency — the most consistently voiced concern, and the one this report — and the ILN, GARP, and CarbonPlan studies alongside it — is fundamentally trying to help address.

Notably, the report’s conclusion cites research by Bingler and Colesanti Senni (2022) to argue that vendors need to improve model transparency, scenario flexibility, and communication of output uncertainty. That citation is almost certainly to their two-author paper “Taming the Green Swan” (Climate Policy, 2022) rather than to the three-author Bingler, Colesanti Senni & Monnin paper on transition risk metric convergence: a distinct but related paper by an overlapping set of authors on the same broad question. Worth keeping the two apart when following the citation trail.

Conclusion

UNEP FI’s technical supplement doesn’t try to crown a winner, and its own limitations section is candid about why a pro bono, single-portfolio exercise with no metric standardization can’t fully isolate methodology from data coverage from definitional choice. But as one data point among several — alongside the ILN’s and GARP’s more tightly controlled vendor benchmarks and CarbonPlan’s transparency-focused test — it adds a distinct and useful piece of evidence: even inside a coalition of 44 banks working collaboratively and transparently with vendors who volunteered their time, the same portfolio still produces physical risk scores that differ by more than 20 points out of 100, and PVaR estimates that differ by more than 35-fold. That’s not evidence that physical climate risk modeling is worthless — it’s evidence that the numbers on their own, without the underlying methodology, aren’t yet directly comparable across vendors, and that buyers who treat them as if they were risk making decisions on a false sense of precision.

Access the Source Material