Verisk released a US Data Center Exposure Database on September 3, covering more than 2,500 facilities with rooftop-level geocoding, building footprints, construction type, floor area, capacity and operational redundancy characteristics, delivered into Synergy Studio and Touchstone. "If you can't identify the exposure, you can't effectively measure or manage it," said chief research officer Jay Guin. That is true, and it is the easier half of the problem.

What Actually Shipped

The product is an exposure file, not a model. It ships as flat-file records, building footprint shapefiles and a 90-metre disaggregation grid, and Verisk positions it for underwriting, exposure management, catastrophe analytics and portfolio accumulation. Rob Newbold, who runs Verisk Catastrophe and Risk Solutions, framed the need in the terms the market already understands: the facilities powering AI growth "represent billions of dollars in concentrated assets." The release cites industry projections putting global data center insurance premiums at roughly $10 billion in 2026 and $23 billion by 2030.

That last figure is a market-size estimate rather than a loss estimate, and the distinction matters for what follows. Premium doubling tells you the industry expects to write more of this risk. It says nothing about how well the risk is understood.

What the database does supply is genuinely useful and genuinely hard to assemble. Rooftop geocoding rather than street or ZIP centroid is the difference between knowing a facility is in a county and knowing whether it sits inside a flood polygon. Construction type and floor area feed the exposure term directly. Redundancy characteristics, which describe how a site is engineered to survive a utility or cooling failure, are the kind of attribute that normally lives in an engineering report rather than a schedule of values.

Exposure Is the Term That Was Already Solvable

A catastrophe model multiplies three things: hazard, exposure and vulnerability. Hazard for these perils is mature, with decades of hurricane, flood, severe convective storm and earthquake work behind it. Exposure is what Verisk just shipped. Vulnerability, the damage ratio a given asset suffers at a given hazard intensity, is the term this asset class cannot yet populate.

Vulnerability curves are fitted to claims. Hyperscale AI campuses in their current form are a 2024-onward phenomenon, which means the loss history available to fit a curve for a multi-billion-dollar facility full of liquid-cooled accelerators is close to empty. The curve has to be built by engineering judgment and analogy to conventional commercial property, and a data center is not conventional commercial property: most of the insured value is equipment rather than structure, the equipment is sensitive to interruptions that leave the building standing, and a large share of the economic loss arrives through delay and business interruption rather than physical damage.

The actuarial consequence is specific, and it is not that the database is unhelpful. It is that model output precision will now improve faster than model output accuracy. Feed better-located, better-attributed exposure into an accumulation run and the average annual loss and probable maximum loss figures that come back will be tighter, more granular and more confident-looking, while resting on a damage function nobody has validated against claims.

A number that narrows from a wide band to a tight one without the underlying uncertainty having changed is the more dangerous input to a rate or a reinsurance structure, because it invites decisions the evidence does not support. The honest treatment is to carry the vulnerability uncertainty explicitly, in the loading or the model-risk margin, rather than let exposure resolution stand in for it. That is a documentation problem as much as a technical one: an accumulation output built on a judgment-based damage curve should say so where a reviewer will see it.

A Partial File Is Not a Neutral One

Verisk attaches a caveat to the release that deserves reading twice: available data fields and levels of detail may vary by facility and source. For a single-risk underwriting decision, a missing attribute is a question to ask the broker. For portfolio accumulation, it is something else, because unpopulated fields get defaults, and defaults are assumptions applied at scale.

The direction of that gap is unlikely to be random. Public documentation is thinnest for the newest and largest builds, where operators disclose least and construction is most recent. If completeness correlates inversely with facility size and recency, the least-attributed records are systematically the largest exposures, and a portfolio run will be most confident exactly where it should be least. The 90-metre disaggregation grid is the tell: it exists because some exposure can be located only to a cell rather than a roof, and grid-spread value behaves differently in a correlated event than value pinned to a footprint.

The peril list is worth noting too. The release names hurricane, flood, severe convective storm and earthquake, which is the set the ILS and reinsurance markets price most efficiently. It is not the set that concentrates data center loss.

Our earlier reporting on Zurich's data-center quota share covered the Swiss Re modelling that puts a quarter of US capacity in high-hail zones and roughly 40% in significant tornado risk, alongside FM Global work attributing about 42% of industrial loss costs to fire against 11% of events. A database built for the modelled perils sharpens the part of the problem the market can already price, and leaves untouched the concentration where severity actually lives.

None of that is an argument against the product. Identifying exposure is a precondition for everything else, and 2,500 geocoded facilities is a real advance over a schedule of values with a city name on it. It is an argument for reading it as what it is: a solved exposure term, arriving well ahead of the vulnerability term it will be multiplied against.

Further Reading

Sources