Forty-eight percent of insurers do not license catastrophe models (Aon survey, 2025), which leaves cat and specialty pricing running on own-company experience fragments. Guidewire's federated machine learning, built on its integration of integrate.ai, trains models by aggregating gradient updates across the carrier ecosystem while every raw exposure record stays local. It solves the data supply problem. It does not solve the filing problem.

Key Takeaways

  • 48% of insurers do not license cat models, and of those that do, nearly 60% run cat teams of five people or fewer, with only 27% maintaining internal teams to assess the models they license.
  • The sparsity is in event count per rating cell, not in loss dollars. Global insured cat losses reached $137 billion in 2024, concentrated in a handful of events, so a single carrier may see a dozen occurrences per peril per decade.
  • Calibration is where the filing is earned. A model trained across 40 carriers can be well calibrated to the population and systematically biased for any carrier with unusual attachment or retention structure.
  • The carrier can audit its own gradient contribution and nothing else. Against an ISO or NCCI pool, where a regulator can be walked through published methodology, that is a new gap in the documentation chain.
  • PSI above 0.20 on a top-importance rating variable is the conventional investigation threshold, and above 0.25 warrants revalidation, before updated global parameters reach new business.

What Sparse Actually Means Here

Credibility standards set the weight given to observed experience against expected. In large-commercial general liability, excess umbrella, inland marine and catastrophe-exposed property, most carriers accumulate too few claims per segment per year to move past blending on their own.

The cat modelling data shows the same shape from the tooling side. Beyond the 48% who license nothing, nearly 60% of those who do operate cat risk teams of five or fewer, often working from broker interpretations, and only 27% maintain internal teams dedicated to assessing the models they license.

The constraint is not the loss total. Global insured natural catastrophe losses reached $137 billion in 2024, but those concentrate in a handful of events. A carrier writing excess earthquake in the Pacific Northwest may accumulate ground-up data from one or two moderate events across fifteen years, while every segment of that book, by attachment point, construction type and occupancy class, carries exposure and almost no claims.

A model fits that data fine. The relativities come out with confidence intervals wide enough to make the model indefensible in a filing, and the traditional remedies all give something up: broad blending with industry data, simplifying the rating plan to reduce cells, or restricting the book to segments where experience happens to be better. None helps when the competitive opportunity sits in the thinnest cells.

Better Training Data Does Not Make a Model Fileable

Federated training splits the process. Each carrier trains a local update on its own data and transmits gradient updates, the mathematical adjustments moving the global model toward better predictions on that data, to an aggregation layer that combines them and redistributes. No exposure records, claim histories or risk characteristics leave any environment. A national carrier with thin Oregon earthquake exposure contributes its sparse gradient and receives a model that has seen more Oregon earthquakes than it ever will.

The technical catch is client drift. When participant data is not identically distributed, gradients from different carriers push the global model in conflicting directions. In insurance that heterogeneity is structural: a coastal homeowners book and a Midwest book have different cat physics, and the aggregate relativities blend both. A January 2026 study on federated parametric insurance index design confirmed a useful common signal can be extracted across portfolios with different underlying physics, with performance depending heavily on how the aggregation protocol handles heterogeneity.

That is why a better holdout Gini does not settle anything.

Actuarial Usability Test What It Checks Federated Model Complication
Calibration Predicted vs. actual loss for own book Global model calibrated to population mean, not to carrier’s own attachment or retention structure
Monotonicity Relativities in directionally defensible direction Portfolio heterogeneity across participants may produce counterintuitive gradient-weighted relativities
Exposure drift Input distribution stability over time Participant mix shifts update the global model without carrier-level notification
Governance trail Documentation of training data, version history Carrier cannot audit full training dataset; only its own contribution is visible
Filing explainability Factor-level attribution of relativities Global gradient aggregation obscures which participants drove a given factor’s direction and magnitude

Calibration is the test the actuary cannot delegate. A model trained across 40 carriers can be well calibrated to the global population and systematically biased for one participant whose attachment structure or claims handling differs. The actual-to-expected ratio on the carrier's own book, using the federated model, is the adjustment that makes the filing possible.

Monotonicity fails differently in this setting. A filing needs relativities that move in defensible directions, and heterogeneity can introduce a non-monotone one: if the carriers dominating gradient updates for a rating variable select differently on that variable, the global model learns a relativity that reflects their underwriting selection rather than loss experience. It is not fileable regardless of holdout performance.

Exposure drift is the operationally hardest. Global updates train on participants' recent experience, so a shift in the participant mix or a material change in any one portfolio moves the model without flagging anything for review at the receiving carrier.

The Audit Chain Runs Through Data the Carrier Cannot See

Existing industry sharing does not work this way. ISO statistical agent programs, NCCI workers compensation data and state rating bureaus collect raw records under statutory authority and publish aggregated experience, so a regulator examining a homeowners filing can request the development and trend methodology and the carrier can point at published documentation.

Federated training buys access to data that would never enter a pool. A carrier with meaningful first-party cyber breach experience will not submit it to a statistical program, because the competitive intelligence inside it is worth more than what comes back. The same holds for parametric triggers in specialty structures and manuscript excess and surplus forms.

The cost of that privacy is the audit trail. The carrier can document its own contribution and the aggregation protocol; it cannot walk a regulator through the training data of the other participants. The NAIC's AI Systems Evaluation Tool pilot, running across 12 states as of March 2026, requires documentation of training data sourcing under Exhibit C, and for a federated model that exhibit resolves to a statement of what the carrier can and cannot audit.

The second gap is cadence. A GLM retrained annually on own data synchronises naturally with the filing cycle. A federated model receiving updated global parameters weekly or monthly can move relativities faster than the filing process tracks them, which forces a question the pilot also asks: is each global update a new version requiring sign-off and possible re-filing, or scheduled maintenance inside the filed specification?

Monitoring is what stands in for an answer until one exists. A Population Stability Index above 0.20 on any top-importance rating variable is the conventional trigger for investigation before updated parameters reach new business, and above 0.25 for revalidation, with segment-level loss ratio deviation as the performance-based complement. With NAIC survey data showing 88% AI adoption among 193 private-passenger auto insurers and 70% among 194 homeowners insurers, most carriers already run machine learning in pricing-adjacent applications. Writing those thresholds down before a federated model goes live is a different position from reconstructing the monitoring narrative once Exhibit C is open.

Further Reading

Sources

  1. Guidewire, “Guidewire and integrate.ai: The Next Frontier of Predictive Modeling with Industry Intelligence” (2025) — Federated machine learning capability, privacy-preserving model training, and Guidewire Intel use cases for sparse and specialty data.
  2. Artemis, “Aon Survey Highlights Critical Gaps in Cat Model Use Among Re/Insurers” (2025) — 48% of insurers lacking cat model licenses, 60% with cat teams of five or fewer, and 27% with dedicated model evaluation capacity.
  3. Global Reinsurance, “Insurers Divided on Cat Model Usage, Aon Study Finds” (2025) — Survey findings on catastrophe model adoption gaps and the broker-dependency pattern in small cat risk teams.
  4. Swiss Re Sigma, Natural Catastrophe Report (2025) — $137 billion in global insured losses from natural catastrophes in 2024, with projections for 2025.
  5. Bhatt et al., “Privacy-Enhancing Collaborative Information Sharing through Federated Learning: A Case of the Insurance Industry” (arXiv, 2024) — Federated learning architecture for insurance claim loss modeling, gradient aggregation methodology, and privacy preservation without raw data sharing.
  6. Federated Learning for Parametric Insurance Index Design Under Heterogeneous Production Losses (arXiv, January 2026) — Evaluation of federated vs. aggregation-based index design for parametric insurance, confirming ability to extract common signal from heterogeneous risk physics.
  7. NAIC, Private Passenger Auto AI/ML Survey Data Call (December 2022) — 88% of 193 responding auto insurers use, plan to use, or are exploring AI/ML models in their operations.
  8. NAIC, Home AI/ML Survey Data Call (August 2023) — 70% of 194 responding homeowners insurers report current or planned AI/ML deployment.
  9. NAIC, Big Data and Artificial Intelligence (H) Working Group (2026) — AI Systems Evaluation Tool pilot running across 12 states as of March 2026, with anticipated adoption at the November 2026 Fall National Meeting.
  10. Datos Insights, cited in Guidewire, 2026 P&C Insurance Trends (2025) — “By 2030, data maturity will make or break insurers,” with legacy data approaches crumbling under AI and analytics demands.