Only 7% of insurance AI initiatives advance beyond the pilot phase, and fewer than half of carriers have deployed AI in a single operational function (BCG, via Simplifai). The attrition is not random. Pilots run on curated data and narrow claim types; portfolio deployment meets model drift, leakage and adverse selection that a controlled test was never built to detect. The gap between a pilot number and a book number is mostly credibility.
Key Takeaways
- 7% reach portfolio scale while 88% of private-passenger-auto carriers use, plan to use or are exploring AI in underwriting and pricing. Exploration and production operation are not points on one continuum.
- More than 81% of insurers spend at least $5 million a year on AI and 14% spend over $50 million, and the spend has continued independently of deployment results.
- A 500-claim pilot carries roughly 30 to 40 percent credibility weight on a pure premium indication at moderate severity. A 10-point improvement measured there indicates 3 to 4 points once weighted.
- 49% of disclosed use cases target speed and cost rather than risk selection, so they cannot reach a combined ratio unless automated disposition accuracy at least matches the manual process it replaced.
- Forty-four percent of insurance executives report governance problems contributed to AI failure and only 24% believe their controls would survive an independent audit.
The Adoption Number and the Deployment Number Measure Different Things
Adoption reads as expansive. Eighty-eight percent of private-passenger-auto carriers use, plan to use or are exploring AI or machine learning in underwriting and pricing, on the NAIC's 2025 survey, and more than 81% of insurers commit at least $5 million a year, with 14% over $50 million.
Production-scale deployment sits in customer service chatbots and document summarization. End-to-end workflow automation in underwriting and claims is the least common form. "Carriers have run hundreds of pilots, spent real money, and in most cases have little production deployment to show for it," per Carrier Management.
The capability side shows the same shape. Twenty of the 30 insurers scored by the 2026 Evident AI Index now report at least one use case with documented outcomes, up eight from the prior year, but 49% of those remain narrow, aimed at speed and process efficiency.
That distinction has an arithmetic edge to it. A carrier processing 40% of personal auto claims without human review saves two to three minutes of adjuster time per claim. It does not move a combined ratio unless automated disposition accuracy on those claims matches or beats the manual process. Establishing that equivalence is a measurement problem, and it is the one most implementations have skipped.
Credibility Is What the Pilot Number Is Missing
Three design features make pilots overstate lift, and the second one is quantifiable.
Data selection is the first. Pilots take the cleanest slice: standardized commercial auto submissions with complete telematics, water damage claims under a severity threshold, personal lines renewals with three or more continuous years. Extending to complex commercial property or catastrophe-exposed homeowners introduces distributional shift, where the training data stops describing the production population and accuracy falls with no model failure in the conventional sense. "A demo that works against sample data is not the same as a solution that works against a Guidewire implementation with 12 years of customization on top of it," Carrier Management noted.
Volume credibility is the second, and it is where the pilot number breaks. Partial credibility on 500 claims of moderate severity assigns roughly 30 to 40 percent weight to the observed data, the rest carried by the complement. A pilot reporting a 10-point loss ratio improvement on that base indicates 3 to 4 points once weighted, so a board figure taken raw overstates the likely portfolio effect by a factor of two to three. The 3,000 exposure units across two full policy years usually cited as the attribution threshold is that credibility weight, not a convention.
Self-selection is the third. Pilot models are built on business existing underwriting rules already accepted, which is a filtered sample biased toward risks the carrier prices confidently. Extension fails disproportionately because it meets exactly what the pilot excluded: the complex commercial submission, the claim with disputed liability, the policy with gaps in coverage history.
| Dimension | Pilot Condition | Portfolio Condition |
|---|---|---|
| Data quality | Curated, complete, standardized | Heterogeneous, missing fields, legacy codes |
| Volume | 200 to 1,000 claims or submissions | Full book, all complexity bands |
| Model drift | None (static training set) | Continuous; requires defined monitoring cadence |
| Override tracking | Typically not captured | Must be tracked and analyzed as diagnostic data |
| Adverse selection | Not present (controlled test) | Present wherever model rejects or re-prices risks |
Getting past this needs a holdout, and the holdout needs to be sized. For an underwriting model claiming a 5-point loss ratio improvement, 85% power at a two-sided alpha of 0.05 takes roughly 1,200 to 1,500 policy years of holdout experience in a personal lines book of typical variance. A 90-day result on a long-tail liability book is a preliminary signal at any sample size, because reopened claims, subrogation and late ALAE have not emerged.
The Model That Works Still Moves Risk Somewhere Else
An AI-driven tightening that improves the loss ratio on the bound book does not remove the risk from the market. Rejected or re-priced risks move to competing carriers or to the assigned risk mechanism, and the effect compounds with concentration: a carrier holding 15% share in a territory that tightens pulls a material volume out of the admitted pool and concentrates the residual across every carrier that did not tighten at the same time.
Override behavior is the diagnostic most carriers are not capturing. An adjuster pool overriding 40% of claims AI recommendations is reporting something specific about model fit. Reading it requires comparing closed-claim outcomes where the recommendation was followed against those where it was overridden, stratified by complexity and severity band, with the lag definition held constant across both.
The ownership question is the one that binds hardest. The NAIC Model Bulletin on the Use of AI Systems, adopted in 24 states and Washington, D.C., requires written governance programs, accountability structures and bias testing, and it does not distinguish proprietary from licensed vendor models. If the model touches a regulated insurance decision, the insurer carries the documentation obligation.
Four records make that obligation operable for any AI-influenced decision under examination: the model version active at decision time, the inputs behind the output, the governance sign-off approving that version for production, and the override record where a human departed from it. None of the four is a default deliverable in a vendor SaaS contract. They are negotiated terms, and the 12-state NAIC AI Systems Evaluation Tool pilot running January through September 2026 is already asking for the governance half of them.
Further Reading on actuary.info
- Capgemini: 42% of P&C Insurers Never Measured AI Outcomes -- the 344-executive survey behind the industry measurement gap, including the 72/28 split between infrastructure and change management spending that explains why deployed AI does not always move the P&L.
- Insurer AI Returns: What the Evident AI Index Measures and What an Actuarial Scorecard Should -- a parallel analysis of the three carriers that have published enterprise-level AI ROI figures and the five metrics that connect AI spending to insurance performance.
- AI Governance Gap in Actuarial Practice -- how governance obligations for AI model validation sit at the boundary between actuarial and technology accountability, and what documentation the actuarial sign-off requires.
- NAIC Pilot Midpoint: What 12 States Found in Insurer AI Inventories -- mid-pilot regulatory findings on how examiners are categorizing insurer AI systems and which documentation gaps are surfacing first.
- Insurer AI Vendor Risk: The 68/18 Accountability Gap -- the governance obligations that carriers cannot transfer to vendors in AI model deployments, and what that means for audit trail requirements under the NAIC bulletin.
- NAIC AI Evaluation Pilot 2026: Industry Pushback and What Examiners Are Actually Looking For -- the industry response to the 12-state examination tool and the specific documentation categories driving examiner attention.
- EXL: 76% of Insurers Think They Lead AI, Only 6% Do -- EXL's finding that most insurers overrate their own AI maturity, with the gap concentrated hardest in actuarial and underwriting deployment.
Sources
- "Insurance AI: What You Won't Read in the Press Releases," Carrier Management, June 19, 2026. carriermanagement.com
- "Insurance industry still stuck in AI pilot phase, report finds," CIO Dive, 2026. ciodive.com
- "2026 Evident AI Index for Insurance Key Findings Report," Evident Insights, June 2026. evidentinsights.com
- NAIC Insurance Topics: Artificial Intelligence, National Association of Insurance Commissioners. Includes the Model Bulletin on the Use of AI Systems (adopted December 2023) and the AI Systems Evaluation Tool pilot details. content.naic.org
- "Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers' Use of AI," Quarles Law Firm, 2026. quarles.com
- "How the NAIC AI model bulletin is evolving and why insurers should prepare now," Plante Moran, March 2026. plantemoran.com