Grant Thornton's 2026 AI Impact Survey, fielded between February 23 and March 18 across 100 insurance respondents, found that 44% of insurance executives say governance or compliance problems contributed to an AI project failing or underperforming. Only 24% are very confident they could pass an independent AI governance review within 90 days.

Carriers can show that their models work. The survey's finding is that most cannot show how they governed the models while they worked, and that is the gap that stops deployments.

Key Takeaways

  • 44% cite governance as a contributor to AI project failure, against 24% who are very confident of passing an independent review inside 90 days, from a survey of 100 insurance respondents among 950 total business leaders.
  • 68% report that AI controls exist but are fragmented across teams and tools, and while boards have widely adopted AI governance policies, only 20% have tested a response plan for an AI failure.
  • 73% are piloting, scaling, or running autonomous AI systems, so wide deployment sits on top of the 76% who acknowledge they could not evidence their controls on demand.
  • Formal inventory exercises surface 20% to 40% more in-scope systems than carriers estimate, which is the single most reliable budget surprise in a governance build.
  • A retrofit runs $4 million to $8 million for a mid-sized carrier, against roughly 15% to 20% of model development cost when governance is instrumented from the start.

What the Survey Found Underneath the Headline

The Grant Thornton figures separate three things that usually get reported as one.

Controls exist. Sixty-eight percent of insurance respondents say AI controls are in place but fragmented across teams and tools: a model risk team holds validation records, compliance holds bias testing results, IT holds data lineage logs, and nothing aggregates them into a record an examiner could read.

Policy exists. Sixty-one percent of boards have established AI governance policies. Set against the 24% audit-confidence figure, the distance between adopting a policy and being able to operate it is most of the survey, usually because the policy was written at a level of abstraction that does not translate into a testable control for a specific underwriting or claims use case.

Testing does not exist. Only 20% have tested their response plan for an AI failure. Tom Puthiyamadam, the Grant Thornton partner overseeing the survey, framed the aggregate pattern as investment that is not correlating with an increase in AI accountability. Meanwhile 73% of respondents are piloting, scaling, or running autonomous systems.

Why the Bill Arrives After the Model Works

The reason governance kills projects rather than delaying them is that the cost is incurred in the wrong order.

A model is built, validated on performance, and moved toward production. Compliance then asks for training data provenance, protected-class impact testing, and model change management records. Those artefacts have to be reconstructed rather than retrieved, because nothing recorded them at ingestion time.

Reconstruction is where the money goes. Building an audit-ready program for a mid-sized carrier with 2,000 to 4,000 in-scope systems runs $4 million to $8 million across personnel, infrastructure, and outside counsel.

WorkstreamTimelineKey Outputs
Model Inventory6-8 weeksComplete registry of all AI/ML systems, risk classification, owner assignment
Bias Testing (First Pass)10-14 weeksStatistical parity, four-fifths rule, proxy screening for all high-risk models
Drift Monitoring Dashboards6-10 weeksAutomated alerts for accuracy degradation, distribution shift, emerging bias
Audit Trail Specification2-4 weeksLogging requirements, retention policy, reproducibility standards
Vendor Contract Amendments2 weeksDisclosure requirements, bias testing access, incident notification SLAs
Organizational Standup8-12 weeksExecutive hire, team build, RACI assignment, reporting cadence

Two numbers govern that estimate. The first is scope: formal inventory exercises turn up 20% to 40% more systems than the carrier expected, once shadow models built inside actuarial teams, vendor-embedded scoring engines, and legacy rule systems that qualify as algorithmic decision-making are counted. A carrier budgeting for 2,000 should provision for 2,400 to 2,800, and the overshoot alone adds $500K to $1M.

The second is the retrofit premium. Instrumenting governance into a new deployment costs roughly 15% to 20% of model development cost. Retrofitting it onto a production pipeline costs multiples of that, because training data provenance has to be rebuilt from version control history and team memory, and change history from deployment logs.

That asymmetry is what makes the decision look like a resource prioritization rather than a governance failure when the project is shelved. It also explains the pattern in the monitoring data: most carriers watch a model for accuracy degradation, because accuracy affects the loss ratio, and few watch the same model for an emerging four-fifths rule violation on a protected class or for input distribution shift, because those affect only the examination that has not happened yet. In a first comprehensive bias pass, 15% to 30% of consumer-facing models get flagged, which is not a finding of discrimination but a finding that the carrier cannot affirmatively evidence its absence.

Four Examinations, One Set of Artefacts

The complication is that the audit layer is not being built for one examiner, and the deadlines land together.

The NAIC's AI Systems Evaluation Tool pilot runs from March 2, 2026 through September 2026 across 12 states, and it is a market conduct examination instrument rather than a survey. Its Exhibit C asks for design, training data and performance on high-risk systems; Exhibit D screens for proxy variables tied to protected characteristics. A carrier in a pilot state that cannot populate them is choosing between rapid remediation and a non-response to its regulator.

The Third-Party Data and Models Working Group advanced its vendor registration framework on March 23, 2026, heading for public exposure in Q3 2026 and adoption consideration in November. Registration creates no safe harbour: the carrier stays accountable for vendor model behaviour, so the registry mainly gives regulators a baseline to compare a carrier's own due diligence file against.

The EU AI Act classifies insurance risk assessment as high-risk under Annex III with enforcement from August 2, 2026, and penalties up to 35 million euros or 7% of worldwide revenue. Colorado's insurance compliance deadline is June 30, 2026, and nearly 25 states have adopted the NAIC Model Bulletin in some form.

Each of those tests the same underlying artefacts: the inventory, the bias testing record, the lineage documentation, the change log. That is the one piece of good news in the arithmetic, because the build is shared. It is also why the 76% figure is not four separate problems that can be sequenced. A carrier that cannot evidence its controls in September cannot evidence them in November either.

Further Reading

Sources