The NAIC's Big Data and Artificial Intelligence (H) Working Group began running its AI Systems Evaluation Tool through a live pilot on March 1, 2026. Twelve state insurance departments are requesting model inventories, governance documentation and high-risk system details from domestic insurers under a framework that did not exist 12 months ago.

Six trade associations had already told regulators the exercise is "one-sided, voluntary for regulators while compulsory for companies." That objection is now being tested inside actual exam rooms.

Key Takeaways

  • 12 states are running the tool from March 2026 through September 2026: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin. They meet monthly to reconcile how the exhibits are being applied.
  • Four exhibits, used in sequence. Exhibit A is a model inventory that screens; B covers the governance program and ORSA integration; C requests detail on models the company itself calls high risk; D traces data lineage.
  • December 5, 2025: six trade associations filed a joint letter raising five objections, including that findings during the pilot could support enforcement before the tool is finalized.
  • 24 states had adopted the Model Bulletin or pursued legislation as of late 2025, with the NAIC reporting 25 by March 2026, yet New York, Colorado, California and Texas each impose requirements the bulletin does not contain.
  • Nearly one-third of health insurers do not regularly test their models for bias, per the NAIC's May 2025 survey. That is the gap Exhibit C's testing questions reach into.

What the Pilot Actually Is

The tool is not a questionnaire but a four-exhibit escalation ladder, and the legal footing under it is ordinary exam authority rather than a new disclosure regime.

Exhibit A is the screen. It asks how many models are in production, how many are new, updated or retired, which carry direct consumer or material financial impact, and whether any have drawn consumer complaints. For a multi-line carrier with models spread across pricing, underwriting, claims triage, fraud detection and marketing, producing a defensible count is the hard part. What Exhibit A returns decides whether a regulator opens the deeper exhibits at all.

Exhibit B asks for the governance program itself, including vendor oversight and how AI considerations feed Enterprise Risk Management and ORSA. Exhibit C asks how each high-risk model was developed, what testing was run, and how much human involvement sits in the loop. Exhibit D covers the data: external against internal sources, third-party licensing, training composition, and the lineage of inputs into consumer-facing decisions. Vendor contracts are where that lineage most often runs out, a gap we mapped in our AI governance gap analysis.

The pilot expanded to 12 states by February 2026 after California joined, having started from ten insurers selected by a smaller group. Information collected is protected by the administering state's confidentiality rules, and the Working Group has said participating regulators will leverage existing exam authority. Counsel advising insurers has read that plainly: Foley and Lardner told clients to treat a request like an early-stage exam inquiry, and noted participation can be effectively mandatory at the state's discretion.

The Company Draws the Perimeter, Then Defends It

The design choice that matters most is in Exhibit C: the insurer sets its own definition of high risk. That is a real concession, and it also makes the classification methodology the thing under examination.

A self-set boundary is only defensible with written criteria, a record of which models were tested against them, and dates for when each was last reviewed. Without that, the boundary is a judgment call made after the request arrives, and the regulator gets to probe where it was drawn. AHIP and several life trade groups have asked the NAIC to anchor a common definition rather than leave each state to interpret it, which would remove the concession and the exposure together.

For actuarial teams the consequence runs through Exhibit B rather than Exhibit C. Folding AI governance into ORSA puts model risk inside the same document that carries the capital assessment, so a pricing or reserving model excluded from the high-risk inventory still has to be accounted for in the risk narrative that supports the capital number. The exclusion has to be explained somewhere.

The bias-testing figure shows how far the documentation gap can run. The NAIC's own May 2025 health survey found nearly one-third of health insurers do not regularly test models for bias, a practice the Model Bulletin already expects. Exhibit C asks for the testing record directly.

Bulletin compliance does not settle the question either. New York's Circular Letter No. 7 requires its own AI risk management framework, Colorado's 2024 Artificial Intelligence Act layers testing procedures onto existing insurance regulation, California's Fair Employment and Housing Act AI rules took effect October 1, 2025, and Texas enacted TRAIGA in June 2025. As Fenwick's tracker sets out, the compliance burden scales with geographic footprint rather than with how much AI a carrier actually uses.

The Objections the Pilot Has Not Answered

The December 5 joint letter raises two objections that go to whether pilot responses are safe to give, and neither has a written answer.

The first is that the same exhibits can be deployed inside a market conduct exam, which looks at consumer treatment, or a financial examination, which looks at solvency and reserving. Those tracks have historically been separate, and combining them means one disclosure can surface in two enforcement contexts. The second is sharper: the associations state that companies can apparently be penalized for negative findings drawn from pilot data before the tool is final. Standard NAIC pilot practice uses results to refine the instrument, and the Working Group has not committed in writing to that limit.

The associations also note the pilot has no binding end date. September 2026 is the stated close, but nothing in the tool holds states to it, and the November 19, 2025 draft went into the field while its own comment period was still open.

Federal action has added a second layer of uncertainty. An executive order signed in early December 2025 sets out a national AI framework aimed at displacing state rules. McCarran-Ferguson has shielded the business of insurance from preemption for decades, and Iowa Commissioner Doug Ommen has argued that blocking commissioners from coordinating AI supervision would depart from a state system that has worked for over 150 years.

Regulators are proceeding as though nothing changed. Carriers deciding how much to invest in a pilot response are weighing an unfinished tool against a federal standard that may reset the frame, and S&P's reporting found NAIC membership itself divided on where the rules should end up.

Further Reading

Sources

  1. NAIC Big Data and Artificial Intelligence (H) Working Group
  2. NAIC Insurance Topics: Artificial Intelligence
  3. NAIC Request for Information: AI Model Law (May 2025)
  4. NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers (December 2023)
  5. NAIC AI Systems Evaluation Tool Draft 4.0
  6. NAIC AI Systems Evaluation Tool Pilot Project Summary
  7. NAIC BDAIWG July 16, 2025 Meeting Minutes
  8. NAIC BDAIWG September 29, 2025 Meeting Minutes
  9. InsuranceNewsNet: NAIC's 2026 AI Evaluation Pilot Moves Ahead as Industry Balks
  10. Fenwick: NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States
  11. Monitaur: NAIC AI Systems Evaluation Tool Pilot, A Guide for Insurers
  12. Foley & Lardner via Mondaq: What To Do If You Receive a NAIC AI Systems Evaluation Tool Pilot Request
  13. Carlton Fields: NAIC Big Data and Artificial Intelligence Working Group Conceptualizes Tools
  14. S&P Global Market Intelligence: NAIC Membership Divided on Developing AI Model Law
  15. American Academy of Actuaries: Comment Letter on AI Model Law RFI
  16. Leader's Edge: Charting the Course of AI Governance
  17. Fenwick: Tracking the Evolution of AI Insurance Regulation
  18. NCOIL: Committee Working Drafts (NCOIL Model Act Regarding Insurers' Use of Artificial Intelligence)
  19. McDermott Will & Emery: An Update on NAIC's Consideration of AI Model Law for Insurers