Three months into the NAIC's 12-state AI Systems Evaluation Tool pilot, the Big Data and Artificial Intelligence (H) Working Group's June 1, 2026 meeting produced the first structured inventory of how insurers actually run AI in production.
The value of that inventory is its provenance. These are regulatory filings made under examination authority, not survey responses, which is why they read differently from every adoption statistic published so far.
Key Takeaways
- 12 states have run the pilot since March 2026: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin.
- A carrier running 200 AI models might see 15 to 30 flagged for deeper review under the proportionality principle, depending on line mix and degree of automation.
- Life insurers keep human review for denials but let AI-driven approvals proceed with less oversight, a structural asymmetry the pilot surfaced and the Working Group has not yet addressed.
- Exhibit C carries the compliance burden, and its high-risk classification is currently self-reported by the insurer, with the four-tier risk taxonomy meant to clarify the boundary still in draft.
- 68% of insurers rely on third-party AI vendors and 18% track vendor model risk systematically, yet those vendors' models still have to appear in the Exhibit A inventory.
What the June 1 Briefing Showed
Pennsylvania Commissioner Michael Humphreys, who chairs the Working Group, opened with a mid-pilot status briefing from the participating states alongside an actuarial panel on governance.
On the P&C side the reported use cases span the value chain. Marketing systems using behavioral and demographic data generally classified as lower risk, though regulators noted that the line between marketing and underwriting blurs when the same signals feeding ad targeting also feed quote eligibility. Underwriting split between renewal evaluation models flagging policies for non-renewal or premium adjustment, and aerial and satellite property inspection, mostly through third-party vendors. Claims covered accident image analysis, ultimate claim estimation and fraud detection, with straight-through processing reaching 70% at some carriers in low-severity auto and property.
Pricing submissions are the ones nearest actuarial work, and the boundary there is not clean. Carriers described gradient-boosted trees and neural architectures generating individual risk scores, but several reported hybrid systems where the ML model produces features that a filed GLM then consumes. Whether the ML component is documented sufficiently for rate filing review is a separate workstream running alongside the pilot.
Life submissions are narrower and higher stakes: accelerated issuance pipelines compressing application-to-bind from weeks to minutes, models contributing to approval and denial outcomes, and AI-assisted risk class assignment. The pilot surfaced one asymmetry worth naming. Most life insurers maintain human review for denials while allowing AI-driven approvals to proceed with less oversight.
Proportionality Decides How Much Gets Examined
Exhibit A, the AI inventory, is the triage layer. A carrier submits model counts by use case, consumer impact and financial materiality, and the reviewing regulator uses that to decide which models escalate into Exhibits B, C or D.
The pilot data suggests regulators are drawing the line along three dimensions: whether the system affects a coverage decision, whether it operates with limited or no human review before its output reaches the consumer, and whether it uses external data whose provenance the insurer does not control. A yes on any two appears sufficient to trigger Exhibit C in most participating states, though no bright-line rule has been formalized.
In practice that means a carrier running 200 models might see 15 to 30 escalated. That is a real reduction against a framework treating every model alike, and those 15 to 30 will face documentation scrutiny most carriers have not met outside a rate filing proceeding.
The classification decision is where this stops being a compliance exercise. Any pricing model consuming ML-generated risk features is a high-risk candidate, and so is a reserving model using AI for severity prediction or IBNR estimation. The boundary between high and medium risk for such a model is a judgment about how much the ML component moves the indication, which is an actuarial question before it is a legal one.
The life side carries the sharper version. If an ML model assigns preferred, standard or substandard risk class, the mortality assumptions behind product pricing and reserving sit downstream of that model. A calibration error there does not stop at the individual applicant: it propagates into pricing margins, reserve adequacy and the experience studies used to validate both, and it does so invisibly because the experience data confirms the classification that produced it. VM-20 prescribed assumptions do not address the case where risk class assignment is itself algorithmically determined.
Exhibit C Carries the Burden, and Its Boundary Is Self-Reported
Exhibit C asks, for each model the insurer classifies as high-risk, how it was developed, what data trained it, what testing including bias testing was performed, the degree of human involvement, the monitoring cadence and the supporting documentation. For a carrier with 20 high-risk models, that is a multi-week project across actuarial, data science, IT, legal and compliance.
The structural problem is that the insurer makes the classification. Classify too few models as high-risk and invite a challenge to the methodology; classify too many and the documentation burden compounds against no regulatory benefit. Carriers have asked for clearer guidance on the high-versus-medium boundary, and the four-tier risk taxonomy proposed at the Spring meeting remains in draft.
The second objection is trade secret exposure. A pricing model's feature set and the variables it weights most heavily are competitive differentiators, and insurers have argued that disclosure creates risk if materials become subject to freedom-of-information requests or move across state lines through NAIC market analysis procedures. The Working Group's answer is that everything collected is protected under the administering state's confidentiality statutes and that the pilot rests on existing examination authority rather than a new disclosure regime. Whether that holds depends on language not yet written.
Sitting underneath both is a capability question the exhibits assume away. The Conning-Datos survey puts 82% of insurers adopting AI in some form against 7% at enterprise scale, and Capgemini found 42% of P&C insurers have never measured AI outcomes. A carrier that has never assessed whether a model performs as intended cannot complete Exhibit C's testing and monitoring sections in a form an examiner will accept, whatever the model is actually doing.
The vendor layer compounds it. AM Best found 68% of insurers rely on third-party AI vendors while 18% track vendor model risk systematically, and those vendor models belong in the Exhibit A count regardless of how little the carrier can see inside them. The pilot states also selected carriers with the largest and most mature AI footprints first, so the mid-pilot picture reflects the top of the distribution rather than the industry median.
Further Reading
- Insurance AI Hits the Pilot-to-Portfolio Wall: why 93% of insurance AI initiatives stall before portfolio scale, with holdout group design, drift monitoring thresholds, and the audit trail the NAIC examination tool will expect to see.
- NAIC AI Evaluation Pilot Launches Amid Industry Pushback: the original pilot structure, the six-association joint letter, and the full four-exhibit breakdown.
- NAIC Targets AI in Claims Handling at Spring 2026: how the Working Group flagged claims for additional scrutiny, the 88% auto insurer AI adoption rate, and state claims AI laws.
- State AI Law Patchwork Forces Carriers Into Four Compliance Regimes: the evaluation pilot as a de facto fourth compliance layer alongside Connecticut SB 5, Colorado SB 26-189, and Texas TRAIGA.
- NAIC Four-Tier AI Risk Taxonomy Redefines Insurer Compliance: the proposed risk classification framework that will anchor the Exhibit C high-risk boundary.
- The AI Risk Evaluation Supplement v5.0 and the Model Inventory Request: what the pilot changed in the document, the materiality threshold a company sets and discloses, and the GLM scope question this data raises.
- Market Conduct Modernization Working Group Targets Exam Frameworks: the new (D) Committee body translating pilot findings into structural exam methodology.
- 68% of Insurers Outsource AI, Only 18% Track Vendor Risk: the AM Best vendor governance gap that Exhibits B and D are designed to surface.
- NAIC Weighs Jump From AI Bulletin to Enforceable Model Law: the 33 RFI comment letters and the fault lines shaping whether the evaluation tool leads to legislation.
- Automated Claims Decisions Face Regulatory Pushback: how the pilot intersects with state unfair claims settlement acts and human-in-the-loop requirements.
- NIST AI Agent Standards Set the Compliance Baseline: the parallel federal governance framework relevant to vendor procurement and agent-based AI architectures.
- Allstate's drift-monitoring patent and claim controls: a carrier-owned example of the dashboards, alerts, second-test verification, and retraining evidence regulators may expect to inspect.
Sources
- NAIC Big Data and Artificial Intelligence (H) Working Group
- NAIC Insurance Topics: Artificial Intelligence
- NAIC AI Systems Evaluation Tool Draft 4.0
- NAIC AI Systems Evaluation Tool Pilot Project Summary
- Fenwick: NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States
- Monitaur: NAIC AI Systems Evaluation Tool Pilot, A Guide for Insurers
- InsuranceNewsNet: NAIC's 2026 AI Evaluation Pilot Moves Ahead as Industry Balks
- Repairer Driven News: Regulators Examine AI Behind Claims Payouts
- NAIC Request for Information: AI Model Law (May 2025)
- NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers (December 2023)
- American Academy of Actuaries: Comment Letter on AI Model Law RFI
- Celent: 2026 Global Insurer GenAI Production Survey
- Capgemini: World Property and Casualty Insurance Report 2025
- Swept AI: NAIC 12-State Pilot Analysis
- Foley & Lardner via Mondaq: What To Do If You Receive a NAIC AI Systems Evaluation Tool Pilot Request