The NAIC's AI Systems Evaluation Tool, a four-exhibit examiner framework now piloting across 12 states through September 2026, turns the 2023 Model Bulletin's disclosure language into a scored market conduct instrument. Ten companies are enrolled.

Adoption is expected at the November 2026 Fall National Meeting. The exhibit that has no equivalent in current rate filing documentation is Exhibit D, and it asks a question a rate filing memorandum was never built to answer.

Key Takeaways

  • Four exhibits, escalating in depth: an AI footprint count, a governance risk assessment, a high-risk system deep dive, and a data sourcing and quality exhibit.
  • 12 pilot states against 24 bulletin-adopting states means exactly half the disclosure-regime jurisdictions are testing whether disclosure can become a scored exam.
  • High-risk is currently self-determined in Exhibit C, and regulators have signaled it is not limited to consumer-facing systems; models affecting solvency sit in the same lens.
  • Exhibit D asks for data lineage, sources through transformations to the values that entered training, which most rate filing workflows do not preserve as a standing artifact.
  • A vendor registry is moving in parallel, with first-state implementation in the same late 2026 to early 2027 window as the tool's adoption vote.

The Four Exhibits and What Each Demands

The tool lets an examiner escalate in stages rather than open every carrier to the same depth. Exhibit A quantifies the AI footprint: how many systems are in production, by function, and which touch a consumer-facing or financially material decision. Exhibit B is a governance risk assessment covering roles, oversight and how AI risk feeds enterprise risk management. Exhibit C narrows to systems the company has classified high-risk, asking for development history, testing results and human-in-the-loop detail. Exhibit D covers data sourcing, quality controls, representativeness, and a field the pilot draft added for reasonable accommodations or policy modifications.

ExhibitWhat it asks forWhat the actuarial/model-risk team must produce
A: AI usage inventoryCount and classify every production AI system by function and impactA queryable model inventory tying each model to owner, use case, and consumer/financial materiality
B: Governance risk frameworkNarrative or checklist description of oversight, vendor management, ERM integrationDocumented model risk policy, escalation paths, and a record of how AI risk is reported into ORSA
C: High-risk system detailDevelopment, testing, and human-in-the-loop detail for self-classified high-risk modelsWritten high-risk classification criteria, applied consistently, plus validation and bias-testing records
D: Data integritySource, lineage, quality checks, and representativeness of training and input dataA traceable lineage from raw source through feature construction to the model's production inputs

The difference from the bulletin is evidentiary rather than topical. The bulletin asked a carrier to describe its governance program in a narrative. The exhibits ask for the artifacts an examiner can cross-check against that narrative.

Proportionality language tells examiners to spend more effort on Exhibits C and D where a system could cause serious consumer or financial harm and less on low-risk back-office automation. That reads as restraint and operates as concentration: it puts the heaviest burden on pricing and underwriting models, which are the systems most likely to be scored high-risk under any reasonable reading.

What Exhibit D Would Demand of a Rating Engine

Take a personal auto insurer running a gradient-boosted model to derive rating factor relativities that feed a GLM base rate structure filed with the states. Current documentation centers on the rate filing memorandum, built to satisfy ASOP No. 12 on risk classification and ASOP No. 56 on modeling, plus whatever internal validation the governance policy requires. That package explains what the model does and why its outputs are actuarially sound.

Exhibit D asks something else: the sources feeding the model, internal policy and claims data alongside purchased data, the quality controls at each stage, and whether the training population represents the population the model is applied to.

For an engine with dozens of candidate features, several from licensed providers such as credit-based insurance scores, motor vehicle records or property attribute vendors, that means answering for every feature where the raw value originated, what transformation produced the feature, and what check confirmed the transformation ran correctly. A rate filing memorandum answers whether a factor is justified. Exhibit D answers whether the data behind it can be proved end to end.

Most workflows do not preserve the second answer. Feature engineering pipelines get rebuilt between model refreshes, vendor contracts rarely specify documentation rights beyond delivery, and lineage tracking, where it exists, lives in a data engineering team's tooling rather than in a form that can be handed to a market conduct examiner. That is a data infrastructure gap, and it is the largest new lift the tool creates relative to ASOP-based filing documentation.

The exam logic behind it is borrowed rather than invented. Financial condition exams have triaged by inherent risk for two decades; market conduct has worked from complaint-and-sampling. The tool imports the financial-exam triage into market conduct through a four-tier severity scale floated at the Spring 2026 meeting, echoing the EU AI Act's classification while staying inside existing state exam authority.

The gap it opens is geographic for now. 24 states and the District of Columbia had adopted the Model Bulletin as of April 2026, with four more applying comparable guidance. In the 12 pilot states a carrier may now face an examiner working through Exhibits C and D with a specific model in hand; in the other bulletin states the lighter regime holds.

The Registry Makes the Exhibit D Answer Checkable

Exhibit D's reach into third-party data meets a second workstream head on. At its March 23, 2026 session the Third-Party Data and Models Working Group sketched a registration regime for vendors supplying AI models and datasets to insurers. Registration means filing information and updating it on a cadence. It is not licensure, and it creates no safe harbor: the insurer remains answerable for the model's behavior in pricing, underwriting, claims and fraud detection.

The consequence is that Exhibit D stops being a question the carrier answers alone. Once the registry exists, an examiner asking for training data lineage can compare the registered vendor's own filed disclosures against what the carrier represents, and any daylight between the two is a finding in itself. A vendor stack that was opaque to regulators by default becomes independently examinable from one side at the same moment the carrier is asked to document it from the other.

The timing lines up closely enough to matter. First-state registry implementation sits in late 2026 or early 2027, the same window in which the evaluation tool would move from pilot to adopted instrument.

The industry's objection so far has been to process rather than to that substance. A joint trade letter filed December 5, 2025 said the industry "remains significantly concerned about the lack of detail and guidance around the proposed pilot," pointing at effectively mandatory participation, no fixed end date, and results potentially informing enforcement before the tool is final. Those may be resolved in the September and October revision cycle. The documentation standard underneath them is the part that gets formalized rather than softened, and the November vote, not the pilot's close, is when it applies.

Further Reading

Sources

  1. NAIC AI Systems Evaluation Tool Pilot Project Summary
  2. NAIC Big Data and Artificial Intelligence (H) Working Group
  3. NAIC Insurance Topics: Artificial Intelligence
  4. Fenwick: NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States
  5. Monitaur: NAIC AI Systems Evaluation Tool Pilot, A Guide for Insurers
  6. InsuranceNewsNet: NAIC's 2026 AI Evaluation Pilot Moves Ahead as Industry Balks
  7. Quarles Law Firm: Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers' Use of AI
  8. Mayer Brown: US NAIC Spring 2026 National Meeting Highlights
  9. Swept AI: The 2026 NAIC Third-Party Model Law, A Vendor Registry Is Coming for Insurance AI