NAIC staff circulated a draft compliance report at the Spring 2026 National Meeting that converts the December 2023 AI Model Bulletin into nine specific disclosure components a carrier completes and attests to. The bulletin is adopted in 24 states and the District of Columbia and has never carried a template for proving compliance.

The Big Data and Artificial Intelligence (H) Working Group takes its next formal step at a July 22, 2026 public meeting. The report needs no Model Law to bite.

Key Takeaways

  • The draft turns the bulletin's five functional domains into nine disclosure components, including a board and senior management attestation that names an executive responsible for the program.
  • The 12-state AI Systems Evaluation Tool pilot runs March through September 2026 inside existing market conduct and financial examinations, not a separate AI track.
  • A carrier nominally running three vendor platforms and two internal models may hold twelve systems in scope once vendor-embedded components are counted.
  • Exhibit C covers third-party vendor models in full, with the carrier owing documentation whether or not the contract obliges the vendor to supply it.
  • Model inventories take 6 to 12 months to build accurately, against a requirement likely arriving in pilot states during 2027.

What the Bulletin Left Out

The bulletin describes what a regulator may ask for and then stops, which is where the enforcement gap lives.

It directs carriers to maintain a written AI program across five domains: governance and accountability, risk management and internal controls, third-party vendor oversight, consumer transparency and notice, and responsiveness to regulatory inquiries. It sets no format, no required fields, no threshold for a sufficiently documented program, and no definition of how a carrier demonstrates compliance rather than asserting it.

That ambiguity is deliberate in principles-based regulation. It lets programs scale to a carrier's size and keeps documentation standards from being frozen against moving technology. The cost is that a carrier can point to a written policy, a quarterly governance committee, and a vendor clause in its procurement templates without producing one data point on how any system performs, what trained it, or whether its outputs have been tested for disparate impact.

The nine components close that. They run from an executive summary establishing scope and geographic reach, through a board and senior management attestation, a models and data sources inventory covering internal training data and external purchased data with selection bias controls disclosed for each, a risk assessment framework classifying systems by tier, structured model cards per in-scope system, a governance narrative tracing the chain from developers to the board, a drift and validation section with monitoring frequency and retraining triggers, protected class inference and bias testing results by system, and consumer complaint disclosure for AI-influenced adverse decisions.

The attestation is the structural change. A senior officer signs, which creates a named chain from the program to a person, where the bulletin left an unsigned policy in a compliance manual.

What the Four Exhibits Actually Ask For

The evaluation tool is the enforcement counterpart, and it is already running inside examinations carriers face for other reasons.

The 12-state pilot runs March through September 2026, folded into market conduct exams, financial exams, and financial analysis rather than a separate track. So a personal auto market conduct exam triggered by complaint volume can pull Exhibit C requests for pricing and underwriting systems, and a life insurer's financial exam on reserve adequacy can pull Exhibits B and D.

Exhibit A quantifies the footprint, and the scope question is harder than it reads. Vendors embed machine learning inside underwriting workbenches, claims triage, fraud detection, and service tools, and few carriers track centrally which of those meet the NAIC definition. The tool reaches vendor-embedded models, automated decision components inside larger platforms, and machine learning features nobody in the organization files under AI. A carrier that believes it runs three vendor platforms and two internal models may hold twelve systems in scope.

Exhibit B separates governance infrastructure from governance theater. A policy describing an AI risk committee, plus quarterly minutes, does not show the committee assesses system risk. The evidence is the output: minutes citing specific performance metrics, escalation records for systems that breached thresholds, documented remediation.

Exhibit C is the demanding one, covering underwriting, pricing, claims handling, fraud detection, and any function where AI materially affects access to or cost of coverage. It asks for model design and architecture, training data composition and vintage, validation procedures and performance metrics, bias testing methodology and results by protected class, and sample case files showing how outputs shaped specific decisions.

That last requirement reaches the modeling pipeline, not the documentation shelf. Disparate impact testing on a gradient-boosted tree or a neural network requires output-based demographic analysis rather than variable-level review, which means building demographic inference and segmented performance reporting into the scoring pipeline itself. A carrier that bought a GLM pricing model in 2021 and a gradient-boosted underwriting score in 2023 owes Exhibit C for both, and only one of them yields to a coefficient review. Exhibit D then runs the same question at the data layer, with aerial imagery, social media, and purchasing behavior data singled out for explicit discrimination risk analysis separate from output-level testing.

The Documentation Is Held by Vendors and Made of Behavior

Two gaps separate what the report asks for from what most carriers can produce, and neither closes on a filing deadline.

The first is contractual. Carrier-vendor agreements signed before 2024 predate all of this and commonly grant no right to model documentation, bias testing results, training data composition, or demographically segmented performance data. When Exhibit C asks about a vendor pricing model, the carrier needs the vendor's material, and where the contract is silent the only route is renegotiation on a multi-quarter timeline. Pilot examinations have surfaced this repeatedly, with carriers citing the vendor as the reason they cannot produce, and pilot states responding that managing the vendor relationship is the carrier's obligation rather than grounds for deference.

The second is that most of the report cannot be written. The attestation, drift, governance narrative, and complaint sections all call for continuous evidence rather than a document. An AI risk policy takes a day to draft. Showing that the governance committee reviewed a model's drift metrics in March and documented a remediation decision in April takes months of changed behavior before the records exist to cite.

The timeline is where those two meet. Model inventories take 6 to 12 months to build accurately, and require actuarial, IT, underwriting, claims, and legal to engage at once because each function's AI is usually owned by a different team with no central steward. Vendor renegotiations run in quarters. Bias testing infrastructure is the deepest lift, because it changes the pipeline rather than the paperwork.

None of it waits on the Model Law question. Thirty-three comment letters on the NAIC's request for information exposed deep disagreement on scope and vendor liability, and any Model Law is years from adoption. The report needs neither: a state that has adopted the bulletin can ask a carrier to complete the form during a market conduct exam under authority it already holds.

Further Reading

Sources

  1. NAIC Big Data and Artificial Intelligence (H) Working Group Committee Page
  2. NAIC AI Systems Evaluation Tool Pilot Project Summary (PDF)
  3. Mayer Brown: NAIC Spring 2026 National Meeting Highlights, Innovation and Technology Committee (April 2026)
  4. Fenwick: NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States (2026)
  5. Foley & Lardner: What To Do If You Receive an NAIC AI Systems Evaluation Tool Pilot Request
  6. Swept AI: NAIC AI Evaluation Tool 12-State Pilot Is Live (2026)
  7. Quarles: Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers' Use of AI
  8. NAIC: Implementation of Model Bulletin State Map