At the Spring 2026 National Meeting on March 24, NAIC staff put a four-tier AI risk taxonomy in front of the Big Data and Artificial Intelligence (H) Working Group, with operational compliance expectations attached to each tier.

That is a change of kind. Everything from the 2020 AI Principles through the December 2023 Model Bulletin was principles-based. The taxonomy is prescriptive, and it arrives alongside a 12-state evaluation pilot and a live request for information on a model law.

Key Takeaways

  • Four tiers, classified by what the AI decides rather than what line it sits in: unacceptable, high, medium and low, with examination resource concentrated on systems that make or materially influence coverage, pricing and claims decisions.
  • The compliance report runs to five sections covering management oversight, data source documentation, drift and validation, protected class inference and bias testing, and consumer complaint procedures.
  • A mid-size carrier running 15 to 25 AI systems has to inventory every one, classify each by tier, and produce a full model card for every high-risk application.
  • A P&C pricing model is high-risk under the NAIC taxonomy and may not fall under the EU AI Act's Annex III provision, which names risk assessment and pricing for natural persons in life and health.
  • The Model Bulletin is guidance, not law, adopted by 24 states and the District of Columbia; the taxonomy gains statutory force only if a model law codifies it.

The Four Tiers

The taxonomy's organizing choice is that a system's tier follows the consequence of its output, not the technology behind it.

Risk Tier NAIC Definition Insurance Use Cases Regulatory Treatment
Unacceptable AI systems using subliminal manipulation or general social scoring Systems exploiting behavioral biases to increase premium acceptance; social-credit-style scoring that conditions coverage on non-insurance behavior Prohibited outright
High Potential for significant harm if failure or misuse occurs Automated underwriting triage with binding authority; claims denial algorithms; pricing models that set final rates without human review; fraud detection systems that trigger coverage rescission Full compliance report, AI model card, bias testing, continuous monitoring, human-in-the-loop requirements
Medium Requires transparency and user disclosure of AI interaction Customer-facing chatbots; sentiment analysis on claims calls; emotion recognition in recorded interactions; AI-assisted (not AI-decided) claims triage Transparency disclosure, periodic review, consumer complaint monitoring
Low Minimal restrictions; deployable without enhanced oversight Spam filters; internal document search; scheduling and workflow automation; data extraction from policy forms Inventory tracking only

The risk taxonomy draws its main line between systems that make or materially influence coverage, pricing and claims decisions and those that support internal operations without touching policyholders. The medium tier holds a growing category of customer-facing tools that do not make final decisions but shape the consumer experience enough to require transparency.

Applied to live deployments the classification is mostly obvious and occasionally not. AIG's Palantir Foundry underwriting platform, processing over 370,000 excess and surplus submissions annually toward a 500,000 target for 2030 and running LLM agents across a $300 million delegated authority portfolio at Lloyd's Syndicate 2479, is high-risk wherever it influences binding or pricing. Internal document search, email routing and scheduling are low-risk and need only an inventory entry.

The agentic claim assistant Travelers launched with OpenAI in February 2026 sits on the boundary. It interacts directly with consumers, which is the medium-tier transparency trigger, and it characterizes damage and cites policy provisions in real time, which is high-tier decision impact.

What a Tier Costs

The taxonomy's practical content is the documentation each tier requires, and the step from medium to high is where the cost is.

The compliance report is the carrier-level document, and per the Mayer Brown summary it covers five areas: management oversight including the full organizational chart of AI governance rather than the bulletin's single designated person, internal and external data source documentation with provenance and representativeness, model drift monitoring with explicit thresholds and revalidation triggers, protected class inference and bias testing with methodology and remediation, and consumer complaint procedures for challenging an AI-influenced decision.

The model card is the system-level document, adapting the concept from Mitchell et al. (2019) for regulatory reporting. For a high-risk system it carries system identification and responsible person, intended use and decision volume, training data summary and representativeness, performance metrics benchmarked to the use case, disparate impact results across protected classes, monitoring cadence and drift triggers, third-party components, and known failure modes.

The arithmetic follows directly from the classification. A regional P&C carrier writing $800 million across six states, three of them in the pilot, running a homeowners pricing model, a claims fraud detection system, a customer service chatbot and an internal document processing tool ends up with two high-risk systems, one medium and one low. That is two full model cards, five compliance report sections with detailed treatment of both high-risk systems, a transparency section for the chatbot, and one inventory line. A carrier running 15 to 25 systems scales that accordingly.

The vendor documentation is where the estimate usually breaks. A vendor-supplied fraud detection model needs training data summary, performance metrics and bias testing results from the vendor, and many carriers have never contracted for the audit rights that would produce them.

The evaluation tool pilot is already operating on this logic across 12 states from March 2, 2026 through September, with weekly regulator coordination calls and information requests to selected carriers. As one analysis of the pilot notes, regulators concentrate on "high-risk AI systems that could cause serious consumer or financial issues, while paying less attention to low-risk back-office systems."

Who Decides the Tier

The complication is that the framework assigning these obligations has no statutory force, and the classification that triggers them is made by the party paying for them.

The Model Bulletin is guidance. It gives state departments a framework to issue, with no independent enforcement mechanism beyond whatever authority existing statutes supply. The taxonomy gains teeth only through codification, and the Working Group has exposed a request for information on a NAIC Model Law on the Use of Artificial Intelligence in the Insurance Industry with a 45-day comment period. The 33 comment letters on the earlier RFI showed the fault lines: scope, vendor liability, and company-size thresholds.

The calendar compounds it. Earliest exposure of a model law draft is the Summer 2026 National Meeting, with adoption no sooner than Spring 2027, and state adoption on individual timelines after that. The Model Bulletin itself took roughly 14 months from adoption to reach 24 states.

Until then the tier assignment is self-classification against definitions with no binding text, and the boundary cases are exactly the systems with the most at stake. Travelers' claim assistant classified as medium carries a transparency disclosure; classified as high it carries a model card with bias testing and drift thresholds. Nothing currently settles which, and the incentive on the classifying party points one way.

The jurisdictional mismatch adds a second unsettled edge for global carriers. The EU AI Act names "AI systems intended to be used for risk assessment and pricing in relation to natural persons in the case of life and health insurance" as high-risk under Annex III. The NAIC taxonomy classifies by the nature of the decision across all lines, so a P&C pricing model is unambiguously high-risk in one regime and outside the named provision in the other, on the same documentation stack.

Further Reading

Sources