At the Spring 2026 National Meeting in San Diego on March 24, the NAIC's Big Data and Artificial Intelligence (H) Working Group gave a full panel to agentic AI. Three risk categories came out of it: cascading errors across autonomous decision chains, accountability gaps where no single person oversees the workflow, and failure modes that model validation does not test for.

The reason this is a break rather than an extension is in the existing text. The December 2023 Model Bulletin requires insurers to designate "a person or persons who are responsible for the AI system," a sentence drafted for one model with one owner.

Key Takeaways

  • The Model Bulletin is adopted by 24 states and the District of Columbia as of the Spring meeting, with four states layering additional insurance-specific AI rules on top, all of it built around defined inputs, identifiable training data and single-model ownership.
  • The 12-state evaluation tool pilot launched March 2, 2026 and its Exhibit C, covering high-risk models with automated decision-making, was designed before agentic AI was named as a distinct category.
  • AIG's Lloyd's Syndicate 2479 went live January 1, 2026 with LLM agents on a delegated authority portfolio managing $300 million of premium, while Lexington processed over 370,000 E&S submissions in 2025 against a 500,000 target by 2030.
  • Travelers launched an agentic voice claim assistant in February 2026, three years and two months after the NAIC adopted the AI Principles that preceded the bulletin now being outgrown.
  • August 2020 to December 2023 was the NAIC's own interval from AI Principles to Model Bulletin, which puts specific agentic guidance no earlier than 2028 on the same cadence.

What the Panel Named

The three categories the working group set out are each a failure the existing framework was not built to catch, rather than a harder version of one it was.

Cascading errors come first. A predictive model produces one output that a human evaluates. An agent chains reasoning steps, calls tools and takes actions, so an error in an early step propagates through everything downstream before any human sees a result.

Accountability is the second, and it is the one with existing text pointing at it. The bulletin's designated-person requirement assumes one model, one owner, one chain. An underwriting agent drawing on a claims model, a pricing algorithm and external data in a single workflow produces a composite output, and the panel noted that no current guidance says which designated person owns it.

Performance limitations are the third. Validation against holdout data, discrimination metrics and stability testing will not surface an agent that performs well on routine work and fails unpredictably at the edge of its training distribution, or one optimizing a proxy rather than the business objective. The panel's word was redesign, not extension.

NAIC staff also presented a four-level risk taxonomy that mirrors the EU AI Act's tiered structure.

Risk Level Description Examples in Insurance
Unacceptable Subliminal manipulation, social scoring Systems that exploit behavioral biases to increase premium acceptance
High Potential for significant harm if failure occurs Automated claims denial, underwriting triage with coverage impact
Medium Requires transparency; chatbots, emotion recognition Customer service AI, sentiment analysis in claims calls
Low Minimal restrictions Spam filters, internal document search, scheduling tools

The Gap Between the Exhibit and the Deployment

The concrete version of the problem is that the instrument regulators are piloting right now assumes the architecture the panel said no longer holds.

The AI Systems Evaluation Tool pilot launched March 2, 2026 across California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia and Wisconsin, and runs to September. Exhibit A quantifies AI usage, Exhibit B assesses governance, Exhibit C gathers detail on high-risk models with automated decision-making, and Exhibit D documents data sources and vendor relationships. Exhibit C is the closest fit for an agentic system and still asks for one model, one decision, one documented input set. A workflow chaining three or four models with tool-calling and branching logic does not resolve into a single response.

What sits on the other side of that gap is already in production. Through Lloyd's Syndicate 2479, launched January 1, 2026, AIG deployed LLM agents against a delegated authority portfolio managing $300 million in premium, using Palantir Foundry ontologies that map entities, risks and relationships, coordinated by an orchestration layer across the enterprise. Lexington processed over 370,000 E&S submissions in 2025 against a 500,000 target by 2030. Travelers, having committed to AI assistants for nearly 10,000 staff, launched an agentic voice claim assistant in February 2026 that consults policies, guides filing decisions and escalates to live agents.

Both carriers name human escalation as the safeguard, and neither has disclosed how the triggers are calibrated or what false-negative rate on escalation they accept. That figure is the whole safeguard. Escalate too rarely and harmful decisions pass unchecked; escalate too often and the system is a conventional workflow with suggestions attached. For anyone signing off on a rate indication or reserve estimate downstream of one of these chains, the escalation false-negative rate is the number that determines whether the human oversight in the governance document exists in the running system, and it is not currently in any exhibit.

The proposed vendor registry, narrowed at the same meeting to pricing and underwriting, has the same shape of limitation. AIG's stack spans Palantir, multiple LLM providers and its own ontology construction; Travelers' spans Anthropic, OpenAI and internal engineering. Cataloging which vendor supplied which model does not capture how they are orchestrated or where autonomy sits, and the NAIC was explicit that the registry "is not intended to relieve insurers of their existing vendor diligence and management obligations."

The Clock Runs at Two Speeds

The constraint is that the interval between recognizing a gap and writing guidance for it is fixed by process, and deployment is not.

The NAIC adopted its AI Principles in August 2020. The Model Bulletin followed in December 2023, more than three years later. The evaluation tool was developed through 2024 and 2025 and reached pilot in March 2026. On that cadence, specific agentic guidance arrives no earlier than 2028, by which point the question is not whether carriers have deployed agentic systems but how deeply embedded they are.

The nearer deadline is tighter still. The pilot closes in September 2026, the tool is updated on pilot feedback through September and October and re-exposed for comment, with formal adoption expected at the Fall 2026 National Meeting. Anything the Spring panel surfaced has to reach the permanent instrument through that window, and the carriers best placed to document where Exhibit C falls short are the pilot participants running agentic systems.

The burden of documenting it is not evenly distributed. Panelists raised scope definition difficulties and the resource disparity between large and small insurers: a carrier with a dedicated AI governance team can assemble full Exhibit B and C responses, while a regional mutual running one vendor-supplied tool may not have the staff, and orchestration documentation widens that difference rather than narrowing it.

The state layer does not close the gap either. Colorado's SB 21-169 tests for algorithmic discrimination, with auto and health insurers facing July 1, 2026 reporting deadlines and life insurers in scope since 2023. The NCOIL model act would require qualified human professionals to make final claims decisions. Neither reaches the intermediate steps of a chain, which is precisely where the panel located the cascading-error risk.

Further Reading

Sources