Standard model risk management frameworks cannot validate large language models running in P&C claims, because those frameworks assume deterministic, repeatable outputs. The same FNOL narrative fed through the same model twice can yield two different coverage conclusions or reserve estimates.

That property has no analog in GLM or gradient-boosted tree validation, and the NAIC's actuarial panel begins probing it directly on July 22, 2026.

Key Takeaways

  • SR 26-2 explicitly carves generative and agentic AI out of scope. The April 2026 revision is the first wholesale update to SR 11-7 since 2011, and the regulator that wrote the standard says it was not built for this technology.
  • Three assumptions break at once. Determinism, measurability against labeled ground truth, and gradual degradation all hold for a GLM and none holds for an LLM reading an unstructured narrative.
  • A misread flood exclusion on one claim in four hundred produces no aggregate signal a quarterly stability report catches, while a GLM overpredicting loss ratios by 15 points trips a dashboard immediately.
  • 24% of 3,000 graded legal AI answers cited or misapplied law that did not support the claim, with every model tested fabricating or misapplying at least one citation.

Three Assumptions That Do Not Hold

Two decades of validation muscle around generalized linear models and gradient-boosted trees assumes three things. The model is deterministic, so identical inputs produce identical outputs. Performance is measurable against a labeled ground truth, so a loss ratio or fraud flag can be scored right or wrong. And degradation is gradual, so a model that scored well last quarter will not fail unpredictably this quarter absent a real shift in the data.

None of the three holds for a language model processing an unstructured claims narrative. Run the identical FNOL description through the identical model twice at any temperature above zero and the outputs can diverge, sometimes only in phrasing, sometimes in the coverage determination or reserve figure recommended. Ask the same coverage question with different word order and the conclusion can shift, a context-dependence with no equivalent in a scorecard model where feature order is irrelevant by construction.

The third break is the hardest to instrument. A GLM that misfires produces an implausible number an actuary flags on inspection. An LLM that misfires produces a confident, well-formatted paragraph reading exactly like a correct one, with no signal in the output distinguishing the two.

Most P&C carriers built their governance around the Federal Reserve's SR 11-7, a banking supervisory letter adopted informally across insurance as the benchmark since its 2011 issuance. In April 2026 the Fed, the OCC and the FDIC issued SR 26-2, the first wholesale revision in over a decade, and it carves generative and agentic AI out of formal scope on the basis that the technology is too novel for the existing validation architecture.

MRM Assumption GLM / Gradient-Boosted Tree Large Language Model
Deterministic output Yes; identical inputs always reproduce identical outputs No; stochastic sampling produces different outputs from identical prompts
Measurable against ground truth Yes; loss ratio, fraud flag, or claim outcome scores as right or wrong Partial; coverage conclusions and reserve narratives often lack a single correct answer to score against
Stable over time absent data shift Yes; performance degrades predictably as population drifts No; a prompt-template change, vendor model update, or fine-tune can silently shift behavior with no data drift at all
Interpretable failure mode Yes; an implausible score is visible on inspection No; a hallucinated but fluent answer is often indistinguishable from a correct one without independent verification

Where the Gap Shows Up in an Exam

Carriers running LLMs in production claims workflows report three recurring failure patterns, and they share a structural feature that matters more than any one of them.

The first is coverage misinterpretation at intake, where a model summarizing an FNOL narrative drops or misstates an exclusion, particularly on claims with overlapping coverages or endorsements requiring the declarations page to be read against the narrative rather than pattern-matched on keywords. The second is inconsistent reserve estimation across narratives describing functionally similar losses, where phrasing alone produces materially different initial recommendations. The third is regulatory language drift, where a model trained on a national corpus applies a coverage standard correct in one state's case law and wrong in another's.

None of them trips a conventional performance-threshold alarm, because there is no stable baseline threshold to trip against. A GLM overpredicting loss ratios by 15 points sets off a monitoring dashboard. A model misreading a flood exclusion on one claim in four hundred, correctly on the rest, produces no aggregate signal a quarterly stability report would catch, and each individual miss is a live coverage decision with a policyholder on the other end of it.

The NAIC's AI Systems Evaluation Tool pilot, running across 12 states from March through September 2026, was not written with language models in mind, but Exhibit C asks for testing evidence and human-in-the-loop protocols for high-risk models. Most carriers can produce that for a triage scoring model built on gradient-boosted trees. Far fewer can produce an equivalent for a model summarizing narratives, because the underlying validation science is not standardized the way GLM backtesting is. The NAIC's March 2026 Issue Brief states that "existing state insurance laws apply regardless of whether" a decision involves a human or an algorithm, so the absence of an LLM-specific standard is not a defense.

Three techniques address the stochastic properties directly rather than forcing an LLM through a deterministic process. Ensemble sampling runs the identical prompt dozens of times and measures variance: a determination returning the same way ninety-eight times out of a hundred is validated differently than one splitting sixty-forty, and that variance statistic is what a validator certifies.

Boundary condition testing feeds deliberately ambiguous coverage language, concurrent-causation losses, stacked endorsements, conflicting policy wording, and checks whether the model's confidence matches its accuracy. Semantic drift monitoring compares output content against a fixed baseline, flagging a shift in how a standard coverage question is interpreted even when no traditional data drift has occurred.

None is a drop-in replacement for backtesting, and none produces the single clean number a GLM sign-off expected. Validating a language model means certifying a distribution and a confidence interval, and most carrier governance documentation still asks for a point estimate.

The Reserving Data Sits Downstream of an Uncharacterized Process

The validation gap has a governance twin. When a technology team builds a claims tool and a claims operations team deploys it, validation responsibility falls between them: technology validates uptime and latency rather than coverage accuracy, and claims lacks the statistical background to design a variance test.

Three artifacts are usually missing as a result. There is no documented output-variance baseline, so nobody has recorded what normal variance looks like for that model and prompt template and nothing exists to compare a suspicious result against. There is no named, credentialed owner for validation as distinct from the technical owner or the deploying department. And there is no semantic drift monitoring running continuously, which means a vendor's silent model update, common given how often providers ship new versions, can change claims behavior with no internal alert until complaints or an examiner's sample review surface it.

Scale is growing faster than the instrumentation. Full-scale AI adoption across the industry grew from 8% to 34% of carriers between 2024 and 2025, and Travelers' early-2026 agentic voice service for auto damage FNOL now makes more than half of eligible claims candidates for fully automated processing.

The adjacent legal profession, whose document-review architecture claims tools often inherit, has the visible failure record. Court sanctions for AI-generated fabrications reached at least $145,000 in the first quarter of 2026, and a study grading 3,000 legal AI answers found 24% cited or misapplied law that did not support the claim.

The reserving consequence follows from the data rather than the model. Any development pattern built on claims where a language model influenced the intake narrative, the coverage read or the initial reserve figure is drawn from a data-generating process whose error behavior has not been statistically characterized. An actuary selecting loss development factors from a period of LLM-influenced settlement activity is selecting from a source the rest of the reserving methodology assumes has been validated, and has not been.

Further Reading

Sources

  1. NAIC Big Data and Artificial Intelligence (H) Working Group
  2. NAIC March 2026 AI Issue Brief
  3. Fenwick: NAIC Expands AI Systems Evaluation Tool Pilot to 12 States
  4. NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers (December 2023)
  5. Crowell & Moring: NAIC Intensifies AI Regulatory Focus
  6. WaterStreet Company: What the NAIC Model Bulletin Means for Insurance AI
  7. Bespoke Mentis: SR 11-7 Guidance Revisited: AI Model Risk in 2026
  8. Stanford RegLab: Hallucinating Law: Legal Mistakes With Large Language Models Are Pervasive