A state rate filing is a point-in-time contract over a model's behavior: the regulator approves the specification submitted, not the class of algorithms it belongs to. A P&C carrier retraining an ML pricing model on a rolling window can therefore be writing new business under relativities no regulator has reviewed. The NAIC's AI Systems Evaluation Tool pilot, live in 12 states since March 2, 2026, is the first instrument built to test whether carriers can document that gap.

Key Takeaways

  • 12 states, from March 2 through September 2026, with revisions in September and October and formal adoption expected at the November 2026 NAIC fall meeting. The group includes California, Florida, Pennsylvania and Virginia.
  • Exhibit C asks for version history, not just the model as deployed: design documentation, training data description, validation procedures, performance metrics and bias testing for high-risk systems, which explicitly include pricing.
  • Materiality is usually defined as a 1% to 3% aggregate rate level move, a standard that a retrain can pass while shifting the relativities charged to identifiable risk segments.
  • PSI above 0.20 on a key rating variable warrants investigation and above 0.25 is a conventional revalidation trigger, but the discipline that matters is writing the threshold down before the retrain, not after.
  • 24 states plus the District of Columbia had adopted the NAIC Model Bulletin by early 2026, so the written governance program the tool tests already exists as an obligation in half the country.

The Version Lock and Where It Breaks

Filed GLMs and traditional actuarial schedules change only when a carrier deliberately amends them, so the approved artifact and the deployed artifact stay identical between filings. Models trained on rolling data windows do not work that way. The model approved in January produces different relativities in July once the training set moves six months.

Take the mechanics. A personal auto carrier files a gradient-boosted pricing model trained on 36 months of loss data ending December 2024, and the regulator approves it in April 2025. The carrier retrains in July 2025 on data through June 2025, a window carrying a sharp inflation spike in replacement parts and a shift in driving patterns. The retrained model assigns materially different relativities to mileage bands and vehicle age segments. No amendment follows, because the internal classification of the retrain is scheduled maintenance within operating parameters.

This is distinct from model underperformance, where the model degrades at its approved version, and from intentional change, where an amendment is filed. It is version control sitting inside the rate-setting workflow, and it became structurally more common as carriers moved from annual model development to quarterly or continuous retrain pipelines between 2021 and 2025.

The Materiality Test the Filing System Uses Is the Wrong One

The default framework in most prior-approval states defines materiality by aggregate rate level: a change beyond roughly 1% to 3% requires a new filing. Applied to an ML pricing model, that test is looking at the wrong number.

A retrain can leave the aggregate rate level flat and still move the relativities charged to identifiable segments. Those segment shifts are changes to the rate classification system whether or not the aggregate premium impact rounds to zero, which is why the actuarial judgment has to compare pre- and post-retrain outputs across the full rating universe rather than the premium total.

That comparison needs metrics with thresholds attached.

Metric Drift Type Detected Conventional Alert Threshold Typical Action
Gini coefficient change Discrimination power loss >3 percentage points vs. validation period Revalidation; materiality review
Population Stability Index Covariate / input distribution shift PSI > 0.20 on any key rating variable Investigate; PSI > 0.25 triggers revalidation
Prediction error by segment Systematic over/underpricing in risk classes Segment loss ratio deviation > 10 points Segment-level materiality analysis; potential amended filing
Lift curve degradation Rank-order performance Top-decile lift drops > 15% vs. development set Retraining trigger; version log entry
K-S statistic Predicted probability distribution shape K-S drop > 0.05 vs. baseline Supplemental validation; escalation review

Gini and lift curves read discrimination power, the K-S statistic reads the shape of the predicted probability distribution, and the Population Stability Index reads whether the input distribution has moved from the training period. PSI above 0.20 on a key rating variable warrants investigation; above 0.25 is the conventional revalidation trigger. The thresholds themselves matter less than fixing them in writing at deployment, because a threshold set after the retrain is a rationalization.

The sign-off is the part automation cannot absorb. A named pricing actuary has to review the materiality analysis, sign the version log entry, and either approve continued production use or escalate for a filing determination, and at carriers where the retrain is automated that review has to gate promotion to production. "Version history matters. Regulators will want to see that governance evolved as your AI use did," Foley & Lardner told carriers receiving pilot requests.

The Drift That Never Triggers a Log Entry

A version log records retrain events. The harder exposure produces no event at all.

Covariate shift changes the statistical distribution of the input variables without any retraining trigger. A homeowners model using aerial imagery features can move when post-wildfire reconstruction patterns change the feature distribution across Western states. A personal auto model on telematics features can move when driving behavior changes after an economic shock. The outputs change, the carrier may not notice until renewal season, and nothing in the change management process fires, because nothing was changed.

That is the gap a version log cannot close on its own, and it is why the monitoring side of the obligation is the substantive half. EIOPA's August 2025 AI Governance Opinion made it explicit for European carriers, requiring performance metrics to detect model drift or data degradation and endorsing SHAP and LIME for identifying which features drive prediction shifts. EIOPA does not reach US carriers, but the pilot is asking the same question with a different instrument.

The pilot response is also not a one-time exercise. "A carrier that produces a weak Exhibit B governance narrative in the pilot now has that narrative on file with a state regulator," Swept AI noted. It becomes the baseline the next market conduct examination starts from, in 2027 or 2028, in a pilot group holding four of the largest insurance markets in the country.

The exposure concentrates where the models do. State Farm, USAA and Allstate together account for roughly 77% of AI patents filed by P&C insurers, so the carriers running the most production models through the most retrain cycles are the ones setting the benchmark the rest of the industry gets measured against.

Further Reading

Sources

  1. NAIC, AI Systems Evaluation Tool Pilot: Pilot Project Summary (March 2026) — Pilot launch date, participating states, timeline, and evaluation dimensions.
  2. Fenwick, “NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States” (March 2026) — Pilot scope, carrier obligations, and adoption timeline for the November 2026 NAIC fall meeting.
  3. Foley & Lardner, “What To Do If You Receive an NAIC AI Systems Evaluation Tool Pilot Request” (2026) — Carrier response strategy and the version-history documentation baseline implication.
  4. Quarles, “Nearly Half of States Have Now Adopted NAIC Model Bulletin on Insurers’ Use of AI” (March 2026) — State adoption count and the written AI program requirement for model change management.
  5. EIOPA, Opinion on AI Governance and Risk Management (EIOPA-BoS-25-360, August 6, 2025) — Performance metrics for drift detection, SHAP/LIME endorsement, and actuarial function accountability for AI controls.
  6. Swept AI, “NAIC AI Systems Evaluation Tool: 12-State Pilot Is Live” (March 2026) — Pilot structure, exhibit breakdown, and the governance narrative baseline risk.
  7. Monitaur, “NAIC AI Systems Evaluation Tool Pilot: A Guide for Insurers” (2026) — Exhibit C documentation requirements for high-risk AI systems, including pricing models.
  8. InsNerds / Insurance Business, “AI Patent Trends by State Farm, USAA, and Allstate Signal Strategic Innovation for P/C Insurers” (2025) — The 77% patent concentration among three carriers and its implications for industry-wide AI governance standards.
  9. Allstate, Machine Learning Monitoring Patent (June 2026), as analyzed on actuary.info — Drift-detection, alerting, and retraining control architecture on the claims side as an operational template for pricing governance.