Edin Imsirovic of AM Best told the NAIC's Big Data and Artificial Intelligence (H) Working Group on August 13 that AI creates no separate risk category for an insurer. The exposures already flow through product and underwriting risk, operational risk, and regulatory and legal risk, and the enterprise risk management question stays whether risk management capability is commensurate with the risk profile (NAIC, August 2026).
He gave the 39 jurisdictions seated in Columbus the routing rule in a single line: where AI affects pricing, risk selection, segmentation or portfolio mix, that points to product and underwriting risk. Later in the session, answering Commissioner Doug Ommen of Iowa, he said AM Best has no official process for gathering data on AI governance, and that most of what it hears arrives in conversation.
Key Takeaways
- Three existing categories, product and underwriting risk, operational risk, and regulatory and legal risk, absorb every AI exposure under AM Best's framing, so a claims tool that corrupts payment or reserve information surfaces as a financial consequence.
- +1 to -4 notches is the enterprise risk management adjustment range in Best's Credit Rating Methodology, set against a combined ceiling of +2 across operating performance, business profile and ERM together.
- 177,436 agent tools reviewed by the UK AI Security Institute with the Bank of England showed action tools climbing from 24% to 65% of monthly downloads, and payment-execution servers going from 46 to more than 1,200 in a year.
- 59% better performance came from raising a frontier model's compute budget from 10 million to 100 million tokens on multistep attack scenarios, with no retraining, which is a capability change no model-keyed revalidation trigger would register.
- 73%, 44% and 33% are the shares of reported cloud, model and data providers held by the top three vendors in each class, the concentration behind the shared-dependency concern Imsirovic raised.
Where AM Best Routes an AI Failure
The presentation carried a disclaimer worth reading literally: it was for discussion, and it changes neither AM Best's criteria and methodology nor its rating guidance (NAIC, August 2026). What it assigned was destinations, which an analyst can act on without new criteria.
Pricing, risk selection, segmentation and portfolio mix go to product and underwriting risk. Data quality, claims-handling errors, cyber exposure, workflow breakdowns and weak oversight of third-party models fall under operational risk and consumer protection. Explainability, litigation and regulatory scrutiny run through the legislative, regulatory, judicial and economic categories.
Imsirovic then walked one use case across all three. A claims AI tool begins as an operational issue, becomes a regulatory issue if consumers are treated inconsistently, and creates financial consequences if its errors reach payments or reserve information. A valuation actuary inherits that last clause, which appears nowhere in the AI governance literature written for boards.
Three working terms sit behind the routing. Predictive AI scores, ranks, classifies or flags. Generative AI produces output that can sound confident when it is wrong. Agentic AI works through a series of steps, uses tools and updates systems, affecting processes directly.
Each creates a different governance problem, and the risks stack: an agentic system still carries the predictive and generative ones underneath its own permission and action risks.
Five factors then set how much governance a deployment needs: the importance of the decision, how much the system can do on its own, explainability, speed of change, and third-party dependence. Proportionality underlies all five, which he tied to the NAIC Model Bulletin and to EIOPA's risk-based approach. The label indicates the type of system; the use case indicates the level of risk.
What the Routing Rule Costs in Notches
The destination decides the size of the consequence, because the building blocks in Best's Credit Rating Methodology do not carry symmetric weight.
| Rating building block | Adjustment to baseline |
|---|---|
| Balance sheet strength | Baseline |
| Operating performance | +2 / -3 notches |
| Business profile | +2 / -2 notches |
| Enterprise risk management | +1 / -4 notches |
| Comprehensive adjustment | +1 / -1 notch |
Source: AM Best, Best's Credit Rating Methodology. Combined lift across the three adjusting blocks is capped at +2.
Read that column as an option payoff and the asymmetry is stark. An excellent AI governance answer is worth at most one notch of lift, and the +2 combined cap can absorb even that. A governance failure is worth four notches of drag. No amount of documented control quality buys back what a single uncontrolled deployment can cost.
The routing rule then does something the ERM range alone does not capture. AM Best's baseline balance sheet assessment names adequacy of reserves as an explicit assessment factor, alongside BCAR, reinsurance quality, liquidity, quality of capital and internal economic capital models. A claims tool whose errors reach reserve information is not adjusting the rating from the ERM block; it is moving the baseline every adjustment starts from.
Business profile widens it again: its named review factors include pricing sophistication and data quality, so an underwriting model the carrier cannot inspect can register in two blocks before ERM is scored.
The capability figure Imsirovic cited is where model validation practice breaks. Testing frontier models on multistep cyberattack scenarios, the AI Security Institute raised the compute budget from 10 million to 100 million tokens and measured a 59% performance improvement with no retraining, no additional fine-tuning and no special expertise (AISI, 2026).
A carrier can hold the same approved model, grant it more time, more attempts, better tools or wider permissions, and materially change its risk. Revalidation triggers written against changes to a model or its data will not fire on any of that. Approving the model, as Imsirovic put it to the working group, is separable from approving the deployment.
The Evidence Sits Behind Exam Confidentiality
Imsirovic named four things a company should show beyond a policy: an AI inventory, monitoring evidence covering what happened when a trigger was hit, a case in which a person challenged a result or stopped a process, and proof it can see a vendor change and return to a safe version.
By his own account AM Best does not systematically collect any of it. Companies typically tell him they do not use AI in underwriting, or that they monitor it if they do. The most consistent answer he receives to a governance question is that a human is in the loop, and it comes through conversation, because AM Best has no official process for gathering AI governance data.
The written record exists; it is being created three feet away. Version 5.0 of the NAIC's AI Risk Evaluation Supplement, exposed August 31 with comments due September 29, asks for a model inventory carrying each model's use case, inherent risk level, consumer impact and financial impact, plus new checklist items on third-party model oversight and on explainability. Twelve states are running the pilot, among them California, Florida, Iowa, Pennsylvania and Wisconsin.
The pilot's own project summary closes that channel to everyone else. Requested information "will be protected under the confidentiality rules of the state administering the exam" (NAIC, 2026). The documentation that would populate an ERM assessment is legally sealed inside the examination that produced it.
Shared dependency is the exposure that gap hides best. Imsirovic raised it alongside the Financial Stability Board's work on third-party dependency and provider concentration, where a Bank of England survey found the top three providers accounting for 73% of reported cloud, 44% of model and 33% of data providers, with one third of all AI use cases running on third-party implementations (FSB, October 2025).
The AISI count of payment-capable agent servers across that same ecosystem went from 46 to more than 1,200 in a year, and finance is where the tools cluster: "high-stakes occupations including financial services agents, accountants, and financial managers, have more action tools than predicted by the overall pattern" (AISI, March 2026). A rating file assembled one balance sheet at a time reads that concentration as twelve hundred separate company decisions.
Further Reading
- NAIC AI Risk Evaluation Supplement v5.0 Makes the Model Inventory an Examination Exhibit
- NAIC Flags Agentic AI as Insurance's Next Governance Gap
- Guidewire's Qusar Puts Agents Inside Core Underwriting and Claims Data
- Why Standard Model Risk Management Cannot Validate LLMs in P&C Claims
- The Vendor Registry Model Law and Third-Party AI in Insurance
- NAIC's July 22 Actuarial Panel Maps AI Governance Duties
- The AI Patent Race in Insurance: How Carriers Are Staking Competing IP Claims
Sources
- NAIC: Big Data and Artificial Intelligence (H) Working Group, minutes of the August 13, 2026 meeting, Columbus, Ohio (drafted August 23, 2026)
- NAIC Big Data and Artificial Intelligence (H) Working Group: charges, exposures and comment instructions
- NAIC: AI Systems Evaluation Tool Pilot, Pilot Project Background and participating states
- NAIC: AI Risk Evaluation Supplement, Summary of Changes from version 4.0 (exposed August 31, 2026)
- AM Best: Best's Credit Rating Methodology, an overview of the rating process and building blocks
- UK AI Security Institute: How are AI agents used? Evidence from 177,000 AI agent tools (March 26, 2026)
- UK AI Security Institute: Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios
- Financial Stability Board: Monitoring Adoption of Artificial Intelligence and Related Vulnerabilities in the Financial Sector (October 2025)
- NAIC Insurance Topics: Artificial Intelligence