AIG chief executive Peter Zaffino put a number on AI governance during the carrier's Q1 2026 earnings call on May 1. A professional claims adjuster reviewed 100 claims for fraud. Anthropic's Claude reviewed the same 100 independently. The two agreed 88% of the time. Most carriers describe AI deployments as "promising results" and "significant efficiency gains." An agreement rate is the rarer thing: a governance figure a regulator, a board member and a reserving actuary can all read without translation.
Key Takeaways
- 88% agreement between Claude and an AIG claims adjuster across 100 fraud determinations, from an out-of-the-box model with no claim-specific tuning. Tuning should move it, which makes 88% a floor rather than a result.
- Plus or minus 6.4 points is the 95% confidence interval on 100 binary observations, so the true rate sits somewhere between roughly 82% and 94%. Reaching plus or minus 2 points takes 1,000 claims.
- 12 claims out of 100 is where the governance question lives. Whether the model over-flagged or under-flagged decides whether the cost lands in claims expense or in incurred losses.
- 12 states are running the NAIC AI Systems Evaluation Tool pilot through September 2026, and its Exhibit C asks carriers to describe performance metrics for high-risk systems without prescribing which.
- 42% of insurers track no AI metrics at all, on Grant Thornton's 2026 survey, so the standardization problem starts well before anyone argues about which metric is best.
What the 88% Was Measured On
Zaffino described the evaluation with unusual specificity for an earnings call. One professional claims adjuster ranked 100 claims as fraudulent or legitimate with documented reasoning. Claude assessed the same 100 with no access to the adjuster's work. The two sets of determinations were then compared.
The load-bearing phrase in his description is "out-of-the-box." The model received no fine-tuning on AIG's proprietary claims data. The 88% is general reasoning applied to insurance fraud indicators: timeline inconsistencies, geolocation mismatches, prior claim patterns, document tampering signals.
Sample size is the first thing to check. For a binary classification on 100 observations, the 95% confidence interval runs roughly plus or minus 6.4 percentage points, putting the true agreement rate between about 82% and 94%. A 1,000-claim sample narrows that to plus or minus 2 points, and 10,000 brings it under 1 point. Reported at that spread, the figure is a direction rather than a measurement.
The rest of what Zaffino disclosed measures something else entirely. AIG Assist in Lexington middle market property produced a 30% improvement in quoting volume, a 55% reduction in time-to-quote and roughly 40% more submissions bound. Those are throughput figures. Only the 88% addresses whether the machine reaches the conclusion a qualified human would.
Why Agreement Beats F1 as a Reported Metric
This section shows why the metric that reads as least rigorous to a data scientist is the one most likely to survive into regulatory reporting.
Accuracy, precision, recall and F1 all score a model against a labeled ground truth. They answer whether the model is right. Claims adjudication rarely offers that luxury: two experienced adjusters can read the same file and disagree in good faith, so the ground truth is itself a judgment call. An agreement rate asks the answerable question instead, and that reframing is why it travels across audiences F1 never reaches.
| Metric | What It Measures | Audience | Governance Utility |
|---|---|---|---|
| Accuracy | Correct predictions / total predictions | Data scientists | Low: misleading with imbalanced classes (fraud rates of 5-10%) |
| Precision | True positives / predicted positives | Data scientists, model ops | Medium: meaningful for false-positive costs but requires technical context |
| Recall | True positives / actual positives | Data scientists, compliance | Medium: captures missed fraud but needs pairing with precision |
| F1 Score | Harmonic mean of precision and recall | Data scientists | Low: composite metric that obscures the tradeoffs boards need to see |
| Raw Agreement Rate | Concordance between AI and expert human | All three audiences | High: intuitive, comparable across functions, audit-friendly |
| Cohen's Kappa | Agreement adjusted for chance concordance | Statisticians, actuaries | High: corrects for base-rate inflation in agreement percentages |
The correction an actuary should insist on is Cohen's Kappa. Raw agreement inflates with the base rate: if 90% of claims are legitimate, a model that calls everything legitimate scores 90% against any adjuster working the same mix. Kappa nets out chance concordance. On Cohen's scale, 0.61 to 0.80 is substantial agreement and above 0.81 is almost perfect. At a 20% fraud prevalence, typical for a flagged-for-review pool, 88% raw would correspond to a Kappa of roughly 0.65 to 0.75. The prevalence in AIG's sample was not disclosed.
The financial content of the metric sits in the 12 claims where the two disagreed, and the direction decides the sign. Claims the model flags and the adjuster clears drive investigation expense and customer friction. Claims the model clears and the adjuster flags are fraud leakage that lands in incurred losses.
Reported without that split, 88% gives a reserving actuary nothing to price. It is the same limitation in the peer disclosures already in circulation: Lemonade's 96% of first notice of loss handled without human intervention and Travelers closing 90% of catastrophe claims within 30 days both measure throughput, not judgment.
Regulators are moving toward the judgment question regardless. The NAIC evaluation tool pilot, launched March 2, 2026 across 12 states and running through September, asks in Exhibit C for performance metrics on high-risk systems without naming any. Treasury's Financial Services AI Risk Management Framework, published February 19, 2026 with 230 control objectives, pushes toward metrics comparable across institutions. A bespoke F1 computed on one carrier's own test set is not one of those.
What Makes an Agreement Rate Gameable
Sampling, thresholds and the human baseline each move the reported number without moving the model.
AIG's 100 claims were not a random draw from the book. Claims routed to fraud review are already a filtered subset with fraud prevalence far above the general population, so 88% on that pool does not generalize to the full claims universe. An aggregate rate that mixes auto glass with complex commercial liability tells a regulator less than a stratified one.
Threshold selection is the sharper problem. Models output probabilities, not verdicts. A claim scored at 0.72 is flagged at a 0.50 threshold and cleared at 0.75. A carrier can lift its reported agreement rate by tuning the threshold toward the base rate of its adjusters' decisions while the underlying model stays exactly as it was. A single agreement rate quoted without a fixed threshold protocol reports a choice, not a performance.
Then there is the baseline. AIG's figure is concordance with one adjuster. If that adjuster flags fraud more aggressively or more conservatively than professional consensus, the metric is measuring an individual. Inter-rater reliability studies, several adjusters scoring the same claims before any model is introduced, would settle it, and published inter-rater data for claims adjudication is scarce.
None of this argues against the disclosure. Zaffino volunteered a figure no regulation required, and Grant Thornton found 42% of insurers tracking no AI metrics at all. The distance between a specific number with known weaknesses and no number is the entire point. But an agreement rate carried without its sample design, its threshold and its evaluator count is a headline, and the 12% is where the analysis has to begin.
Further Reading
- Hartford's Algorithmic Impact Assessment Sets the Carrier Transparency Bar - How Hartford operationalized qualitative AI governance documentation ahead of the NAIC evaluation pilot.
- NAIC Four-Tier AI Risk Taxonomy: What the Compliance Framework Means for Insurers - Mapping each tier of the proposed taxonomy to specific insurance use cases and documentation requirements.
- Carrier AI Projects Fail at the Audit Layer, Not the Tech - The 44/24 governance gap and what independent AI audits actually test in carrier deployments.
- NAIC Flags Agentic AI as Insurance's Next Governance Gap - How autonomous AI agent cycles push the boundaries of existing governance frameworks.
- AI Governance Gap in Actuarial Practice - ASOP 56 compliance requirements and model risk management for AI systems in actuarial workflows.
- Verisk Fraud Study Reveals Generational Moral Hazard in AI Claims - How 55% Gen Z willingness to alter claims intersects with AIG’s 88% agreement rate benchmark and the detection confidence gap across the industry.
- Adversarial Self-Critique Rewrites AI Underwriting Governance - How a three-agent architecture from arXiv 2602.13213 builds agreement-rate monitoring into the underwriting pipeline itself, reducing hallucination from 11.3% to 3.8% and producing the decision audit trail EU AI Act Article 9 requires.
- Sixfold’s AI Underwriter Turns Carrier Expertise Into Machine Memory - The case for applying agreement-rate monitoring to straight-through bind output from Sixfold’s June 2026 AI Underwriter, with the ASOP 56 validation structure that governs a self-updating carrier-walled model.
Sources
- AIG Q1 2026 Earnings Call Transcript (May 1, 2026) - Motley Fool
- AI Advancing Faster Than Expected as AIG Builds Multi-Agentic Solution (2026) - Reinsurance News
- NAIC Spring 2026 National Meeting Highlights: H Committee Update (April 2026) - Mayer Brown
- NAIC Expands AI Systems Evaluation Tool Pilot to 12 States (2026) - Fenwick
- NAIC Spring 2026 Regulatory Update (April 2026) - Sidley Austin
- U.S. Treasury Financial Services AI Risk Management Framework (February 2026) - U.S. Department of the Treasury
- FDA AI/ML Device Clearances: Clinical Validation Analysis (2025) - JAMA Network Open
- Interrater Reliability: The Kappa Statistic - PMC / Annals of Family Medicine
- 2026 AI Impact Survey Report: Insurance Edition (2026) - Grant Thornton
- Agentic AI for Actuarial Workflows Research Initiative (2026) - Society of Actuaries
- CAS AI Primer: A Practical Guide for Actuaries (2026) - Casualty Actuarial Society
- AI Collaboration and Risk Management (2026) - The Hartford
- Lemonade 2025 Annual Report / 10-K - SEC Filings