AIG chief executive Peter Zaffino put a number on AI governance during the carrier's Q1 2026 earnings call on May 1. A professional claims adjuster reviewed 100 claims for fraud. Anthropic's Claude reviewed the same 100 independently. The two agreed 88% of the time. Most carriers describe AI deployments as "promising results" and "significant efficiency gains." An agreement rate is the rarer thing: a governance figure a regulator, a board member and a reserving actuary can all read without translation.

Key Takeaways

  • 88% agreement between Claude and an AIG claims adjuster across 100 fraud determinations, from an out-of-the-box model with no claim-specific tuning. Tuning should move it, which makes 88% a floor rather than a result.
  • Plus or minus 6.4 points is the 95% confidence interval on 100 binary observations, so the true rate sits somewhere between roughly 82% and 94%. Reaching plus or minus 2 points takes 1,000 claims.
  • 12 claims out of 100 is where the governance question lives. Whether the model over-flagged or under-flagged decides whether the cost lands in claims expense or in incurred losses.
  • 12 states are running the NAIC AI Systems Evaluation Tool pilot through September 2026, and its Exhibit C asks carriers to describe performance metrics for high-risk systems without prescribing which.
  • 42% of insurers track no AI metrics at all, on Grant Thornton's 2026 survey, so the standardization problem starts well before anyone argues about which metric is best.

What the 88% Was Measured On

Zaffino described the evaluation with unusual specificity for an earnings call. One professional claims adjuster ranked 100 claims as fraudulent or legitimate with documented reasoning. Claude assessed the same 100 with no access to the adjuster's work. The two sets of determinations were then compared.

The load-bearing phrase in his description is "out-of-the-box." The model received no fine-tuning on AIG's proprietary claims data. The 88% is general reasoning applied to insurance fraud indicators: timeline inconsistencies, geolocation mismatches, prior claim patterns, document tampering signals.

Sample size is the first thing to check. For a binary classification on 100 observations, the 95% confidence interval runs roughly plus or minus 6.4 percentage points, putting the true agreement rate between about 82% and 94%. A 1,000-claim sample narrows that to plus or minus 2 points, and 10,000 brings it under 1 point. Reported at that spread, the figure is a direction rather than a measurement.

The rest of what Zaffino disclosed measures something else entirely. AIG Assist in Lexington middle market property produced a 30% improvement in quoting volume, a 55% reduction in time-to-quote and roughly 40% more submissions bound. Those are throughput figures. Only the 88% addresses whether the machine reaches the conclusion a qualified human would.

Why Agreement Beats F1 as a Reported Metric

This section shows why the metric that reads as least rigorous to a data scientist is the one most likely to survive into regulatory reporting.

Accuracy, precision, recall and F1 all score a model against a labeled ground truth. They answer whether the model is right. Claims adjudication rarely offers that luxury: two experienced adjusters can read the same file and disagree in good faith, so the ground truth is itself a judgment call. An agreement rate asks the answerable question instead, and that reframing is why it travels across audiences F1 never reaches.

MetricWhat It MeasuresAudienceGovernance Utility
AccuracyCorrect predictions / total predictionsData scientistsLow: misleading with imbalanced classes (fraud rates of 5-10%)
PrecisionTrue positives / predicted positivesData scientists, model opsMedium: meaningful for false-positive costs but requires technical context
RecallTrue positives / actual positivesData scientists, complianceMedium: captures missed fraud but needs pairing with precision
F1 ScoreHarmonic mean of precision and recallData scientistsLow: composite metric that obscures the tradeoffs boards need to see
Raw Agreement RateConcordance between AI and expert humanAll three audiencesHigh: intuitive, comparable across functions, audit-friendly
Cohen's KappaAgreement adjusted for chance concordanceStatisticians, actuariesHigh: corrects for base-rate inflation in agreement percentages

The correction an actuary should insist on is Cohen's Kappa. Raw agreement inflates with the base rate: if 90% of claims are legitimate, a model that calls everything legitimate scores 90% against any adjuster working the same mix. Kappa nets out chance concordance. On Cohen's scale, 0.61 to 0.80 is substantial agreement and above 0.81 is almost perfect. At a 20% fraud prevalence, typical for a flagged-for-review pool, 88% raw would correspond to a Kappa of roughly 0.65 to 0.75. The prevalence in AIG's sample was not disclosed.

The financial content of the metric sits in the 12 claims where the two disagreed, and the direction decides the sign. Claims the model flags and the adjuster clears drive investigation expense and customer friction. Claims the model clears and the adjuster flags are fraud leakage that lands in incurred losses.

Reported without that split, 88% gives a reserving actuary nothing to price. It is the same limitation in the peer disclosures already in circulation: Lemonade's 96% of first notice of loss handled without human intervention and Travelers closing 90% of catastrophe claims within 30 days both measure throughput, not judgment.

Regulators are moving toward the judgment question regardless. The NAIC evaluation tool pilot, launched March 2, 2026 across 12 states and running through September, asks in Exhibit C for performance metrics on high-risk systems without naming any. Treasury's Financial Services AI Risk Management Framework, published February 19, 2026 with 230 control objectives, pushes toward metrics comparable across institutions. A bespoke F1 computed on one carrier's own test set is not one of those.

What Makes an Agreement Rate Gameable

Sampling, thresholds and the human baseline each move the reported number without moving the model.

AIG's 100 claims were not a random draw from the book. Claims routed to fraud review are already a filtered subset with fraud prevalence far above the general population, so 88% on that pool does not generalize to the full claims universe. An aggregate rate that mixes auto glass with complex commercial liability tells a regulator less than a stratified one.

Threshold selection is the sharper problem. Models output probabilities, not verdicts. A claim scored at 0.72 is flagged at a 0.50 threshold and cleared at 0.75. A carrier can lift its reported agreement rate by tuning the threshold toward the base rate of its adjusters' decisions while the underlying model stays exactly as it was. A single agreement rate quoted without a fixed threshold protocol reports a choice, not a performance.

Then there is the baseline. AIG's figure is concordance with one adjuster. If that adjuster flags fraud more aggressively or more conservatively than professional consensus, the metric is measuring an individual. Inter-rater reliability studies, several adjusters scoring the same claims before any model is introduced, would settle it, and published inter-rater data for claims adjudication is scarce.

None of this argues against the disclosure. Zaffino volunteered a figure no regulation required, and Grant Thornton found 42% of insurers tracking no AI metrics at all. The distance between a specific number with known weaknesses and no number is the entire point. But an agreement rate carried without its sample design, its threshold and its evaluator count is a headline, and the 12% is where the analysis has to begin.

Further Reading

Sources