Mike McGavick, the former XL Group CEO now chairing mea Platform's operating board, told the CAS Seminar on Reinsurance on June 5, 2026 that smaller insurance-trained models beat general-purpose LLMs on insurance documents, and that OpenAI's and Anthropic's models "lose their way" with the terminology.
The claim is testable, and the published numbers mostly support it. What they do not support is the conclusion carriers are drawing about how much review the accuracy gain removes.
Key Takeaways
- $32 billion a year is Accenture's estimate of operational inefficiency in insurance, which McGavick sizes as 12 to 14 cents of every premium dollar not going to loss costs or rate relief.
- 0.5 to 3 combined ratio points is what mea Platform claims for carrier clients, a range wide enough to reorder competitive position in any line if it holds.
- 0.6% hallucination is the published ceiling, from the INS-S1 insurance model family, against an 11.3% baseline measured for general-purpose models in underwriting decision support.
- 30% accuracy gain and 30% lower inference cost is what EXL reports for its fine-tuned Insurance LLM against generic models on claims and underwriting tasks.
- 42% of P&C insurers lack the measurement infrastructure to track model performance over time, which is the capability a domain model requires most and a general model requires least.
What McGavick Actually Argued
The economic case rests on an Accenture estimate that operational inefficiency costs the sector roughly $32 billion annually, which McGavick puts at 12 to 14 cents of every premium dollar absorbed by manual data handling, duplicative workflows and systems that cannot talk to each other.
That figure is an expense load, and expense load is a direct input to a rate indication. Any systematic reduction flows to the combined ratio, which is why mea's claimed 0.5 to 3 percentage points is the number worth interrogating rather than the accuracy statistics.
His argument for why domain training specifically captures it is about document form rather than model size. Submissions arrive as "napkins with drawings, Excel sheets, photos." A general model trained on internet text parses clean documents well and degrades on the mixed-format, terminology-dense material that constitutes daily insurance operations. A model trained for five years on submissions, ACORDs, loss runs, wordings, bordereaux and claims files builds representations of those concepts rather than approximating them from context.
He framed the industry's pattern as three phases with any transformative technology: exclude it, harness it to cut cost, then get greedy and write new products for risks the technology makes insurable. The current argument sits entirely in phase two.
What the Accuracy Numbers Buy, and What They Do Not
The empirical claim holds up. The question is what a validated actuary can do with it.
The strongest published result is the INS-S1 insurance model family, which reports a 0.6% hallucination rate while outperforming DeepSeek-R1 and Gemini-2.5-Pro on domain tasks, alongside INSEva, a benchmark of more than 39,000 insurance-specific test samples. Against that, Roy and Singh measured 11.3% baseline hallucination for general-purpose models in underwriting decision support, cut to 3.8% by a domain-aware critic agent that challenges the primary agent before human review. That architecture moved decision accuracy from 92% to 96% across 500 expert-validated underwriting cases.
| Vendor | Architecture | Domain Focus | Reported Advantage |
|---|---|---|---|
| Mea Platform | From-scratch dsLM + Knowledge Graph | Full insurance operations | Weeks to production; 0.5-3 pts combined ratio improvement |
| Akur8 | Transparent ML + agentic AI | Actuarial pricing & reserving | End-to-end actuarial workflow; 50%+ ARR growth |
| EXL | Fine-tuned LLM (NVIDIA) | Claims & underwriting | 30% accuracy gain; 30% lower cost vs. generic LLMs |
| FIS | Embedded GenAI in existing tooling | Risk modeling | 24/7 actuarial guidance; multilingual support |
The vendor split reflects different bets on where domain knowledge belongs. Mea trains from scratch. Akur8 narrows to the actuarial pricing and reserving lifecycle and optimizes only for that. EXL fine-tunes on 25 years of medical records processing for bodily injury, workers' compensation and general liability, reporting 30% accuracy gains and 30% lower cost against generic models. FIS embeds an assistant in tooling actuarial teams already license.
Here is where the numbers stop translating. For a reserve estimate or a rate indication, the acceptable count of fabricated data points is zero, not 0.6%. Moving from 11.3% to 0.6% does not make an output signable; it changes the volume of errors a human reviewer has to catch, which is what makes an AI-assisted workflow economically viable rather than merely possible. The reviewer is still in the loop at 0.6%, and the expense saving has to be modelled net of that reviewer rather than as a replacement for them.
That is the arithmetic that matters against the 12 to 14 cent expense load. A domain model that removes most of the manual extraction while requiring the same actuarial sign-off yields a fraction of the 0.5 to 3 combined ratio points on offer.
Where the Domain Bet Erodes
Three things work against domain specificity, and none of them is a defect in the models.
The first is maintenance. A model fine-tuned on 2024 claims data does not reflect 2026 litigation trends, social inflation patterns or new regulatory guidance, so the total cost of ownership includes perpetual retraining, data curation and drift monitoring. General-purpose providers absorb that cost themselves. Capgemini found 42% of P&C insurers lack the measurement infrastructure to track model performance over time, which means a large share of the market is buying the architecture that needs the most governance while lacking the capability to supply it.
The second is the improvement curve. Capabilities that needed fine-tuning 18 months ago are reachable through prompting on current model versions. A carrier committing to a domain architecture is betting that insurance language and regulatory constraint create a durable gap rather than a temporary one. The evidence splits by task: structured extraction and compliance documentation favor domain models, cross-functional and generative work favors general ones.
The third is breadth. A carrier writing personal auto, commercial property, specialty and reinsurance may cover more use cases with one well-prompted general model than with several domain models each carrying separate procurement, integration and governance.
The CAS is funding the counterweight at up to $80,000 across two tracks, with deliverables required as open source under the Mozilla Public License 2.0. That produces shared infrastructure and shared baselines, which is the part the vendor market will not supply on its own. It does not produce a model anyone will run in production. McGavick's own closing line concedes where the burden lands regardless of architecture: "If I use the word model, it's the actuary in the end that's going to be telling the rest of us what can be trusted and what cannot."
Further Reading on actuary.info
- CAS Funds Research to Fine-Tune LLMs for P&C Actuarial Reasoning
- Akur8's Matrisk Acquisition Builds the First End-to-End Actuarial AI Platform
- Three Carrier AI Architectures: Platform, Partnership, and Proprietary
- EXL's Insurance LLM Patents Reveal Domain AI Strategy
- Synthetic Data Wins CAS Ratemaking Prize: Privacy-Safe Pricing With KDE
Sources
- Carrier Management, "Exclude It, Harness It, Get Greedy: McGavick's Take on Insurers' AI Playbook," June 5, 2026 - carriermanagement.com
- CAS, "2026 Request for Proposals: Adapting Large Language Models (LLMs) for Specialized P&C Actuarial Reasoning," February 2026 - casact.org
- EXL, "EXL Launches Specialized Insurance Large Language Model Leveraging NVIDIA AI Enterprise," September 2024 - exlservice.com
- EXL, "AI-Powered Insurance Workflows: Operationalizing LLMs with EXL Insurance LLM" - exlservice.com
- Fintech Global, "How Akur8 Is Building an End-to-End Actuarial Platform for the Next Era of Insurance," March 2026 - fintech.global
- FIS, "FIS Launches 24/7 AI Assistant to Ease Risk Models Management," February 2026 - fisglobal.com
- mea Platform, "Agentic AI for (Re)Insurance Operations" - meaplatform.com
- mea Platform, "mea Platform Announces New AI Products to Replace Core Insurance Industry Workflows," October 2025 - businesswire.com
- Reinsurance News, "HDI Global Partners with mea Platform to Expand AI-Driven Input Management," 2026 - reinsurancene.ws
- Roy, J. and Singh, S., "Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique," arXiv 2602.13213, January 2026 - arxiv.org
- "An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control" (INS-S1), arXiv 2603.14463, March 2026 - arxiv.org
- NAIC, "Artificial Intelligence and State Insurance Regulation," March 2026 - naic.org
- Fenwick, "NAIC Expands AI Systems Evaluation Tool Pilot Program to 12 States," 2026 - fenwick.com
- Plante Moran, "How the NAIC AI Model Bulletin Is Evolving and Why Insurers Should Prepare Now," March 2026 - plantemoran.com
- Celent, "GenAI in Insurance" - celent.com
- Datos Insights, "Insurance Leaders Gathered in Boston to Define the New Insurance Carrier Operating Model for AI," April 2026 - datos-insights.com
- Cogent, "Domain-Specific Language Models (DSLMs): The End of the General-Purpose LLM Hype in 2026" - cogentinfo.com
We are seeking feedback on how to improve the site and deliver high-quality content relevant to actuaries. Help us make it better.
Stay ahead with daily actuarial intelligence - news, analysis, and career insights delivered free.
Subscribe to Actuary Brew Browse All Insights