Roots Automation relaunched its insurance AI platform as Bevaya on May 28, 2026, with two numbers doing the positioning work: 300 million proprietary insurance documents in InsurGPT's training corpus, and 115 live carrier deployments including 3 of the top 5 P&C carriers by premium.

The two answer different questions. The corpus figure is a claim about domain fluency on submission parsing and coverage analysis. The deployment figure is a claim about production evidence, and only one of them is verifiable today.

Key Takeaways

  • 300 million documents is the corpus claim, against the 20 to 30 million insurance documents a single large carrier typically holds across policy administration, claims, and underwriting systems.
  • 115 live deployments including 3 of the top 5 P&C carriers give the platform a production track record that a newly fine-tuned in-house model cannot match on day one.
  • InsurGPT is an ensemble of specialized sub-models with an orchestration layer, not a single fine-tuned general-purpose architecture, which changes what version stability means in a rate filing cycle.
  • The NAIC AI Systems Evaluation Tool pilot reached 12 states by mid-2026 and requires carriers to document how third-party models were trained and what governs model updates.
  • AM Best found 68% of roughly 150 rated carriers and MGAs use third-party AI while only 18% cite third-party model risk as a challenge.

What the Two Numbers Claim

The architecture matters more than the rebrand, because it determines what the corpus figure buys.

A fine-tuned general model, typically a GPT-4 class architecture pre-trained on broad internet text, learns language from that general pre-training and then learns insurance from a smaller adaptation layer. A purpose-built domain model learns language itself from domain data, so its internal representation of a term like occurrence policy or loss development factor is shaped by insurance usage rather than by every English sentence that mentions insurance.

InsurGPT sits closer to the second. Roots Automation describes an ensemble of specialized models rather than one fine-tuned architecture: separate sub-models for submission parsing, coverage determination, and claims classification, combined by an orchestration layer. Ensembles are more robust than any single component, which is the technical reason for the design.

On structured tasks the distinction may not pay. Extracting loss run data from a standardized ACORD form is a task where a fine-tuned general model and a domain model can land at near-identical accuracy, a point developed in the domain-trained versus general LLM debate from the CAS Seminar on Reinsurance. The gap opens on ambiguous policy language, jurisdiction-specific coverage interpretations, and non-standard claims documents, where a model either has encountered the pattern in training or is interpolating.

The Moat Is the Accumulation Rate, Not the Corpus Size

A static 300 million documents is a weaker asset than the number suggests, and a stronger one for a reason the press release does not lead with.

Breadth is the honest half of the claim. A single large carrier holding 20 to 30 million documents has a real training set, but it is one book: its policy forms, its claim types, its coverage language. A corpus assembled across multiple carriers spans form variations and claim patterns that single-carrier data never contains, and edge cases are exactly where a domain model that has seen too little fails on coverage determination.

Recency is where a static corpus decays. Commercial auto documents gathered in 2024 do not represent the coverage disputes now forming around autonomous vehicle liability endorsements, telematics-based rating modifications, or loss types that standard forms address ambiguously. The 115 active deployments matter here more than the 300 million does, because production use keeps adding training signal that a fixed dataset cannot. That is the same shape as Sedgwick Omni's claims data advantage: the accumulation rate outlasts the head start.

The pricing consequence sits in that same feedback loop. If sub-models are independently retrained and updates ship across the installed base, the system generating outputs inside a carrier's workflow this quarter is not necessarily the one that produced the outputs supporting an in-flight rate filing. An ensemble makes this sharper than a single model would, because one component can shift while the others hold and the orchestration layer absorbs the change.

The contractual answer is narrow and specific: the right to freeze a named model version for the length of an annual pricing cycle, notification on every version change, and rollback if an update degrades a previously validated task. Without those terms the platform is a moving input under a filing that assumes a fixed one.

The Documentation Burden Falls Asymmetrically

Regulatory documentation is where the insurance-native claim stops being marketing, and it does not land evenly across vendor types.

The NAIC's AI Systems Evaluation Tool pilot expanded to 12 states by mid-2026 and requires carriers to document how third-party AI models were trained, what datasets were used, and what governance controls apply to model updates. The proposed Third-Party Model Vendor Registry would go further, requiring vendors serving the insurance market to register models and submit training data descriptions before their platforms could be used by carriers subject to market conduct examination. It is proposed, not adopted, and its timing tracks how fast states take up the evaluation tool.

The asymmetry is structural. A carrier running a GPT-4 class model fine-tuned on its own documents through an enterprise API cannot document the base model's training composition, because that composition is not the carrier's to disclose and foundation model labs do not publish it in detail. A vendor whose entire corpus is its own has no such gap. For Bevaya specifically, a registry process is an upgrade rather than a threat: it is the mechanism that would move 300 million documents from an assertion in a launch announcement to a figure someone has checked.

The market has not priced that work in. AM Best's survey of roughly 150 rated carriers and MGAs found 68% using third-party AI solutions and only 18% naming third-party model risk as a challenge, an accountability gap that a market conduct examiner in a pilot state will reach before anything else in an AI governance file.

And the corpus advantage has a ceiling that neither side states plainly. Publicly available insurance text, regulatory filings, NAIC publications, coverage decisions, and actuarial literature, is substantial. Policy language in force, loss run formats, adjuster notes, and underwriting questionnaire responses are not public at scale, which is what keeps an open-source alternative from closing the gap on the tasks where domain data pays. The moat is narrower than 300 million implies and wider than open-source convergence arguments allow, and which of those is true depends entirely on the task.