On May 28, 2026, Liberty Mutual became the first major US carrier to embed its own rating engine inside a foundation model platform, quoting personal auto through natural language conversation in Arizona, Kentucky, Ohio, Missouri, New Mexico, Utah, and Wisconsin, with more than 40 states planned by year-end.
Unlike aggregator apps that show ranges and hand off, this routes through Liberty Mutual's own algorithm to a single personalized premium. The engine was calibrated on structured fields. It is now fed free-text conversation.
Key Takeaways
- The app produces a single personalized premium from Liberty Mutual's own rating algorithm, not an estimated range, with the bind completed on LibertyMutual.com.
- Territorial factors across the 30-plus five-digit ZIP codes of greater Phoenix can vary 25 to 40 percent between the most and least favorable segments.
- Vehicle symbol drives collision base rates that vary 10 to 25 percent between trims of the same model year and nameplate, decoded from a 17-character VIN.
- The interpretation layer that converts conversation into rated fields appears in no state rate filing and in none of the actuarial memoranda supporting the territorial or symbol exhibits.
- Channel-level loss development needs 12 to 24 months, and the program reaches 40-plus states within roughly seven of them.
What the Rating Engine Expects at Input
Personal auto rating plans are built on fields with defined formats and filed lookup tables, and each one carries actuarial support developed at that precision.
Garaging ZIP code sets the territorial factor, the primary geographic variable in most state plans. Greater Phoenix alone spans more than 30 five-digit ZIP codes, and territorial factors across them can differ by 25 to 40 percent. A garaging ZIP is not a city or a neighborhood. It is a five-digit code resolving to one factor in a filed table.
Vehicle symbol works the same way off a 17-character VIN, which decodes to year, make, model, series, and body style. A 2019 Toyota Camry carries LE, SE, SE Nightshade, XLE, XSE, and TRD trims, each with its own symbol assignment, and symbol-to-symbol collision base rate variation within a single model year and nameplate typically runs 10 to 25 percent. Date of birth applies in single-year increments for young drivers and two- to five-year bands above that, so mid-thirties and thirty-four are not the same rated value where single-year factors apply. Prior continuous insurance needs carrier, dates, and coverage type. Prior claims need counts, dates, and coverages.
Each was described in the supporting actuarial memorandum as a defined, structured field, and the filed rates are defensible in rate review because that precision held through the whole data pipeline. The conversational interface collects all of them as natural language.
Where the Premium Moves
The gap is not in the rating algorithm. It is in the layer that now sits in front of it.
Converting "I live in the west suburbs of Cleveland" into a specific five-digit garaging ZIP is a translation the rating engine was never designed to perform, and the model performing it is not in any Liberty Mutual state filing. It does not appear in the territorial factor exhibit or the symbol exhibit. Its error rate against ambiguous geography, incomplete vehicle descriptions, and self-reported history that conflicts with motor vehicle records is not validated to the standard that governs the algorithm downstream of it.
The consequence is directional, not merely noisy. Territorial boundaries in dense metropolitan markets do not follow intuitive geographic lines, and a consumer's own description of where they live tends to resolve toward the recognizable suburb rather than the higher-rated urban tier. A quote in Columbus, Ohio that lands on a suburban Dublin ZIP instead of a downtown Columbus one produces a premium below what the filed plan requires for the actual garaging location. The same asymmetry applies to trim: a consumer describing "the sporty Camry from a few years ago" invites a symbol assignment that a VIN would have settled.
The seven launch states are the mild version of this problem. California, New York, New Jersey, Massachusetts, and Florida carry the most heterogeneous territorial structures and the most demanding filing review; Arizona, Kentucky, Ohio, Missouri, New Mexico, Utah, and Wisconsin have moderate territorial variance. The 40-state expansion by year-end brings the high-complexity markets into the channel before the interpretation model has a full policy year of loss experience from the easy ones.
There is a live regulatory question underneath. Most states require prior approval or notification for material changes to a rating algorithm and its input definitions. Whether an input-parsing step counts as part of the rating algorithm has not been answered in published state guidance for this configuration. The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted December 2023 and implemented by 24 states as of mid-2026, asks for validation and testing of the quality and integrity of data used in AI system inputs. If the interpretation model participates in the rating decision, it sits inside that program.
Two Records That Do Not Match
A web form leaves one record. A conversation leaves two, and the space between them is where the exposure sits.
Structured quoting produces field entry timestamps, validation error sequences, override events, and the final payload handed to the rating engine, which is what a market conduct examiner expects when a consumer disputes a premium. A conversational transaction produces a chat transcript of what the consumer said and a separate structured object of what the model derived. Reconstructing a quote requires the transcript, the parsed output at each field, every disambiguation decision the model made, and the final payload, linked through one transaction identifier. Whether that full chain is captured and retained in an examinable format is not addressed in the launch materials.
The monitoring problem has the same shape. Channel-level adverse selection cannot be evaluated without 12 to 24 months of development, and channel preference is a behavioral signal that correlates with risk in both directions. Consumers who abandoned the web form because of its friction, including those anticipating a surcharge, may find the conversational route easier. Highly engaged platform users may skew younger with cleaner records. Either could dominate, and neither will be visible in the data until well after the book has been written.
Liberty Mutual has margin for that. Its 2025 full-year combined ratio was 88.4 percent against 95.9 percent in 2024, on consolidated net written premium of $43.6 billion. What it does not have is a containable pilot. A channel that starts in seven states and reaches more than 40 rate filings before its first claims year develops is an expanding share of new business origination, priced on the assumption that a translation layer nobody filed is getting the ZIP code right.
Further Reading on actuary.info
- Three Carrier AI Architectures: Platform, Partnership, and Proprietary Models Compared - State Farm chose OpenAI Frontier, Travelers deployed Anthropic via TravAI, and Allstate built ALLIE in-house: a comparative framework with actuarial implications for ASOP No. 56 model validation and rate filing documentation across carrier AI strategies.
- Allstate Patents Turn Road Risk Into Rating Evidence - How Allstate’s 2026 paired patents close the sensor-to-premium pipeline and raise ASOP 56 modeling obligations and state rate reviewer expectations for input-derived rating variables.
- Hartford’s Algorithmic Impact Assessment Sets the Carrier Transparency Bar - The first voluntary Algorithmic Impact Assessment from a top-20 carrier, covering bias audits and governance documentation mapped to NAIC 12-state AI evaluation pilot requirements.
- OpenAI Sits in 90% of Carrier AI Stacks: The Vendor Concentration Risk Nobody Is Pricing - Survey data on platform dependency across carrier AI deployments, with analysis of switching costs and correlated model-drift exposure that applies directly to ChatGPT-based distribution infrastructure.
- How Insurance-Native AI Platforms Reframe the Carrier Build-vs-Buy Decision - The data moat argument for insurance-specific AI and the governance implications for carriers choosing between foundation model platforms and proprietary rating infrastructure.
Sources
- Liberty Mutual Group, Official Product Announcement: Liberty Mutual Insurance Launches First-of-its-Kind Carrier-Backed Conversational AI Quoting App in ChatGPT for Auto Insurance (May 28, 2026).
- PR Newswire, Liberty Mutual Insurance Launches First-of-its-Kind Carrier-Backed Conversational AI Quoting App in ChatGPT for Auto Insurance (May 28, 2026).
- Property Casualty 360, Liberty Mutual launches ChatGPT-based auto quoting app (June 2, 2026).
- Fintech Global, Liberty Mutual launches AI car insurance quoting via ChatGPT (June 3, 2026).
- eMarketer, Liberty Mutual brings auto insurance quoting to ChatGPT.
- Insurance Business, Liberty Mutual becomes first major US carrier to offer insurance quotes inside ChatGPT.
- PR Newswire, Liberty Mutual Insurance Reports Fourth Quarter and Full Year 2025 Results.
- NAIC, Model Bulletin: Use of Artificial Intelligence Systems by Insurers (December 4, 2023).
- Actuarial Standards Board, ASOP No. 56, Modeling (adopted October 2019, effective October 2020).
- Actuarial Standards Board, ASOP No. 23, Data Quality (adopted December 2016).
We are seeking feedback on how to improve the site and deliver high-quality content relevant to actuaries. Help us make it better.
Stay ahead with daily actuarial intelligence - news, analysis, and career insights delivered free.
Subscribe to Actuary Brew Browse All Insights