Machine learning tools trained on venue-level verdict databases, judge disposition histories and third-party litigation funding records are now supplementing commercial auto reserve triangles, because the traditional loss development method cannot see a jury award coming.

$31.3B
2024 Nuclear Verdict Total, Up 116% YoY
$51M
2024 Median Nuclear Verdict (vs. $21M in 2020)
104.4 → 106.3
Projected Commercial Auto Combined Ratio, 2026-2029
14
Consecutive Years of Commercial Auto Underwriting Losses

Nuclear verdicts jumped 52% in 2024 to 135 total, worth $31.3 billion, up 116% year over year, with a median award of $51 million against $21 million in 2020.

Key Takeaways

  • $2.7 billion of a $4 to $5 billion industry reserve deficiency sits in accident years 2021 and later, so the shortfall is a live mispricing of hard-market vintages rather than a legacy tail.
  • A venue-and-judge score exists the day counsel and a court are assigned, months or years before trial, which is the one thing a backward-looking average cannot supply.
  • 8% to 10% flagged against roughly 1% converting. A scenario reserve built off flagged-claim counts is a conversion-rate problem, and that rate is the parameter most exposed to stale calibration.

Why the Curve Cannot See a Verdict Coming

Chain-ladder and Bornhuetter-Ferguson share a foundational assumption: that the ratio of losses at one maturity to the next is reasonably stable across accident years. That holds for a claims population where severity is unimodal, clustering around a central tendency with a thinning tail.

Commercial auto bodily injury no longer looks like that. A growing share of claims resolve either in a settlement well below the plaintiff's demand or in a verdict at or above it, with comparatively little mass between. Gen Re's litigation analytics research notes 89 nuclear verdicts in 2023 alone totaling $14.5 billion, the highest in 15 years at the time, before 2024 more than doubled it.

A bimodal distribution breaks the link-ratio assumption in a specific way. The average development factor at a given maturity blends the settlement mode and the trial mode, weighted by how many claims in that cohort land in each. When the trial-mode share creeps up, the historical average understates the expected value for the next cohort, because it was estimated when trial-mode claims were a smaller share.

The triangle is not wrong about the past. It cannot extrapolate a shifting mixture weight forward, which is how the median nuclear verdict climbed from $21 million in 2020 to $51 million in 2024 while link ratios moved far more slowly.

The effect surfaces in reserve development rather than the initial pick. 2024 statutory results put commercial auto liability at an 87.6 loss ratio, the highest in 11 years, on a record $6.4 billion liability underwriting loss, while physical damage posted a best-ever $1.5 billion profit: a 24.6-point gap between two coverages inside one combined line.

AM Best estimates a $4 to $5 billion industry reserve deficiency with $2.7 billion concentrated in accident years 2021 and later. Christopher Graham of AM Best notes that "with claims remaining open longer, insurers have more direct costs in attorney fees and expert witnesses as cases are negotiated before trial", which extends exactly the maturities where triangle methods are least reliable.

Two Signals That Arrive Before the Triangle Does

What the analytics layer changes is not accuracy but timing, and timing is what the reserving method lacks.

Products such as CLARA Analytics and Lex Machina mine judicial records and settlement histories to score open claims by jurisdiction, judge assignment and injury severity. Lex Machina's dataset, built from PACER filings and state court records, surfaces judge-specific settlement rates, letting a model flag a venue-judge combination that resolves a higher share of comparable claims by trial. CLARA reports a 2% to 5% reduction in total incurred losses on scored claims.

A traditional case reserve is an adjuster's judgment, updated through discovery, appearing in the aggregate triangle at the next valuation date. A venue-and-judge score exists the day the claim is assigned counsel and a court, because jurisdiction and judge assignment are known facts rather than projections.

Third-party litigation funding is the second layer, with industry estimates putting $16.1 billion committed to US litigation in 2024. A funded claim carries information no triangle holds: institutional capital does not deploy against a claim it expects to settle quickly and cheaply, so funding is a third party's underwriting judgment that the claim is worth a protracted, expensive path to trial. Reliance Partners describes institutional investors "backing high-stakes lawsuits using sophisticated data analytics, fueling protracted litigation and increasing both the frequency and severity of claims".

Funding decisions are typically made in the first 12 to 18 months after a complaint, well before a trial date three to five years out on a congested docket. That asymmetry is the actuarial value.

Input Traditional Triangle ML Litigation Model
Signal timing Emerges at next valuation date Available at claim/venue assignment
Distribution shape assumed Unimodal, smooth decay Bimodal, settlement vs. trial modes
Primary data source Own historical paid/incurred losses Venue/judge records, TPLF activity
Failure mode Understates accelerating tail Stale training data on fast-moving target

The output enters reserving as a supplement, in three forms. A percentile-based IBNR ladder segments open claims by risk score and treats the high-score cohort as its own severity distribution with a fatter right tail, mirroring the segmentation already applied to liability against physical damage. A scenario reserve sits outside the triangle-derived central estimate. And a supplemental tail factor applies to the 36-to-72-month band where AM Best's $2.7 billion has emerged.

The scenario reserve is where the arithmetic gets delicate. Scoring tools flag roughly 8% to 10% of open commercial auto bodily injury claims as elevated risk against a base rate where perhaps 1% ultimately reach that tier. The reserve is therefore a conversion-rate calculation calibrated to a carrier's own flagged-to-actual history, not a claim count.

The Model Is Calibrated on the Slower World

The objection writes itself, and it deserves stating as sharply as the case for the tools: a model trained on verdict distributions from 2018 through 2023 is being pointed at 2025 and 2026, precisely the period when the underlying inflation is accelerating.

If frequency grew 52% in a single year and the dollar total grew 116%, a model calibrated on the prior five years is calibrated on a slower-moving target. The parallel to trend selection is exact. The structural-shift problem that makes a long-term historical severity trend inadequate for commercial auto rate filings, addressed by weighting recent experience more heavily, applies with equal force to any tail model fitted on a rolling historical verdict window.

That is a limitation rather than an engineering defect, and it has a direction. A venue-scoring model unrefreshed since verdict inflation accelerated in 2024 is arguably worse than no model, because it lends precision to an estimate stale in the direction that matters most: understating severity, the same failure the triangle already exhibits. The question worth putting to a vendor is how often the verdict database and model weights refresh, and whether out-of-sample performance can be shown on the 2023-to-2024 acceleration specifically rather than on a longer, calmer window.

Relying on a vendor's model for tail development also creates a disclosure obligation a triangle-based selection does not. The actuary depends on proprietary training data, a refresh cadence and a conversion-rate calibration none of which is directly auditable. The opinion should document which claims population the model was applied to, the vendor's stated refresh frequency, and a reconciliation showing how much of the held reserve traces to the ML-informed adjustment against the traditional indication. That reconciliation is what separates a supplemental estimate from an opaque override, and it tracks the same concerns raised in our work on what actuaries owe when AI tools touch a reserve opinion.

It matters most when reserves move. If held reserves rise materially and the increase is attributed to a flagged cohort, a reader needs to know whether that reflects new information, a claim newly flagged or funding newly detected, or a re-calibration of the model after a period of understatement. Those are different events with different implications for future volatility, and conflating them understates how much of the reserve position now depends on a vendor's maintenance schedule rather than the carrier's own claims experience.

Further Reading

Sources