Machine learning reserving models detect development trend shifts two to four quarters ahead of traditional methods. Gradient-boosted trees flag severity movement in long-tail casualty before it reaches standard triangle diagnostics; recurrent networks catch seasonality in short-tail auto physical damage that Bornhuetter-Ferguson smooths over.

ASOP No. 43, adopted in 2007, and ASOP No. 56, effective October 2020, both assume the actuary can explain each step. On a 500-tree ensemble trained on 200 features from individual claim records, compliance becomes translation.

Key Takeaways

  • ASOP 43 section 3.6 requires disclosure of assumptions with a material effect, and a gradient-boosted model's "assumptions" are hyperparameters: learning rate, tree depth, estimator count, feature selection criteria and loss function.
  • A 3% to 5% RMSE improvement is what a gradient-boosted model might deliver over a well-selected chain-ladder on personal auto physical damage with 20 years of stable development, which may not cover the compliance overhead.
  • Long-tail training data is stale by construction: a model fitted on accident years 2015 through 2020 does not carry the post-pandemic litigation surge behind $12.5 billion of other-liability-occurrence deficiencies concentrated in 2021 to 2024.
  • The NAIC Model Bulletin's AIS Program enumerates six functions and reserving is not among them, while the 12-state evaluation pilot prioritizes high-risk systems, which a model driving carried reserves plainly is.
  • The SOA's own commentary concedes that where a model appears as a black box, "all efforts to comply with ASOP requirements may be hampered if it is not possible to peer into the black box."

Where the Accuracy Actually Is

The technology is in production rather than under evaluation, and its advantage is uneven enough to be worth locating precisely.

Milliman's Arius now carries an Advanced Analytics module putting gradient-boosted and random forest output alongside chain-ladder and Bornhuetter-Ferguson results in one environment. WTW's Radar 5, launched October 2025, spans pricing, portfolio, claims and underwriting, with early tests showing fraud detection rates up more than 100%, which reaches case reserve accuracy where fraud drives development. The academic base has moved with it, from Kevin Kuo's DeepTriangle in 2019 through interpretability work in Variance and the CAS E-Forum.

Dimension Chain-Ladder / BF GLMs ML (GBM, Neural Nets)
Accuracy (stable lines) Strong; well-calibrated with adequate history Comparable; better for heterogeneous portfolios Marginal improvement; 2-5% RMSE reduction typical
Accuracy (volatile lines) Degrades with pattern changes Moderate; limited nonlinear capture Significant edge; catches trend shifts 2-4 quarters earlier
Interpretability Fully transparent; each step auditable Coefficients interpretable; interactions less so Requires post-hoc explainability (SHAP, PDP, LIME)
Regulatory acceptance Universal Broadly accepted; some states require additional disclosure Limited; no state has explicitly approved ML-only statutory opinions
Documentation burden Standard actuarial report format Moderate; coefficient tables and diagnostics High; requires model inventory, validation reports, drift monitoring
ASOP 43 alignment Direct; standards written for these methods Compatible with standard disclosures Requires significant interpretation and supplemental documentation

The advantage is not uniform. On personal auto physical damage with 20 years of consistent development, a gradient-boosted model might cut root mean squared error 3% to 5% against a well-selected chain-ladder, an improvement that may not justify the documentation it brings. High claim volumes and granular data favor ML, which is why personal auto and workers' compensation lead adoption.

Long-tail casualty is where the case is strongest and the risk highest. General liability and commercial auto develop over five to fifteen years, and models trained on claim-level features such as attorney involvement, injury type, jurisdiction and litigation status detect pattern shifts faster than aggregate triangles. Cat reserving is the opposite case: low-frequency high-severity data overfits, and ML's contribution there is ALAE estimation and subrogation recovery.

What Counts as an Assumption

The friction is not that the models are opaque in principle. It is that the standards ask for a specific artifact these models do not produce.

ASOP 43 section 3.6 requires disclosure of assumptions with a material effect on the unpaid claim estimate. For chain-ladder those are well understood: loss development factor selections, tail factors, expected loss ratios, each defensible in a sentence. An ML model's equivalent inputs are the learning rate, tree depth, number of estimators, feature selection criteria and the loss function being optimized. Explaining to a regulator why a learning rate of 0.05 with 800 trees was selected, and how that choice materially moves the reserve, is a different act from explaining why a 5-year weighted development factor was chosen over a 3-year.

Section 3.1's data requirement compounds it. Triangle selection, exclusions and adjustments document in a page. A model trained on 200 features drawn from claim records, medical bill line items and external feeds runs to dozens.

Section 3.8, on multiple methods, is where the judgment actually lands. Most practitioners run ML alongside traditional methods rather than instead of them, which fits the standard well. The unresolved part is weighting: the basis for giving the ML indication 30% weight rather than 50% is traditional actuarial judgment applied to a method the actuary cannot decompose into intuitive components, and the disclosure has to carry that.

ASOP 56 pushes on the same seam from the other side. Section 3.2 asks for reasonable efforts to confirm that model structure, data, assumptions, governance and testing suit the intended purpose. A three-hidden-layer network with dropout regularization has a structure that is mathematically precise and practically opaque: the architecture is describable, and why it produces a particular estimate for a particular claim segment needs interpretability tools that sit outside the model. Section 3.6.2(b)'s hold-out testing requirement is the one place the two frameworks align cleanly, since cross-validation and out-of-time testing are built into any competent pipeline.

Limitations are the harder disclosure. Chain-ladder's are catalogued: it assumes stable development, BF needs a reliable expected loss ratio, both struggle on immature years. A gradient-boosted model may perform well in aggregate and produce unreliable estimates for a small jurisdiction thin in the training data, or capture nonlinear interactions that improve accuracy overall while destabilizing the tail. Enumerating those requires understanding behavior across operating conditions, not architecture, which is the gap the Academy's own modeling commentary describes.

The Retrain and the Opinion

The complication is that the appointed actuary signs a point-in-time attestation on a model that is not a point-in-time object.

The Statement of Actuarial Opinion covers carried reserves as of the statement date, and the appointed actuary must be able to explain the basis to regulators, boards and potentially in litigation. An ML indication is a prediction from one trained instance, and it can move materially when the model retrains. If a retrain happened between Q3 and year-end, the opinion has to document how the retrain moved the indication and why the movement was reasonable, which means the model version becomes part of the reserve record.

Underneath that sits a data problem the accuracy comparison does not show. The lines where ML detects change fastest are the long-tail lines whose training sets are furthest out of date, because five to fifteen years of development means the fitted history predates the current environment by construction. A model trained on accident years 2015 through 2020 has no exposure to the post-pandemic litigation surge or the social inflation behind $12.5 billion of other-liability-occurrence deficiencies concentrated in 2021 through 2024. Early detection of a shift the model has never seen is the specific thing it is worst at.

The regulatory frame has not caught up either, in a way that cuts against carriers rather than for them. Over half the states have adopted or substantially replicated the NAIC Model Bulletin on AI from December 2023, which requires a written AI System Program covering product development, marketing, underwriting, rating and pricing, claim administration and fraud detection. Reserving is not enumerated.

The 12-state evaluation tool pilot running March through September 2026 applies proportionality and prioritizes high-risk systems, and a model that influences carried reserves and the appointed actuary's opinion is high-risk on any reading. So the use case most exposed to examiner questions is the one the governing document does not list, and the practical answer is to put reserving models in the AIS inventory before anyone asks.

Further Reading on actuary.info

Sources

  1. ASOP No. 43: Unpaid Claim Estimates, Actuarial Standards Board (2007)
  2. ASOP No. 56: Modeling, Actuarial Standards Board (2020)
  3. Kuo, Kevin. DeepTriangle: A Deep Learning Approach to Loss Reserving. Risks, Vol. 7, No. 3 (2019)
  4. Generalized DeepTriangle: A Flexible Architecture for Loss Reserving. Risks, Vol. 12, No. 1 (2024)
  5. Assured Research: P/C Industry Loss Reserves Redundant by $20.7 Billion, Carrier Management (March 2026)
  6. AM Best: Premium Slowdown, Higher Combined Ratio Expected for 2026, Insurance Journal (February 2026)
  7. NAIC Expands AI Evaluation Tool Pilot to 12 States, Fenwick (2026)
  8. Milliman Arius: Machine Learning for Reserving, Azure ML Partnership Brief
  9. WTW Introduces Radar 5 Insurance Analytics Platform, Reinsurance News (October 2025)
  10. SOA Modeling Section: On ASOP 56 and Modeling (April 2021)
  11. Lorentzen, C. and Mayer, M. Peeking into the Black Box: An Actuarial Case Study for Interpretable Machine Learning, SSRN
  12. Towards Explainability of Machine Learning Models in Insurance Pricing, Variance Journal
  13. Machine Learning in Insurance, CAS E-Forum
  14. Actuarial Modeling Through a New Lens, American Academy of Actuaries
  15. Travelers Reports First Quarter 2026 Results, Travelers Investor Relations
  16. Chubb Reports Q1 2026 Results, Insurance Journal (April 2026)
  17. Deloitte 2026 Global Insurance Outlook
Feedback

We are seeking feedback on how to improve the site and deliver high-quality content relevant to actuaries. Help us make it better.

Submit feedback

Stay ahead with daily actuarial intelligence - news, analysis, and career insights delivered free.

Subscribe to Actuary Brew Browse All Insights