CMS named Milliman the winner of its Crushing Fraud Chili Cook-Off Competition on December 15, 2025, for a glass-box fraud detection tool: a deterministic algorithm rooted in actuarial science whose every flag can be deconstructed into the inputs that produced it. The competition did not ask for the highest detection rate. It asked for reasoning a program integrity team could read, and that requirement is the finding.
Key Takeaways
- More than 250 submissions arrived in Phase 1 and CMS named ten finalists on October 20, 2025. The winning criterion was explainability at the provider level, not detection accuracy.
- $28.83 billion at a 6.55% rate was Medicare Fee-for-Service improper payments in FY 2025, down from $31.70 billion at 7.66% the prior year, which is the base a payment integrity credit is drawn against.
- Roughly 250 models per day run at CMS's Center for Program Integrity, which is why a system that flags 10% of legitimate providers exhausts investigative capacity regardless of its sensitivity.
- $230 million in recoverable claims, three times its legacy system on the same dataset, is what Milliman identified in its Mastercard Healthcare Solutions partnership, alongside 2,700 providers flagged as high risk.
- More than $2 billion saved since March 2025 by the Fraud Defense Operations Center, which moves the actuarial question from how much can be recovered to how much was prevented, a quantity nobody observes directly.
What CMS Chose to Judge
The competition ran in two phases. Phase 1 took proposals from August 19 through September 19, 2025, drawing more than 250 submissions from consultancies, academic teams and technology vendors. CMS named ten finalists on October 20, including Milliman, MindPetal, a joint Stanford and UCSF team and Visual Connections.
Phase 2 gave finalists CMS Limited Data Sets covering Medicare Fee-for-Service claims in Hospice, Part B and Durable Medical Equipment, three of the highest-risk categories for improper billing, and asked for findings plus a scalable policy solution.
The evaluation criteria say more than any policy statement. CMS required that solutions be explainable, stating that pattern detection alone was insufficient and that insights had to be transparent and accessible to program integrity teams, regulators and policymakers. A flag is only operationally useful if an investigator can articulate why a billing pattern is anomalous, both to triage it and to defend it in an administrative proceeding.
Milliman's tool synthesizes three anomaly dimensions into one composite provider score that can be taken back apart:
| Anomaly Category | What It Measures | Why It Matters for Fraud Detection |
|---|---|---|
| Behavioral | Billing patterns, service frequency, procedure code usage relative to specialty norms | Identifies providers whose clinical behavior deviates significantly from peers serving similar patient populations |
| Network | Referral relationships, shared patient clusters, geographic service patterns | Detects coordinated billing schemes involving multiple providers or facilities |
| Financial | Total billing volume, cost per beneficiary, reimbursement rate patterns | Flags providers billing at volumes or rates inconsistent with practice size and specialty |
False Positives Are What the Requirement Is Actually Pricing
Black-box models are good at sensitivity and poor at specificity, and in Medicare claims that asymmetry has a name attached to it.
A deep learning model trained to flag high-cost providers will flag oncologists treating advanced cancers, hospice providers managing complex end-of-life care, and specialists in underserved areas carrying referral volumes out of proportion to practice size. None of them are committing fraud. They are managing sicker and more expensive populations.
Each false positive costs investigative hours, a provider dispute process, a possible payment suspension that interrupts care, and network relations. With roughly 250 models per day running at the Center for Program Integrity, a system that catches 95% of fraudulent providers while flagging 10% of legitimate ones does not produce enforcement, it produces a queue.
The glass-box design attacks specificity rather than sensitivity. Because the algorithm is deterministic and statistically grounded, it normalizes for patient acuity, specialty billing distributions and regional cost variation before scoring. An oncologist billing $4 million a year for chemotherapy infusions in a high-prevalence region scores differently from a DME supplier billing $4 million for catheter kits with no clinical justification, even at identical volume.
That is where the actuarial consequence sits. A payment integrity program's PMPM credit in a rate filing is a function of positive predictive value, not detection rate. A high-false-positive system overstates recoverable dollars at projection and underdelivers at recovery, because investigating a legitimately expensive provider yields no recoupment. Milliman's $230 million in recoverable claims, three times its legacy system on the same data, with 2,700 providers flagged, is a benchmark precisely because both the numerator and the flag count are stated.
Prevention Moves the Number Somewhere Nobody Can Audit
The enforcement apparatus around the competition is running on the opposite design principle, and it complicates the measurement.
The Fraud Defense Operations Center, launched March 2025, replaced pay-and-chase with pre-payment detection and has saved more than $2 billion, identifying $2.6 billion in Medicare overpayments across 3,262 providers in 2025. CMS Deputy Administrator Kim Brandt described it as a Netflix-type algorithm labeling applicants as potentially high risk, with the validation step left to people. In one category CMS recorded a 99% decrease in skin and tissue substitute billing after detecting impossibly high claim levels.
A 99% drop in a billing category is not the same measurement as a recovery. Recoveries are observed: a dollar is identified, disputed and returned. Prevented payments are inferred from a counterfactual, and the counterfactual includes whatever legitimate billing in that category stopped alongside the fraudulent billing. Medicare FFS improper payments fell to $28.83 billion at 6.55%, from $31.70 billion at 7.66%, but that rate is computed on claims that were paid, so blocking claims upstream moves it whether or not the blocked claims were improper.
The enforcement environment is expanding on the same basis. The February 2026 CRUSH request for information proposes identity verification, Medicare Advantage preclusion list strengthening and AI-assisted coding oversight. CMS deferred roughly $259.5 million in federal Medicaid matching funds from Minnesota and threatened more than $1 billion, freezing enrollment across 13 provider categories. DOJ's 2025 takedown charged 324 defendants over $14.6 billion in alleged fraud.
For a health plan actuary, the explainability standard is therefore doing double duty. It makes a flag defensible in an investigation, and it is the only thing that makes a prevented-payment estimate auditable at all.
Further Reading
- MHPAEA 2026: Health Actuaries Must Now Prove Parity Holds – Technical walkthrough of MHPAEA compliance and the data-driven comparative analysis framework, with parallels to the documentation and transparency standards CMS is requiring for AI-powered fraud detection programs.
- CMS Star Ratings Overhaul: $18.6B in MA Quality Bonuses at Stake – How CMS is restructuring the Star Ratings methodology that drives Medicare Advantage quality bonuses, providing broader context for the agency’s data-driven approach to program management.
- CMS 2027 MA Final Rule: 2.48% Rate Increase vs. 0.09% NPRM – Component decomposition of the Medicare Advantage payment environment that shapes the financial incentives around program integrity investment.
- CBO Part D Spending Forecasts and the $500B Actuarial Gap – The Medicare Part D spending trajectory that makes fraud prevention in pharmaceutical and DME claims increasingly urgent.
- The AI Governance Gap in Actuarial Practice – How the actuarial profession is building governance frameworks for AI tools, directly relevant to the model validation and ASOP No. 56 compliance requirements for payment integrity AI.
Sources
- GovCIO Media & Research: CMS Uses Explainable AI to Strengthen Medicare Fraud Detection
- CMS: Crushing Fraud Chili Cook-Off Competition
- Morningstar/Business Wire: Milliman Wins CMS Crushing Fraud Chili Cook-Off Competition (December 2025)
- Nextgov: CMS Seeks to Expand Tech-Driven Fight Against Medicaid Fraud (March 2026)
- Nextgov: CMS Saved $2 Billion by Using AI to Fight Fraud (February 2026)
- Foley & Lardner: New Federal Focus on Fraud, Waste and Abuse May Signal Changes for the Health Care Industry (April 2026)
- Milliman: Payment Integrity and AI Express
- Milliman: Effective Claims Auditing for Healthcare Fraud, Waste, and Abuse
- U.S. Department of Justice: 2025 National Health Care Fraud Takedown (324 Defendants, $14.6 Billion)
- CMS: Fiscal Year 2025 Improper Payments Fact Sheet
- Morgan Lewis: CMS Announces Sweeping Anti-Healthcare Fraud Initiatives (February 2026)