The CMS CY 2027 final rule (CMS-4207-F), published April 2, 2026 and effective June 1, removes 11 Star Rating measures on which the industry already averaged above 94%. CMS did not change the weighting methodology. It shrank the denominator, and the arithmetic pushes clinical outcomes and patient experience from roughly half of the overall score to about 65%. For the 40% of MA-PD contracts sitting between 3.5 and 4.0 stars, that changes which investments earn the 5% bonus.

Key Takeaways

  • 11 measures removed, phased across 2028 and 2029, drawn predominantly from administrative and process categories where average industry scores exceeded 94%.
  • Clinical outcomes plus patient experience move from roughly 50% to about 65% of overall weight, without any change to the measure-level weighting rules.
  • Triple-weighted outcome measures gain the most: Controlling Blood Pressure, Diabetes HbA1c Poor Control above 9.0%, and Plan All-Cause Readmissions.
  • CAHPS measures at 1.5x weight become the second-largest block, and their contract-level sampling noise can move a plan half a star between measurement years on variance alone.
  • The new Depression Screening measure averages two rates: a screening rate and a 30-day follow-up rate, so 95% screening against 55% follow-up scores 75%.

The Denominator Shrank

Before CY 2027, MA-PD contracts were scored on up to 43 measures across clinical outcomes, CAHPS patient experience, pharmacy quality and administrative process. The administrative and process block, covering call center availability, appeals timeliness, complaints and member retention, carried roughly 25 to 30% of overall weight.

The scoring machinery is unchanged. A clustering algorithm still assigns each measure to a one-through-five star band, and the measure-level weights still run 1x for process, 1.5x for intermediate outcomes and patient experience, and 3x for outcomes, before the overall score is computed.

What changed is what remains in the pool. Most of the 11 removed measures carried a 1x process weight, so removal takes roughly 11 weight units out of the denominator. CMS added back the Depression Screening and Follow-Up measure and three clinical measures for the 2027 ratings, being Care for Older Adults Functional Status Assessment, Concurrent Use of Opioids and Benzodiazepines, and Polypharmacy with multiple anticholinergic medications, each at a weight of 1. The additions do not replace what was removed, so the share held by 1.5x and 3x measures rises.

Measure CategoryPre-Removal SharePost-Removal ShareDirection
Clinical Outcomes (HEDIS, 3x weight)~25-28%~30-35%Gains most
Patient Experience (CAHPS, 1.5x weight)~25-28%~28-32%Moderate gain
Pharmacy Quality~12-15%~15-18%Moderate gain
Administrative/Process (1x weight)~25-30%~12-18%Largest reduction

The practical reading is about where past spending went. A plan that built call center capacity, appeals throughput and complaint reduction was earning star credit for work that no longer scores. A plan that built HEDIS performance and provider quality now gets more score per dollar from the same programs.

The Weight Moves Toward the Measures That Are Hardest to Forecast

The redistribution is not neutral across the remaining measures, and the ones that gain most are the ones a bid model can least confidently project.

Triple-weighted Part C outcome measures gain disproportionately, because removing 1x process measures raises every 3x measure's share of the total. Three matter most: Controlling Blood Pressure for members 18 to 85 with hypertension, Diabetes HbA1c Poor Control above 9.0% where a lower rate is better, and Plan All-Cause Readmissions on a risk-adjusted 30-day basis. Unlike the removed measures, all three show genuine spread across plans, which is precisely why they differentiate.

The second block is the problem. CAHPS patient experience measures at 1.5x weight now constitute the second-largest share, and they are survey instruments. At contract level the confidence intervals are wide enough that a plan's score can move half a star between measurement years from sampling variation alone, with no change in what the plan did.

That noise now carries more weight in the overall score, and the financial consequence is specific. With 40% of MA-PD contracts between 3.5 and 4.0 stars and a 5% quality bonus on the county benchmark at stake, the variance in the star forecast is the variance in bid revenue. A contract expected to land just under 4.0 stars, carrying half a star of survey-driven error, does not support a single revenue assumption. It supports a two-outcome one, and the bid has to be priced against the distribution rather than the point estimate.

Nothing in the rule changed how much a star is worth. It changed how much of the star is determined by measurement error.

The New Measure Averages Two Rates, and One of Them Is Not the Plan's to Control

The Depression Screening and Follow-Up measure is the hardest item in the rule to model, for two reasons that compound.

It scores two distinct rates that CMS averages. The screening rate is the share of eligible members aged 12 and over screened with a standardized instrument, typically the PHQ-2 or PHQ-9, during the measurement year. The follow-up rate is the share of members who screened positive and received documented follow-up within 30 days, whether a behavioral health referral, a documented treatment plan, pharmacotherapy or further evaluation.

Averaging removes the option of winning on volume. A plan screening 95% of eligible members but achieving 55% follow-up within 30 days averages 75%, which after clustering is unlikely to reach four or five stars. Screening is an operational process a plan can drive through its own workflows. Follow-up depends on behavioral health provider availability, appointment capacity and member engagement in a clinical area with historically high no-show rates. The half of the measure that is harder to move is the half that caps the score.

The second problem is that there is nothing to forecast from. Plans hold no historical performance on this measure or on the three other new clinical additions. The clustering algorithm sets star bands from the observed distribution across contracts, so the first year's thresholds are determined by how everyone else performs, and a plan's placement will not correlate with its historical performance on the measures being removed.

That is the position the rule leaves plans in. The weight has moved toward outcome measures with real spread, toward survey measures with real noise, and toward a new measure with no history and a rate the plan only partly controls, on a scoring distribution where 40% of contracts sit within half a star of the bonus threshold.

Further Reading on actuary.info

Sources