We tested the flag. Here is everything it did and did not do.
The backtest: 14,583 English care homes, three baseline dates, every one of them followed to CQC's 1 July 2026 ratings file. The finding holds. It is also narrower than the copy this site used to carry, and the comparisons that narrowed it are printed here rather than left in a folder.
Analysis run 30 July 2026 · Regulatory data: Care Quality Commission, Open Government Licence v3 · Full results table
Over 4 years, 1 July 2022 to 1 July 2026. Every figure on this page is a row count in the published table.
One sentence, and every number in it is a row count.
Of the 803 English care homes CQC rated Good or Outstanding overall but Requires improvement or Inadequate on Well-led at 1 July 2022, 106 (13.2 per cent) held a lower overall rating at 1 July 2026, against 996 of 11,077 (9.0 per cent) whose Well-led was Good or Outstanding. Risk ratio 1.47 [1.22-1.77], p = 1.5e-4.
In plain terms: about one in eight flagged homes lost the rating inside 4 years, against about one in eleven of the clean ones. That is the whole of it. It is a real difference and it is a difference in odds, not a forecast about any particular home. Working a flagged list, most of your calls will still be to homes that were never going to be downgraded. There will just be fewer of them than if you worked the register in alphabetical order.
The outcome being counted is the commercial one. Of the 106 falls in the flagged arm, 98 went to Requires improvement and 8 to Inadequate. Not one is an Outstanding home slipping quietly to Good, and that is structural rather than lucky. The flagged arm contains 0, 0, 0 Outstanding homes across the three windows, because CQC does not rate a home Outstanding overall while marking its Well-led below Good. Every fall this flag can ever produce is a Good home losing Good.
The control arm is not built that way. It holds 638 Outstanding homes in the 4-year window and 109 of its 996 falls are Outstanding homes dropping to Good, which is not the outcome a consultant sells against. Count only falls that land in Requires improvement or Inadequate and the risk ratio goes up rather than down, to 1.61 [1.34-1.95]. The headline uses the smaller number.
How the cohort was drawn, and what a fall is.
The sample. Every row in CQC's 01 July 2022 Latest ratings.ods file marked as a care home, taking the whole-location rating rather than a per-service breakdown, in the 9 English regions, carrying an overall rating that can be ranked. That is 14,583 homes out of the 14,603 in the file. The 20 excluded are 16 with no overall row and 4 carrying an overall value that is not one of the four rating words. Those counts are printed in the results table for each window so the filter can be checked rather than believed.
The arms. Flagged is Good or Outstanding overall with Well-led at Requires improvement or Inadequate, which is the FirstFlag test exactly as the method page states it. Clean is Good or Outstanding overall with Well-led at Good or Outstanding, which is to say Well-led was not marked down. Every Good or Outstanding home in the baseline file is in exactly one of those two arms and 0 are in neither. A third arm, used further down, is Good overall with Well-led Good or Outstanding and exactly one of the other four key questions at Requires improvement.
An earlier version of this page described the clean arm as Well-led Good or Outstanding while the code behind it kept only Well-led Good. That dropped 585 homes out of the comparison in the Jul 2022 window, 95 of which fell, and it pushed the headline up from 1.47 to 1.54. The wide arm is now both what the code does and what this page says, and it is the smaller number. The narrow version is still computed in every window and printed in the results table under the heading Sensitivity 1, so the choice can be checked: Jul 2022 1.47 wide against 1.54 narrow, Jul 2023 1.38 wide against 1.45 narrow, Jul 2024 1.42 wide against 1.48 narrow.
A fall. Ratings are ordered Outstanding, Good, Requires improvement, Inadequate. A home fell if its overall rating in the 1 July 2026 file sits lower in that order than it did at baseline. A re-publication at the same level is not a fall. A re-worded judgement is not a fall. The single assessment framework introduced values that are not ratings at all, and those are counted in their own column and never as a fall.
160 of the flagged homes and 1,985 of the clean ones are missing from the Jul 2026 file because they closed. All of them stay in the denominator as non-falls. Dropping closed homes from one arm and keeping them in the other is how a result of this shape gets manufactured, and dropping them from both still bends it, because the arms do not close at the same rate. Closures were reconciled against CQC's 01_July_2026_Deactivated_Locations.ods file, 64,880 rows. A further 3 flagged and 16 clean homes are absent from the endpoint file without appearing in the closure file, meaning still registered but no longer carrying a published rating. They are kept as non-falls too.
What was fixed in advance, and what cannot be shown. Three baselines, 1 July 2022, 3 July 2023, 1 July 2024, all three reported below. They are July anniversaries of the one endpoint file and they were set in the extract step before the join ran. That is a design description, not a registration. There is no timestamped artefact on this site or in the repository that proves the windows were chosen before the results were seen, so this page does not claim there is one. An earlier version did.
What can be checked instead: no window was dropped, and no comparison was run and left out. The results table counts every test in it, including the ones that go against us, and the count is at the foot of that file. The 4-year window carries the headline because it is the only one long enough for a home's next inspection to have plausibly happened. Take that as a stated reason and read the other two rows yourself.
The two ways this could have been an artefact.
Both of these would produce the gap above without any home actually being worse run. Both were tested before the result was written up.
The flagged ratings were just older.
A home whose last rating is four years old has waited longer for re-inspection than one rated last spring, and a rating nobody revisits cannot fall. If the flagged arm were the staler of the two, the gap would be an inspection-queue effect wearing the clothes of a quality effect.
It is not, and the raw numbers run the other way. Median age of the baseline overall rating was 1,029 days in the flagged arm against 1,141 in the clean one, so the flagged homes had been looked at more recently. Stratifying on rating age in quintiles computed across the whole cohort, the risk ratio moves from 1.47 to 1.49 [1.24-1.80]. It holds in all three windows, which is the column marked age-adjusted in the table below.
The flagged homes closed instead of falling.
If flagged homes shut at a higher rate, the ones left at the endpoint would be a healthier selection and the comparison would be measuring who survived rather than what happened to them. This is the survivorship version of the same trick the denominator rule above is there to prevent.
The honest version is weaker than a flat no. 19.9% of flagged homes closed against 17.9% of clean ones: risk ratio 1.11 [0.96-1.28], p = 0.1542. The interval crosses 1 in all three windows, so the difference is not distinguishable from zero. It is also above 1 in all three windows, 1.11 over 4 years, 1.17 over 3 years, 1.18 over 2 years, and the gap is widest in the shortest window rather than the longest. That is a direction, and it is worth saying which way it points.
It points against the flag, not for it. A home that has closed cannot be marked down, so excess closures among flagged homes remove falls we would otherwise have counted. The survivorship version of this objection needs the bias running the other way. Here it does not.
Closure is a competing risk and the rule for handling it is a choice, so all three defensible rules are printed rather than the one that suits us. Keeping closed homes in the denominator as non-falls, which is what this page publishes, gives 1.47 [1.22-1.77]. Dropping them from both arms gives 1.50 [1.25-1.81]. Counting a closure as a bad outcome in its own right gives 1.23 [1.11-1.36]. The published rule sits between the other two and every one of the three intervals excludes 1.
Some of this is the inspection queue, and the mix moves with the horizon.
A fall needs two separate things to happen. CQC has to come back, and the rating has to drop when it does. Splitting the age-adjusted effect across those two channels says which one is carrying it, and the answer is not the same at two years as it is at four.
| Window | CQC came back | It dropped when they did | Share that is deterioration |
|---|---|---|---|
| 1 July 2022 | 1.16 | 1.31 | 65% |
| 3 July 2023 | 1.19 | 1.19 | 51% |
| 1 July 2024 | 1.33 | 1.08 | 22% |
Risk ratios, flagged against clean, age-adjusted, split on the log scale.
Read plainly, that says over 4 years the flag is mostly telling you the home is genuinely more likely to be marked down. Over the shortest window it is mostly telling you CQC is coming back to this home sooner. Both are useful to a consultant and they are not the same product, so we are not going to describe one as the other.
Including the two that flatter us less.
| Baseline | Flagged fell | Clean fell | Risk ratio | Age-adjusted | Flagged vs one other domain |
|---|---|---|---|---|---|
| 1 July 20224 years | 13.2%106 / 803 | 9.0%996 / 11,077 | 1.47[1.22-1.77] | 1.49[1.24-1.80] | 1.30[1.01-1.67] |
| 3 July 20233 years | 9.2%76 / 825 | 6.7%727 / 10,907 | 1.38[1.10-1.73] | 1.42[1.13-1.78] | 1.03[0.77-1.39] |
| 1 July 20242 years | 7.1%59 / 832 | 5.0%525 / 10,533 | 1.42[1.10-1.84] | 1.46[1.12-1.90] | 0.91[0.65-1.28] |
Percentages are the unconditional probability of a fall. Intervals are 95%. The last column is the whole flagged arm against control B, one estimator in all three rows.
The three windows share most of their homes. Overlap between the flagged arms, measured on CQC location ids, runs from 0.51 to 0.78 by Jaccard: 634 homes shared between Jul 2022 and Jul 2023, 555 homes shared between Jul 2022 and Jul 2024, 727 homes shared between Jul 2023 and Jul 2024. So these are three views of one cohort, not three independent replications. They must never be pooled, and three intervals excluding 1 is not three times the evidence of one. They are printed together because a single window would be a choice, and the choice would be ours.
Shared membership understates it. What the three rows really share is the falls themselves. Added up across the windows the flagged arm produced 241 falls, and those come from 130 distinct homes: 62 of the 76 falls in the Jul 2023 window are homes that also fell in the Jul 2022 one, 38 of the 59 falls in the Jul 2024 window are homes that also fell in the Jul 2022 one, 49 of the 59 falls in the Jul 2024 window are homes that also fell in the Jul 2023 one. Read the table as one set of events counted up to three times, not as a result that replicated.
The last column is the flag measured against homes whose single below-Good key question is one of the other four. It is the comparison that decides whether Well-led is special. Both arms are the full arms, so the numerator is the same flagged arm as every other column here; the version that also strips the flagged arm down to homes carrying exactly one below-Good key question is in the results table, and it does not change the story. The direction reverses across the windows, which is the first entry in the section below.
One more thing before those entries, because it changes how every interval on this page should be read. The results table runs 70 distinct risk ratios with a significance value, plus 27 age-adjusted estimates reported as intervals. No correction for multiple comparisons is applied to any single line, so a Bonferroni threshold across them is 0.05 over 70, or 7.14e-4. The headline sits at 1.5e-4 and clears it, along with 8 other comparisons, most of which are that same contrast restated or stress-tested and are therefore not extra evidence for it. The two shorter windows do not clear it. Neither does any single domain comparison below, nor the flagged arm against control B in any window. Those are directions worth a purpose-built study, not findings to sell, and this page will not present them as more than that.
Caring has the highest fall rate of the five, in every window.
The comparison above pools four key questions into one control arm, which is convenient and which hides something. Below is the same cohort unpooled: Good overall, exactly one key question at Requires improvement, none Inadequate, split by which key question it is. The number of marks is held at one, so the only thing varying is which one.
| Key question at Requires improvement | Jul 2022 | Jul 2023 | Jul 2024 |
|---|---|---|---|
| Safe | 9.8%56 / 572 | 8.3%44 / 527 | 7.2%36 / 499 |
| Effective | 7.1%18 / 252 | 7.6%18 / 237 | 7.1%15 / 212 |
| Caring | 24.1%7 / 29 | 21.4%6 / 28 | 23.8%5 / 21 |
| Responsive | 12.9%26 / 201 | 10.3%18 / 174 | 8.3%13 / 157 |
| Well-led | 13.1%105 / 800 | 9.1%75 / 820 | 7.0%58 / 828 |
Unconditional probability of a fall by the 1 July 2026 file, with the numerator and denominator of every cell.
Caring is the highest of the five in all three windows, and it is not close in the shortest one. Against homes whose single mark is Well-led, the risk ratios for Caring are 1.84 [0.94-3.59] in the Jul 2022 window, 2.34 [1.12-4.92] in the Jul 2023 window, 3.40 [1.52-7.60] in the Jul 2024 window. 2 of those three intervals exclude 1.
| Against Well-led, same cohort | Jul 2022 | Jul 2023 | Jul 2024 |
|---|---|---|---|
| Safe | 0.75[0.55-1.01] p 0.0615 | 0.91[0.64-1.30] p 0.6941 | 1.03[0.69-1.54] p 0.9122 |
| Effective | 0.54[0.34-0.88] p 0.0095 | 0.83[0.51-1.36] p 0.5165 | 1.01[0.58-1.75] p 1.0000 |
| Caring | 1.84[0.94-3.59] p 0.0970 | 2.34[1.12-4.92] p 0.0428 | 3.40[1.52-7.60] p 0.0154 |
| Responsive | 0.99[0.66-1.47] p 1.0000 | 1.13[0.69-1.84] p 0.6669 | 1.18[0.66-2.10] p 0.6129 |
Now the caveats, which are real and which are not a way of taking it back. The Caring cells are 29, 28, 21 homes. Not one of those risk ratios clears the multiplicity threshold set out above. Almost no home in England carries a Caring mark on its own, which is why the cells are that size, and a rate computed on 29 homes moves a long way if two of them go the other direction.
It is printed at full size anyway. Further down this page there is a small-n null that happens to suit us, the 9 homes carrying two marks, and it is used to decline a claim we would otherwise like to make. A page that prints the small cell when the small cell is convenient and pools the small cell away when it is not is doing the opposite of what this page claims to do. So: on this evidence the sharpest single key question is Caring, at 24.1% against Well-led at 13.1% over 4 years, and it is not the one we filter on.
What the study does not show.
Below are 7 sentences the data will not carry. They are here in full size, above the sales argument rather than under it, because a finding is only worth something if you can see what sits next to it. Every one of them is in the published table, so the only question was whether you read them here or found them yourself.
- 01Not shown
“Well-led is the earliest or the strongest warning of a downgrade.”
Against homes whose single below-Good key question is one of the other four, the flagged arm ran at a risk ratio of 1.30 in the four-year window, 1.03 in the three-year, and 0.91 in the two-year. The direction reverses. Whatever raises the risk, it is a below-Good key question rather than that key question.
- 02Not shown
“Well-led is the sharpest of the five key questions.”
Caring is, in all three windows. Holding the number of below-Good key questions at exactly one, homes whose single mark is Caring fell at 24.1%, 21.4% and 23.8% against Well-led's 13.1%, 9.1% and 7.0%. Those cells are 21 to 29 homes and none of them clears the multiplicity threshold, so it is a direction rather than a finding. It is printed because we print the small-n null that suits us two entries down, and printing one without the other would be a choice about which small cell the reader sees.
- 03Not shown
“You could not build a better list yourself.”
In the two-year window you could. Every Good-rated home with any key question below Good is 1,741 homes and holds 22.6% of all the falls, against our 832 homes and 10.1%. That filter costs nothing and the file it comes from is linked at the bottom of this page.
- 04Not shown
“Two problems are worse than one.”
Probably, but this cohort cannot show it. Only nine Good-rated homes in the 2022 file carried two key questions below Good, and two of them fell. Nine homes is not a finding, and the entry above it is the same size of cell pointing the other way.
- 05Not shown
“The flagged homes closed at the same rate as the rest.”
Not quite. Closure risk ratios were 1.11 over four years, 1.17 over three and 1.18 over two, every one of them above 1 and the gap widest in the shortest window. None is distinguishable from no difference, and none is evidence of a survivorship artefact, because the direction runs the wrong way for that: a home that has closed cannot be marked down, so any excess closure among flagged homes removes falls we would otherwise have counted.
- 06Not shown
“The Well-led mark caused the fall.”
Nothing here is causal. These are two groups of homes that differed at baseline and were followed forward, not an experiment. The most that can be said is that the flag sorts homes into a higher-risk and a lower-risk group four years ahead of the outcome.
- 07Not shown
“A flagged home is going to lose its rating.”
Roughly seven in eight did not, over four years. The flag moves the odds from about one in eleven to about one in eight. That is a targeting instrument, not a prophecy, and any consultant who opens a call as though it were a prophecy will be corrected by the person who answers.
In the most recent window, a filter you can build for free beat ours.
Density is only half of what a list has to answer. The other half is reach: not how concentrated the list is in future falls, but how many of them it contains at all. The pool for the 1 July 2024 window is 11,365 English care homes rated Good or Outstanding overall, of which 584 went on to fall.
| List | Homes | Falls in list | Share of all falls | Hit rate | Lift |
|---|---|---|---|---|---|
| FirstFlag flag: Well-led below Good | 832 | 59 | 10.1% | 7.1% | 1.38 |
| One key question below Good, not Well-led | 889 | 69 | 11.8% | 7.8% | 1.51 |
| Any key question below Good | 1,741 | 132 | 22.6% | 7.6% | 1.48 |
Our list is 832 homes and holds 10.1% of the falls in the pool. Every Good-rated home with any key question below Good is 1,741 homes and holds 22.6%, at a slightly better lift as well. The middle row beat us too: homes whose one below-Good key question is not Well-led are more numerous, hold more of the falls than we do, and carry the highest lift in the table. If the number you care about is how many of the sector's coming downgrades you can get in front of, the wider filter wins in this window, and it is free.
Why we filter on Well-led anyway.
Start from what the backtest actually establishes, which is narrower than either the old pitch or the objection above. A home rated Good or Outstanding overall that already carries a below-Good key question is more likely to be downgraded than one that carries none. That is the finding. It does not identify which key question, and on this evidence it is not Well-led in particular.
| Baseline | No key question below Good | Exactly one below Good | Exactly two below Good |
|---|---|---|---|
| 1 July 2022 | 8.9%n = 10,015 | 11.4%n = 1,854 | 9homes |
| 3 July 2023 | 6.4%n = 9,922 | 9.0%n = 1,786 | 20homes |
| 1 July 2024 | 4.7%n = 9,624 | 7.4%n = 1,718 | 19homes |
Probability of a fall by the number of key questions below Good at baseline, Good and Outstanding homes only. The last column is a headcount rather than a rate: too few homes carry two below-Good key questions for a percentage to mean anything.
So the risk is real whichever key question is carrying the mark, and choosing Well-led is therefore not a claim about prediction. It is a claim about what a consultant can sell and then deliver, which is a different question and one the backtest was never going to answer.
Well-led is the domain a governance engagement moves. Audit trails, records, statutory notifications, who is accountable for what, how long the manager has been in post, whether anybody read the last action plan back and closed it out. That is scoped work with a document trail at the end of it. It can start in a week, it can finish inside one engagement, and the evidence of it is written down where the next inspector will look.
A Safe mark is usually staffing levels or medicines competency. Fixing it means recruiting in a market that has no staff in it, or retraining, or spending money on the building. A Caring mark is culture, which is the slowest thing in a home to shift and the hardest to evidence when it has shifted. Both are real work. Neither is a proposal a consultant can scope on a first call and close in a fortnight.
That cuts both ways and the Caring table above is where it bites. Caring is the sharpest single mark in this cohort and it is also the one a consultant can do least about inside one engagement, so the domain that predicts best is the domain that sells worst. We are not going to resolve that by pretending the table says something else.
That is the whole argument, and notice that it does not need Well-led to be the sharpest signal. It is not the sharpest signal. It is the one where the phone call has somewhere to go, sitting on top of a risk that is real either way.
If what you want is the widest net rather than the most workable one, take the wider filter. It is one pass over CQC's monthly spreadsheet, it costs nothing, and the file is linked below. Saying that here is cheaper than you working it out in month three.
This site used to imply that Well-led goes first, and that the flag was a warning the other domains do not give. That was an assumption nobody had tested. It has now been tested and it did not survive, so it is gone from the copy. The comparisons that killed it are the last column of the three-window table and the five-key-question table under it.
The age of the Well-led mark matters more than anything else here.
One split in this dataset changes the size of the effect more than any other, and it was not on this page until now. Take the same two arms and separate them on how old the baseline Well-led judgement already was. A mark published in the last two years is an inspector's current finding. An older one may already have been worked through, with the home carrying a label that no longer describes it.
| Baseline | Well-led mark under 2 years old | Well-led mark 2 years or older |
|---|---|---|
| 1 July 2022 | 2.19[1.62-2.96] | 1.19[0.93-1.52] |
| 3 July 2023 | 1.69[1.16-2.46] | 1.28[0.96-1.70] |
| 1 July 2024 | 1.97[1.25-3.09] | 1.28[0.93-1.76] |
Risk ratios, flagged against clean, computed inside each stratum. Both arms split on the same cut. Intervals are 95%.
In the 4-year window the recent-mark stratum is 43 falls out of 208 flagged homes against 169 out of 1,788 clean ones, a risk ratio of 2.19 [1.62-2.96]. The older-mark stratum is 1.19 [0.93-1.52], and its interval touches 1 in all three windows. The whole of the headline effect is concentrated in homes whose Well-led mark is still fresh.
Only the 4-year recent-mark figure clears the multiplicity threshold above, so this is one cut on one dataset and not a law. It is published because it is the most useful thing in the study for anyone actually working a list: a flagged home whose Well-led mark went up last year is a different prospect from one whose mark has sat there since before the pandemic, and nothing on this site said so before today.
Seven in eight flagged homes did not fall.
Nothing here is causal. Two groups of homes differed on a public record at a baseline date and were followed forward. Nobody assigned anything, so the study cannot say the Well-led mark caused the downgrade, only that homes carrying it were downgraded more often.
And the base rate is the base rate. Over 4 years, 697 of the 803 flagged homes did not lose their overall rating: roughly seven in eight. The flag moves the odds from about one in eleven to about one in eight. That is a targeting instrument. Anyone who opens a call as though it were a prophecy will be corrected by the person who picks up, and deservedly.
The table, the code and the files are all published.
None of this needs to be taken on trust. The results table every figure on this page is drawn from is here, so are the scripts that produced it, and so is the public archive they read.
The spreadsheets are CQC's own, published under the Open Government Licence. Three baselines and one endpoint, plus the closure file the disappearances were reconciled against. The baselines come out of the archived monthly folder linked above. The other two are current downloads from CQC.
- 01 July 2022 Latest ratings.ods
- 03 July 2023 Latest ratings.ods
- 01 July 2024 Latest ratings.ods
- 01_July_2026_Latest_ratings.ods
- 01_July_2026_Deactivated_Locations.ods
The analysis. Joins the baselines to the endpoint, reconciles the disappearances against the closure file, runs every comparison on this page and writes the results table.
Turns one CQC spreadsheet into a per-location snapshot. Applies the care-home filter and no region filter, so the region tally stays auditable.
The spreadsheet reader, behaving exactly as the production refresh does, so the backtest cannot be parsing the files differently from the live product.
Lists and downloads the archived monthly files from the public folder, anonymously.
Dumps the raw rows for any location id out of all four snapshots, so a single home can be traced through the join by hand.
Prints sheet names and header rows of a spreadsheet, which is how the column names were confirmed stable across vintages before anything was joined.
To re-run it: pull the four spreadsheets, run extract.mjs over each one, then analyse.mjs, and it writes the same table. To check a single home rather than the whole cohort, verify.mjs prints its raw rows out of all four snapshots so you can follow one location id through the join by hand. If a number here does not survive that, we want to hear about it, and the number changes.
The list is the cheap half.
You can build the flag list yourself out of the same monthly file, and the method page sets out, item by item, what it will not come with. Those items are what we charge for. This page is only here to tell you what the sorting underneath them is worth, before you decide whether the rest is.