We tested the flag. Here is everything it did and did not do.
FirstFlag has no customers and no track record, so there is no hit rate on this site and there will not be one until there is something real to count. What there is now is a backtest: 14,583 English care homes, three baseline dates fixed before a line of analysis ran, every one of them followed to CQC's 1 July 2026 ratings file. The finding holds. It is also narrower than the copy this site used to carry, and the comparisons that narrowed it are printed here rather than left in a folder.
Analysis run 30 July 2026 · Regulatory data: Care Quality Commission, Open Government Licence v3 · Full results table
Over 4 years, 1 July 2022 to 1 July 2026. Every figure on this page is a row count in the published table.
One sentence, and every number in it is a row count.
Of the 803 English care homes CQC rated Good or Outstanding overall but Requires improvement or Inadequate on Well-led at 1 July 2022, 106 (13.2 per cent) held a lower overall rating at 1 July 2026, against 901 of 10,492 (8.6 per cent) whose Well-led was Good. Risk ratio 1.54 [1.27-1.86], p = 0.000027.
In plain terms: about one in eight flagged homes lost the rating inside 4 years, against about one in twelve of the clean ones. That is the whole of it. It is a real difference and it is a difference in odds, not a forecast about any particular home. Working a flagged list, most of your calls will still be to homes that were never going to be downgraded. There will just be fewer of them than if you worked the register in alphabetical order.
The outcome being counted is the commercial one. Of the 106 falls in the flagged arm, 98 went to Requires improvement and 8 to Inadequate. Not one of them is an Outstanding home slipping quietly to Good.
How the cohort was drawn, and what a fall is.
The sample. Every row in CQC's 01 July 2022 Latest ratings.ods file marked as a care home, taking the whole-location rating rather than a per-service breakdown, in the 9 English regions, carrying an overall rating that can be ranked. That is 14,583 homes out of the 14,603 in the file. The 20 excluded are 16 with no overall row and 4 carrying an overall value that is not one of the four rating words. Those counts are printed in the results table for each window so the filter can be checked rather than believed.
The arms. Flagged is Good or Outstanding overall with Well-led at Requires improvement or Inadequate, which is the FirstFlag test exactly as the method page states it. Clean is Good or Outstanding overall with Well-led at Good or Outstanding. A third arm, used further down, is Good overall with Well-led Good and exactly one of the other four key questions at Requires improvement.
A fall. Ratings are ordered Outstanding, Good, Requires improvement, Inadequate. A home fell if its overall rating in the 1 July 2026 file sits lower in that order than it did at baseline. A re-publication at the same level is not a fall. A re-worded judgement is not a fall. The single assessment framework introduced values that are not ratings at all, and those are counted in their own column and never as a fall.
160 of the flagged homes and 1,914 of the clean ones are missing from the Jul 2026 file because they closed. All of them stay in the denominator as non-falls. Dropping closed homes from one arm and keeping them in the other is how a result of this shape gets manufactured, and dropping them from both still bends it, because the arms do not close at the same rate. Closures were reconciled against CQC's 01_July_2026_Deactivated_Locations.ods file, 64,880 rows. A further 3 flagged and 15 clean homes are absent from the endpoint file without appearing in the closure file, meaning still registered but no longer carrying a published rating. They are kept as non-falls too.
Pre-specification. Three baselines, 1 July 2022, 3 July 2023, 1 July 2024, all chosen before the analysis ran, and all three reported below. The 4-year window carries the headline because it is the only one long enough for a home's next inspection to have plausibly happened, and that reason was written down before the numbers came back. Had we picked it afterwards for being the friendliest, this page would be the thing it claims not to be.
The two ways this could have been an artefact.
Both of these would produce the gap above without any home actually being worse run. Both were tested before the result was written up.
The flagged ratings were just older.
A home whose last rating is four years old is more overdue for re-inspection than one rated last spring, and a rating nobody revisits cannot fall. If the flagged arm were the staler of the two, the gap would be an inspection-queue effect wearing the clothes of a quality effect.
It is not, and the raw numbers run the other way. Median age of the baseline overall rating was 1,029 days in the flagged arm against 1,140 in the clean one, so the flagged homes had been looked at more recently. Stratifying on rating age in quintiles computed across the whole cohort, the risk ratio moves from 1.54 to 1.56 [1.29-1.88]. It holds in all three windows, which is the column marked age-adjusted in the table below.
The flagged homes closed instead of falling.
If flagged homes shut at a higher rate, the ones left at the endpoint would be a healthier selection and the comparison would be measuring who survived rather than what happened to them. This is the survivorship version of the same trick the denominator rule above is there to prevent.
They did not. 19.9% of flagged homes closed against 18.2% of clean ones: risk ratio 1.09 [0.95-1.26], p = 0.24. That difference is not distinguishable from zero, and the interval crosses 1 in all three windows.
Some of this is the inspection queue, and the mix moves with the horizon.
A fall needs two separate things to happen. CQC has to come back, and the rating has to drop when it does. Splitting the age-adjusted effect across those two channels says which one is carrying it, and the answer is not the same at two years as it is at four.
| Window | CQC came back | It dropped when they did | Share that is deterioration |
|---|---|---|---|
| 1 July 2022 | 1.14 | 1.38 | 71% |
| 3 July 2023 | 1.17 | 1.26 | 59% |
| 1 July 2024 | 1.31 | 1.12 | 30% |
Risk ratios, flagged against clean, age-adjusted, split on the log scale.
Read plainly, that says over 4 years the flag is mostly telling you the home is genuinely more likely to be marked down. Over the shortest window it is mostly telling you CQC is coming back to this home sooner. Both are useful to a consultant and they are not the same product, so we are not going to describe one as the other.
Including the two that flatter us less.
| Baseline | Flagged fell | Clean fell | Risk ratio | Age-adjusted | Against other domains |
|---|---|---|---|---|---|
| 1 July 20224 years | 13.2%106 / 803 | 8.6%901 / 10,492 | 1.54[1.27-1.86] | 1.56[1.29-1.88] | 1.29[1.00-1.67] |
| 3 July 20233 years | 9.2%76 / 825 | 6.4%659 / 10,339 | 1.45[1.15-1.81] | 1.48[1.17-1.86] | 1.03[0.77-1.39] |
| 1 July 20242 years | 7.1%59 / 832 | 4.8%477 / 9,971 | 1.48[1.14-1.92] | 1.51[1.16-1.97] | 0.91[0.65-1.28] |
Percentages are the unconditional probability of a fall. Intervals are 95%.
The three windows share most of their homes. Overlap between the flagged arms, measured on CQC location ids, runs from 0.51 to 0.78 by Jaccard: 634 homes shared between Jul 2022 and Jul 2023, 555 homes shared between Jul 2022 and Jul 2024, 727 homes shared between Jul 2023 and Jul 2024. So these are three views of one cohort, not three independent replications. They must never be pooled, and three intervals excluding 1 is not three times the evidence of one. They are printed together because a single window would be a choice, and the choice would be ours.
The last column is the flag measured against homes whose single below-Good key question is one of the other four. It is the comparison that decides whether Well-led is special. The direction reverses across the windows, which is the first entry in the section below.
What the study does not show.
Below are 5 sentences the data will not carry. They are here in full size, above the sales argument rather than under it, because a finding is only worth something if you can see what sits next to it. Every one of them is in the published table, so the only question was whether you read them here or found them yourself.
- 01Not shown
“Well-led is the earliest or the strongest warning of a downgrade.”
Against homes whose single below-Good key question is one of the other four, the flagged arm ran at a risk ratio of 1.29 in the four-year window, 1.03 in the three-year, and 0.91 in the two-year. The direction reverses. Whatever raises the risk, it is a below-Good key question rather than that key question.
- 02Not shown
“You could not build a better list yourself.”
In the two-year window you could. Every Good-rated home with any key question below Good is 1,741 homes and holds 22.6% of all the falls, against our 832 homes and 10.1%. That filter costs nothing and the file it comes from is linked at the bottom of this page.
- 03Not shown
“Two problems are worse than one.”
Probably, but this cohort cannot show it. Only nine Good-rated homes in the 2022 file carried two key questions below Good, and two of them fell. Nine homes is not a finding.
- 04Not shown
“The Well-led mark caused the fall.”
Nothing here is causal. These are two groups of homes that differed at baseline and were followed forward, not an experiment. The most that can be said is that the flag sorts homes into a higher-risk and a lower-risk group four years ahead of the outcome.
- 05Not shown
“A flagged home is going to lose its rating.”
Roughly seven in eight did not, over four years. The flag moves the odds from about one in twelve to about one in eight. That is a targeting instrument, not a prophecy, and any consultant who opens a call as though it were a prophecy will be corrected by the person who answers.
In the most recent window, a filter you can build for free beat ours.
Density is only half of what a list has to answer. The other half is reach: not how concentrated the list is in future falls, but how many of them it contains at all. The pool for the 1 July 2024 window is 11,365 English care homes rated Good or Outstanding overall, of which 584 went on to fall.
| List | Homes | Falls in list | Share of all falls | Hit rate | Lift |
|---|---|---|---|---|---|
| FirstFlag flag: Well-led below Good | 832 | 59 | 10.1% | 7.1% | 1.38 |
| One key question below Good, not Well-led | 889 | 69 | 11.8% | 7.8% | 1.51 |
| Any key question below Good | 1,741 | 132 | 22.6% | 7.6% | 1.48 |
Our list is 832 homes and holds 10.1% of the falls in the pool. Every Good-rated home with any key question below Good is 1,741 homes and holds 22.6%, at a slightly better lift as well. The middle row beat us too: homes whose one below-Good key question is not Well-led are more numerous, hold more of the falls than we do, and carry the highest lift in the table. If the number you care about is how many of the sector's coming downgrades you can get in front of, the wider filter wins in this window, and it is free.
Why we filter on Well-led anyway.
Start from what the backtest actually establishes, which is narrower than either the old pitch or the objection above. A home rated Good or Outstanding overall that already carries a below-Good key question is more likely to be downgraded than one that carries none. That is the finding. It does not identify which key question, and on this evidence it is not Well-led in particular.
| Baseline | No key question below Good | Exactly one below Good | Exactly two below Good |
|---|---|---|---|
| 1 July 2022 | 8.9%n = 10,015 | 11.4%n = 1,854 | 9homes |
| 3 July 2023 | 6.4%n = 9,922 | 9.0%n = 1,786 | 20homes |
| 1 July 2024 | 4.7%n = 9,624 | 7.4%n = 1,718 | 19homes |
Probability of a fall by the number of key questions below Good at baseline, Good and Outstanding homes only. The last column is a headcount rather than a rate: too few homes carry two below-Good key questions for a percentage to mean anything.
So the risk is real whichever key question is carrying the mark, and choosing Well-led is therefore not a claim about prediction. It is a claim about what a consultant can sell and then deliver, which is a different question and one the backtest was never going to answer.
Well-led is the domain a governance engagement moves. Audit trails, records, statutory notifications, who is accountable for what, how long the manager has been in post, whether anybody read the last action plan back and closed it out. That is scoped work with a document trail at the end of it. It can start in a week, it can finish inside one engagement, and the evidence of it is written down where the next inspector will look.
A Safe mark is usually staffing levels or medicines competency. Fixing it means recruiting in a market that has no staff in it, or retraining, or spending money on the building. A Caring mark is culture, which is the slowest thing in a home to shift and the hardest to evidence when it has shifted. Both are real work. Neither is a proposal a consultant can scope on a first call and close in a fortnight.
That is the whole argument, and notice that it does not need Well-led to be the sharpest signal. It is not the sharpest signal. It is the one where the phone call has somewhere to go, sitting on top of a risk that is real either way.
If what you want is the widest net rather than the most workable one, take the wider filter. It is one pass over CQC's monthly spreadsheet, it costs nothing, and the file is linked below. Saying that here is cheaper than you working it out in month three.
This site used to imply that Well-led goes first, and that the flag was a warning the other domains do not give. That was an assumption nobody had tested. It has now been tested and it did not survive, so it is gone from the copy, and the comparison that killed it is the last column of the table above.
Seven in eight flagged homes did not fall.
Nothing here is causal. Two groups of homes differed on a public record at a baseline date and were followed forward. Nobody assigned anything, so the study cannot say the Well-led mark caused the downgrade, only that homes carrying it were downgraded more often.
And the base rate is the base rate. Over 4 years, 697 of the 803 flagged homes did not lose their overall rating: roughly seven in eight. The flag moves the odds from about one in twelve to about one in eight. That is a targeting instrument. Anyone who opens a call as though it were a prophecy will be corrected by the person who picks up, and deservedly.
The table, the code and the files are all published.
None of this needs to be taken on trust. The results table every figure on this page is drawn from is here, so are the scripts that produced it, and so is the public archive they read.
The spreadsheets are CQC's own, published under the Open Government Licence. Three baselines and one endpoint, plus the closure file the disappearances were reconciled against. The baselines come out of the archived monthly folder linked above. The other two are current downloads from CQC.
- 01 July 2022 Latest ratings.ods
- 03 July 2023 Latest ratings.ods
- 01 July 2024 Latest ratings.ods
- 01_July_2026_Latest_ratings.ods
- 01_July_2026_Deactivated_Locations.ods
The analysis. Joins the baselines to the endpoint, reconciles the disappearances against the closure file, runs every comparison on this page and writes the results table.
Turns one CQC spreadsheet into a per-location snapshot. Applies the care-home filter and no region filter, so the region tally stays auditable.
The spreadsheet reader, behaving exactly as the production refresh does, so the backtest cannot be parsing the files differently from the live product.
Lists and downloads the archived monthly files from the public folder, anonymously.
Dumps the raw rows for any location id out of all four snapshots, so a single home can be traced through the join by hand.
Prints sheet names and header rows of a spreadsheet, which is how the column names were confirmed stable across vintages before anything was joined.
To re-run it: pull the four spreadsheets, run extract.mjs over each one, then analyse.mjs, and it writes the same table. To check a single home rather than the whole cohort, verify.mjs prints its raw rows out of all four snapshots so you can follow one location id through the join by hand. If a number here does not survive that, we want to hear about it, and the number changes.
The list is the cheap half.
You can build the flag list yourself out of the same monthly file, and the method page sets out, item by item, what it will not come with. Those items are what we charge for. This page is only here to tell you what the sorting underneath them is worth, before you decide whether the rest is.