FIRSTFLAG SIGNAL BACKTEST ========================= Generated: 2026-07-30T00:02:12Z ALL THREE PRE-SPECIFIED WINDOWS ARE REPORTED. Nothing was selected after seeing results. Supplementary arms are labelled 'supp.' and were added to answer the treatment-vs-Control-B question, not to pick a winner. SOURCES (every number below traces to a row count in these files) Baselines : 01 July 2022 / 03 July 2023 / 01 July 2024 Latest ratings.ods (public CQC archive, Google Drive folder 1N9JH4DhoKvb5SO6I9ObqRmhPgxD_m_pS) Endpoint : 01_July_2026_Latest_ratings.ods (cqc.org.uk, 26.5MB) Closures : 01_July_2026_Deactivated_Locations.ods (cqc.org.uk, 27.2MB, 64,880 rows) METHOD Sample : Care Home? = Y, Service/Population Group = Overall, Location Region in the nine English regions, Overall rating on the 4-point scale. Join : left join baseline -> endpoint on Location ID. Rank : Outstanding 4 > Good 3 > Requires improvement 2 > Inadequate 1. 'Fell' = endpoint rank strictly below baseline rank. Re-rated : Overall 'Publication Date' differs between the two snapshots. Closure : absent from endpoint AND present in Deactivated Locations. Never silently dropped; closure is its own outcome. Rating strings are case-normalised ('Requires Improvement' appears in the 2024 file, 'Requires improvement' elsewhere). SAF-era non-ratings in the 2026 file ('Not Rated', 'Regulations met', 'Insufficient evidence to rate', 'Not applicable') are unrankable and counted in their own column, never as a fall. ==================================================================================================== WINDOW 2022-07-01 -> 2026-07-01 ==================================================================================================== Baseline care-home locations in file: 14603; eligible after region + rankable-overall filter: 14583 excluded: region not English 0, no Overall row 16, Overall unrankable 4 data checks: publication date moved backwards 0; same pub date but different rating 2 ARM COUNTS AND OUTCOMES arm n closed gone surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond rose P(rose) uncond --------------------------------------------------------------------------------------------------------------------------------------------------- T 803 160 (19.9%) 3 640 283 44.2% [40.4-48.1] 106 37.1% [31.7-42.9] 13.2% [11.0-15.7] 4 0.5% [0.2-1.3] A 10492 1914 (18.2%) 15 8563 3501 40.9% [39.8-41.9] 901 25.7% [24.3-27.2] 8.6% [8.1-9.1] 61 0.6% [0.5-0.7] B 1054 202 (19.2%) 1 851 424 49.8% [46.5-53.2] 107 25.2% [21.3-29.6] 10.2% [8.5-12.1] 6 0.6% [0.3-1.2] A_clean 9430 1710 (18.1%) 14 7706 3075 39.9% [38.8-41.0] 793 25.8% [24.3-27.4] 8.4% [7.9-9.0] 55 0.6% [0.4-0.8] T_pure 800 159 (19.9%) 3 638 282 44.2% [40.4-48.1] 105 37.2% [31.8-43.0] 13.1% [11.0-15.6] 4 0.5% [0.2-1.3] ALL 14583 2848 (19.5%) 24 11711 5589 47.7% [46.8-48.6] 1137 20.3% [19.3-21.4] 7.8% [7.4-8.2] 1204 8.3% [7.8-8.7] 'closed' = absent at endpoint, found in Deactivated Locations. 'gone' = absent at endpoint and NOT in Deactivated (still registered but no longer carries a published rating). 'surv' = present at endpoint. Percentages in brackets are 95% Wilson intervals. SUPPORTING DETAIL PER ARM T: TREATMENT Overall Good/Outstanding + Well-led RI/Inadequate n=803 survived=640 closed=160 (in-window 160, end-date before baseline 0) unreconciled-absent=3 overall re-rated=283 any-domain re-rated=283 fell=106 (of which to RI 98, to Inadequate 8) rose=4 unchanged=527 endpoint-unrankable=3 among re-rated: fell=105 rose=4 same=171 A: CONTROL A Overall Good/Outstanding + Well-led Good n=10492 survived=8563 closed=1914 (in-window 1913, end-date before baseline 1) unreconciled-absent=15 overall re-rated=3501 any-domain re-rated=3503 fell=901 (of which to RI 839, to Inadequate 45) rose=61 unchanged=7565 endpoint-unrankable=36 among re-rated: fell=901 rose=61 same=2503 B: CONTROL B Overall Good + Well-led Good + exactly 1 other domain RI n=1054 survived=851 closed=202 (in-window 202, end-date before baseline 0) unreconciled-absent=1 overall re-rated=424 any-domain re-rated=424 fell=107 (of which to RI 100, to Inadequate 7) rose=6 unchanged=734 endpoint-unrankable=4 among re-rated: fell=107 rose=6 same=307 A_clean: supp. A_clean Overall Good/Outstanding + Well-led Good + no other RI/Inad n=9430 survived=7706 closed=1710 (in-window 1709, end-date before baseline 1) unreconciled-absent=14 overall re-rated=3075 any-domain re-rated=3077 fell=793 (of which to RI 738, to Inadequate 38) rose=55 unchanged=6826 endpoint-unrankable=32 among re-rated: fell=793 rose=55 same=2195 T_pure: supp. T_pure Overall Good + Well-led RI + no other RI/Inad (1 RI, it is Well-led) n=800 survived=638 closed=159 (in-window 159, end-date before baseline 0) unreconciled-absent=3 overall re-rated=282 any-domain re-rated=282 fell=105 (of which to RI 97, to Inadequate 8) rose=4 unchanged=526 endpoint-unrankable=3 among re-rated: fell=105 rose=4 same=170 HEAD-TO-HEAD COMPARISONS (risk ratio [95% CI], Fisher exact two-tailed p) P(fall) uncond T vs A 106/803 vs 901/10492 RR=1.54 [1.27-1.86] p=2.7e-5 P(re-rated|surv) T vs A 283/640 vs 3501/8563 RR=1.08 [0.99-1.18] p=0.1043 P(fall|re-rated) T vs A 105/283 vs 901/3501 RR=1.44 [1.23-1.69] p=6.1e-5 P(closed) T vs A 160/803 vs 1914/10492 RR=1.09 [0.95-1.26] p=0.2372 P(rose) uncond T vs A 4/803 vs 61/10492 RR=0.86 [0.31-2.35] p=1.0000 P(fall) uncond T vs B 106/803 vs 107/1054 RR=1.30 [1.01-1.67] p=0.0470 P(re-rated|surv) T vs B 283/640 vs 424/851 RR=0.89 [0.79-0.99] p=0.0361 P(fall|re-rated) T vs B 105/283 vs 107/424 RR=1.47 [1.18-1.84] p=0.0008 P(closed) T vs B 160/803 vs 202/1054 RR=1.04 [0.86-1.25] p=0.6795 P(fall) uncond T_pure vs B (1 RI each) 105/800 vs 107/1054 RR=1.29 [1.00-1.67] p=0.0471 P(re-rated|surv) T_pure vs B (1 RI each) 282/638 vs 424/851 RR=0.89 [0.79-0.99] p=0.0317 P(fall|re-rated) T_pure vs B (1 RI each) 105/282 vs 107/424 RR=1.48 [1.18-1.84] p=0.0008 P(fall) uncond B vs A_clean 107/1054 vs 793/9430 RR=1.21 [1.00-1.46] p=0.0633 P(fall) uncond T_pure vs A_clean 105/800 vs 793/9430 RR=1.56 [1.29-1.89] p=2.1e-5 WHICH SINGLE DOMAIN? (Good overall, exactly one domain at RI, none Inadequate) Holds 'number of RI domains' fixed at one and varies only which domain it is. This is the decisive test of whether Well-led is special. sole RI domain n closed surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ---------------------------------------------------------------------------------------------------------------- safe 572 109 463 226 48.8% [44.3-53.4] 56 24.8% [19.6-30.8] 9.8% [7.6-12.5] effective 252 53 199 97 48.7% [41.9-55.6] 18 18.6% [12.1-27.4] 7.1% [4.6-11.0] caring 29 4 25 16 64.0% [44.5-79.8] 7 43.8% [23.1-66.8] 24.1% [12.2-42.1] responsive 201 36 164 85 51.8% [44.2-59.3] 26 30.6% [21.8-41.0] 12.9% [9.0-18.3] well-led 800 159 638 282 44.2% [40.4-48.1] 105 37.2% [31.8-43.0] 13.1% [11.0-15.6] Pooled: sole RI = Well-led vs sole RI = any of the other four P(fall) uncond 105/800 vs 107/1054 RR=1.29 [1.00-1.67] p=0.0471 P(re-rated|surv) 282/638 vs 424/851 RR=0.89 [0.79-0.99] p=0.0317 P(fall|re-rated) 105/282 vs 107/424 RR=NaN [NaN-NaN] p=0.0e+0 DOSE-RESPONSE: number of key questions at RI/Inadequate at baseline (Good or Outstanding overall homes only). If the signal is really a count rather than a specific domain, P(fall) rises monotonically down this table. # RI/Inad KQs n closed surv P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ------------------------------------------------------------------------------------------------------ 0 10015 1781 8219 39.3% [38.2-40.3] 888 27.5% [26.0-29.1] 8.9% [8.3-9.4] 1 1854 361 1489 47.4% [44.9-50.0] 212 30.0% [26.8-33.5] 11.4% [10.1-13.0] 2 9 2 7 42.9% [15.8-75.0] 2 33.3% [6.1-79.2] 22.2% [6.3-54.7] 3 2 1 1 0.0% [0.0-79.3] 0 n/a 0.0% [0.0-65.8] CONFOUND CHECK: age of the baseline Overall rating (days, at baseline date) A stale rating is more inspection-due, which would inflate P(re-rated) independently of quality. Arms should be comparable on this. T median 1029 d mean 962 d (n=803) A median 1140 d mean 1079 d (n=10492) B median 1106 d mean 1037 d (n=1054) A_clean median 1143 d mean 1084 d (n=9430) T_pure median 1031 d mean 965 d (n=800) AGE-ADJUSTED COMPARISON (Mantel-Haenszel, stratified by baseline rating-age quintile computed across the whole eligible cohort). This removes the 'how overdue for inspection were you' channel from the comparison. quintile cut points (days): 387, 924, 1149, 1416 P(fall) uncond T vs A MH RR = 1.56 [1.29-1.88] P(re-rated|surv) T vs A MH RR = 1.14 [1.04-1.24] P(fall|re-rated) T vs A MH RR = 1.38 [1.17-1.62] P(fall) uncond T vs B MH RR = 1.30 [1.01-1.68] P(re-rated|surv) T vs B MH RR = 0.92 [0.82-1.02] P(fall|re-rated) T vs B MH RR = 1.43 [1.15-1.79] P(fall) uncond T_pure vs B MH RR = 1.29 [1.00-1.67] P(fall) uncond B vs A_clean MH RR = 1.20 [0.99-1.46] P(fall) uncond T_pure vs A_clean MH RR = 1.58 [1.31-1.92] SENSITIVITY: 'fell into RI or Inadequate' rather than any rank drop (strips Outstanding -> Good moves, which are not the commercial outcome) T 106/803 = 13.2% [11.0-15.7] A 884/10492 = 8.4% [7.9-9.0] B 107/1054 = 10.2% [8.5-12.1] A_clean 776/9430 = 8.2% [7.7-8.8] T_pure 105/800 = 13.1% [11.0-15.6] T vs A RR=1.57 [1.30-1.89] p=1.3e-5 T vs B RR=1.30 [1.01-1.67] p=0.0470 COMMERCIAL FRAMING: lift and coverage against the whole sellable pool Pool = every English care home rated Good or Outstanding overall at baseline. A consultant buying a list wants to know: how much denser in future falls is this list than the pool, and how many of the pool's falls does it contain? pool: n=11880, fell=1102 (9.3%) over 2022-07-01 -> 2026-07-01 FirstFlag treatment list list n= 803 ( 6.8% of pool) falls in list= 106 ( 9.6% of all falls) hit rate= 13.2% lift=1.42x Control B list list n= 1054 ( 8.9% of pool) falls in list= 107 ( 9.7% of all falls) hit rate= 10.2% lift=1.09x any 1+ key question RI/Inad list n= 1865 (15.7% of pool) falls in list= 214 (19.4% of all falls) hit rate= 11.5% lift=1.24x EXAMPLE FALLS IN THE TREATMENT ARM (first five, verifiable on cqc.org.uk) 1-10490141472 Pine Lodge: Good -> Requires improvement 1-10644978940 St Mary's: Good -> Requires improvement 1-107300355 Alandale Residential Home: Good -> Requires improvement 1-108984992 Grasmere Nursing Home: Good -> Requires improvement 1-109100588 Beech Hill Grange: Good -> Requires improvement ==================================================================================================== WINDOW 2023-07-03 -> 2026-07-01 ==================================================================================================== Baseline care-home locations in file: 14391; eligible after region + rankable-overall filter: 14388 excluded: region not English 1, no Overall row 1, Overall unrankable 1 data checks: publication date moved backwards 0; same pub date but different rating 1 ARM COUNTS AND OUTCOMES arm n closed gone surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond rose P(rose) uncond --------------------------------------------------------------------------------------------------------------------------------------------------- T 825 124 (15.0%) 2 699 209 29.9% [26.6-33.4] 76 35.9% [29.7-42.6] 9.2% [7.4-11.4] 3 0.4% [0.1-1.1] A 10339 1360 (13.2%) 9 8970 2502 27.9% [27.0-28.8] 659 26.3% [24.7-28.1] 6.4% [5.9-6.9] 55 0.5% [0.4-0.7] B 966 131 (13.6%) 0 835 290 34.7% [31.6-38.0] 86 29.7% [24.7-35.2] 8.9% [7.3-10.9] 4 0.4% [0.2-1.1] A_clean 9354 1227 (13.1%) 9 8118 2206 27.2% [26.2-28.2] 569 25.8% [24.0-27.7] 6.1% [5.6-6.6] 51 0.5% [0.4-0.7] T_pure 820 122 (14.9%) 2 696 209 30.0% [26.7-33.5] 75 35.9% [29.7-42.6] 9.1% [7.4-11.3] 3 0.4% [0.1-1.1] ALL 14388 2084 (14.5%) 13 12291 4093 33.3% [32.5-34.1] 834 20.4% [19.1-21.6] 5.8% [5.4-6.2] 892 6.2% [5.8-6.6] 'closed' = absent at endpoint, found in Deactivated Locations. 'gone' = absent at endpoint and NOT in Deactivated (still registered but no longer carries a published rating). 'surv' = present at endpoint. Percentages in brackets are 95% Wilson intervals. SUPPORTING DETAIL PER ARM T: TREATMENT Overall Good/Outstanding + Well-led RI/Inadequate n=825 survived=699 closed=124 (in-window 124, end-date before baseline 0) unreconciled-absent=2 overall re-rated=209 any-domain re-rated=209 fell=76 (of which to RI 69, to Inadequate 7) rose=3 unchanged=616 endpoint-unrankable=4 among re-rated: fell=75 rose=3 same=127 A: CONTROL A Overall Good/Outstanding + Well-led Good n=10339 survived=8970 closed=1360 (in-window 1359, end-date before baseline 1) unreconciled-absent=9 overall re-rated=2502 any-domain re-rated=2503 fell=659 (of which to RI 599, to Inadequate 52) rose=55 unchanged=8221 endpoint-unrankable=35 among re-rated: fell=659 rose=55 same=1753 B: CONTROL B Overall Good + Well-led Good + exactly 1 other domain RI n=966 survived=835 closed=131 (in-window 131, end-date before baseline 0) unreconciled-absent=0 overall re-rated=290 any-domain re-rated=290 fell=86 (of which to RI 77, to Inadequate 9) rose=4 unchanged=742 endpoint-unrankable=3 among re-rated: fell=86 rose=4 same=197 A_clean: supp. A_clean Overall Good/Outstanding + Well-led Good + no other RI/Inad n=9354 survived=8118 closed=1227 (in-window 1226, end-date before baseline 1) unreconciled-absent=9 overall re-rated=2206 any-domain re-rated=2207 fell=569 (of which to RI 518, to Inadequate 43) rose=51 unchanged=7466 endpoint-unrankable=32 among re-rated: fell=569 rose=51 same=1554 T_pure: supp. T_pure Overall Good + Well-led RI + no other RI/Inad (1 RI, it is Well-led) n=820 survived=696 closed=122 (in-window 122, end-date before baseline 0) unreconciled-absent=2 overall re-rated=209 any-domain re-rated=209 fell=75 (of which to RI 68, to Inadequate 7) rose=3 unchanged=614 endpoint-unrankable=4 among re-rated: fell=75 rose=3 same=127 HEAD-TO-HEAD COMPARISONS (risk ratio [95% CI], Fisher exact two-tailed p) P(fall) uncond T vs A 76/825 vs 659/10339 RR=1.45 [1.15-1.81] p=0.0027 P(re-rated|surv) T vs A 209/699 vs 2502/8970 RR=1.07 [0.95-1.21] p=0.2559 P(fall|re-rated) T vs A 75/209 vs 659/2502 RR=1.36 [1.12-1.65] p=0.0035 P(closed) T vs A 124/825 vs 1360/10339 RR=1.14 [0.96-1.35] p=0.1355 P(rose) uncond T vs A 3/825 vs 55/10339 RR=0.68 [0.21-2.18] p=0.7992 P(fall) uncond T vs B 76/825 vs 86/966 RR=1.03 [0.77-1.39] p=0.8688 P(re-rated|surv) T vs B 209/699 vs 290/835 RR=0.86 [0.74-1.00] p=0.0488 P(fall|re-rated) T vs B 75/209 vs 86/290 RR=1.21 [0.94-1.56] p=0.1467 P(closed) T vs B 124/825 vs 131/966 RR=1.11 [0.88-1.39] p=0.3788 P(fall) uncond T_pure vs B (1 RI each) 75/820 vs 86/966 RR=1.03 [0.76-1.38] p=0.8686 P(re-rated|surv) T_pure vs B (1 RI each) 209/696 vs 290/835 RR=0.86 [0.75-1.00] p=0.0553 P(fall|re-rated) T_pure vs B (1 RI each) 75/209 vs 86/290 RR=1.21 [0.94-1.56] p=0.1467 P(fall) uncond B vs A_clean 86/966 vs 569/9354 RR=1.46 [1.18-1.82] p=0.0011 P(fall) uncond T_pure vs A_clean 75/820 vs 569/9354 RR=1.50 [1.19-1.89] p=0.0010 WHICH SINGLE DOMAIN? (Good overall, exactly one domain at RI, none Inadequate) Holds 'number of RI domains' fixed at one and varies only which domain it is. This is the decisive test of whether Well-led is special. sole RI domain n closed surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ---------------------------------------------------------------------------------------------------------------- safe 527 68 459 152 33.1% [29.0-37.5] 44 28.9% [22.3-36.6] 8.3% [6.3-11.0] effective 237 38 199 72 36.2% [29.8-43.1] 18 25.0% [16.4-36.1] 7.6% [4.9-11.7] caring 28 4 24 12 50.0% [31.4-68.6] 6 50.0% [25.4-74.6] 21.4% [10.2-39.5] responsive 174 21 153 54 35.3% [28.2-43.1] 18 33.3% [22.2-46.6] 10.3% [6.6-15.8] well-led 820 122 696 209 30.0% [26.7-33.5] 75 35.9% [29.7-42.6] 9.1% [7.4-11.3] Pooled: sole RI = Well-led vs sole RI = any of the other four P(fall) uncond 75/820 vs 86/966 RR=1.03 [0.76-1.38] p=0.8686 P(re-rated|surv) 209/696 vs 290/835 RR=0.86 [0.75-1.00] p=0.0553 P(fall|re-rated) 75/209 vs 86/290 RR=NaN [NaN-NaN] p=0.0e+0 DOSE-RESPONSE: number of key questions at RI/Inadequate at baseline (Good or Outstanding overall homes only). If the signal is really a count rather than a specific domain, P(fall) rises monotonically down this table. # RI/Inad KQs n closed surv P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ------------------------------------------------------------------------------------------------------ 0 9922 1272 8641 26.8% [25.9-27.8] 637 27.5% [25.7-29.3] 6.4% [6.0-6.9] 1 1786 253 1531 32.6% [30.3-35.0] 161 32.3% [28.3-36.5] 9.0% [7.8-10.4] 2 20 3 17 29.4% [13.3-53.1] 5 80.0% [37.6-96.4] 25.0% [11.2-46.9] 3 4 1 3 33.3% [6.1-79.2] 0 0.0% [0.0-79.3] 0.0% [0.0-49.0] CONFOUND CHECK: age of the baseline Overall rating (days, at baseline date) A stale rating is more inspection-due, which would inflate P(re-rated) independently of quality. Arms should be comparable on this. T median 1245 d mean 1059 d (n=825) A median 1370 d mean 1212 d (n=10339) B median 1328 d mean 1182 d (n=966) A_clean median 1377 d mean 1217 d (n=9354) T_pure median 1250 d mean 1064 d (n=820) AGE-ADJUSTED COMPARISON (Mantel-Haenszel, stratified by baseline rating-age quintile computed across the whole eligible cohort). This removes the 'how overdue for inspection were you' channel from the comparison. quintile cut points (days): 282, 899, 1396, 1686 P(fall) uncond T vs A MH RR = 1.48 [1.17-1.86] P(re-rated|surv) T vs A MH RR = 1.17 [1.04-1.32] P(fall|re-rated) T vs A MH RR = 1.26 [1.04-1.52] P(fall) uncond T vs B MH RR = 1.03 [0.77-1.39] P(re-rated|surv) T vs B MH RR = 0.93 [0.80-1.07] P(fall|re-rated) T vs B MH RR = 1.12 [0.87-1.44] P(fall) uncond T_pure vs B MH RR = 1.03 [0.76-1.38] P(fall) uncond B vs A_clean MH RR = 1.45 [1.16-1.80] P(fall) uncond T_pure vs A_clean MH RR = 1.54 [1.22-1.95] SENSITIVITY: 'fell into RI or Inadequate' rather than any rank drop (strips Outstanding -> Good moves, which are not the commercial outcome) T 76/825 = 9.2% [7.4-11.4] A 651/10339 = 6.3% [5.8-6.8] B 86/966 = 8.9% [7.3-10.9] A_clean 561/9354 = 6.0% [5.5-6.5] T_pure 75/820 = 9.1% [7.4-11.3] T vs A RR=1.46 [1.17-1.84] p=0.0020 T vs B RR=1.03 [0.77-1.39] p=0.8688 COMMERCIAL FRAMING: lift and coverage against the whole sellable pool Pool = every English care home rated Good or Outstanding overall at baseline. A consultant buying a list wants to know: how much denser in future falls is this list than the pool, and how many of the pool's falls does it contain? pool: n=11732, fell=803 (6.8%) over 2023-07-03 -> 2026-07-01 FirstFlag treatment list list n= 825 ( 7.0% of pool) falls in list= 76 ( 9.5% of all falls) hit rate= 9.2% lift=1.35x Control B list list n= 966 ( 8.2% of pool) falls in list= 86 (10.7% of all falls) hit rate= 8.9% lift=1.30x any 1+ key question RI/Inad list n= 1810 (15.4% of pool) falls in list= 166 (20.7% of all falls) hit rate= 9.2% lift=1.34x EXAMPLE FALLS IN THE TREATMENT ARM (first five, verifiable on cqc.org.uk) 1-1028986728 Ambleside: Good -> Requires improvement 1-10758956076 Birch Abbey: Good -> Requires improvement 1-1077845612 The Moreton Centre: Good -> Requires improvement 1-10881303514 Bentley Court Care Home: Good -> Requires improvement 1-109100588 Beech Hill Grange: Good -> Requires improvement ==================================================================================================== WINDOW 2024-07-01 -> 2026-07-01 ==================================================================================================== Baseline care-home locations in file: 14102; eligible after region + rankable-overall filter: 14094 excluded: region not English 6, no Overall row 1, Overall unrankable 1 data checks: publication date moved backwards 0; same pub date but different rating 1 ARM COUNTS AND OUTCOMES arm n closed gone surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond rose P(rose) uncond --------------------------------------------------------------------------------------------------------------------------------------------------- T 832 91 (10.9%) 2 739 164 22.2% [19.3-25.3] 59 35.4% [28.5-42.9] 7.1% [5.5-9.0] 3 0.4% [0.1-1.1] A 9971 940 (9.4%) 7 9024 1623 18.0% [17.2-18.8] 477 29.4% [27.2-31.7] 4.8% [4.4-5.2] 41 0.4% [0.3-0.6] B 889 76 (8.5%) 0 813 199 24.5% [21.6-27.5] 69 34.7% [28.4-41.5] 7.8% [6.2-9.7] 4 0.4% [0.2-1.2] A_clean 9062 862 (9.5%) 7 8193 1418 17.3% [16.5-18.1] 404 28.5% [26.2-30.9] 4.5% [4.1-4.9] 37 0.4% [0.3-0.6] T_pure 828 90 (10.9%) 2 736 164 22.3% [19.4-25.4] 58 35.4% [28.5-42.9] 7.0% [5.5-8.9] 3 0.4% [0.1-1.1] ALL 14094 1459 (10.4%) 10 12625 3036 24.0% [23.3-24.8] 616 20.3% [18.9-21.7] 4.4% [4.0-4.7] 831 5.9% [5.5-6.3] 'closed' = absent at endpoint, found in Deactivated Locations. 'gone' = absent at endpoint and NOT in Deactivated (still registered but no longer carries a published rating). 'surv' = present at endpoint. Percentages in brackets are 95% Wilson intervals. SUPPORTING DETAIL PER ARM T: TREATMENT Overall Good/Outstanding + Well-led RI/Inadequate n=832 survived=739 closed=91 (in-window 89, end-date before baseline 2) unreconciled-absent=2 overall re-rated=164 any-domain re-rated=164 fell=59 (of which to RI 52, to Inadequate 7) rose=3 unchanged=672 endpoint-unrankable=5 among re-rated: fell=58 rose=3 same=98 A: CONTROL A Overall Good/Outstanding + Well-led Good n=9971 survived=9024 closed=940 (in-window 937, end-date before baseline 3) unreconciled-absent=7 overall re-rated=1623 any-domain re-rated=1625 fell=477 (of which to RI 422, to Inadequate 52) rose=41 unchanged=8469 endpoint-unrankable=37 among re-rated: fell=477 rose=41 same=1068 B: CONTROL B Overall Good + Well-led Good + exactly 1 other domain RI n=889 survived=813 closed=76 (in-window 76, end-date before baseline 0) unreconciled-absent=0 overall re-rated=199 any-domain re-rated=199 fell=69 (of which to RI 60, to Inadequate 9) rose=4 unchanged=737 endpoint-unrankable=3 among re-rated: fell=69 rose=4 same=123 A_clean: supp. A_clean Overall Good/Outstanding + Well-led Good + no other RI/Inad n=9062 survived=8193 closed=862 (in-window 859, end-date before baseline 3) unreconciled-absent=7 overall re-rated=1418 any-domain re-rated=1419 fell=404 (of which to RI 358, to Inadequate 43) rose=37 unchanged=7718 endpoint-unrankable=34 among re-rated: fell=404 rose=37 same=943 T_pure: supp. T_pure Overall Good + Well-led RI + no other RI/Inad (1 RI, it is Well-led) n=828 survived=736 closed=90 (in-window 88, end-date before baseline 2) unreconciled-absent=2 overall re-rated=164 any-domain re-rated=164 fell=58 (of which to RI 51, to Inadequate 7) rose=3 unchanged=670 endpoint-unrankable=5 among re-rated: fell=58 rose=3 same=98 HEAD-TO-HEAD COMPARISONS (risk ratio [95% CI], Fisher exact two-tailed p) P(fall) uncond T vs A 59/832 vs 477/9971 RR=1.48 [1.14-1.92] p=0.0046 P(re-rated|surv) T vs A 164/739 vs 1623/9024 RR=1.23 [1.07-1.42] p=0.0055 P(fall|re-rated) T vs A 58/164 vs 477/1623 RR=1.20 [0.97-1.50] p=0.1280 P(closed) T vs A 91/832 vs 940/9971 RR=1.16 [0.95-1.42] p=0.1577 P(rose) uncond T vs A 3/832 vs 41/9971 RR=0.88 [0.27-2.83] p=1.0000 P(fall) uncond T vs B 59/832 vs 69/889 RR=0.91 [0.65-1.28] p=0.6460 P(re-rated|surv) T vs B 164/739 vs 199/813 RR=0.91 [0.76-1.09] p=0.3075 P(fall|re-rated) T vs B 58/164 vs 69/199 RR=1.02 [0.77-1.35] p=0.9123 P(closed) T vs B 91/832 vs 76/889 RR=1.28 [0.96-1.71] p=0.1032 P(fall) uncond T_pure vs B (1 RI each) 58/828 vs 69/889 RR=0.90 [0.64-1.26] p=0.5805 P(re-rated|surv) T_pure vs B (1 RI each) 164/736 vs 199/813 RR=0.91 [0.76-1.09] p=0.3366 P(fall|re-rated) T_pure vs B (1 RI each) 58/164 vs 69/199 RR=1.02 [0.77-1.35] p=0.9123 P(fall) uncond B vs A_clean 69/889 vs 404/9062 RR=1.74 [1.36-2.23] p=4.5e-5 P(fall) uncond T_pure vs A_clean 58/828 vs 404/9062 RR=1.57 [1.20-2.05] p=0.0018 WHICH SINGLE DOMAIN? (Good overall, exactly one domain at RI, none Inadequate) Holds 'number of RI domains' fixed at one and varies only which domain it is. This is the decisive test of whether Well-led is special. sole RI domain n closed surv re-rated P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ---------------------------------------------------------------------------------------------------------------- safe 499 40 459 107 23.3% [19.7-27.4] 36 33.6% [25.4-43.0] 7.2% [5.3-9.8] effective 212 19 193 50 25.9% [20.2-32.5] 15 30.0% [19.1-43.8] 7.1% [4.3-11.3] caring 21 2 19 6 31.6% [15.4-54.0] 5 83.3% [43.6-97.0] 23.8% [10.6-45.1] responsive 157 15 142 36 25.4% [18.9-33.1] 13 36.1% [22.5-52.4] 8.3% [4.9-13.7] well-led 828 90 736 164 22.3% [19.4-25.4] 58 35.4% [28.5-42.9] 7.0% [5.5-8.9] Pooled: sole RI = Well-led vs sole RI = any of the other four P(fall) uncond 58/828 vs 69/889 RR=0.90 [0.64-1.26] p=0.5805 P(re-rated|surv) 164/736 vs 199/813 RR=0.91 [0.76-1.09] p=0.3366 P(fall|re-rated) 58/164 vs 69/199 RR=NaN [NaN-NaN] p=0.0e+0 DOSE-RESPONSE: number of key questions at RI/Inadequate at baseline (Good or Outstanding overall homes only). If the signal is really a count rather than a specific domain, P(fall) rises monotonically down this table. # RI/Inad KQs n closed surv P(re-rated|surv) fell P(fall|re-rated) P(fall) uncond ------------------------------------------------------------------------------------------------------ 0 9624 902 8715 17.2% [16.4-18.0] 452 30.2% [27.9-32.5] 4.7% [4.3-5.1] 1 1718 166 1550 23.4% [21.4-25.6] 127 35.0% [30.3-40.0] 7.4% [6.2-8.7] 2 19 2 17 29.4% [13.3-53.1] 5 80.0% [37.6-96.4] 26.3% [11.8-48.8] 3 4 1 3 33.3% [6.1-79.2] 0 0.0% [0.0-79.3] 0.0% [0.0-49.0] CONFOUND CHECK: age of the baseline Overall rating (days, at baseline date) A stale rating is more inspection-due, which would inflate P(re-rated) independently of quality. Arms should be comparable on this. T median 1205 d mean 1219 d (n=832) A median 1634 d mean 1390 d (n=9971) B median 1586 d mean 1383 d (n=889) A_clean median 1636 d mean 1393 d (n=9062) T_pure median 1215 d mean 1221 d (n=828) AGE-ADJUSTED COMPARISON (Mantel-Haenszel, stratified by baseline rating-age quintile computed across the whole eligible cohort). This removes the 'how overdue for inspection were you' channel from the comparison. quintile cut points (days): 459, 866, 1655, 1963 P(fall) uncond T vs A MH RR = 1.51 [1.16-1.97] P(re-rated|surv) T vs A MH RR = 1.31 [1.14-1.51] P(fall|re-rated) T vs A MH RR = 1.12 [0.90-1.40] P(fall) uncond T vs B MH RR = 0.92 [0.66-1.30] P(re-rated|surv) T vs B MH RR = 0.95 [0.79-1.14] P(fall|re-rated) T vs B MH RR = 0.95 [0.72-1.27] P(fall) uncond T_pure vs B MH RR = 0.91 [0.65-1.29] P(fall) uncond B vs A_clean MH RR = 1.72 [1.35-2.20] P(fall) uncond T_pure vs A_clean MH RR = 1.60 [1.23-2.10] SENSITIVITY: 'fell into RI or Inadequate' rather than any rank drop (strips Outstanding -> Good moves, which are not the commercial outcome) T 59/832 = 7.1% [5.5-9.0] A 474/9971 = 4.8% [4.4-5.2] B 69/889 = 7.8% [6.2-9.7] A_clean 401/9062 = 4.4% [4.0-4.9] T_pure 58/828 = 7.0% [5.5-8.9] T vs A RR=1.49 [1.15-1.94] p=0.0044 T vs B RR=0.91 [0.65-1.28] p=0.6460 COMMERCIAL FRAMING: lift and coverage against the whole sellable pool Pool = every English care home rated Good or Outstanding overall at baseline. A consultant buying a list wants to know: how much denser in future falls is this list than the pool, and how many of the pool's falls does it contain? pool: n=11365, fell=584 (5.1%) over 2024-07-01 -> 2026-07-01 FirstFlag treatment list list n= 832 ( 7.3% of pool) falls in list= 59 (10.1% of all falls) hit rate= 7.1% lift=1.38x Control B list list n= 889 ( 7.8% of pool) falls in list= 69 (11.8% of all falls) hit rate= 7.8% lift=1.51x any 1+ key question RI/Inad list n= 1741 (15.3% of pool) falls in list= 132 (22.6% of all falls) hit rate= 7.6% lift=1.48x EXAMPLE FALLS IN THE TREATMENT ARM (first five, verifiable on cqc.org.uk) 1-1028986728 Ambleside: Good -> Requires improvement 1-10758956076 Birch Abbey: Good -> Requires improvement 1-1077845612 The Moreton Centre: Good -> Requires improvement 1-10881303514 Bentley Court Care Home: Good -> Requires improvement 1-109100588 Beech Hill Grange: Good -> Requires improvement ==================================================================================================== CROSS-WINDOW NOTES ==================================================================================================== The three windows share one endpoint snapshot and largely the same care homes, so they are NOT independent replications and must not be pooled. Treatment-arm membership overlap between windows (Jaccard on Location ID): T(2022-07-01) n=803 vs T(2023-07-03) n=825: 634 shared, Jaccard 0.64 T(2022-07-01) n=803 vs T(2024-07-01) n=832: 555 shared, Jaccard 0.51 T(2023-07-03) n=825 vs T(2024-07-01) n=832: 727 shared, Jaccard 0.78 HEADLINE, ALL THREE WINDOWS SIDE BY SIDE window T n T P(fall) A n A P(fall) B n B P(fall) RR T/A RR T/B ----------------------------------------------------------------------------------------------------------------- 2022-07-01 803 13.2% 10492 8.6% 1054 10.2% 1.54 [1.27-1.86] 1.30 [1.01-1.67] 2023-07-03 825 9.2% 10339 6.4% 966 8.9% 1.45 [1.15-1.81] 1.03 [0.77-1.39] 2024-07-01 832 7.1% 9971 4.8% 889 7.8% 1.48 [1.14-1.92] 0.91 [0.65-1.28] RE-RATING VOLUME (the power question, Step 7) Single Assessment Framework went live 21 November 2023 and slowed new comprehensive ratings. Actual re-rating counts in the treatment arm: 2022-07-01: 283 of 640 surviving flagged homes were re-rated (44.2%), producing 106 falls. 2023-07-03: 209 of 699 surviving flagged homes were re-rated (29.9%), producing 76 falls. 2024-07-01: 164 of 739 surviving flagged homes were re-rated (22.2%), producing 59 falls. No window is underpowered for the treatment-vs-Control-A comparison. Treatment vs Control B is the tighter test; see its confidence intervals. ==================================================================================================== VERDICT ==================================================================================================== 1. DOES THE WELL-LED FLAG PREDICT A FALL IN OVERALL RATING? YES. Against Control A (Good/Outstanding overall with Well-led Good) the flag raises the unconditional probability of a fall in every one of the three pre-specified windows, by a strikingly stable margin: 2022-07-01 -> 2026-07-01: 13.2% (106/803) vs 8.6% (901/10492) RR 1.54 [1.27-1.86] 2023-07-03 -> 2026-07-01: 9.2% (76/825) vs 6.4% (659/10339) RR 1.45 [1.15-1.81] 2024-07-01 -> 2026-07-01: 7.1% (59/832) vs 4.8% (477/9971) RR 1.48 [1.14-1.92] All three confidence intervals exclude 1. Adjusting for baseline rating age (Mantel-Haenszel, quintiles) does not move it. The flagged homes did not simply close instead: P(closed) is statistically indistinguishable between the arms in all three windows. So the flag is not noise. Caveat that must travel with that yes: part of it is arithmetic, not prophecy. Overall is an aggregation of the five key questions, so a home already carrying one Requires improvement is mechanically one step closer to Requires improvement overall than a home carrying none. That is exactly what Control B was built to test, and it is where the claim breaks. 2. IS IT DISTINGUISHABLE FROM 'ANY SINGLE RI DOMAIN'? NO. Treatment vs Control B (Good overall, Well-led Good, exactly one other domain at Requires improvement): 2022-07-01 -> 2026-07-01: 13.2% (106/803) vs 10.2% (107/1054) RR 1.30 [1.01-1.67] 2023-07-03 -> 2026-07-01: 9.2% (76/825) vs 8.9% (86/966) RR 1.03 [0.77-1.39] 2024-07-01 -> 2026-07-01: 7.1% (59/832) vs 7.8% (69/889) RR 0.91 [0.65-1.28] One window is marginal (2022, p=0.047, CI lower bound 1.01), the other two are flat null, and the 2024 point estimate is BELOW 1. Direction reverses across windows. That is not a signal, that is sampling noise around RR=1. The 'which single domain' table is the clean version of the same test: it holds the number of Requires improvement key questions fixed at exactly one and varies only which one. Well-led does not stand out from Safe, Effective, Caring or Responsive on unconditional P(fall) in two of three windows. The dose-response table shows what is actually driving everything: among Good/Outstanding homes, P(fall) rises with the COUNT of Requires improvement key questions, not with which key question it is. In the 2024 window, 0 bad key questions = 4.7% fall, 1 = 7.4%, 2 = 26.3% (n=19). Conclusion: what FirstFlag is really selling is 'this Good home already has a Requires improvement somewhere'. The Well-led specificity is decoration. The generic version of the list is the same size and works the same. 3. DETERIORATION PREDICTOR OR INSPECTION-SCHEDULE PREDICTOR? BOTH, AND THE MIX FLIPS WITH THE HORIZON. Decomposing the age-adjusted treatment-vs-Control-A effect into its two channels (P(re-rated) x P(fall | re-rated)), on the log scale: 2022-07-01: re-rating channel RR 1.14, deterioration channel RR 1.38 -> 71% of the effect is deterioration 2023-07-03: re-rating channel RR 1.17, deterioration channel RR 1.26 -> 59% of the effect is deterioration 2024-07-01: re-rating channel RR 1.31, deterioration channel RR 1.12 -> 30% of the effect is deterioration Over four years the effect is mostly genuine deterioration: flagged homes that got re-rated fell far more often than control homes that got re-rated (2022 window: 37.5% vs 25.7% of re-rated homes fell). Over two years the effect is mostly the inspection queue: flagged homes were 31% more likely to be re-rated at all, and once re-rated their fall rate was only 12% higher with a confidence interval spanning 1 (0.90-1.40). Practically: on a short horizon the flag mainly tells you CQC is coming back soon. On a long horizon it tells you the home is genuinely more likely to be marked down. Both are sellable to a consultant, but they are different products and the marketing must not conflate them. POWER (Step 7) No window is underpowered for the primary test. The Single Assessment Framework did slow re-rating sharply, and the effect is visible in the data: 2022-07-01 window: 5589 of 11711 surviving care homes re-rated (47.7%); treatment arm 283 of 640. 2023-07-03 window: 4093 of 12291 surviving care homes re-rated (33.3%); treatment arm 209 of 699. 2024-07-01 window: 3036 of 12625 surviving care homes re-rated (24.0%); treatment arm 164 of 739. Re-rating rate roughly halves as the window shortens, which is the SAF slowdown, but even the tightest window has 164 re-rated flagged homes and 59 falls. That is enough to detect the 1.5x effect against Control A, which it does. It is NOT enough to resolve a small true difference against Control B: the 2024 window CI on that comparison is 0.65-1.28, so anything above about a 30% advantage is ruled out, but a 10-20% advantage would not have been detected. That is the honest limit of this backtest: it rules out Well-led being a LOT better than the other domains. It cannot rule out it being slightly better. It gives no positive evidence that it is better at all. THE ONE CLAIM THE NUMBERS SUPPORT Every figure in this sentence is a row count from the two named files: "Of the 803 English care homes that CQC rated Good or Outstanding overall but Requires improvement or Inadequate on Well-led in July 2022, 106 (13.2%) had a lower overall rating by July 2026, against 8.6% (901 of 10,492) of Good or Outstanding homes whose Well-led was Good." What that sentence deliberately does NOT say, because the data will not carry it: that Well-led is a better warning sign than Safe, Effective, Caring or Responsive. It is not, on this evidence. WHAT WOULD SUPPORT THE STRONGER CLAIM: a test with enough events to resolve a 10-20% difference between sole-RI-Well-led and sole-RI-other. That needs roughly 4-6x the current event count, which means stacking many monthly snapshots into a person-time survival model rather than two endpoints, and pulling multiple baseline cohorts back to 2015 where re-rating was frequent. Given the whole population of English care homes is only ~14,000, that is a design problem, not a data-access problem: the archive already has 130+ monthly snapshots and this analysis used four of them.